教程 / 实战普通
Your AI guardrail is green. It's also catching nothing.
内容摘要
The scariest security failure isn't the guardrail that's down, or the one that's weak. It's the one that's running, passing every health check, and quietly configured to catch nothing — green by construction. I found one in a benchmark of 629 real agent attacks: a famous prompt-injection model catching 1%, not because it's bad, but because its default threshold was ~50x too high. Here's the anatomy of a guardrail that has no symptom.