教程 / 实战普通
Meta's prompt-injection detector caught 1% of real agent attacks. One config change made it 99%. That's the problem.
内容摘要
I threw 629 real AgentDojo attacks at 10 open-source prompt-injection detectors — buried inside ordinary tool output, the way an agent firewall actually sees them. Most are smoke alarms that either sleep through the fire or scream at your toast. Then I tuned the thresholds and the whole leaderboard flipped upside down. Here's the reproducible benchmark, and why it means your agent needs something other than a text classifier.