AIQB
TutorialsOrdinary

Meta's prompt-injection detector caught 1% of real agent attacks. One config change made it 99%. That's the problem.

Source: DEV Community·

Summary

I threw 629 real AgentDojo attacks at 10 open-source prompt-injection detectors — buried inside ordinary tool output, the way an agent firewall actually sees them. Most are smoke alarms that either sleep through the fire or scream at your toast. Then I tuned the thresholds and the whole leaderboard flipped upside down. Here's the reproducible benchmark, and why it means your agent needs something other than a text classifier.
TierOrdinary
Published
Indexed by AIQB
SourceDEV Community
AIQB record IDintel-3123e40560319e0043e87cda