GuardianAgentBench: Where Agents Fail and How to Guard Them
Provides a rigorous benchmark and demonstrates that execution-time guardrails outperform prompt-based defenses for safe LLM agent deployment.
GuardianAgentBench evaluates LLM agent safety across 580 scenarios, finding that even the best model achieves only 74.8% accuracy, with guardrails recovering 19.9% of failures at a 0.5% false positive rate.
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three production-ready frameworks: LangChain, LlamaIndex, and Vectara. The benchmark incorporates rigorous multi-stage validation and five adversarial attack modes. Experiments with six state-of-the-art models reveal that even the strongest configuration achieves only 74.8% overall accuracy and expose two distinct failure regimes: stronger models under-call required tools, while weaker models mis-select and over-call tools. Performance degrades monotonically with both tool-set size and sequential turn depth, with long-horizon planning proving the steeper bottleneck. Our guardrail implementation consistently outperforms system-prompt-based defenses across all models, recovering 19.9% of failures at a false positive rate of just 0.5%. These results demonstrate that execution-time structural intervention improves safety without disrupting correct agent behavior.