Tool use / function calling
LlamaFirewall
LlamaFirewall: An open source guardrail system for building secure AI agents
Superseded baseline#38 of 55 most-superseded · first seen May 6, 2025
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites LlamaFirewall as a baseline.
Unlike existing guardrail framework such as LlamaFirewall, which abort tasks when prompt injection or unsafe behaviors are detected, our approach monitors each tool invocation in real time and provides feedback before execution, guiding safety-aware tool invocation reasoning in LLM-based agents.
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.