Tool use / function calling

LlamaFirewall

LlamaFirewall: An open source guardrail system for building secure AI agents

Superseded baseline#38 of 55 most-superseded · first seen May 6, 2025

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites LlamaFirewall as a baseline.

Unlike existing guardrail framework such as LlamaFirewall, which abort tasks when prompt injection or unsafe behaviors are detected, our approach monitors each tool invocation in real time and provides feedback before execution, guiding safety-aware tool invocation reasoning in LLM-based agents.
ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.