Living systematic review
Tool use / function calling
Teaching LLMs to call external tools and APIs — function-calling, tool selection/retrieval, and tool-augmented agents.
52 papers79 critique receipts186 benchmark resultsupdated Jun 18, 2026
Most-superseded baselines
Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.
- 1ReAct
ReAct: Synergizing Reasoning and Acting in Language Models
6 critique · 4 beaten on benchmarks
- 2ToolLLM
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
4 critique · 3 beaten on benchmarks
- 3ToolACEin ReAct
ToolACE: Winning the Points of LLM Function Calling
2 critique · 4 beaten on benchmarks
- 4API-Bankin ReAct
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
5 critique · 0 beaten on benchmarks
- 5Gorillain ToolLLM
Gorilla: Large Language Model Connected with Massive APIs
3 critique · 1 beaten on benchmarks
- 6GRPO
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
2 critique · 1 beaten on benchmarks
- 7StepTool
2 critique · 1 beaten on benchmarks
- 8Toolformer
Toolformer: Language Models Can Teach Themselves to Use Tools
2 critique · 1 beaten on benchmarks
- 9APIGenin ReAct
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
3 critique · 0 beaten on benchmarks
- 10xLAMin ReAct
xLAM: A Family of Large Action Models to Empower AI Agent Systems
0 critique · 2 beaten on benchmarks
- 11ART
ART: Automatic multi-step reasoning and tool-use for large language models
1 critique · 1 beaten on benchmarks
- 12ExpeLin ART
ExpeL: LLM Agents Are Experiential Learners
1 critique · 1 beaten on benchmarks
The competition
Methods that fight on the same benchmarks cluster into distinct sub-problems.
ToolLLM10 methods
ToolLLM · Gorilla · ToolAlpaca · Less-is-More · TinyAgent · ToolPlanner
Probe&Prefill7 methods
Probe&Prefill · When2Tool / ToolReadable · Tool-identity steering · Tool-identity · NexusRaven · Functionary
AgentAuditor6 methods
AgentAuditor · AGrail · GuardAgent · LlamaFirewall · ShieldAgent · ToolSafe
GRPO5 methods
GRPO · SAGE · Reflexion · Reinforced Agent · RC-GRPO
Toolformer4 methods
Toolformer · CAST · CostBench · ToolAlign
ART3 methods
The frontier
Recent methods not yet superseded in the knowledge base.
- May 28, 2026
- May 27, 2026
- May 20, 2026
- May 14, 2026
- May 8, 2026
- May 8, 2026
- Apr 29, 2026
- Apr 22, 2026
- Apr 22, 2026
- Apr 20, 2026
- Apr 19, 2026
- Apr 10, 2026