Living systematic review

Tool use / function calling

Teaching LLMs to call external tools and APIs — function-calling, tool selection/retrieval, and tool-augmented agents.

52 papers79 critique receipts186 benchmark resultsupdated Jun 18, 2026

Most-superseded baselines

Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.

  1. 1
    ReAct

    ReAct: Synergizing Reasoning and Acting in Language Models

    6 critique · 4 beaten on benchmarks

  2. 2
    ToolLLM

    ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

    4 critique · 3 beaten on benchmarks

  3. 3
    ToolACEin ReAct

    ToolACE: Winning the Points of LLM Function Calling

    2 critique · 4 beaten on benchmarks

  4. 4
    API-Bankin ReAct

    API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

    5 critique · 0 beaten on benchmarks

  5. 5
    Gorillain ToolLLM

    Gorilla: Large Language Model Connected with Massive APIs

    3 critique · 1 beaten on benchmarks

  6. 6
    GRPO

    Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

    2 critique · 1 beaten on benchmarks

  7. 7
    StepTool

    2 critique · 1 beaten on benchmarks

  8. 8
    Toolformer

    Toolformer: Language Models Can Teach Themselves to Use Tools

    2 critique · 1 beaten on benchmarks

  9. 9
    APIGenin ReAct

    APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets

    3 critique · 0 beaten on benchmarks

  10. 10
    xLAMin ReAct

    xLAM: A Family of Large Action Models to Empower AI Agent Systems

    0 critique · 2 beaten on benchmarks

  11. 11
    ART

    ART: Automatic multi-step reasoning and tool-use for large language models

    1 critique · 1 beaten on benchmarks

  12. 12
    ExpeLin ART

    ExpeL: LLM Agents Are Experiential Learners

    1 critique · 1 beaten on benchmarks

The competition

Methods that fight on the same benchmarks cluster into distinct sub-problems.

ReAct24 methods

ReAct · ToolACE · API-Bank · APIGen · xLAM · StableToolBench

ToolLLM10 methods

ToolLLM · Gorilla · ToolAlpaca · Less-is-More · TinyAgent · ToolPlanner

StepTool9 methods

StepTool · ToolRL · CodeAct · Search-R1 · Search-R1 PPO · R2IF

Probe&Prefill7 methods

Probe&Prefill · When2Tool / ToolReadable · Tool-identity steering · Tool-identity · NexusRaven · Functionary

AgentAuditor6 methods

AgentAuditor · AGrail · GuardAgent · LlamaFirewall · ShieldAgent · ToolSafe

GRPO5 methods

GRPO · SAGE · Reflexion · Reinforced Agent · RC-GRPO

Toolformer4 methods

Toolformer · CAST · CostBench · ToolAlign

ART3 methods

ART · ExpeL · Stepwise Experience Recall (SEER)

ReTool4 methods

ReTool · SWiRL · CoCoDA · SPaRK

Mem04 methods

Mem0 · NLSI · PEToolLLM · PRefine

The frontier

Recent methods not yet superseded in the knowledge base.