Tool use / function calling
ReTool
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Superseded baseline#15 of 55 most-superseded · first seen Apr 15, 2025
Superseded — cited as a baseline and beaten by newer methods
1 papers critique it · 1 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites ReTool as a baseline.
While these works advance our understanding of exploration in language models, none simultaneously learn exploration policies that select among heterogeneous tools while rewarding both answer quality and trajectory diversity.
Beaten on benchmarks
Head-to-head results where a newer method reports beating ReTool. Values are copied from the source paper's tables — verify against the cited paper.
CoCoDA beats ReTool
11.24 vs 10.55
FinQA · [0.6B student]
CoCoDA: Co-evolving Compositional DAG for Tool-Augmented Agents
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 8, 2026