Tool use / function calling

Search-R1

Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Superseded baseline#47 of 55 most-superseded · first seen Mar 12, 2025

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Search-R1 as a baseline.

Because the reward is shared across all segments, the contribution of any single tool call is hard to isolate, and unnecessary tool calls on easy questions still receive positive reinforcement when the episode succeeds.
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.