LLM reasoning / chain-of-thought
GRPO
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Superseded baseline#10 of 772 most-superseded · first seen Feb 5, 2024
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
5 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites GRPO as a baseline.
methods based on RLVR predominantly rely on GRPO, which provides only coarse, outcome-level supervision and lacks the fine-grained signals necessary for improving complex, step-by-step reasoning
“online RL algorithms represented by GRPO often lead to unsatisfactory training results due to insignificant differences in reward signals”
“This introduces the credit assignment problem, as the scalar signal at step T fails to distinguish the contribution of each token s_t.”
“Inadequate exploration, as independent sampling strategy struggles to produce diverse trajectories with sufficient exploration due to structural inefficiency”
“reward design is intrinsically noisy. Unlike multi-hop search with near-unique references, many practical tool-use tasks admit multiple valid outputs (e.g., recommendations). As a result, outcome-only rewards induce high-variance gradients and provide weak incentives for reasoning, even when augmented with LLM-as-a-judge or learned reward models”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 27, 2026
- Apr 27, 2026
- Jan 14, 2026