LLM reasoning / chain-of-thought
Math-Shepherd
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Heavily superseded — a standard baseline that newer methods routinely beat
5 papers critique it · 8 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Math-Shepherd as a baseline.
However, its reliance on large-scale sampling makes it computationally expensive and primarily confined to mathematical reasoning.
“Math-Shepherd proposes an automated method for estimating intermediate step correctness using Monte Carlo estimation, though it generates some incorrect labels and demands extensive computational resources.”
“training PRMs requires stepwise human annotation, which is often infeasible for open-source communities”
“these approaches are typically operated and optimized based on the target policy and focus on finding the first error location, which can restrict the versatility of the PRM in evaluating a wide range of policies, experience performance degradation when applied to out-of-distribution (OOD) policies, and reduce usability in optimizing subsequent RL algorithms that require complete process rewards along the trajectory”
“Although this strategy is effective, it remains both computationally expensive and indirect.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Math-Shepherd. Values are copied from the source paper's tables — verify against the cited paper.
GenPRM-7B (Maj@8) beats Math-Shepherd
80.5 vs 31.5
Qwen2.5-Math-PRM-7B beats Math-Shepherd
73.5 vs 31.5
Avg. F1 · [ProcessBench, 7B+ PRMs]
The Lessons of Developing Process Reward Models in Mathematical ReasoningQwen2.5-Math-7B-NAIT beats Math-Shepherd
57.4 vs 31.5
Avg. F1 · [MCE-based PRMs]
Towards Robust Process Reward Modeling via Noise-aware LearningProcess Reward Models for Reflective Mathematical Reasoning beats Math-Shepherd
0.167 vs 0.100
PRM@8-step · [AIME2024]
Beyond the First Error: Process Reward Models for Reflective Mathematical ReasoningGroundedPRM beats Math-Shepherd
39.7 vs 31.5
F1 score · [auto-labeled supervision at 40K samples]
GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Jun 3, 2026
- May 2, 2026
- Apr 19, 2026
- DC-W2SDC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological ReasoningMar 9, 2026
- Feb 9, 2026
- Jan 29, 2026
- Noise-Aware Iterative Training (NAIT)Towards Robust Process Reward Modeling via Noise-aware LearningJan 19, 2026
- GroundedPRMGroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level ReasoningOct 16, 2025