LLM reasoning / chain-of-thought

Math-Shepherd

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Heavily superseded#5 of 772 most-superseded · first seen Dec 14, 2023

Heavily superseded — a standard baseline that newer methods routinely beat

5 papers critique it · 8 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Math-Shepherd as a baseline.

However, its reliance on large-scale sampling makes it computationally expensive and primarily confined to mathematical reasoning.
Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs
Math-Shepherd proposes an automated method for estimating intermediate step correctness using Monte Carlo estimation, though it generates some incorrect labels and demands extensive computational resources.
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
training PRMs requires stepwise human annotation, which is often infeasible for open-source communities
Entropy-Regularized Process Reward Model
these approaches are typically operated and optimized based on the target policy and focus on finding the first error location, which can restrict the versatility of the PRM in evaluating a wide range of policies, experience performance degradation when applied to out-of-distribution (OOD) policies, and reduce usability in optimizing subsequent RL algorithms that require complete process rewards along the trajectory
AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification
Although this strategy is effective, it remains both computationally expensive and indirect.
Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards

Beaten on benchmarks

Head-to-head results where a newer method reports beating Math-Shepherd. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.