LLM reasoning / chain-of-thought

PRM

Let's Verify Step by Step

Superseded baseline#6 of 772 most-superseded · first seen May 31, 2023

Superseded — cited as a baseline and beaten by newer methods

7 papers critique it · 3 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites PRM as a baseline.

although PRMs can validate the output of LLMs more accurately than ORMs, they often require high quality annotated data
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
However, key challenges remain, such as the difficulty of obtaining high-quality labels and the limited effectiveness of current PRM approaches
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
In our experiments, we find, however, that even state-of-the-art PRMs can be miscalibrated, assigning overly optimistic scores---particularly on challenging, out-of-distribution problems.
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
PRMs poorly approximate state values and reliability degrades with reasoning depth, suggesting credit assignment issues
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
Process Reward Models (PRMs) are bound by prohibitive annotation costs, while verifier-free proxies frequently yield sparse signals that lack awareness of the intermediate reasoning process.
Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language Models
PRM-based approaches require substantial computational resources for training step-level reward models and conducting multi-step inference processes
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
obtaining high-quality step-by-step annotations is challenging: current efforts relying on human annotation, Monte Carlo sampling, or LLM-as-a-judge are either costly or noisy.
Uncertainty-Aware Step-wise Verification with Generative Reward Models

Beaten on benchmarks

Head-to-head results where a newer method reports beating PRM. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.