LLM reasoning / chain-of-thought
Outcome Reward Models
Superseded baseline#92 of 772 most-superseded
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
2 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Outcome Reward Models as a baseline.
Traditional outcome-only verifiers (Outcome Reward Models) are limited, evaluating only the final answer and often missing intermediate errors that compromise the reasoning trajectory Wang2024.
“most applications of RLVR to date focus on outcome-level verification, where the model receives a scalar reward only if the final answer is correct. While such outcome rewards improve overall performance, they provide little guidance on the internal reasoning process itself. As a result, a model may reach a correct conclusion through unsound, inconsistent, or opaque reasoning traces, limiting interpretability and trustworthiness.”
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.