LLM reasoning / chain-of-thought
visualprm
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Superseded baseline#752 of 772 most-superseded · first seen Mar 13, 2025
Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites visualprm as a baseline.
visualprm relied primarily on MC-score-based dataset construction, which strongly influenced PRM performance. Our central question is whether superior performance can be achieved by using stronger VLMs to judge each reasoning step, and subsequently employing these judgments as supervision signals for training VL-PRMs.
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.