LLM reasoning / chain-of-thought

visualprm

VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Superseded baseline#752 of 772 most-superseded · first seen Mar 13, 2025

Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here

1 papers critique it · 0 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites visualprm as a baseline.

visualprm relied primarily on MC-score-based dataset construction, which strongly influenced PRM performance. Our central question is whether superior performance can be achieved by using stronger VLMs to judge each reasoning step, and subsequently employing these judgments as supervision signals for training VL-PRMs.
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.