Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
Least supersededMost
What papers say
Verbatim critique sentences, each from a paper that cites LLaVA-o1 as a baseline.
“
However, these approaches, even advanced models like GPT-5 or Gemini, perform CoT in pure text space. Once visual features are initially encoded, they cannot be re-accessed during reasoning.