Cited as a baseline — critiqued by newer work, not yet beaten on a benchmark here
1 papers critique it · 0 beat it on benchmarks
Least supersededMost
What papers say
Verbatim critique sentences, each from a paper that cites Prompted self-correction as a baseline.
“
our experiments show this baseline achieves only 19.8% on MATH-500 with Llama-3-8B, lower than standard AR (28.8%), consistent with Huang2024, who show that LLMs cannot reliably self-correct without external feedback