LLM reasoning / chain-of-thought
Self-Refine
Self-Refine: Iterative Refinement with Self-Feedback
Superseded baseline#9 of 772 most-superseded · first seen Mar 30, 2023
Superseded — cited as a baseline and beaten by newer methods
4 papers critique it · 3 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Self-Refine as a baseline.
Though rewriting techniques like Self-Refine~Madaan2023SelfRefineIR can help relieve, this tendency can still lead to misleading or wrong outcomes in real-world complex mathematical problems.
“rather than relying on post-hoc reflection or global templates”
“In all of these methods the guidance signal---critique, scoring model, abstraction prompt---is produced by an untrained, prompted LLM.”
“Sometimes LLMs can directly provide correct answers to questions, but after applying CoT-like methods, it brings extra reasoning paths to models, causing their answers to be wrong.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Self-Refine. Values are copied from the source paper's tables — verify against the cited paper.
RIDERS beats Self-Refine
65.3 vs 52.4
Co-ReAct beats Self-Refine
36.92 vs 34.24
DeepResearchBench Average · [Qwen3-14B]
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct AgentsPASR beats Self-Refine
61.7 vs 57.5
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- May 27, 2026
- Tree-of-ThoughtsTree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design PatternsMay 27, 2026
- May 22, 2026
- May 22, 2026
- Novelty-based Tree-of-Thought SearchNovelty-based Tree-of-Thought Search for LLM Reasoning and PlanningMay 7, 2026
- Decoding-Time Debiasing via Process Reward ModelsDecoding-Time Debiasing via Process Reward Models: From Controlled Fill-in to Open-Ended GenerationMay 4, 2026
- Apr 27, 2026
- Apr 22, 2026
- CoT-PoT ensemblingSelf-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM ReasoningApr 19, 2026
- AtroposAtropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model HotswapApr 16, 2026
- Apr 1, 2026
- Learning When to SampleLearning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought ReasoningMar 17, 2026