LLM reasoning / chain-of-thought

Self-Refine

Self-Refine: Iterative Refinement with Self-Feedback

Superseded baseline#9 of 772 most-superseded · first seen Mar 30, 2023

Superseded — cited as a baseline and beaten by newer methods

4 papers critique it · 3 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Self-Refine as a baseline.

Though rewriting techniques like Self-Refine~Madaan2023SelfRefineIR can help relieve, this tendency can still lead to misleading or wrong outcomes in real-world complex mathematical problems.
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
rather than relying on post-hoc reflection or global templates
Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
In all of these methods the guidance signal---critique, scoring model, abstraction prompt---is produced by an untrained, prompted LLM.
Co-ReAct: Rubrics as Step-Level Collaborators for ReAct Agents
Sometimes LLMs can directly provide correct answers to questions, but after applying CoT-like methods, it brings extra reasoning paths to models, causing their answers to be wrong.
Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense Reasoning

Beaten on benchmarks

Head-to-head results where a newer method reports beating Self-Refine. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.