Living systematic review
LLM reasoning / chain-of-thought
Eliciting multi-step reasoning from LLMs — chain/tree-of-thought, self-consistency, process reward models, self-refinement, and test-time compute scaling.
391 papers892 critique receipts2,230 benchmark resultsupdated Jun 18, 2026
Most-superseded baselines
Ranked by how many distinct papers critique or beat each method — the standard baselines newer work routinely measures against.
- 1Chain-of-Thought
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
25 critique · 13 beaten on benchmarks
- 2Self-Consistencyin Chain-of-Thought
Self-Consistency Improves Chain of Thought Reasoning in Language Models
24 critique · 9 beaten on benchmarks
- 3ToTin Chain-of-Thought
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
9 critique · 5 beaten on benchmarks
- 4ORM
10 critique · 4 beaten on benchmarks
- 5Math-Shepherdin ORM
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
5 critique · 8 beaten on benchmarks
- 6PRMin Chain-of-Thought
Let's Verify Step by Step
7 critique · 3 beaten on benchmarks
- 7ReAct
ReAct: Synergizing Reasoning and Acting in Language Models
5 critique · 3 beaten on benchmarks
- 8Best-of-Nin Chain-of-Thought
3 critique · 4 beaten on benchmarks
- 9Self-Refinein Chain-of-Thought
Self-Refine: Iterative Refinement with Self-Feedback
4 critique · 3 beaten on benchmarks
- 10GRPO
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
5 critique · 0 beaten on benchmarks
- 11H-CoT
H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking
2 critique · 2 beaten on benchmarks
- 12MCTS
2 critique · 2 beaten on benchmarks
The competition
Methods that fight on the same benchmarks cluster into distinct sub-problems.
Chain-of-Thought178 methods
Chain-of-Thought · Self-Consistency · ToT · PRM · Best-of-N · Self-Refine
ORM99 methods
ORM · Math-Shepherd · EurusPRM-Stage2 · VersaPRM · Monte Carlo estimation · EurusPRM-Stage1
ReAct33 methods
ReAct · Transformer · Block Universal Transformer · Data Interpreter · HRM · AutoGen
MCTS23 methods
MCTS · PRIME · Qwen2.5-7B-Instruct · external CoT monitoring · Integration (post-hoc merging) · Monte Carlo Tree Search (MCTS) / MCTSr
single-teacher CoT distillation18 methods
single-teacher CoT distillation · Mixture-of-Agents and LLM-Blender (runtime ensembles) · MoT (Merge of Thought) · multi-teacher CoT aggregation · parameter merging frameworks · pruning-based CoT refinement
Outcome Reward Models17 methods
Outcome Reward Models · MCTS-based scoring · Neural Process Reward Models · visualprm · Gemma3-27B (Baseline) · Qwen-2.5-VL-32B (Baseline)
Diable16 methods
Diable · LUNA · SPACE-3 · GNNs · RAG and KG prompting · semantic parsing
VL-Rethinker-7B15 methods
VL-Rethinker-7B · LLaVA-CoT · Insight-V · DPO (Direct Preference Optimization) · LLaVA-o1 · PPO (Proximal Policy Optimization)
DIN-SQL14 methods
DIN-SQL · DAIL-SQL GPT-4 · ROUTE Qwen2.5-7B · predict SQL-only Llama-3.1-8B-Instruct · STaR-SQL · DTS-SQL Mistral-7B
The frontier
Recent methods not yet superseded in the knowledge base.
- Jun 9, 2026
- Jun 8, 2026
- Jun 7, 2026
- Jun 3, 2026
- Jun 3, 2026
- May 29, 2026
- May 28, 2026
- May 27, 2026
- May 27, 2026
- May 27, 2026
- May 27, 2026
- May 25, 2026