Retrieval-augmented generation
IRCoT
Superseded — cited as a baseline and beaten by newer methods
4 papers critique it · 19 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites IRCoT as a baseline.
While these approaches improve evidence coverage, multi-round retrieval introduces unavoidable computational overhead, limiting their use in latency-sensitive applications.
“Since the LLMs are prone to hallucinate, the generated CoT sentences may be inaccurate~luo2023reasoning,nguyen2024direct and lead to suboptimal performance.”
“However, these models all rely on LLM-generated thoughts, making them prone to hallucination.”
“While suitable for multi-step reasoning tasks, its rigid structure may limit performance in scenarios requiring parallel information aggregation.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating IRCoT. Values are copied from the source paper's tables — verify against the cited paper.
GlobalRAG beats IRCoT
6.63 vs 0.09
Avg F1 · [Qwen2.5-14B-Instruct]
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level ReasoningCoE-8B (Phase II) beats IRCoT
87.5 vs 28.5
Chain-Acc · [SlideVQA (Complex Layouts)]
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented GenerationEVO-RAG beats IRCoT
51.8 vs 19.9
MuSiQue EM · [Flan-T5-XXL backbone]
Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented GenerationTSSS beats IRCoT
14.5 vs 6.5
LcRL beats IRCoT
40.6 vs 20.3
fEM · [Qwen3-4B, MKQA]
Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented GenerationGFM-RAG beats IRCoT
0.060 vs 3.441
Time (s) · [2Wiki]
GFM-RAG: Graph Foundation Model for Retrieval Augmented GenerationSIM-RAG beats IRCoT
46.1 vs 23.5
EM · [GPT-4/BM25 (Group b, Learned)]
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-PracticingPAR-RAG beats IRCoT
0.60 vs 0.35
ANCHOR beats IRCoT
37.85 vs 22.60
Average EM · [Qwen2.5-7B-Instruct]
Graph-Anchored Knowledge Indexing for Retrieval-Augmented GenerationTaSR-RAG beats IRCoT
35.9 vs 22.6
EM · [Qwen2.5-7B-Instruct]
TaSR-RAG: Taxonomy-guided Structured Reasoning for Retrieval-Augmented GenerationRA-ISF beats IRCoT
46.0 vs 34.0
Avg. · [Llama-2_13b With Retrieval]
RA-ISF: Learning to Answer and Understand from Retrieval Augmentation via Iterative Self-FeedbackInSemRAG beats IRCoT
57.42 vs 44.51
HotpotQA EM · [GPT]
Efficient RAG with Intent-Aware Retrieval and Semantics-Preserving Chunking
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Narrative Knowledge WeaverNarrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text UnderstandingJun 4, 2026
- Jun 4, 2026
- May 30, 2026
- May 27, 2026
- LegalGraphRAGLegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal ReasoningMay 27, 2026
- In-Context Optimization for RAGIn-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent PerspectiveMay 25, 2026
- EfficientGraph-RAGEfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented GenerationMay 25, 2026
- May 22, 2026
- May 12, 2026
- May 7, 2026
- Chain of Evidence (CoE)Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented GenerationMay 2, 2026
- CERTA"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented GenerationMay 1, 2026