Retrieval-augmented generation
MMGraphRAG
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
Superseded baseline#52 of 1,179 most-superseded · first seen Jul 28, 2025
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 2 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites MMGraphRAG as a baseline.
such approaches typically convert visual content into textual graph nodes via MLLMs, effectively reducing multimodal structure to text-centric representations. As a result, fine-grained visual evidence may be abstracted away, limiting faithful cross-modal reasoning.
“MMGraphRAG links scene graphs with textual representations but suffers from structural blindness—treating tables and formulas as plain text without proper entity extraction, losing structural information for reasoning”
Beaten on benchmarks
Head-to-head results where a newer method reports beating MMGraphRAG. Values are copied from the source paper's tables — verify against the cited paper.
MG²-RAG beats MMGraphRAG
38.15 vs 0.39
InfoSeek Unseen-E · [Qwen2.5-VL-7B backbone]
MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented GenerationRAGAnything beats MMGraphRAG
42.8 vs 37.7
Accuracy · [Overall]
RAG-Anything: All-in-One RAG Framework
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Apr 7, 2026
- Graph-to-Frame RAGGraph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video ReasoningApr 6, 2026
- Apr 4, 2026
- AutoThinkRAGAutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text InteractionMar 17, 2026
- Feb 27, 2026
- VimRAGVimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory GraphFeb 13, 2026
- Feb 5, 2026
- Feb 1, 2026
- Oct 8, 2025