Retrieval-augmented generation
Video-RAG
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
Superseded baseline#55 of 1,179 most-superseded · first seen Nov 20, 2024
Superseded — cited as a baseline and beaten by newer methods
2 papers critique it · 2 beat it on benchmarks
What papers say
Verbatim critique sentences, each from a paper that cites Video-RAG as a baseline.
appending textual context such as ASR, OCR, or descriptions to the prompt
“This contrasts with existing agent- or retrieval-augmented generation-based methods~VideoAgent, Video-Agent, Video-RAG, DrVideo, which rely heavily on external tools for frame-level information extraction, limiting their capacity to respond to diverse queries due to the inherent constraints of these tools.”
Beaten on benchmarks
Head-to-head results where a newer method reports beating Video-RAG. Values are copied from the source paper's tables — verify against the cited paper.
G2F-RAG beats Video-RAG
57.0 vs 48.5
WildVideo · [LLaVA-Video 7B]
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video ReasoningVideoStir beats Video-RAG
60.3 vs 58.7
Overall · [LLaVA-Video (7B)]
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.
- Apr 7, 2026
- Graph-to-Frame RAGGraph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video ReasoningApr 6, 2026
- Apr 4, 2026
- AutoThinkRAGAutothinkRAG: Complexity-Aware Control of Retrieval-Augmented Reasoning for Image-Text InteractionMar 17, 2026
- Feb 27, 2026
- VimRAGVimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory GraphFeb 13, 2026
- Feb 5, 2026
- Feb 1, 2026
- Oct 8, 2025