Retrieval-augmented generation

Video-RAG

Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension

Superseded baseline#55 of 1,179 most-superseded · first seen Nov 20, 2024

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites Video-RAG as a baseline.

appending textual context such as ASR, OCR, or descriptions to the prompt
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
This contrasts with existing agent- or retrieval-augmented generation-based methods~VideoAgent, Video-Agent, Video-RAG, DrVideo, which rely heavily on external tools for frame-level information extraction, limiting their capacity to respond to diverse queries due to the inherent constraints of these tools.
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

Beaten on benchmarks

Head-to-head results where a newer method reports beating Video-RAG. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.