Retrieval-augmented generation
Video-LLaMA 2
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Superseded baseline#338 of 1,179 most-superseded · first seen Jun 11, 2024
Superseded — cited as a baseline and beaten by newer methods
0 papers critique it · 1 beat it on benchmarks
Beaten on benchmarks
Head-to-head results where a newer method reports beating Video-LLaMA 2. Values are copied from the source paper's tables — verify against the cited paper.
AffectAgent beats Video-LLaMA 2
42.74 vs 35.99
Mean · [all MLLM backbones - Video-LLaMA 2]
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
What to use instead
Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.