CVJul 2

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

arXiv:2607.0173715.0
Predicted impact top 19% in CV · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of inefficient frame sampling in long-form video QA for multimodal LLM users, offering a plug-and-play solution that improves performance under fixed token budgets.

ReQuest introduces a question-adaptive keyframe selection pipeline for long-form video QA that uses uncertainty-driven selective computation, achieving consistent accuracy gains on Video-MME, MLVU, and LongVideoBench without modifying the underlying MLLM.

Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for evidence localization. We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation. ReQuest integrates (i) a lightweight question-aware selector distilled from MLLM-generated supervision, (ii) Re-thinking Routing that triggers additional inference only when the model is uncertain with a length-adaptive criterion, and (iii) uncertainty-guided adaptive non-maximum suppression that selects temporally diverse frames while adjusting spacing based on question difficulty. As a plug-andplay method, ReQuest improves long-video QA without modifying or fine-tuning the underlying MLLM. Experiments on Video-MME, MLVU, and LongVideoBench demonstrate consistent accuracy gains with competitive computational cost, with particularly strong improvements in medium and long video regimes.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes