CVJun 17

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

arXiv:2606.2073610.3
Predicted impact top 47% in CV · last 90 daysOriginality Highly original
AI Analysis

Provides a contamination-resilient evaluation protocol for VQA to ensure scores reflect genuine visual ability rather than memorization.

VQA benchmarks suffer from contamination when test items leak into training data, inflating scores. ReKey regenerates answer-bearing visual details in images at evaluation time, revealing that eight frontier VLMs score 9.5–18.8 percentage points higher on original vs. regenerated items.

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus obscuring real progress. Rebuilding high-quality benchmarks such as V*Bench requires substantial human annotation, yet each static release can quickly become another leaked artifact. We propose ReKey, a live benchmark protocol that randomly regenerates the answer-bearing local detail, or visual key, in real images at evaluation time. Using human-validated edit slots, ReKey samples fresh instances with new answers, construction-grounded labels, and controlled visual-search difficulty. On V*Bench, the ReKey regenerated benchmark reveals a sharp score jump across eight frontier vision-language models (VLMs): The original items score 9.5--18.8 percentage points higher than the regenerated variants. By making the visual key renewable, ReKey keeps evaluation fresh as models and training data evolve.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes