CVAIJun 2

A Dataset for Dynamic Human Preferences for Vision Language Models

arXiv:2606.076536.9h-index: 2
Predicted impact top 71% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For researchers evaluating VLMs in human-interactive settings, this benchmark addresses the gap in assessing real-time preference adaptation, though it is an incremental contribution.

This work introduces a benchmark for evaluating Vision Language Models' ability to adapt to dynamic human preferences provided in-context at inference time, and evaluates state-of-the-art models on it.

Given the increased adoption of Vision Language Models (VLMs) in human-interactive settings, it is important that we evaluate how well these models can adapt to real-time preferences for different users. While an increasing number of vision-language benchmarks have recently been introduced, they focus largely on evaluating static capabilities and generally-held preferences learned from extensive training data. This work introduces a new benchmark for evaluating the ability of VLMs to understand dynamic human-preferences, i.e. preferences that are passed in-context at inference time. We provide an automated pipeline for generating this benchmark with variations on image dependence, a dynamic multi-modal human-preference dataset, and evaluations of state-of-the-art models on the novel benchmark.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes