CVLGJun 11

Self-Evolving Visual Questioner

CMU
arXiv:2606.1392920.5h-index: 25
Predicted impact top 12% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in vision-language AI, this provides a method to autonomously improve question generation, reducing reliance on costly human-annotated data.

This work introduces a self-evolving framework for vision-language models to generate high-quality, diverse, and visual-centric questions without external supervision, achieving substantial improvements in question quality and difficulty over static training data.

Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high-quality training data or the cost of curating them. We show that a VLM can continuously improve itself as a visual questioner without any external supervision. We propose a self-evolving framework that uses a VLM itself as both a proposer and a filter to produce harder, more informative, and visual-centric questions, while maintaining their exploration diversity to avoid training collapse. These questions are then used to train the VLM in both questioner and answerer modes. To evaluate the questioner, we introduce an agentic protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across various backbone VLMs show that our method substantially enhances the quality and substantially expands the difficulty boundary of autonomous question generation. Under the same budget, our self-supervision is more effective than training on the static source data. Moreover, the self-evolving questioner remains a competitive or even better answerer.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes