CVAIJul 5

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

arXiv:2607.0416313.8
Predicted impact top 23% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For practitioners deploying LVLMs, this work offers a lightweight method to mitigate hallucinations without retraining, though the gains are incremental over existing decoding-stage interventions.

SeeMe proposes a training-free framework that restructures visual tokens to reduce hallucinations in large vision-language models, achieving consistent improvements across MME, POPE, and AMBER benchmarks.

Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain susceptible to hallucinations, generating content that is inconsistent with the actual visual input. Existing methods primarily intervene at the decoding stage, while overlooking a critical source of hallucinations: irrelevant or noisy visual tokens that mislead the decoding process. To address this issue, we propose SeeMe, a training-free framework that introduces the concept of feature engineering from traditional machine learning into LVLMs. SeeMe restructures visual tokens through a three-stage token engineering process to suppress hallucination sources while preserving informative visual evidence. Experiments on MME, POPE, and AMBER benchmarks across four LVLMs demonstrate that SeeMe consistently reduces hallucinations and improves output consistency, providing a novel perspective for mitigating hallucinations in LVLMs.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes