CVAICLAug 2

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

arXiv:2608.0102110.41 citations
Predicted impact top 33% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

This work provides a method for creating more stable and model-independent hallucination benchmarks for vision-and-language models, which is important for evaluating rapidly evolving models.

The authors investigate whether human-written hallucination samples can replace model-generated ones for benchmarking vision-and-language hallucination detection. They create a dataset of 1,600 human-written samples across four languages and 18,400 model-generated samples, finding that human samples yield higher agreement, better control, and similar distributional properties, suggesting they are a viable substitute.

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600 human-written samples, spanning four languages (Chinese, English, French, Italian), and 18,400 samples from five vision-and-language models, all annotated for hallucinations using a fine-grained span-level labeling scheme. We find that human-written samples result in higher agreement and allow greater control of dataset contents, while remaining distributionally similar to samples derived from vision-and-language samples and providing a reasonable portrayal of detection capabilities - suggesting that human data is a viable substitute for model-based hallucination benchmarks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes