When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs
This work addresses the problem of understanding how LLMs process discourse coherence, revealing that they fall for similar illusions as humans, which is important for improving model reliability in language understanding.
The study investigates whether Dutch language models exhibit coherence illusions—where an incoherent discourse seems coherent due to a matching distractor—and finds that models show reduced surprisal for incoherent continuations when a distractor is present, mirroring human behavior. Attention entropy and energy metrics reveal shared mechanisms across experiments.
Psycholinguistics studies show that human readers fall for coherence illusions: an incoherent discourse can seem coherent simply because a distractor matches what comes next. We investigate whether Dutch language models (6 monolingual and 4 multilingual) show the same behavior on texts that link back to earlier context with words such as 'again' and 'too'. First, we find that surprisal at the critical word tracks human acceptability judgments and eye-tracking data. Models are more surprised by incoherent continuations, but a matching distractor in the prior context reduces this surprisal. Second, attention entropy at the critical position identifies heads that behave differently under coherence vs. incoherence. We find that ablating these heads shows transfer effects across experiments, suggesting a shared mechanism. Third, we introduce energy from the associative-memory literature as a metric to quantify discourse coherence. Taken together, our results show that coherence illusions arise in Dutch LLMs, with entropy and energy exposing mechanisms that operate across settings.