AIJun 21

Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs

arXiv:2606.226337.7
Predicted impact top 79% in AI · last 90 daysOriginality Incremental advance
AI Analysis

For researchers studying LLM behavior and alignment, this work links behavioral conflict resolution to internal model uncertainty, suggesting potential interventions for improving reliability.

The paper studies how LLMs resolve conflicts between their prior outputs and new evidence, introducing Trust Elasticity (TE) to measure persuadability. Across four models, TE varies substantially but is near-zero for clearly false claims, and this variation correlates with internal uncertainty indicators like Confidence Miscalibration and Internal Uncertainty Change.

Large language models (LLMs) frequently encounter inputs that disagree with their prior outputs, through user pushback, retrieved documents, or web search results. While the way they resolve such conflicts -- a process we frame as cognitive dissonance resolution -- has been characterized behaviorally, its connection to internal model uncertainty is not well understood. To study this systematically, we vary persuasion attempts along two dimensions, source authority and evidence quality, across 12 health-science claims of stratified epistemic status. Dissonance can be resolved through persuasion, backfire, or immunity. We introduce Trust Elasticity (TE), an econometrics-inspired measure of how readily a model is persuaded toward conflicting evidence. Across four LLMs, TE varies substantially, while clearly false claims elicit near-zero TE across all models. On two open-weight models, we further find that this variation is associated with two complementary internal uncertainty indicators, Confidence Miscalibration in Qwen and Internal Uncertainty Change in Llama. These results link cross-model behavioral variation to a measurable internal property and point to interventions targeting internal uncertainty as future work.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes