LGNCJul 7

Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia

arXiv:2607.0662618.6h-index: 21
Predicted impact top 5% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For AI alignment and neuroscience, this work provides a causal mechanistic link between model internals and reward valuation deficits, but the approach is incremental as it applies known perturbation methods to a new domain.

The paper identifies reward-anticipatory units in vision-language models analogous to the Nucleus Accumbens in humans, and shows that perturbing these units induces anhedonia-like behavior (shifts to low-effort, low-reward options) while preserving non-reward task performance, mirroring clinical anhedonia scales.

Recent Vision-Language Models capture increasingly complex aspects of human cognition. Here we ask whether this alignment extends to reward valuation, which we assess in a mechanistic framework built on clinical tests that were developed to evaluate anhedonia and motivational deficits in major depressive disorder. In the brain, anhedonia is frequently linked to dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in vision language models, and test their causal role via targeted perturbations. Perturbing NAc-selective units induces behavioral effects that mirror human anhedonia: the model shifts toward low-effort, low-reward options in effort-based decision-making tasks. Crucially, our results reflect a specific deficit in reward valuation and anticipation rather than a loss of task capability: the perturbed model maintains baseline performance when reward-based choice is removed. This induced vulnerability further aligns with clinical anhedonia and motivation scales, including DARS and MAP-SR. Taken together, these results reveal reward valuation circuits in AI models that parallel those in humans.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes