MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias
For practitioners deploying MLLMs in visually grounded tasks, this work reveals a previously unknown source of bias and offers a simple fix without retraining.
The paper identifies a 'late-layer textual override' in MLLMs where correct vision-based predictions formed in intermediate layers are overridden by textual bias in final outputs. The proposed training-free method, CALRD, recovers these suppressed predictions, achieving up to 9.4% absolute improvements on conflict benchmarks.
When vision contradicts text, multimodal large language models (MLLMs) consistently favor text, even when images provide clear evidence otherwise. This bias poses risks for applications requiring visual grounding, yet its cause remains unclear. In this paper, we uncover a surprising finding: models often get it right initially, forming correct vision-based predictions in their intermediate layers, before changing their minds and favoring text in the final output. We call this "late-layer textual override". The visual information is encoded, it simply does not survive to the output. More intriguingly, we find that how predictions change reveals whether they're correct: 85% of failures shift toward text, while 89% of successes shift toward vision. This directional signature enables a simple but powerful intervention: when we detect a confident visual prediction being suppressed, we restore it. We propose CALRD (Conflict-Aware Layer Reference Decoding), a training-free method that recovers overridden predictions at inference time. Experiments across five MLLMs of varying architectures demonstrate up to 9.4% absolute improvements on conflict benchmarks while largely preserving standard performance, without training or external knowledge. It recovers what the model already knew but failed to preserve.