Response drift across frontier large language models
For developers and users of frontier LLMs, this work establishes that response drift is universal and cannot be captured by automated metrics, highlighting the necessity of human evaluation.
All frontier LLMs exhibit response drift from expert-validated references, with eight models showing 78-81% deviation and two achieving 47-49% deviation, as measured by 29,140 human assessments across 62 questions. Drift is domain- and question-dependent, and automated metrics explain less than 2% of human-judged variance.
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation. Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments. Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%). Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85. Automated similarity metrics explain less than 2% of variance in human judgements. These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.