Brent Mittelstadt

h-index1
3papers
9citations

3 Papers

10.1CYMay 15
AI-Mediated Communication Can Steer Collective Opinion

Stratis Tsirtsis, Kai Rawal, Chris Russell et al.

Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LLMs) now polish users' posts on LinkedIn and provide context for content shared on X. While prior work has shown that AI can express biased opinions and shape individuals' opinions during human-AI interactions, less attention has been paid to its influence on collective opinion formation when mediating human-to-human communication. We address this gap via a combination of empirical and theoretical analyses. We show empirically that LLMs from multiple popular families introduce directional biases when instructed to edit human-written texts on contested topics, for example, nudging texts in favor of gun control and against atheism. Building on this observation, we introduce a mathematical model of opinion dynamics in which an AI system sits between users on a social network, transforming the opinions they express and perceive. By analytically characterizing the equilibrium of this model and performing simulations on real social network data, we show that biases introduced by AI in human-to-human communication can be amplified through the network and shift collective opinion in their direction. In light of these findings, we investigate whether such biases are controllable by online platforms. We audit the "Explain this post" feature on X and find evidence of pro-life bias in Grok's outputs on abortion-related content, which we trace back to specific design choices. We conclude with a discussion of the broader implications of our findings in relation to ongoing legislative efforts in the European Union.

25.8CLJun 27
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Will Hawkins, Kaivalya Rawal, Jonathan Rystrøm et al.

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning with non-adversarial data. We present the first comprehensive empirical study of this phenomenon in multilingual settings by fine-tuning Llama-3.2, Qwen3, and Gemma-3 models using benign data translated across nine languages. We find that safety outcomes are highly sensitive to both the choice of fine-tuning language and the evaluation language, with adversarial compliance rates increasing four-fold in some settings. Multilingual safety drift is decoupled from general capability metrics, and occurs heterogeneously across languages and models. Fine-tuning in non-English languages often induces smaller internal representational drifts than English, but these shifts lead models to default to either exaggerated compliance or refusal. As such, assessing fine-tuning impacts solely in English provides inadequate assurance for deployment. To facilitate further research into these cross-lingual safety blind spots, we release the Multilingual-Benign-Tune dataset and the SORRY-Bench-Multilingual evaluation suite.

CYJun 13
The Fallacy of Sustainable Generative AI: Limitations in EU Environmental Regulation of Data Centres and Paths Forward

Daria Onitiu, Sandra Wachter, Brent Mittelstadt

In the age of Artificial Intelligence (AI), Large Language Models, Generative AI and larger frontier AI models, data centres create a significant environmental burden on electricity grids and fresh water resources. Requiring data centre operators and Big Tech under the recast Energy Efficiency Directive (recast EED) to quantify, report and disclose the facility-level energy and water impacts seems to be a step into the right direction towards more transparency and accountability. Yet when two recast EED approved benchmarks - the Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) - can be skewed to create a false sense on efficiency gains, current EU policy pushing for sustainable hyperscale data centre expansion appears misplaced. This paper argues that current PUE and WUE reporting frameworks illustrate what we term the "efficiency paradox," according to which positive scores require retrofitting larger AI data centres at the expense of energy supply and people's water access. Countering this efficiency paradox requires a new strategy for data centre operators and the EU Commission to demonstrate the ecological gains of optimising for efficiency through individual reporting and additional policy interventions. We make three policy proposals to show how this strategy can be formalised in practice: (i) measures to reveal and certify efficiency improvements, (ii) documentation of trade-offs in PUE and WUE improvements and, (iii) a monitoring framework of their diminishing returns and countereffects over time. Implementing these measures will ensure that the recast EED common rating scheme is fit-for-purpose, balancing sustainability with AI innovation for local communities.