CLAIJun 9

When Roleplaying, Do Models Believe What They Say?

arXiv:2606.11502v116.7h-index: 4
Predicted impact top 57% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers studying AI alignment and truthfulness, this work distinguishes between superficial role-playing and deeper belief internalization, showing that persona adoption primarily affects outputs while Emergent Misalignment affects internal representations.

The paper investigates whether role-playing historical personas changes language models' internal truth representations or only their outputs. Using linear probes, they find that persona adoption suppresses false claims consistent with the persona less than other false claims, but these claims remain classified as false internally, contrasting with Emergent Misalignment where false claims shift toward the true region.

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models operate, with models constantly selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question with linear truth probes, applying them to LLMs role-playing historical personas whose likely beliefs differ from modern consensus. For each persona, we compare false claims the persona would likely have endorsed (*era-believed*) with topic-matched false claims they would not have endorsed (*era-false*). Across prompting, in-context learning, and supervised fine-tuning, persona induction suppresses era-believed statements less than equally false alternatives, yet they remain classified as false overall. Role-play therefore shifts what these models say more than what they internally represent as true. We contrast this with models trained on harmful advice that exhibit Emergent Misalignment (EM). Across three model families (Qwen 2.5 14B, Qwen 3 8B, and Llama 3.3 70B), their false claims move substantially toward the true region of probe space, are defended under challenge roughly half the time versus about a sixth for role-play, and are used in downstream reasoning. Role-play and Emergent Misalignment thus are points on a spectrum of belief internalization, where role-play changes what a model says with little representational change, while Emergent Misalignment shifts the internal representation of false claims without fully marking them as true.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes