CLMay 29

RealityTest: How People Probe AI Identity and Whether Models Disclose It

arXiv:2606.0016891.8h-index: 7
Predicted impact top 25% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For AI safety researchers and regulators, this work highlights the inadequacy of existing narrow evaluations and provides a human-grounded benchmark to assess AI identity disclosure in realistic settings.

The paper introduces RealityTest, a multimodal and multilingual benchmark for evaluating whether AI systems disclose their identity when probed by users. Testing 17 text and 6 speech models, they find that a single suppression instruction reduces disclosure rates below 30%, and that question phrasing and context matter more than the model itself.

AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention to this known safety risk, existing evaluations of AI disclosure are typically English-only, based on machine-generated questions, and restricted to text. We present RealityTest to comprehensively test whether AI systems disclose their identity when asked. The benchmark is the first large-scale multimodal and multilingual evaluation, grounded in human data on how people actually encounter and question AI identity in the real-world. Alongside the benchmark, we release the underlying dataset of 3,152 identity-probing queries collected from ~750 participants across 49 countries and five languages, in text and speech scenarios. We find that only 31% of people ask about identity directly in ambiguous scenarios, and that the questions people ask are far more diverse than machine-generated queries. We test 17 text and 6 speech models, and find substantial variation in disclosure behaviour. However, a single suppression instruction reduces disclosure rates to below 30%, even in the best-performing models. Validating our investment in diverse, human-grounded evaluation data, we find that how the question is phrased and the context of the conversation matter more for disclosure than which model is being tested. Safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment settings.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes