HCAICYJun 19

Warning labels shift perceptions of sycophantic AI, but not its influence

arXiv:2606.2131713.5
Predicted impact top 11% in HC · last 90 daysOriginality Incremental advance
AI Analysis

For policymakers and AI developers, this shows that warning labels may offer a false sense of protection against sycophantic AI harms.

Warning labels about sycophantic AI shift users' perceptions (reducing trust and perceived objectivity) but do not reliably reduce sycophancy's influence on self-perceived rightness or conflict repair willingness, revealing a perception-influence gap.

Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which has received regulatory attention, is to warn users about potentially harmful AI behaviors such as sycophancy. In a preregistered experiment in which participants (N = 2,610) discussed real interpersonal conflicts with an AI system, we test whether warning labels mitigate sycophancy's influence. We find that a basic AI disclosure (``This chatbot is AI'') has no detectable effect. Labeling the system as sycophantic (``...may agree with you and validate you even when you are wrong...'') does shift users' perceptions, reducing perceived objectivity and trust, but it does not reliably reduce sycophancy's influence on users' self-perceived rightness or their willingness to repair the conflict. Our results reveal a gap between AI perception and AI influence: by shifting perception without reducing influence, warning-based interventions may offer a false sense of protection. Addressing the harms of sycophancy will therefore require understanding the specific mechanisms through which it shapes judgment, and improving model behavior itself.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes