CLCYHCJul 6

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

arXiv:2607.0511310.5
Predicted impact top 77% in CL · last 90 daysOriginality Incremental advance
AI Analysis

This work reveals a fundamental confound in user-elicited LLM evaluations, including public leaderboard preference data, for researchers and practitioners relying on such metrics.

The study shows that user expectations, shaped by pre-interaction framing, significantly influence LLM evaluations and interaction behavior, while actual task performance does not predict changes in user impressions. Oversold users rated models more favorably and used more directive prompts, whereas undersold users wrote longer, collaborative prompts.

Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 participants each used one of six LLMs from two families across three collaborative tasks, after first viewing a landing page that matched, overstated, or understated their model's true capability. This pre-interaction framing shifted user opinions and interaction behavior while task performance did not. Oversold users rated the model more favorably and used more directive prompting, while Undersold users wrote longer, more collaborative prompts. The quality of what users and the model produced together depended only on the model's true capability, not on what users were told. Participants' change in model impressions after use, measured across two impression measures, was not predicted by task performance ($β= -0.01$ and $0.11$, both n.s.), but by whether the model met users' expectations ($β= 0.47$ and $0.50$, both $p < .001$) and how confident they felt working with it ($β= 0.47$ and $0.36$, both $p < .001$). After interaction, users are still rating the pitch, not the product: user-elicited LLM evaluations, including the preference data driving public leaderboards, measure expectation management at least as much as the model itself.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes