ROJun 22

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure

arXiv:2606.233397.4
Predicted impact top 59% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For HRI researchers, this work identifies a critical failure mode (strangeness inversion) in using LLMs as interaction evaluators, showing they cannot access internal affective states needed for certain constructs.

The paper proposes using LLM-powered robots to self-evaluate HRI interactions via standard questionnaires, finding moderate-to-strong agreement with human ratings on engagement (r up to .72) but systematic inversion on comfort/strangeness (r = -.44 to -.67), which persisted across models and live deployment. This establishes boundary conditions for LLM-based self-evaluation in HRI.

Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We propose an \textit{inverted evaluation}, in which LLM-powered robots complete the same standardized instruments from their own perspective, and test whether these ratings agree with human ground truth. In Study~1, five LLMs completed HRI-CUES, Godspeed, and RoSAS questionnaires for 25~interactions ($N = 1{,}522$ evaluations) from the HRI-CUES dataset. LLMs achieved moderate-to-strong agreement on engagement dimensions (satisfaction $r$ up to $.65$ and enjoyment $r$ up to $.72$) with excellent test-retest reliability (ICC $\geq .82$), but \textit{systematically inverted} the comfort/strangeness dimension ($r = -.44$ to $-.67$, all $p < .05$), conflating engagement with comfort. In Study~2, a Nao robot running Claude~Sonnet~4.5 replicated these patterns in live interactions ($N = 4$), including real-time turn-by-turn assessment. The strangeness failure persisted across five models, synthetic controls, and embodied deployment for two participants. We argue that current LLM-based robots lack access to the internal affective states needed to assess constructs like strangeness, and that inverted evaluation requires supplementary modalities (e.g., physiological signals, gaze, proxemics) to move beyond behavioral proxies. These findings establish boundary conditions for using LLMs as interaction evaluators in HRI.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes