Model Releases
When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure
arXiv:2606.23339v1 Announce Type: new Abstract: Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We pro
arXiv:2606.23339v1 Announce Type: new Abstract: Human-robot interaction (HRI) evaluation relies almost exclusively on human-completed questionnaires, leaving the robot's perspective unexamined. We propose an extit{inverted evaluation}, in which LLM-powered robots complete the same standardized instruments from their own perspective, and test whether these ratings agree with human ground truth. In Study1, five LLMs completed HRI-CUES, Godspeed, and RoSAS questionnaires for 25interactions (N = 1{,}522 evaluations) from the HRI-CUES dataset. LLMs achieved moderate-to-strong agreement on engagement dimensions (satisfaction r up to .65 and enjoyment r up to .72) with excellent test-retest reliability (ICC geq .82), but extit{systematically inverted} the comfort/strangeness dimension (r = -.44 to -.67, all p < .05), conflating engagement with comfort. In Study2, a Nao robot running ClaudeSonnet~4.5 replicated these patterns in live interactions (N = 4), including real-time turn-by-turn assessment. The strangeness failure persisted across five models, synthetic controls, and embodied deployment for two participants. We argue that current LLM-based robots lack access to the internal affective states needed to assess constructs like strangeness, and that inverted evaluation requires supplementary modalities (e.g., physiological signals, gaze, proxemics) to move beyond behavioral proxies. These findings establish boundary conditions for using LLMs as interaction evaluators in HRI.
Source: arXiv cs.RO | 2026-06-23