ROJun 29

REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

arXiv:2606.299375.4
Predicted impact top 60% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For the HRI community, this benchmark provides a standardized framework to evaluate robot failures and build adaptive recovery systems, addressing limitations of prior work that treated failures as independent and binary.

REPAIR-Bench introduces a benchmark with 214 trials from 41 participants for evaluating robot failure perception and recovery, spanning three tasks: longitudinal failure detection, visual failure-type classification, and user-centered recovery prediction. Baseline results show hierarchical recurrent modeling improves failure detection (F1: 0.80 vs. 0.68) and a fine-tuned Mistral-7B achieves Hit@5=0.76 for recovery prediction.

Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i) often treats failures as independent events, (ii) emphasizes binary failure detection, (iii) with rule-based recovery modeling. We present REPAIR-Bench, built on 214 interaction trials from 41 participants, the benchmark spans four induced failure types and provides synchronized facial action units, head pose, speech transcripts, and post-interaction affect and recovery reports. The benchmark spans three novel evaluation tasks that jointly capture the lifecycle of failure in human-robot interaction (HRI): (i) failure detection over inter-dependent interaction sessions, modeling longitudinal user adaptation across repeated failures; (ii) visual failure-type classification beyond binary success/failure formulations; and (iii) user-centered recovery prediction, inferring users' preferred recovery strategies from interaction context rather than relying on manually designed or rule-based strategies. In baseline experiments, hierarchical recurrent modeling improved failure detection over a single-session model (strict F1: 0.80 vs. 0.68), achieved a failure localization mean signed error of -0.51 s, median absolute error of 2.97 s and, for recovery prediction, a QLoRA-tuned Mistral-7B reached Hit@5=0.76 and F1@5=0.32. REPAIR-Bench provides both the HRI and Medical HRI communities with a standardized framework for (1) evaluating robot failures and (2) building transparent, adaptive, and trustworthy recovery systems.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes