ROJul 16

Risk-Aware Preference Learning for Stochastic Outcomes

arXiv:2607.154832.7h-index: 5
Predicted impact top 84% in RO · last 90 daysOriginality Incremental advance
AI Analysis

For human-robot interaction researchers, it demonstrates that ignoring human risk sensitivity leads to suboptimal reward learning in safety-critical stochastic settings.

This paper shows that modeling human risk sensitivity via Cumulative Prospect Theory (CPT) significantly improves reward learning from preferences over stochastic outcomes in social robot navigation, reducing regret compared to expected utility models.

Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting loss aversion. We study the consequences of this mismatch in social robot navigation, where safety-critical outcomes (e.g., collisions) are rare but highly consequential. We compare EU with Cumulative Prospect Theory (CPT), a nonlinear model of human decision-making, within a Bradley-Terry preference learning framework. Our preliminary experiments show that when preferences are generated by risk-sensitive users, CPT-based learners recover reward functions with substantially lower regret compared to EU-based learners. Our results highlight the importance of modeling human risk sensitivity when learning rewards from preferences over stochastic robot outcomes.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes