LGAIJul 10

Risk-Aware General-Utility Markov Decision Processes

arXiv:2607.092982.8h-index: 2
Predicted impact top 89% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a framework and algorithm for risk-aware decision-making in general-utility MDPs, enabling risk-averse optimization for a broad class of objectives.

The paper introduces risk-aware general-utility Markov decision processes (GUMDPs) to trade off expected performance with risk aversion using the entropic risk measure, and proposes a Monte Carlo Tree Search (MCTS) method to solve them with provable accuracy, demonstrating success across diverse tasks.

We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective values, where the objective function depends on the frequency of visitation of states induced by the agent's policy. First, we motivate, propose, and formalize risk-aware GUMDPs, which enable agents and decision makers to trade off expected performance by risk aversion while benefiting from the rich set of objectives that can be cast under the framework of GUMDPs. We focus our attention on the entropic risk measure (ERM). Second, we show how we can solve risk-aware GUMDPs with ERM objectives by resorting to online planning techniques. In particular, we propose an approach based on Monte Carlo Tree Search (MCTS) to provably solve risk-aware GUMDPs up to any desired accuracy. Third, we provide a set of experimental results showcasing that our approach is successful when optimizing for a spectrum of risk-aware behaviors in the context of GUMDPs under diverse tasks (standard MDPs, maximum state entropy exploration, imitation learning, and multi-objective MDPs).

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes