3.4SYJul 7
On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ gamesShengyuan Huang, Xiaoguang Yang, Yifen Mu et al.
In infinite-horizon discrete-time linear-quadratic (LQ) dynamic games, computing feedback Nash equilibria (FNEs) remains computationally challenging. Motivated by this, we study a finite-horizon strategy for approximating one of the infinite-horizon FNEs. The finite-horizon strategy is as follows. Each player $i$ has an individual prediction horizon $T^i$. In the infinite-horizon game, at each stage, each player $i$ computes its control in the following way: player $i$ envisions an auxiliary $T^i$-stage game in which the same set of players play, computes the unique FNE of the auxiliary game using a standard method, and implements only the first-stage control. Our main result is, under suitable conditions, the total cost under these finite-horizon strategies converges to that under one of the infinite-horizon FNEs when all players' prediction horizons tend to infinity. Moreover, we derive an explicit cubic-polynomial upper bound on this cost gap with respect to the distance between the corresponding strategy matrices. This strategy is tractable and implementable, as it avoids the direct solution of the coupled algebraic Riccati equations (CARE) of infinite-horizon LQ games.
6.6GTMar 29
Decentralized MARL for Coarse Correlated Equilibrium in Aggregative Markov GamesSiying Huang, Yifen Mu, Ge Chen
This paper studies the problem of decentralized learning of Coarse Correlated Equilibrium (CCE) in aggregative Markov games (AMGs), where each agent's instantaneous reward depends only on its own action and an aggregate quantity. Existing CCE learning algorithms for general Markov games are not designed to leverage the aggregative structure, and research on decentralized CCE learning for AMGs remains limited. We propose an adaptive stage-based V-learning algorithm that exploits the aggregative structure under a fully decentralized information setting. Based on the two-timescale idea, the algorithm partitions learning into stages and adjusts stage lengths based on the variability of aggregate signals, while using no-regret updates within each stage. We prove the algorithm achieves an epsilon-approximate CCE in O(S Amax T5 / epsilon2) episodes, avoiding the curse of multiagents which commonly arises in MARL. Numerical results verify the theoretical findings, and the decentralized, model-free design enables easy extension to large-scale multi-agent scenarios.
6.9THJun 30
Robust Aggregation of Calibrated ForecastsXinxiang Guo, Yingkai Li, Yifen Mu
Decision-makers often rely on multiple probabilistic forecasts that are individually calibrated but need not be fully informative. We develop a framework for aggregating such forecasts when the decision-maker knows only that experts satisfy calibration. We show that the joint distribution of calibrated forecasts can contain decision-relevant information that is unavailable from any single expert, so the standard optimal-in-hindsight (OIH) benchmark may substantially understate attainable performance. To formalize this idea, we introduce a robust max-min benchmark: the best payoff a decision-maker can guarantee against all profile-wise conditional-mean mappings compatible with calibration. This benchmark is tractable, admits a linear-programming formulation, and dominates the OIH benchmark up to calibration error. It can nevertheless be strictly below the Bayesian benchmark, clarifying the value of knowing experts' information structures. Finally, we provide online algorithms that attain the robust benchmark under forecast-only feedback and stronger contextual benchmarks under state feedback.