Autocorrelation effects in a stochastic-process model for solving two-armed bandit problems

arXiv:2603.055596.6h-index: 29
Predicted impact top 52% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work provides a theoretical explanation for experimental observations in photonic chaotic decision-making systems, offering guidance for designing ultrafast reinforcement learning systems in wireless communications and robotics.

The authors analyze a stochastic-process model for the two-armed bandit problem and find that negative autocorrelation is optimal in reward-rich environments, while positive autocorrelation is optimal in reward-poor environments, with performance independent of autocorrelation when winning probabilities sum to one.

Decision makers exploiting photonic chaotic dynamics obtained by semiconductor lasers provide an ultrafast approach to solving multi-armed bandit problems by using a temporal optical signal as the driving source for sequential decisions. In such systems, the sampling interval of the chaotic waveform shapes the temporal correlation of the resulting time series, and experiments have reported that decision accuracy depends strongly on this autocorrelation property. However, it remains unclear whether the benefit of autocorrelation can be explained by a minimal mathematical model. Here, we analyze a stochastic-process model for solving the two-armed bandit problem based on time series, where the threshold and a two-valued Markov signal evolve jointly. Numerical results reveal an environment-dependent structure: negative (positive) autocorrelation is optimal in reward-rich (reward-poor) environments. These findings show that negative autocorrelation of the time series is advantageous when the sum of the winning probabilities is more than one, whereas positive autocorrelation is useful when the sum of the winning probabilities is less than one. Moreover, the performance is independent of autocorrelation if the sum of the winning probabilities equals one, which is mathematically clarified. This study paves the way for solving the two-armed bandit problems for reinforcement learning applications in wireless communications and robotics.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes