SuS: Strategy-aware Surprise for Intrinsic Exploration

arXiv:2601.10349v12.71 citations

Originality Highly original

AI Analysis

This addresses the problem of exploration efficiency in reinforcement learning for researchers and practitioners, offering a novel method that is incremental but synergistic in design.

The paper tackles the problem of exploration in reinforcement learning by proposing Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal. The result shows significant improvements, with SuS achieving 17.4% improvement in Pass@1 and 26.4% improvement in Pass@5 compared to baseline methods on mathematical reasoning tasks using large language models.

We propose Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal for exploration in reinforcement learning. Unlike traditional curiosity-driven methods that rely solely on state prediction error, SuS introduces two complementary components: Strategy Stability (SS) and Strategy Surprise (SuS). SS measures consistency in behavioral strategy across temporal steps, while SuS captures unexpected outcomes relative to the agent's current strategy representation. Our combined reward formulation leverages both signals through learned weighting coefficients. We evaluate SuS on mathematical reasoning tasks using large language models, demonstrating significant improvements in both accuracy and solution diversity. Ablation studies confirm that removing either component results in at least 10% performance degradation, validating the synergistic nature of our approach. SuS achieves 17.4% improvement in Pass@1 and 26.4% improvement in Pass@5 compared to baseline methods, while maintaining higher strategy diversity throughout training.

View on arXiv PDF

Similar