LGAICLJul 3

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning

Jingchu Wang, Bingbing Xu, Yige Yuan, Dan Zhang, Bin Xie, Xiaoqian Sun, Huawei Shen
arXiv:2601.119609.6h-index: 12Has Code
Predicted impact top 11% in LG · last 90 daysOriginality Highly original
AI Analysis

Improves reinforcement learning for LLM reasoning by addressing the entanglement of training and inference policies, benefiting researchers working on reasoning tasks.

R^2PO decouples training trajectories from inference responses in LLM reasoning, achieving 3.4% accuracy gain on MATH-500 and 1.3% on APPS with more diverse rollouts and reduced length bias.

Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincide with the one producing inference responses. We argue that this is a misleading inductive bias: the optimization-optimal trajectory distribution favors informative gradients, whereas the inference-optimal response distribution emphasizes accuracy and consistency. Forcing both into a single policy entangles their gradients and suppresses exploration. We propose R$^2$PO (Residual Rollout Policy Optimization), which attaches a lightweight Residual Rollout-Head atop the policy to decouple training trajectories from inference responses, diversifying rollouts during training while keeping inference generation intact. Experiments show that R$^2$PO consistently outperforms baselines, with average accuracy gains of 3.4% on MATH-500 and 1.3% on APPS, alongside more diverse rollouts and reduced length bias. Our code is available at https://github.com/RRPO-ARR/Code.

Code Implementations1 repo
Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes