LGJun 29

When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon

arXiv:2606.3044512.0
Predicted impact top 17% in LG · last 90 daysOriginality Incremental advance
AI Analysis

Provides a principled understanding of when online imitation learning helps in LLM post-training, clarifying the role of realizability for practitioners.

The paper challenges the view that error accumulation explains online imitation learning's advantage in LLM post-training, showing instead that benefits depend on realizability. Under non-realizability, offline IL faces an information-theoretic bottleneck even at horizon 1, while online IL provably achieves high performance under certain misspecification structures.

Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-training approach, often outperforming offline supervised fine-tuning (SFT). Yet a principled understanding of when and why online interaction helps remains unclear. In this work, we challenge the view that error accumulation is the main source of online IL's advantage, and instead show that the benefits of online interaction depend critically on whether the setting is realizable, i.e., whether the student policy class can represent the expert policy. Under realizability, we empirically find that offline IL already matches expert performance. In contrast, in non-realizable (misspecified) settings, we prove that offline IL encounters an information-theoretic bottleneck even when horizon $H=1$, and propose a structural characterization of misspecification relative to the reward, under which online IL provably achieves high performance despite a large distributional mismatch between the expert and student policies.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes