LGJun 18

Learning through Internalization

arXiv:2606.2093711.2
Predicted impact top 33% in LG · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in mechanistic interpretability and learning theory, this work provides theoretical foundations for understanding how models transition from explicit reasoning to direct computation, though the analysis is limited to a simplified setting.

The paper studies how transformers internalize explicit computational procedures (e.g., chain-of-thought tokens) into their weights, showing that internalization can facilitate learning of computationally hard tasks like parity. It provides the first provable analysis of successful internalization in a simplified one-layer transformer.

We study internalization processes, by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how they facilitate learning. We investigate how transformers internalize the simulation of semiautomata by internalizing chain-of-thought (CoT) tokens, which classes of semiautomata are harder to internalize, and expose the flip side of internalization, that is, a progressive degradation of out-of-distribution performance. We then provide the first provable analysis of successful internalization: for the task of learning parities, we show that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, learning to directly compute the parity. This task is computationally hard to learn from data without CoT supervision. Finally, we discuss how learning through internalization relates to the \textit{Positive Distribution Shift} phenomenon recently introduced by~\citet{Med+26}.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes