CLJun 10

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models

arXiv:2606.12689v114.8
Predicted impact top 69% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For interpretability researchers, this work demonstrates that observable patterns in latent reasoning models are insufficient to infer internal mechanisms, requiring matched controls and causal tests.

Latent reasoning models (LRMs) like Coconut and CODI exhibit patterns (BFS frontiers, decodable arithmetic) also present in controls lacking key components, and these patterns do not always causally affect behavior. Causal interventions show latent-thought utilization is graded, with behavioral influence concentrated in low-rank directions.

Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decodable arithmetic computation, as evidence for internal reasoning mechanisms. Evaluating two LRMs (Coconut and CODI) against controls lacking the proposed recurrence or curriculum, we find these patterns also appear in the controls and do not always causally affect behavior. Causal interventions reveal that latent-thought utilization is not binary but graded, scaling with a thought's causal effect on model behavior. Geometric analyses reveal this effect concentrates in low-rank directions whose step-to-step geometry grows more structured as their behavioral influence increases. Latent thoughts should therefore be treated as hidden computation, not hidden explanation: decodability, attention, or static structure alone cannot establish mechanism. LRM interpretability thus requires matched controls and causal tests.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes