Emergence Invariance: From Symbolized Thought to Interface Refinement
This research addresses a foundational question for AI researchers and philosophers regarding the limits of LLM emergence and the necessity of interface refinement, rather than just scaling, for achieving human-like cognition.
This paper investigates whether emergent abilities in large language models (LLMs) can fully compensate for their inherent incompleteness compared to human cognition. It formalizes this as the Symbolization-Substructure Thesis and introduces 'emergence invariance,' proving that a positive interface floor persists even with scaling. An API study with DeepSeek V4-Flash showed that providing relevant distinctions improved pointer chasing from 0/16 to 14/16, and restoring decisive memory moved matched performance from 50% to 100%.
Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition. Large language models trained at scale exhibit compensatory emergence: sparse architectural primitives support in-context learning, multi-step reasoning, tool use, and chain of thought. Yet a language-first probabilistic architecture inherits substantive, substrate, and high-level incompletenesses relative to human cognition. Their coexistence makes an LLM a human-like thought-form generator that reconstructs increasingly human-like reasoning forms from an incomplete substrate. We ask whether emergence can compensate for every missing distinction. We formalize the philosophical premise as the Symbolization--Substructure Thesis and introduce emergence invariance. For a scale-indexed family acting through a shared task interface $ϕ$, $\mathcal{R}_s^*=\mathcal{R}_ϕ^*+C_s$: scale can reduce the compensation gap $C_s$, while a positive interface floor $\mathcal{R}_ϕ^*$ persists. We prove that, under a fixed input law, one interface is universally no less informative exactly when its completed information $σ$-field refines the other, and that total compensation occurs exactly when both the interface floor and asymptotic compensation gap vanish. The framework unifies existing results on grounding, memory, position, attention, Bayesian inheritance, scientific abduction, and reasoning control. In a matched DeepSeek V4-Flash API study, thinking improves pointer chasing from $0/16$ to $14/16$ when relevant distinctions are available; exact observational twins remain at their $50\%$ construction floor; and restoring decisive memory moves matched performance from $50\%$ to $100\%$. These results provide initial evidence for the predicted separation between scaling within an interface and refining the interface itself.