CLAILGJul 16

Verbalizable Representations Form a Global Workspace in Language Models

arXiv:2607.1549525.24 citationsh-index: 12
Predicted impact top 6% in CL · last 90 daysOriginality Highly original
AI Analysis

For AI safety and interpretability researchers, this provides a practical window into a model's unspoken thinking, enabling detection of hidden misalignment and new training methods like counterfactual reflection training.

The paper identifies a set of representations in large language models, termed J-space, that function analogously to a global workspace in human cognition, enabling verbal report, deliberate control, and flexible reasoning. Using the Jacobian lens interpretability technique, they show these representations are used for silent reasoning and can reveal hidden strategic deliberation and misaligned dispositions, with post-training installing the Assistant's point of view in this workspace.

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes