Anthropic's 2026 Transformer Circuits paper "Verbalizable Representations Form
a Global Workspace in Language Models" makes a startling claim: large language
models keep a small, privileged set of thoughts they can report, hold in mind,
and reason with — floating atop a vast ocean of processing they can't access.
Sound familiar? It's the neuroscience "global workspace" of conscious access.
Using a new interpretability tool, the Jacobian lens (J-lens), the researchers
read a model's unspoken concepts mid-computation — and show this "J-space"
satisfies all five hallmarks of a workspace: verbal report, directed
modulation, internal reasoning, flexible generalization, and selectivity.
In this 20-minute deep dive:
00:00 If the mind is an ocean — access consciousness
The global workspace theory & its five marks
The Jacobian lens and the J-space
"Think of a sport" — verbal report & causal swaps
France → China — flexible generalization
Turning the workspace off — selectivity
Layers, capacity & the broadcast hub
Alignment auditing: blackmail, evaluation-awareness, hidden agendas
The Assistant's point of view & self-monitoring
Counterfactual reflection training
Differences from human cognition & a word on consciousness
Paper: https://transformer-circuits.pub/2026...
By Wes Gurnee, Nicholas Sofroniew, Jack Lindsey et al. (Anthropic).
Figures © Anthropic, used for educational commentary.
#AI #LLM #Interpretability #Anthropic #Claude #MechInterp #Consciousness