
Anthropic research identifies a global-workspace-like representation region inside Claude models
Original: Anthropic Found a Mind Hiding Inside Their Language Model
Short summary
Anthropic's Transformer Circuits Thread team published a paper showing that Claude and similar LLMs contain a small set of internal vector representations that behave like a cognitive global workspace — the brain-theory concept for reportable, conscious-access thinking. Using a new interpretability tool called the Jacobian Lens, they identified a 'J space' where concept vectors activate even when the model never verbalizes them, and demonstrated that swapping vectors in this space predictably changes the model's output. The findings pass five functional tests of a global workspace, offering a new window into what models are 'thinking' before they speak.
- •Anthropic found a 'J space' in Claude models that behaves like a cognitive global workspace for reportable thoughts
- •New Jacobian Lens tool decodes intermediate layer representations by averaging how perturbations affect output across many prompts
- •Swapping active concept vectors in J space predictably alters the model's verbalized answers, confirming functional properties
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



