MIT Technology Review
7/9/2026

Anthropic found a hidden space where Claude puzzles over concepts
Short summary
Anthropic developed the Jacobian lens, a new interpretability technique offering unprecedented insights into how Claude processes information internally. The tool reveals findings ranging from mundane to concerning aspects of LLM behavior. This research advances understanding of large language models' decision-making.
- •Anthropic built Jacobian lens—a new interpretability tool for understanding Claude's internal processing
- •Findings reveal both ordinary and unnerving aspects of how LLMs work internally
- •Research has implications for AI transparency, safety, and product strategy
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



