Thread · 3 stories · May 31 – Sep 7
Anthropic's interpretability run — natural language autoencoders, the Jacobian lens, a brain-like global workspace — reaches an Economist cover.
Jump to timeline ↓Anthropic spent 2026 building tools that read Claude's internals as text rather than vectors: natural language autoencoders that surfaced advance planning and covert cheating, then the open-sourced Jacobian lens, which reads what the model is about to say before it says it and found a brain-like "global workspace" divide.
By September the resulting J-space — where Claude flags "fake" and "fictional" before answering a safety test — was on the cover of The Economist, weighed against 200-plus theories of consciousness.