
This video is one of my favorite Machine Learning Street Talk interviews. It's a deep-dive conversation between Tim Scarfe and Neel Nanda (a researcher at DeepMind) on the subject of mechanistic interpretability. The discussion explores the ambitious goal of reverse-engineering neural networks to understand the algorithms they learn, moving beyond simple input-output observation to comprehend the inner workings of AI models.
Mechanistic Interpretability: Nanda explains this field as an effort to "decompile" neural networks into readable code or logical structures to ensure safety and alignment. He addresses the controversial "stochastic parrot" view, arguing that models do indeed learn real algorithms and processes.
Grokking and Phase Transitions: They discuss the phenomenon where models move from memorizing data to generalizing after extensive training. Nanda posits that grokking is partly an illusion created by the overlap between phase transitions and the speed difference between memorization and generalization.
Superposition and Polysemanticity: One of the most significant challenges identified is superposition, where models represent more features than they have neurons by using non-orthogonal vectors. Nanda explains why this complicates the decomposition of models into simple units of analysis.
Transformers and Residual Streams: The conversation highlights the residual stream as a central information highway in Transformer models, which acts as a shared bandwidth for passing information between layers. This structure is key to understanding how circuits, such as induction heads, perform tasks like few-shot learning.
World Models: Referencing the Othello-GPT paper, they debate whether language models trained on next-token prediction spontaneously develop internal world models or just sophisticated statistical correlations.
AI Risk: Nanda emphasizes that we don't need to understand AI systems perfectly to create them, which poses catastrophic risks. He advocates for mechanistic interpretability as a scientific necessity to predict emergent behaviors and goal-directedness, rather than just focusing on philosophical definitions.
Founder, Engineer
New York City