Matthieu Wyart, a statistical physicist, joined Tim Scarfe on Machine Learning Street Talk to answer a question that sounds simple and is not: why can deep networks discover abstractions that shallow models miss?
Wyart's answer is that the structure is in the data, not the architecture. Language and images are built from parts within parts — a hidden hierarchy — and depth is what lets a network recover those coarse-grained variables. That recovery is his account of how deep nets escape the curse of dimensionality, and it draws directly on the Random Hierarchy Model work.
Two consequences follow that are worth arguing with:
The conversation runs from jamming transitions and rough loss surfaces — Wyart's home turf — through Chomsky and context-free grammars, to where current systems fall short of genuine scientific invention rather than recombination. It closes on diffusion models and neural scaling laws, including a phase transition in diffusion models that reveals the hierarchical nature of data and work deriving scaling laws from the statistics of natural language, and on the limits of physics-inspired theory generally.
The last question is a personal one: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?
"Depth helps because data is hierarchical" is a claim with teeth — it predicts which problems depth should help on, and it implies the objective we train against is leaving sample efficiency on the table. That is a more falsifiable position than most theory-of-deep-learning arguments get to.
Sources:
Founder, Engineer
New York City