Thread · 2 stories · Jun 16 – Jul 17 · concluded
Apple argued reasoning models don't really reason; the rebuttals said the failures were context limits and missing tools, not missing thought.
Jump to timeline ↓Apple's paper claimed o1- and o3-class models collapse on puzzles like Tower of Hanoi, and that what looks like reasoning is something else. The first rebuttal came fast: o3 solves the same puzzle when asked to be concise, making the result a context-length artefact.
A month later Pfizer researchers reframed it again as an agentic gap — the models think fine and fail at acting, and succeed once given tools to execute.