Thread · 4 stories · Mar 30 – Sep 3
The benchmark humans aced and machines scored under 1% on fell within six months — first to scaffolding, then to Astra.
Jump to timeline ↓ARC-AGI-3 launched as the cleanest capability gap in the field: humans at 100%, AI systems below 1%, measuring skill acquisition rather than recall.
What closed it was mostly not new models. Compaction tricks tripled GPT-5.6 Sol's score without touching the weights, an open harness took Best@1 from 30% to 95.5%, and Astra's headline number turned out to depend on which harness you ran it in.