Author
Ang Li
By and about Ang Li
PoLar: Dynamically skipping or looping transformer layers boosts math accuracy by 60.9%
Researchers introduce Program-of-Layers (PoLar), a training-free framework that dynamically skips, keeps, or repeats transformer layer segments per input, achieving up to 87.8% on DART-Math with Qwen2.5-3B, without modifying base model weights.
newsNYU fits a pretraining–RL scaling law on chess, with RL's optimal share rising to 28%
A team from NYU, Modal Labs, UCLA, UIUC and Columbia trained 10 chess models from 5M to 1B parameters to fit a joint pretraining-RL scaling law, with the compute-optimal RL share rising from about 19-20% at 50-80M parameters to 28% at 680M.
blogAI Socratic March 2026
Top AI updates from Jan 15 to Feb 15 2026
newsAlibaba Qwen 3.5 Small Model Series
Alibaba released four Qwen 3.5 small models ranging from 0.8B to 9B parameters with native multimodal capabilities, then saw its lead researcher Junyang Lin and three others unexpectedly resign a day later.