Author
A team from NYU, Modal Labs, UCLA, UIUC and Columbia trained 10 chess models from 5M to 1B parameters to fit a joint pretraining-RL scaling law, with the compute-optimal RL share rising from about 19-20% at 50-80M parameters to 28% at 680M.
We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy