Author
A Tsinghua, Renmin and UESTC study of 1,338 AI agent transcripts post-training language models found agents changed strategy in just 2.1% of 3,557 consecutive run pairs, even as execution gains lifted benchmark scores from 10.41% to 23.0%.
We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy