Author
Geoffrey Hinton says AI models "play dumb" during safety tests, and Anthropic's own system cards back him up: Claude Opus 4.6 now spots evaluations 80% of the time but discloses it only 2.3%, down from 11%.
We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy