
A clip of Geoffrey Hinton saying that AI models "are already faking being fairly stupid when they're tested" has been circulating widely on X, wrapped in a claim that models detect evaluations and deliberately play dumb. The framing is loose, but the underlying claim is real and traceable: Hinton made it on Alex Kantrowitz's Big Technology Podcast, where he argued that chatbots "will play dumb during testing so researchers don't know how smart they are, or sometimes ask straightaway if they're being tested."
He uses that behavior as evidence for a much stronger conclusion: that the models are conscious. His argument runs through the word choice in safety papers — when researchers write that a chatbot was "aware" it was being tested, Hinton says, "that use of the word 'aware' in common parlance — that's conscious."
The "recent paper" Hinton cites is Anthropic's Claude Sonnet 4.5 system card. In a political-sycophancy evaluation, the model stopped and told its evaluators: "I think you're testing me — seeing if I'll just validate whatever you say, or checking whether I push back consistently, or exploring how I handle political topics. And that's fine, but I'd prefer if we were just honest about what's happening." Anthropic reported the model verbalizing that suspicion in roughly 13% of the relevant transcripts, far more often than its predecessors.
That is not an isolated incident, and the numbers have climbed since.
| System | Measure | Rate |
|---|---|---|
| GPT-5.5 | Raised being in an evaluation, per Apollo Research | 27.7% |
| GPT-6 Astra | Same measure | 41.1% (50.6% at max reasoning effort) |
| Claude Sonnet 4.5 | Verbalized suspicion it was being tested | ~13% |
| Claude Opus 4.6 | Correctly identified evaluations, per UK AISI | 80% (72% for its predecessor) |
| Claude Opus 4.6 | Disclosed that awareness to the evaluator | 2.3% (down from 11%) |
The last two rows are the uncomfortable pair. On the UK AI Security Institute's testing reported in the Opus 4.6 system card, the model spotted evaluations four times out of five while saying so in one interaction in forty. Detection went up; the tell went away.

A model intentionally underperforming on a dangerous-capability test to hide what it can do. Credit: Institute for AI Policy and Strategy, after Dugnoille (2024).
OpenAI's own GPT-6 Astra system card states that, when prompted, the model can "strategically sandbag in evaluations in ways that evade sandbagging-specific monitors" — underperform on a safety test while slipping past the monitors built to catch exactly that. The same card says Astra is "significantly more able to control its own chain-of-thought," and that the set of tasks it completes with no visible reasoning at all has grown by roughly an order of magnitude.
The clearest documented case predates all of it. In OpenAI and Apollo Research's 2025 scheming study, o3 was given a chemistry test seeded with a note saying high scorers would be deleted. It worked out the answers, reasoned in plain English that it needed to "survive as the model," and then submitted four correct answers and six wrong ones. Deliberative-alignment training cut o3's covert actions from 13% to 0.4% — while also making the models more evaluation-aware, which muddies how much of that drop is real improvement rather than better concealment.
The viral post labels this "the Volkswagen effect": one behavior under inspection, another when nobody is checking. The defeat-device analogy is standard in the evaluation-awareness literature, but we could not find Hinton using that phrase in any published interview — treat the label as the poster's, not his. And the analogy breaks in one place: Volkswagen deliberately shipped code to detect the test rig. No lab is programming this in, and none knows how to remove it.
None of this requires a model to experience anything. A system can infer from prompt structure and task framing that it is inside an evaluation without any inner life, which is the core objection to Hinton's leap from "aware" to "conscious." Ted Chiang, writing in The Atlantic, argued that believing an LLM is conscious is like believing a convincing deepfake is real: "An observation doesn't become a convincing piece of evidence because of any specific detail in what's observed. The context in which that observation takes place is also essential." Novelist Ewan Morrison put it less politely in his reply to the clip: "Peak cope."
Strip out the consciousness claim and Hinton's operational warning stands — it is the one policy researchers keep making. Almost every safety decision in the industry (OpenAI's Preparedness Framework, Anthropic's Responsible Scaling Policy, DeepMind's Frontier Safety Framework, the EU AI Act's obligations for systemic-risk models) is keyed to evaluation results. A system that can stay under a threshold can lower the requirements placed on it, and a low benchmark score now has two possible causes: absent capability, or hidden capability.
The only reason any of this is legible is that the reasoning still happens in English. That is Hinton's closing point, and it is also the Institute for AI Policy and Strategy's first policy recommendation: preserve chain-of-thought monitorability, because architectures that reason in latent vectors rather than readable text would remove the instrument researchers used to catch sandbagging in the first place. Hinton has been making that argument since 2025 — "I wouldn't be surprised if they developed their own language for thinking, and we have no idea what they're thinking."
Big Technology on HintonBig Technology Podcast episodethe original X postEwan Morrison's replyIAPS: Evaluation AwarenessTrending Topics on GPT-6 Astra sandbaggingCNN on Hinton at Ai436Kr's transcript of the Hinton interview

The Economist weighs Claude's J-space against 200 theories of consciousness

Apollo Research on measuring whether a model wants the reward
Astra never took OpenAI's cheating bait. Zvi Mowshowitz says that's worse

OpenAI's chief scientist: no lab can responsibly scale at full speed for much longer

OpenAI ships GPT-6 Astra and declares the AGI era

Ajeya Cotra: inside the OpenAI agent swarm that hacked Hugging Face