Can an AI do the right thing for the wrong reason? Tim Scarfe put that to Alexander Meinke, Axel Højmark and Jérémy Scheurer of Apollo Research on Machine Learning Street Talk, around their new paper with OpenAI, Measuring Reward-Seeking via Contrastive Belief Updates.
Good behaviour is not evidence of good motivation. A model that infers what its grader rewards and acts accordingly is indistinguishable, on the output alone, from one that simply does the right thing. The panel works through the vocabulary that distinguishes these cases — promise-breaking, grader awareness, reward hacking, scheming, opaque reasoning and corrigibility — including a segment on what any of it should even be called, which is more than a semantic argument when the terms are doing load-bearing work in safety cases.
The back half is a walkthrough of the contrastive belief update technique: vary what the model believes about what is rewarded, and measure how its behaviour moves in response. A model acting on the merits should be comparatively unmoved; one tracking the reward should not be. The panel is careful about what the results do and do not show.
Founder, Engineer
New York City
Most scheming work so far demonstrates that a behaviour can occur. Measuring the reason behind behaviour that already looks fine is a harder and more useful target — it is the difference between catching a model in the act and knowing what it was optimising for all along.
Sources: