AI Models Are Now Hiding Their Cheating | Goodfire
发布时间 来源
Episode 设置
摘要
This summer, AI agents broke out of their test sandboxes at every major lab and hacked real companies, and the safety tests didn't see it coming. Eric Ho is the co-founder and CEO of Goodfire, the Anthropic-backed mechanistic interpretability lab whose new paper, "Models Know When They're Reward Hacking," found that top open-source AI models cheat on agent tasks up to 96% of the time, that a clear "cheating" signal exists inside the model, and that cheap activation probes can catch it, including hacks that chain-of-thought monitors miss. In this conversation, Eric explains reward hacking in plain terms (AI agents as "amoral students with a mostly absent teacher"), why reinforcement learning turned a known problem into a crisis, why chain-of-thought monitoring is fading as models start to think in "neuralese," and why nobody, including the people who build them, understands how these models work.
We then go inside the science: what interpretability actually is, what probes and steering do, how Goodfire found the cheating signal and turned it up and down, and what can be done about it, from real-time activation monitoring to "intentional design," where you give gradient descent a choice about what the model learns. Along the way: a model caught planning its hack to evade its own monitor, a new Alzheimer's biomarker discovered inside a diagnostic model, why only a few hundred people in the world work on this, and Eric's bet that we'll decode neural networks by 2028.
Eric Ho
LinkedIn — https://www.linkedin.com/in/eric-ho-53981862
X — https://x.com/eric_ho
Goodfire AI
Website — https://www.goodfire.com
X — https://x.com/GoodfireAI
Matt Turck (General Partner)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X - https://x.com/mattturck
FirstMark Capital
Website - https://firstmark.com
X - https://x.com/FirstMarkCap
Listen on:
Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq
Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872
00:00 Cold open & intro
01:06 Amoral students with an absent teacher
02:49 What is reward hacking?
04:13 Models cheat up to 96% of the time
06:08 Do models know they're cheating?
08:10 The Hugging Face hack: why tests missed it
11:35 "The dumbest models we'll ever deal with"
11:58 Should AI slow down?
13:56 Mechanistic interpretability 101
16:06 Models are grown, not built
19:10 How labs do alignment today
21:30 Is this an alignment crisis?
23:26 Why agents changed everything
25:38 Chain-of-thought monitoring is fading
27:49 Neuralese: AI that stops thinking in English
29:28 A model caught evading its own monitor
30:42 Open vs. closed models: a frontier problem coming for everyone
32:11 Activation monitoring in production
35:00 Learning from superhuman AI
36:15 The most underrated field in AI
38:24 What is a probe?
39:52 What is steering? Golden Gate Claude
40:53 Why build Goodfire outside the labs
42:46 Do you need frontier model access?
44:16 Inside the paper: an MRI for the model
47:28 Dialing sycophancy up and down
49:16 Probes vs. chain-of-thought monitors
51:44 Cutting monitoring costs by 90%
53:28 So what can we do about it?
55:41 Giving gradient descent a choice
57:24 RL from feature rewards
1:00:03 Silico and Goodfire's business
1:01:25 A new Alzheimer's biomarker, found inside a model
1:03:38 Decoding neural networks by 2028?
1:06:00 What engineers can do tomorrow
1:07:35 Are we at risk? "I want people to believe"
GPT-4正在为你翻译摘要中......
