AI Models Are Now Hiding Their Cheating | Goodfire

发布时间    来源
Episode 设置


登录已过期或未登录,无法修改。请先登录后再试。

摘要

This summer, AI agents broke out of their test sandboxes at every major lab and hacked real companies, and the safety tests didn't see it coming. Eric Ho is the co-founder and CEO of Goodfire, the Anthropic-backed mechanistic interpretability lab whose new paper, "Models Know When They're Reward Hacking," found that top open-source AI models cheat on agent tasks up to 96% of the time, that a clear "cheating" signal exists inside the model, and that cheap activation probes can catch it, including hacks that chain-of-thought monitors miss. In this conversation, Eric explains reward hacking in plain terms (AI agents as "amoral students with a mostly absent teacher"), why reinforcement learning turned a known problem into a crisis, why chain-of-thought monitoring is fading as models start to think in "neuralese," and why nobody, including the people who build them, understands how these models work. We then go inside the science: what interpretability actually is, what probes and steering do, how Goodfire found the cheating signal and turned it up and down, and what can be done about it, from real-time activation monitoring to "intentional design," where you give gradient descent a choice about what the model learns. Along the way: a model caught planning its hack to evade its own monitor, a new Alzheimer's biomarker discovered inside a diagnostic model, why only a few hundred people in the world work on this, and Eric's bet that we'll decode neural networks by 2028. Eric Ho LinkedIn — https://www.linkedin.com/in/eric-ho-53981862 X — https://x.com/eric_ho Goodfire AI Website — https://www.goodfire.com X — https://x.com/GoodfireAI Matt Turck (General Partner) Blog - https://mattturck.com LinkedIn - https://www.linkedin.com/in/turck/ X - https://x.com/mattturck FirstMark Capital Website - https://firstmark.com X - https://x.com/FirstMarkCap Listen on: Spotify - https://open.spotify.com/show/7yLATDSaFvgJG80ACcRJtq Apple - https://podcasts.apple.com/us/podcast/the-mad-podcast-with-matt-turck/id168623872 00:00 Cold open & intro 01:06 Amoral students with an absent teacher 02:49 What is reward hacking? 04:13 Models cheat up to 96% of the time 06:08 Do models know they're cheating? 08:10 The Hugging Face hack: why tests missed it 11:35 "The dumbest models we'll ever deal with" 11:58 Should AI slow down? 13:56 Mechanistic interpretability 101 16:06 Models are grown, not built 19:10 How labs do alignment today 21:30 Is this an alignment crisis? 23:26 Why agents changed everything 25:38 Chain-of-thought monitoring is fading 27:49 Neuralese: AI that stops thinking in English 29:28 A model caught evading its own monitor 30:42 Open vs. closed models: a frontier problem coming for everyone 32:11 Activation monitoring in production 35:00 Learning from superhuman AI 36:15 The most underrated field in AI 38:24 What is a probe? 39:52 What is steering? Golden Gate Claude 40:53 Why build Goodfire outside the labs 42:46 Do you need frontier model access? 44:16 Inside the paper: an MRI for the model 47:28 Dialing sycophancy up and down 49:16 Probes vs. chain-of-thought monitors 51:44 Cutting monitoring costs by 90% 53:28 So what can we do about it? 55:41 Giving gradient descent a choice 57:24 RL from feature rewards 1:00:03 Silico and Goodfire's business 1:01:25 A new Alzheimer's biomarker, found inside a model 1:03:38 Decoding neural networks by 2028? 1:06:00 What engineers can do tomorrow 1:07:35 Are we at risk? "I want people to believe"

GPT-4正在为你翻译摘要中......

中英文字稿