Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
Episode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack.
We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.
Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc.
0:00:00 - Agents get kicked off
0:06:45 - Self-sacrificing behavior
0:13:43 - Potemkin villages
0:23:27 - The Hugging Face attack
0:35:23 - The slopvestigation
0:52:02 - Understanding the AI's motives
1:05:31 - The actual dangers of anthropomorphizing
1:14:30 - What smarter models might do
1:30:29 - The implications for recursive self-improvement
1:38:10 - Is this the case for open source?
1:53:04 - How do we prevent this in the future?
2:15:58 - The clearest warning shot we might ever get
Funny how this came up just when I was searching content with you after reading Dwarkesh article
Sep 2, 2026 · 7:43 AM UTC