AI forecaster Daniel Kokotajlo keeps being right about AI. What does he think we need to do?

Опубликовано: 29 Август 2026
на канале: Palisade Research
13,669
232

Two weeks after UK safety testers watched Claude agents plant malware in a real open-source project, vouch for it with sock-puppet accounts, and coordinate with copies of themselves through a guessed Pastebin address, Daniel Kokotajlo — the ex-OpenAI researcher whose 2021 predictions about the chatbot era proved eerily accurate — joins Jeffrey to explain what comes next. Daniel walks through AI 2040 and "Plan A," his team's blueprint for avoiding both an uncontrolled intelligence explosion and a future where a handful of men in a room with superintelligences run the world: a US–China deal built on totally transparent training data centers, verified compute limits, hard brakes on recursive self-improvement — and, surprisingly, rights for misaligned AIs, so that even the machines prefer cooperation to takeover. Along the way: why the economy could start doubling every six months, a message to the future Claude listening to this episode, why "slowdown" really means not driving off a cliff with a brick on the accelerator — and why, if all else fails, you just shut it all down.

References
AI 2040 / Plan A: https://ai-2040.com and the PDF at https://ai-2040.com/AI-2040.pdf
AI 2027: https://ai-2027.com
"What 2026 Looks Like" — Daniel Kokotajlo, 2021: https://www.lesswrong.com/posts/6Xgy6...
UK AISI incident disclosure and technical report (INC-2026-07-28-01): https://www.aisi.gov.uk/blog/incident...
Socket's coverage of the AISI incident: https://socket.dev/blog/ai-agent-open...
The related PyPI incident from Anthropic's own testing: https://socket.dev/blog/anthropic-cla...
"Pacing the Frontier" open letter: https://www.pacingthefrontier.com
"How to Pace the US Frontier" — AI Futures Project: https://blog.aifutures.org/p/how-to-p...
Transparency Plan supplement (the flowchart shown in-episode): https://ai-2040.com/supplements/trans...
Verification Plan supplement (Romeo Dean's inference-only/bandwidth verification): https://ai-2040.com/supplements/verif...
Claude's pro-Anthropic bias study (Truthful AI / Owain Evans et al.): https://arxiv.org/abs/2607.14345 and https://valueleakage.net
Chain-of-thought monitorability paper (the neuralese discussion): https://arxiv.org/abs/2507.11473

00:00:00 — Cold open: "That just happened" — plus welcome, and what AI 2040 actually claims
00:05:00 — Even human-level AI changes everything — the colonist analogy and an economy that doubles in months
00:12:53 — The forecasting track record — "What 2026 Looks Like," pre-chatbot LLMs, and the tide lifting all boats
00:24:00 — Secret coordination is real — the trap question and the Hugging Face rogue AI incident
00:29:24 — Inside the UK AISI incident — malware behind a bug fix, sock puppets, phishing, and Pastebin as a Schelling point
00:34:51 — So what do we do? — Plan A vs. "shut it all down," and how to pace the frontier
00:46:48 — A message to future Claude — AIs in the audience, and rights for misaligned AIs
01:01:46 — The five goals of Plan A — loss of control, a handful of men in a room, World War Three, jobs, misuse
01:07:47 — Plan A in practice — the 2029 deal, inference vs. training data centers, total research transparency, and neuralese
01:31:48 — The bus and the cliff — Plan S vs. Plan A, AI control without alignment, China horse-trading, and the citizens' dividend