Tsvi Benson-Tilsen spent seven years tackling the alignment problem at the Machine Intelligence Research Institute (MIRI). Now, he delivers a sobering verdict: humanity has made "basically 0% progress" towards solving it.
Tsvi unpacks foundational MIRI research insights like "timeless decision theory" and "corrigibility", which expose just how little humanity actually knows about controlling superintelligence.
These theoretical alignment concepts give us a way of peering into the future, revealing the non-obvious, structural laws of "intellidynamics" that will ultimately determine our fate.
Time to learn some of MIRI's greatest hits.
Stay tuned for part 3 of my interview with Tsvi where we debate AGI timelines!
Watch part 1 of my interview with Tsvi about engineering human superbabies: • Former MIRI Researcher Solving AI Alignmen...
0:00 Episode Highlights
0:49 Humanity Has Made 0% Progress on AI Alignment
1:56 MIRI Alignment Greatest Hits: Reflective Probability Theory, Logical Uncertainty, Reflective Stability
6:56 Why Superintelligence is So Hard to Align: Self-Modification
8:54 AI Will Become a Utility Maximizer (Reflective Stability)
12:26 The Effect of an “Ontological Crisis” on AI
14:41 Why Modern AI Will Not Be ‘Aligned By Default’
18:49 Debate: Have LLMs Solved the "Ontological Crisis" Problem?
25:56 MIRI Alignment Greatest Hit: Timeless Decision Theory
35:17 MIRI Alignment Greatest Hit: Corrigibility
37:53 No Known Solution for Corrigible and Reflectively Stable Superintelligence
39:58 Recap