Neel Nanda is a researcher at Google DeepMind working on mechanistic interpretability.
In this conversation, we discuss Neel's background, research methodology, his advice for people who want to get started, but also his research around superposition, toy models of universality and grokking, among other things.
Transcript & Audio: https://theinsideview.ai/neel
Patreon supporters:
Ana Paiva
Tassilo Neubauer
MonikerEpsilon
Alexey Malafeev
Jack Seroy
JJ Hepburn
Max Chiswick
William Freire
Edward Huff
Gunnar Höglund
Ryan Coppolo
Cameron Holmes
Emil Wallner
Jesse Hoogland
Jacques Thibodeau
Vincent Weisser
OUTLINE
00:00 Intro
01:34 Why Neel Started Doing Walkthroughs Of Papers On Youtube
08:36 Induction Heads, Or Why Nanda Comes After Neel
12:56 Detecting Induction Heads In Basically Every Model
15:12 How Neel Got Into Mechanistic Interpretability
16:59 Neel's Journey Into Alignment
22:46 Enjoying Mechanistic Interpretability And Being Good At It Are The Main Multipliers
25:26 How Is AI Alignment Work At DeepMind?
26:23 Scalable Oversight
29:07 Most Ambitious Degree Of Interpretability With Current Transformer Architectures
31:42 To Understand Neel's Methodology, Watch The Research Walkthroughs
33:00 Three Modes Of Research: Confirming, Red Teaming And Gaining Surface Area
35:35 You Can Be Both Hypothesis Driven And Capable Of Being Surprised
37:28 You Need To Be Able To Generate Multiple Hypothesis Before Getting Started
38:32 All the theory is bullshit without empirical evidence and it's overall dignified to make the mechanistic interpretability bet
40:48 Mechanistic interpretability is alien neuroscience for truth seeking biologists in a world of math
42:49 Actually, Othello-GPT Has A Linear Emergent World Representation
45:45 You Need To Use Simple Probes That Don't Do Any Computation To Prove The Model Actually Knows Something
48:06 The Mechanistic Interpretability Researcher Mindset
50:26 The Algorithms Learned By Models Might Or Might Not Be Universal
52:26 On The Importance Of Being Truth Seeking And Skeptical
54:55 The Linear Representation Hypothesis: Linear Representations Are The Right Abstractions
58:03 Superposition Is How Models Compress Information
01:00:52 The Polysemanticity Problem: Neurons Are Not Meaningful
01:06:19 Superposition and Interference are at the Frontier of the Field of Mechanistic Interpretability
01:08:10 Finding Neurons in a Haystack: Superposition Through De-Tokenization And Compound Word Detectors
01:09:40 Not Being Able to Be Both Blood Pressure and Social Security Number at the Same Time Is Prime Real Estate for Superposition
01:15:39 The Two Differences Of Superposition: Computational And Representational
01:18:44 Toy Models Of Superposition
01:26:16 How Mentoring Nine People at Once Through SERI MATS Helped Neel's Research
01:32:02 The Backstory Behind Toy Models of Universality
01:35:56 From Modular Addition To Permutation Groups
01:39:29 The Model Needs To Learn Modular Addition On A Finite Number Of Token Inputs
01:42:31 Why Is The Paper Called Toy Model Of Universality
01:46:53 Progress Measures For Grokking Via Mechanistic Interpretability, Circuit Formation
01:53:22 Getting Started In Mechanistic Interpretability And Which WalkthroughS To Start With
01:56:52 Why Does Mechanistic Interpretability Matter From an Alignment Perspective
01:59:18 How Detection Deception With Mechanistic Interpretability Compares to Collin Burns' Work
02:01:57 Final Words From Neel