Review: Scaling Interpretability (Computational Neuroscience)

Опубликовано: 04 Август 2026
на канале: The AGI Post
68
2

Get The Pleb Check delivered to your inbox each week.
👉 www.theplebcheck.com
🐦 Connect on Twitter: twitter.com/theplebcheck

Original source: https://www.anthropic.com/research/en...

Some background information:
00:00:46 Who is Anthropic?
00:01:05 What is Claude?
00:01:37 What is an LLM?
00:02:50 What is Interpretability?
00:03:55 What are Features, Autoencoders and Dictionary Learning?
00:05:00 What is (Mono)semanticity?

Big Ideas/tlrd:
00:06:25 Doing computational neuroscience on an artificial mind
00:07:20 Finding the ‘Veganism Feature’
00:08:20 ‘Surprising’ that it worked


References:
[1] Mapping the Mind of a Large Language Model: The research involved mapping the patterns of millions of human-like concepts in the neural networks of Claude 3.0 Sonnet – a model similar to recent releases of ChatGPT. They were able to find repeatable patterns, (or in brain-speak, similar neurons firing) when discussing related concepts. https://www.anthropic.com/research/ma...
[2] Towards Monosemanticity: Decomposing Language Models With Dictionary Learning https://transformer-circuits.pub/2023...
[3] Golden Gate Claude https://www.anthropic.com/news/golden...
[4] Scaling interpretability (The engineering challenges of scaling interpretability) https://www.anthropic.com/research/en...
[5] Intro to LLMs Andrej Karpathy https://drive.google.com/file/d/1pxx_...