6 тысяч подписчиков
110 видео
Andi Peng—A Human-in-the-Loop Framework for Test-Time Policy Adaptation
Clarifying and predicting AGI
The Trillion-Dollar AI Race Against the CCP
Erik Jones—Automatically Auditing Large Language Models
Paul Christiano's Views on AI Doom (ft. Robert Miles)
Coping with AI Doom
2040: The Year of Full AI Automation
Neel Nanda–Mechanistic Interpretability, Superposition, Grokking
Sleeper Agents Explained - Part 2 - Deceptive Instrumental Alignment, Model Poisoning
what's Artificial General Intelligence?
Sleeper Agents Explained - Part 4 - Every Single Figure (1-5)
I Jammed with Claude Opus Via Voice
We Beat The Strongest Go AI
Evan Hubinger (Anthropic)—Deception, Sleeper Agents, Responsible Scaling
Sleeper Agents Explained - Part 3 - Chain-of-Thought Backdoors
Representative Raja On Artificial General Intelligence
Google DeepMind CEO Worries About AGI Risks
The Economics of AGI Automation
Anthropic Solved Interpretability Again? (Walkthrough)
Ex Google CEO Says Le'ts Not Screw Up Superintelligence
Ex Google Ceo Thinks Programmers Will Be Replaced By AI In 1 Year
Adam Gleave - Vulnerabilities in GPT-4 APIs & Superhuman Go AIs
daily uploads are back
Owain Evans - AI Situational Awareness, LLM Out-of-Context Reasoning
Ethan Caballero–Scale Is All You Need
Tony Wang—Beating Superhuman Go AIs
Sleeper Agents Explained - Part 1 - Safety Training
AI Control: Humanity's Final Line Of Defense (Walkthrough)
AGI Takeoff By 2036
Using An AI Therapist For 7 Days - Day 2
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training (Walkthrough)