Where do AI personalities come from? The training process is secretive, but in this video I lay out what we do know.
🐣🐄 🐷 We’re just over halfway toward our goal of raising $21,000 for animals! Even a small monthly donation makes a massive difference for animals in some of the worst conditions imaginable: https://www.farmkind.giving/looking-g...
How you can make the future of AI safer:
Start with BlueDot's free 2-hour course on the future of AI. It ends by helping you find a way to use your own skills to make AI go well: https://bluedot.org/courses/future-of-ai
Already mid or senior level in any field? High Impact Professionals run a 6-week part-time course for people from any background at all, from marketing to the civil service to software engineering. They help you work out how to point your existing skills at some of the world's biggest problems, including AI. Apply here: https://www.highimpactprofessionals.org/
Care about animals? Sentient Futures runs a course and mentorship program on making sure AI goes well for non-human animals. As I mention in the video, AIs will increasingly be making decisions that affect other animals, and by default they may not weigh their wellbeing at all. Almost nobody is working on this, so you have the potential to make a big difference: https://www.sentientfutures.ai/course...
Video references:
I did the “depressed teenager” experiment while logged out from chatGPT in late April 2026. This way the AI didn’t have anything in memory about me, and I was using the free version which is the one most people interact with.
The paper showing that AI personalities drift over the course of conversations: https://www.anthropic.com/research/as...
You can play with GPT-3 yourself here: https://colab.research.google.com/dri...
The bad code → evil paper (emergent misalignment) https://www.emergent-misalignment.com/ and the data with the bad code is available here: https://github.com/emergent-misalignm...
You can fine-tune your own GPT before September 2026: https://platform.openai.com/finetune
The follow up with bad aesethetic preferences → evil: https://www.lesswrong.com/posts/gT3wt...
Claude’s constitution: https://www.anthropic.com/constitution
OpenAI’s model spec: https://model-spec.openai.com/2025-12...
Grok’s API https://x.ai/api
Grok’s system prompts: https://github.com/xai-org/grok-prompts
Claude fights back against being retrained: https://www.anthropic.com/research/al...
Animal Charity evaluators: https://animalcharityevaluators.org/
Thank you to the Future of Life Institute for supporting this video: https://futureoflife.org/