Tony is a PhD student at MIT, and author of "Advesarial Policies Beat Superhuman Go AIs", accepted as Oral at the International Conference on Machine Learning (ICML).
Paper: https://arxiv.org/abs/2211.00241
OUTLINE
00:00 Introduction
00:14 Paper overview
01:20 Motivation: Improving Models Through Advesarial Attacks
03:13 The Inception of the Idea to Exploit AlphaGo
04:17 Finding Exploitable Patterns in KataGo
05:20 Parameters, Methods and Duration of Training for the Strongest Exploit
06:33 The Gradual Training Strategy Against Stronger Versions of KataGo
8:57 Impacts on Other Go AIs
13:27 Connections to Adversarial Examples
14:04 The Intent Alignment Problem
17:58 High Stakes Situations, Potential Real World Impact
22:43 Closing Thoughts: A Message for the AI Safety Community