In a groundbreaking collaboration, Meta and NYU present a paradigm shift in AI training. Proposing Self-Rewarding Language Models, the study leverages Language Model-as-a-Judge prompting during training, enabling models to generate their own rewards. Through Iterative DPO training, improvements in instruction following and the capacity to self-generate high-quality rewards are demonstrated. Fine-tuning Llama 2 70B over three iterations yields a model that surpasses benchmarks on the AlpacaEval 2.0 leaderboard, outperforming Claude 2, Gemini Pro, and GPT-4 0613. This work hints at the potential for models to continually advance in both axes, pushing the boundaries of AI's evolutionary trajectory.