Microsoft Phi-3

Опубликовано: 19 Февраль 2026
на канале: Data Science Gems
561
11

Phi-3-mini is a 3.8 billion parameter language model trained on 3.3 trillion tokens, with overall performance that rivals models such as Mixtral 8x7B and GPT-3.5. Despite its small size, it achieves 69% on MMLU and 8.38 on MT-bench. The innovation lies in the dataset for training, which is a scaled-up version of the one used for phi-2, composed of heavily filtered web data and synthetic data. The model is also aligned for robustness, safety, and chat format. Additionally, initial parameter-scaling results are provided for a 7B and 14B models trained for 4.8T tokens, called phi-3-small and phi-3-medium, both of which are significantly more capable than phi-3-mini. Respectively, they achieve 75% and 78% on MMLU, and 8.7 and 8.9 on MT-bench.

In this video, I talk about the following: How is Phi-3 trained? How does Phi-3 perform?

For more details, please look at https://azure.microsoft.com/en-us/blo... and https://arxiv.org/pdf/2404.14219

Abdin, Marah, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach et al. "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone." arXiv:2404.14219 (2024).