Microsoft VASA-1: Lifelike Audio Driven Talking Faces Generated in Real Time

Опубликовано: 12 Июнь 2026
на канале: Data Science Gems
391
10

VASA is a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. VASA-1 is capable of not only producing lip movements that are exquisitely synchronized with the audio, but also capturing a large spectrum of facial nuances and natural head motions that contribute to the perception of authenticity and liveliness. The core innovations include a diffusion-based holistic facial dynamics and head movement generation model that works in a face latent space, and the development of such an expressive and disentangled face latent space using videos. VASA-1 significantly outperforms previous methods along various dimensions comprehensively. VASA-1 delivers high video quality with realistic facial and head dynamics and also supports the online generation of 512×512 videos at up to 40 FPS with negligible starting latency. It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors.

In this video, I talk about the following: What can VASA-1 do? What is the architecture of VASA-1 and how it is trained? How does the Diffusion Model in VASA-1 work? How does VASA-1 perform in comparison with other face video generation models?

For more details, please look at https://arxiv.org/pdf/2404.10667 and https://www.microsoft.com/en-us/resea...

Xu, Sicheng, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. "VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time." arXiv:2404.10667 (2024).