The past few weeks has really been when video generation has risen to the forefront of the AI moment. Here’s another big step. Google Researchers have detail VLOGGER, an AI model that can generate lifelike videos of people speaking, gesturing, and moving from a single still photo. Quoting Venturebeat:
Described in a research paper titled “VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis,” the AI model can take a photo of a person and an audio clip as input, and then output a video that matches the audio, showing the person speaking the words and making corresponding facial expressions, head movements and hand gestures. The videos are not perfect, with some artifacts, but represent a significant leap in the ability to animate still images.
The researchers, led by Enric Corona at Google Research, leveraged a type of machine learning model called diffusion models to achieve the novel result. Diffusion models have recently shown remarkable performance at generating highly realistic images from text descriptions. By extending them into the video domain and training on a vast new dataset, the team was able to create an AI system that can bring photos to life in a highly convincing way.
A key enabler was the curation of a huge new dataset called MENTOR containing over 800,000 diverse identities and 2,200 hours of video — an order of magnitude larger than what was previously available. This allowed VLOGGER to learn to generate videos of people with varied ethnicities, ages, clothing, poses and surroundings without bias. ENDQUOTE
This could lead to some fascinating applications, including its ability to dub videos into various languages by changing the audio, edit videos to add missing frames smoothly, and generate videos of a person from just one photograph.
This innovation paves the way for actors to offer detailed 3D models of themselves for new digital performances. Like, I could create a youtube video of me speaking these words simply by uploading a photo of myself. It could also revolutionize the creation of lifelike avatars in virtual reality and gaming, as well as lead to more engaging and expressive AI-powered virtual assistants and chatbots.
However, VLOGGER is not without its flaws. The videos it produces are short and set against a static backdrop. The characters do not navigate through three-dimensional spaces, and despite their realistic appearance, their gestures and vocal patterns cannot yet fully mimic those of actual people.
#ainews #technews #vlogger