Lumiere is the most recent Text to Video (T2V) model released by Google. The model addresses the problem of temporal inconsistency prevalent in video generation models today. For this, it introduces a novel space-time UNet architecture and MultiDiffusion design. This video goes into the details.
⌚️ ⌚️ ⌚️ TIMESTAMPS ⌚️ ⌚️ ⌚️
0:00 - Intro
1:06 - Temporal Consistency
2:36 - Text to Video generation Pipeline
4:27 - UNet Overview
5:08 - Space-Time UNet or STUNet
7:03 - MultiDiffusion
8:12 - Qualitative Results
9:00 - Quantitative Results
9:30 - Conclusion
MY KEY LINKS
YouTube: / @aibites
Twitter: / ai_bites
Patreon: / ai_bites
Github: https://github.com/ai-bites
WHO AM I?
I am a Machine Learning Researcher / Practioner who has seen the grind of academia and start-ups equally. I started my career as a software engineer 15 years back. Because of my love for Mathematics (coupled with a glimmer of luck), I graduated with a Master's in Computer Vision and Robotics in 2016 when the now happening AI revolution just started. Life has changed for the better ever since.
#machinelearning #deeplearning #aibites