How Does Text-to-Video Work? A 3Blue1Brown-Style Teardown of MiniMax-H3

Опубликовано: 16 Сентябрь 2026
на канале: LLM Internals
43
1

You type one sentence — and a few minutes later you get a video, with sound. What happens in between?

This episode takes the open-source video generation model MiniMax-H3 completely apart: how the text encoder turns one sentence into a row of vectors; why the VAE "shrinks the picture before doing the work"; how feature maps get sliced into a queue of 63,700 positions; why the DiT's two 50s are entirely different things; how attention keeps the sprite in frame 139 looking like frame 1; and how the decoder turns it all back into pictures.

Architecture and defaults come from the official MiniMax-H3 implementation; sequence lengths are measured.

This is episode 1 of the "Make Tech Clear" series. Next up: how this model gets made faster, step by step — a real performance-engineering campaign.

中文原版 (Chinese original):    • Video  

Chinese version (our own Bilibili channel 大模型成长之路 — same team, self-made, not a re-upload): https://www.bilibili.com/video/BV1VBg...