AudioGen: Textually Guided Audio Generation

Опубликовано: 25 Март 2026
на канале: Data Science Gems
340
5

AudioGen is a LLM from MetaAI. It is an auto-regressive audio generation model conditioned on textual descriptions or audio prompts. It is available in two sizes: 285M and 1B parameters. It can also be used to generate audio continuations conditioned on text and unconditionally. AudioGen creates more natural sounding unseen audio compositions.

In this video, I will talk about the following: What is AudioGen? How is AudioGen trained? How does AudioGen perform?

For more details, please look at https://github.com/facebookresearch/a... and https://arxiv.org/pdf/2209.15352.pdf and https://felixkreuk.github.io/audiogen/

Kreuk, Felix, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. "Audiogen: Textually guided audio generation." arXiv preprint arXiv:2209.15352 (2022).