Generative Modeling Meets Deep Learning: A Modern Statistical Approach to Music Signal Analysis.

Опубликовано: 24 Июль 2026
на канале: C4DM - Centre for Digital Music
132
5

Generative Modeling Meets Deep Learning: A Modern Statistical Approach to Music Signal Analysis.

Author: Kazuyoshi Yoshii (Kyoto University)


MIP-Frontiers final workshop October 2021


Abstract
I will present a modern statistical approach to automatic audio-to-score transcription based on an effective combination of an acoustic model, a language model, and an inference model. These models are implemented with classical probabilistic models (e.g., HMM) or with deep neural networks (e.g., LSTM) for richer expression capabilities if necessary. As subtasks of music transcription based on this approach, I will introduce music structure analysis and automatic transcription of singing voice, chords, keys, drums for popular music.
I will also introduce the state-of-the-art automatic piano transcription system that can yield decent symbolic piano scores.

Bio
Kazuyoshi Yoshii received the M.S. and Ph.D. degrees in informatics from Kyoto University, Kyoto, Japan, in 2005 and 2008, respectively.
He is an Associate Professor at the Graduate School of Informatics, Kyoto University, and concurrently the Leader of the Sound Scene Understanding Team, Center for Advanced Intelligence Project (AIP), RIKEN, Tokyo, Japan. His research interests include music information processing, audio signal processing, and statistical machine learning.


More information on Kazuyoshi Yoshii: http://sap.ist.i.kyoto-u.ac.jp/member...


More information on MIP-Frontiers: https://mip-frontiers.eu/