2020 Intelligent Sensing Summer School
Audio-visual modalities fusion for efficient emotion recognition with active learning capabilities
Antonios Gasteratos, DUT
Speech enhancement is the task of extracting clean speech from a noisy mixture. While the sound signal may suffice in low-noise conditions, the visual input can bring complementary information to help the enhancement process. In this tutorial, two strategies based on variational auto-encoders will be described and compared: (i) the systematic use of both audio and video against (ii) the automatic inference of the optimal mixture between audio and video.
As part of the 2020 Intelligent Sensing Summer School: http://cis.eecs.qmul.ac.uk/school2020...