59494 подписчиков
676 видео
INTERSPEECH 2021 Acoustic Echo Cancellation Challenge - (Oral presentation)
Fast Text-Only Domain Adaptation of RNN-Transducer Prediction Network - (Oral presentation)
SponSpeech: Adaptive Text to Speech for Spontaneous Style - (3 minutes introduction)
Intonation Transcription and Modelling in Research and Speech Technology Applications
A Benchmark of Dynamical Variational Autoencoders applied to Speech Spectrogram Modeling - (Oral...
EfficientSing: A Chinese Singing Voice Synthesis System Using Duration-Free Acoustic Model and H...
Tied & Reduced RNN-T Decoder - (3 minutes introduction)
Automated Detection of Voice Disorder in the Saarbruecken Voice Database: Effects of Pathology S...
SC-GlowTTS: an Efficient Zero-Shot Multi-Speaker Text-To-Speech Model - (3 minutes introduction)...
Token-Level Supervised Contrastive Learning for Punctuation Restoration - (3 minutes introductio...
Deep audio-visual speech separation based on facial motion - (3 minutes introduction)
Cross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis - (3 minutes int...
Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings - (3 minutes introduction)
Graph-based Label Propagation for Semi-Supervised Speaker Identification - (3 minutes introducti...
Digital Einstein Experience: Fast Text-to-Speech for Conversational AI - (3 minutes introduction...
VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis - (Oral p...
Bridging the gap between streaming and non-streaming ASR systems by distilling ensembles of CTC ...
NeMo (Inverse) Text Normalization: From Development To Production - (longer introduction)