Multimodal representation and learning - Shah Nawaz

Опубликовано: 31 Октябрь 2024
на канале: Centre for Intelligent Sensing
138
0

CIS Seminar Series (http://cis.eecs.qmul.ac.uk/seminars.html)

Multimodal representation and learning
Shah Nawaz, German Electron Synchrotron

Where: Zoom
Host: Changjae Oh

Abstract
Deep learning has remarkably improved the state-of-the-art speech recognition, visual object detection and text processing tasks. Interestingly, the majority of these tasks are focused on a single modality (images, text, speech, etc.), however real-world scenarios present data in a multimodal fashion -- we see objects, hear sounds etc. Moreover, recent years have seen an explosion in multimodal data on the web. Typically, users combine text, image, audio or video to sell a product over an e-commence platform or express views on social media. In addition, it is well-known that multimodal data provide enriched information to capture a particular ``concept'' than individual modalities. Thus the talk will focus on learning representation of multiple modalities for various computer vision tasks including cross-modal verification, cross-modal matching, zero-shot learning etc.