Join us for this AI webinar where we delve into the cutting-edge realm of multimodal large language models. From this webinar you will understand the significance, challenges, and state-of-the-art solutions associated with developing AI systems that can seamlessly integrate visual, auditory, and textual data in real time.
Topics include:
Enhancing human-AI interaction at the edge: Learn about live, camera-based instruction, a new emerging use-case for AI at the edge.
How real-time instruction can benefit from end-to-end training through multi-modal streaming models.
Fitness coaching as a challenging testbed application, where AI models need to provide real-time, context-aware feedback by recognizing and responding to complex human actions.
Improving visual grounding: explore how training multi-modal models on low-level tasks like object detection and tracking can improve their ability to perform high-level reasoning over videos, leading to groundbreaking performance in visual reasoning tasks.
Speaker: Roland Memisevic
Senior Director of Engineering at Qualcomm Canada ULC
Roland Memisevic is senior director of engineering at Qualcomm Canada ULC. He joined Qualcomm in 2021 after the acquisition of TwentyBN, an AI startup that he founded in 2015. Roland received a PhD in Computer Science from the University of Toronto in 2008 doing research on neural networks. He continued his work as a post-doc at the University of Toronto and at ETH Zurich and has been on faculty at the University of Frankfurt and at the MILA institute at the University of Montreal. Roland’s research interests are in end-to-end learning and the emergence of human-like common sense in neural networks.
Moderated by Armina Stepan (Qualcomm Technologies Netherlands).