AI with multimodal sensing at the Embedded Vision Summit 2023

Опубликовано: 15 Август 2026
на канале: AI with Sohini
495
7

The Embedded Vision Summit is the premier conference for innovators and engineers incorporating computer vision and visual/perceptual AI in products. Conference Registration Link: https://embeddedvisionsummit.swoogo.c...

Unlike academic conferences, the Summit is relentlessly focused on practical computer vision and edge AI. It’s focused on people who are building products now, or in the next year or two, and who need working solutions to real-world problems. The Summit has two major parts: first, there are great presentations and panels about practical edge AI and CV; second are Technology Exhibits, where you can meet with 70+ exhibitors, see and touch their products and, critically, talk directly to the people who developed them.

First major topic is the keynote by Prof. Kristen Grauman from UT Austin and Facebook Research Team on Multimodal sensing – that is, using data from different types of sensors (and even non: https://www.cs.utexas.edu/users/grauman/. Here we will dive into the multi-modality of data to combine audio, language, and vision. We will learn about Ego4D, a massive new open-sourced multimodal egocentric dataset that captures the daily-life activity of people around the world. (   • EGO4D - Around the World in 3,000 Hours of...  ).This data can be layered with natural language queries (“Where did I last see X? Did I leave the garage door open?”), injecting semantics from text and speech into powerful video representations, and learning audio-visual models to understand a camera wearer’s physical environment or augmenting their hearing in busy places.

Other Summit talks on this theme include: "Tracking and Fusing Diverse Risk Factors to Drive a SAFER Future” (by Tahmida Mahmud and Stefan Heck of Nauto). This work used 1 billion miles of real-life driving data to find out, feeding it into a model fusing 26 road context, driver action/attention and vehicle dynamics factors to predict collisions, near misses and safe driving. This work found that drivers can handle most single risks, but multiple simultaneous risks can be deadly. https://embeddedvisionsummit.com/2023...

It is important to note that all these updates in AI including the chatGPT capabilities have been fueled by enhanced hardware performances to process large volumes of heterogeneous data. We learn about these hardware capabilities in another talk by Pulin Desai of Cadence titled “Tensilica Processor Cores Enable Sensor Fusion for Robust Perception” https://embeddedvisionsummit.com/2023...

Which shows some of the heterogeneous sensor combinations and sensor fusion approaches that are gaining adoption in applications such as driver assistance and mobile robots. You will get to see how the Cadence Tensilica ConnX and Vision processor IP core families and their associated software tools and libraries support sensor fusion applications with high performance, efficiency and ease of development.

The conference is in person in Santa Clara, California. If you can make it, please register using promo code SUMMIT23-SOHINI to save 15% on registration until April 28 (and 10% after that).