Teaching cars to see at scale - Computer Vision at Motional - Dr. Holger Caesar

Опубликовано: 23 Июль 2026
на канале: 2d3d.ai
790
21

Autonomous vehicles have an enormous potential to save lives and reduce greenhouse gas emissions, as well as making transportation more flexible and comfortable. Machine learning is a key tool that enables autonomous vehicles to perceive their environment and learn from a constantly growing body of driving data.
In this talk I present how we develop perception systems at Motional. Besides presenting our perception algorithms (PointPillars, PointPainting) and public benchmark datasets (nuScenes, nuImages), I discuss how to build real-world machine learning solutions. A particular focus will be on the aspects that academia cannot solve for us: selecting the right data using Active Learning, defining what to annotate and scaling the pipeline up to previously unseen quantities of data.

Lecture references:   / lecture_references_teaching_cars_to_see_at...  

00:00 Intro
01:48 What is Motional?
06:20 Algorithms - PointPillars
08:50 Algorithms - PointPainting
11:30 CoverNet: Multimodal Behavior Prediction using Trajector
12:23 Network architecture
17:28 nuScenes: friends and followers
25:47 Data-driven development
29:19 Active Learning
33:17 nuScenes meets SiaSearch
34:29 Data annotation - Algorithms be changed, data is there to stay
38:41 NN's will learn what you ask them to learn
39:27 What is the impact of annotation errors?
48:37 Annotation error analysis
50:15 Conclusion
49:44 Discussion

[Chapters were auto-generated using our proprietary software - contact us if you are interested in access to the software]

The talk is based on the papers:
nuScenes: A multimodal dataset for autonomous driving (CVPR 2020)
arxiv: https://arxiv.org/abs/1903.11027
git: https://github.com/nutonomy/nuscenes-...

PointPainting: Sequential Fusion for 3D Object Detection (CVPR 2020)
https://arxiv.org/abs/1911.10150
git: https://github.com/rshilliday/painting

PointPillars: Fast Encoders for Object Detection from Point Clouds
(CVPR 2019)
arxiv: https://arxiv.org/abs/1812.05784
git: https://github.com/nutonomy/second.py...

Presenter BIO:

Dr. Holger Caesar is a Senior Research Scientist at Motional, formerly known as nuTonomy. Working in Singpore on the Machine Learning Team, his job is to make autonomous vehicles perceive and understand their environment. At Motional, he leads the Data-Curation team, whose goal it is to find scalable and cost-efficient approaches to annotate vast amounts of data. Holger is the project lead for the nuScenes autonomous driving dataset and contributed to the PointPillars method for object detection from lidar. He previously did his PhD in Computer Vision under the supervision of Prof. Vittorio Ferrari at the University of Edinburgh and ETH Zurich. Holger released the COCO-Stuff dataset and several methods for fully and weakly supervised segmentation and detection. He co-organized numerous workshops for COCO and WAD at ECCV, ICCV, CVPR, ICRA, IROS and NIPS.

More information about Dr. Holger Caesar and his research can be found at https://www.it-caesar.com/

-------------------------
Find us at:

Newsletter for updates about more events ➜ http://eepurl.com/gJ1t-D
Sub-reddit for discussions ➜   / 2d3dai  
Discord server for, well, discord ➜   / discord  
Blog ➜ https://2d3d.ai

AI consultancy Abelians ➜ https://abelians.com/