Exploring Vision-Language-Action (VLA) Models: From LLMs to Embodied AI

Опубликовано: 12 Июнь 2026
на канале: Voxel51
5,185
103

This talk will explore the evolution of foundation models, highlighting the shift from large language models (LLMs) to vision-language models (VLMs), and now to vision-language-action (VLA) models. We'll dive into the emerging field of robot instruction following—what it means, and how recent research is shaping its future. I will present insights from my 2024 work on natural language-based robot instruction following and connect it to more recent advancements driving progress in this domain.

Shreya Sharma is a Research Engineer at Reality Labs, Meta, where she works on photorealistic human avatars for AR/VR applications. She holds a bachelor’s degree in Computer Science from IIT Delhi and a master’s in Robotics from Carnegie Mellon University. Shreya is also a member of the inaugural 2023 cohort of the Quad Fellowship. Her research interests lie at the intersection of robotics and vision foundation models.

Linkedin:   / shreya-sharma99  

#computervision #ai #artificialintelligence #machinelearning #datascience

00:00 - Intro
00:13 - Defining Embodied AI
01:51 - Language Instruction Understanding in Robotics
03:45 - Grounding Instructions with Environment Perception
05:01 - Early Approaches: Human Demonstration and Goal Prediction
06:30 - GoalNet and Classical Planning Integration
08:01 - Leveraging Early LLMs for Common Sense Reasoning
10:00 - Generalizing Robotic Tasks with Language Models
11:24 - Real-World Task Examples and Long-Horizon Planning
11:59 - Off-Road Navigation Challenges in Robotics
13:54 - Simulation to Real-World Transfer
15:00 - Using LLMs to Learn Reward Functions
16:22 - Introduction to Vision-Language-Action (VLA) Models
17:49 - Multimodal Perception and Instruction Grounding
19:00 - Advantages of VA Models in Complex Environments
20:34 - VA Models in Practical Robotic Applications
21:00 - LLM-Powered High-Level Planning
23:20 - Few-Shot Prompting for Instruction Grounding
24:01 - Future Directions in VA Model Research
25:54 - Conclusion and Q\&A