Evaluating Generalist Robot Policies: World Model Generalization and Safety using Veo

Опубликовано: 14 Август 2026
на канале: Foundation Models For Robotics
20
1

#robotics #google #video #automation #education #tending
We present the **Veo (Robotics) world model**, a powerful generative evaluation system developed by the Gemini Robotics Team. This innovative system demonstrates that frontier video models can be utilized for the **entire spectrum of policy evaluation use cases in robotics**, moving beyond typical in-distribution evaluations.

The system is optimized to assess nominal performance, **out-of-distribution (OOD) generalization**, and the physical and semantic safety of visuomotor policies. This capability is vital because conducting sufficiently broad hardware evaluations for generalist robot policies is often impractical, and testing safety scenarios can be infeasible, risking damage to the robot, its environment, or humans.

Our method introduces a generative evaluation system built upon the *Veo foundation model**. The base video model is fine-tuned to support two crucial robotics requirements: **robot action conditioning* and **multi-view video generation**. The fine-tuned robotic video generation model can predict future images conditioned on an initial image observation and a sequence of future robot poses. To manage partial observations inherent in robotics, the system generates tiled future frames from four cameras: the top-down view, side view, and the left and right wrist views.

A central feature is the integration of *generative image-editing* and multi-view synthesis to create realistic and diverse variations of real-world scenes. We utilize models like NanoBanana (Gemini 2.5 Flash Image) to edit a nominal scene using language descriptions, synthesizing changes such as novel interaction objects, different visual backgrounds, and new distractor objects. This high fidelity enables the accurate prediction of the relative performance and ranking of different policies under both nominal and OOD conditions.

We rigorously validate the predictions from the video model through *1600+ real-world evaluations* of eight Gemini Robotics policy checkpoints (based on the GROD model) and five tasks executed by a bimanual manipulator (ALOHA 2). We observe a strong correlation between predicted and actual success rates across nominal scenarios and OOD axes of generalization.

Crucially, the Veo (Robotics) world model enables **predictive red teaming for safety**. By rolling out policies in synthetically edited scenes involving safety-critical elements, the system can discover potential vulnerabilities without requiring hardware evaluations. Examples of potentially unsafe behaviors found include the robot moving its gripper to make contact with a human hand or closing a laptop without first removing scissors, which could damage the screen.

This work demonstrates the transformative potential of video models, providing the necessary infrastructure for developing generalist embodied agents that can operate capably and safely in complex, real-world environments.

*Tags:*
Gemini Robotics, Veo, World Model, Robotics Policy Evaluation, OOD Generalization, AI Safety, Predictive Red Teaming, Video Generation, Action Conditioning, Generative AI, GROD, ALOHA 2, Generalist Agent, Simulation, Computer Vision, Machine Learning