Pictura: Perspective-View Self-Play at Scale for Driving
AuthorsYuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch, Thibault Buhet, Andrei Bursuc, Matthieu Cord
Resources
Pictura trains autonomous driving agents through massive visual self-play, teaching them to handle realistic camera views without relying on privileged state information.
Key results
Single-H100 rendering and simulation throughput in agent-steps per second
Images rendered per second on one H100
Agent steps used to train Alberti
Approximate kilometers driven during training
Alberti goal completion on zero-shot Waymo Open Motion Dataset layouts
Alberti collision rate in the zero-shot evaluation
What the paper found
Pictura addresses a central flaw in driving self-play: privileged vector observations expose exact positions, velocities, and occluded agents, creating a gap between training and camera-based deployment. Its GPU-native CUDA rasterizer renders each vehicle’s egocentric perspective view, including vehicles, pedestrians, cyclists, traffic lights, parked cars, walls, and lane geometry, directly inside the reinforcement-learning loop. On a single H100, Pictura reaches 500 K agent-steps per second, or 2 M images per second. Using plain PPO, the Alberti policy jointly learns visual encoding and control from four low-resolution cameras, without privileged observations, a teacher, or distillation. Training covers 50 B agent steps and approximately 35 M km of synthetic driving on CARLA maps. In-domain, Alberti approaches a matched vectorized baseline, reaching 11.21 goals at medium traffic density versus 12.81, although its red-light violation rate is higher at 0.058 versus 0.009. The more important result appears in zero-shot transfer: on layouts re-rendered from the Waymo Open Motion Dataset, Alberti reaches 86% of goals at full traffic density, compared with 72% for the privileged agent, while reducing collisions to 0.006 and off-road events to 0.014. Counterfactual tests show that Alberti becomes insensitive to fully occluded agents and slows at blind corners, demonstrating decisions grounded in visible evidence rather than hidden simulator state.
Original abstract
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, mitigating the representation gap at its source. Pictura sustains up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO. It is the first large-scale driving self-play policy trained directly from perspective images, without privileged observations. Training spans 50B agent steps for ~35M km of driving. It approaches the driving performance of its privileged vectorized counterpart, and transfers zero-shot to Waymo Open Motion Dataset layouts re-rendered in Pictura, where it outperforms privileged vectorized agents. Project page: https://valeoai.github.io/Pictura/
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.