NTH

Pictura: Perspective-View Self-Play at Scale for Driving

AuthorsYuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut, Victor Bares, Yihong Xu, Éloi Zablocki, Alexandre Boulch, Thibault Buhet, Andrei Bursuc, Matthieu Cord

August 10, 2026 2 min read
Watch on YouTube
The one-line take

Pictura trains autonomous driving agents through massive visual self-play, teaching them to handle realistic camera views without relying on privileged state information.

Key results

500K
Pictura throughput

Single-H100 rendering and simulation throughput in agent-steps per second

2M
Rendered image throughput

Images rendered per second on one H100

50B
Self-play training scale

Agent steps used to train Alberti

35M
Synthetic driving distance

Approximate kilometers driven during training

86%
WOMD full-density goal completion

Alberti goal completion on zero-shot Waymo Open Motion Dataset layouts

0.006
WOMD full-density collision rate

Alberti collision rate in the zero-shot evaluation

What the paper found

Pictura addresses a central flaw in driving self-play: privileged vector observations expose exact positions, velocities, and occluded agents, creating a gap between training and camera-based deployment. Its GPU-native CUDA rasterizer renders each vehicle’s egocentric perspective view, including vehicles, pedestrians, cyclists, traffic lights, parked cars, walls, and lane geometry, directly inside the reinforcement-learning loop. On a single H100, Pictura reaches 500 K agent-steps per second, or 2 M images per second. Using plain PPO, the Alberti policy jointly learns visual encoding and control from four low-resolution cameras, without privileged observations, a teacher, or distillation. Training covers 50 B agent steps and approximately 35 M km of synthetic driving on CARLA maps. In-domain, Alberti approaches a matched vectorized baseline, reaching 11.21 goals at medium traffic density versus 12.81, although its red-light violation rate is higher at 0.058 versus 0.009. The more important result appears in zero-shot transfer: on layouts re-rendered from the Waymo Open Motion Dataset, Alberti reaches 86% of goals at full traffic density, compared with 72% for the privileged agent, while reducing collisions to 0.006 and off-road events to 0.014. Counterfactual tests show that Alberti becomes insensitive to fully occluded agents and slows at blind corners, demonstrating decisions grounded in visible evidence rather than hidden simulator state.

Original abstract

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, mitigating the representation gap at its source. Pictura sustains up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO. It is the first large-scale driving self-play policy trained directly from perspective images, without privileged observations. Training spans 50B agent steps for ~35M km of driving. It approaches the driving performance of its privileged vectorized counterpart, and transfers zero-shot to Waymo Open Motion Dataset layouts re-rendered in Pictura, where it outperforms privileged vectorized agents. Project page: https://valeoai.github.io/Pictura/

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis