NTH

Robostral Navigate

AuthorsArjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal

July 28, 2026 2 min read
Watch on YouTube
The one-line take

Robostral Navigate uses an 8B vision-language model and only monocular RGB images to achieve state-of-the-art robot navigation while dramatically reducing training costs.

Key results

8B
Navigation model size

Vision-language model predicting image-space navigation waypoints

2.4M
Simulation trajectories

Expert trajectories generated for training

350k
Simulation scenes

Indoor and outdoor simulated scenes used for data generation

22
Training-token reduction

Reduction factor from prefix caching and tree attention

77.4%
R2R-CE success rate

Validation-unseen success rate using one RGB camera

75.1%
RxR-CE success rate

English-only validation-unseen success rate

What the paper found

Robostral Navigate, presented by Mistral, is an 8B vision-language navigation model designed to work across robot types with minimal sensing: it uses only a single monocular RGB camera and predicts image-space waypoints, avoiding robot-specific metric coordinates and camera recalibration. Its training pipeline generated 2.4M simulated trajectories across 350k scenes, then used prefix caching with a tree-based attention mask to remove redundant history encoding and reduce training-token processing by 22 times. A 121M diffusion policy converts the model’s coarse waypoints into 10-hertz trajectories, while CISPO reinforcement learning on difficult failure cases improves exploration and recovery, adding roughly 4% success. The same learned weights were deployed on the substantially different Galaxea R1 and Hiwonder JetAuto platforms. On the R2R-CE validation-unseen benchmark, Robostral Navigate achieved 77.4% success, exceeding the best monocular baseline by 10.5 points and the strongest depth- or multi-camera system by 5.3 points. On RxR-CE validation-unseen, it reached 75.1% success and 68.7% SPL, demonstrating both strong completion and path efficiency without depth, LiDAR, panoramic cameras, or pre-built maps.

Original abstract

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis