Robostral Navigate
AuthorsArjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
Resources
Robostral Navigate uses an 8B vision-language model and only monocular RGB images to achieve state-of-the-art robot navigation while dramatically reducing training costs.
Key results
Vision-language model predicting image-space navigation waypoints
Expert trajectories generated for training
Indoor and outdoor simulated scenes used for data generation
Reduction factor from prefix caching and tree attention
Validation-unseen success rate using one RGB camera
English-only validation-unseen success rate
What the paper found
Robostral Navigate, presented by Mistral, is an 8B vision-language navigation model designed to work across robot types with minimal sensing: it uses only a single monocular RGB camera and predicts image-space waypoints, avoiding robot-specific metric coordinates and camera recalibration. Its training pipeline generated 2.4M simulated trajectories across 350k scenes, then used prefix caching with a tree-based attention mask to remove redundant history encoding and reduce training-token processing by 22 times. A 121M diffusion policy converts the model’s coarse waypoints into 10-hertz trajectories, while CISPO reinforcement learning on difficult failure cases improves exploration and recovery, adding roughly 4% success. The same learned weights were deployed on the substantially different Galaxea R1 and Hiwonder JetAuto platforms. On the R2R-CE validation-unseen benchmark, Robostral Navigate achieved 77.4% success, exceeding the best monocular baseline by 10.5 points and the strongest depth- or multi-camera system by 5.3 points. On RxR-CE validation-unseen, it reached 75.1% success and 68.7% SPL, demonstrating both strong completion and path efficiency without depth, LiDAR, panoramic cameras, or pre-built maps.
Original abstract
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.