NTH

ABot-N1: Toward a General Visual Language Navigation Foundation Model

AuthorsRuiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu

July 21, 2026 2 min read
Watch on YouTube
The one-line take

ABot-N1 is a navigation foundation model that turns language and vision into pixel-level goals before executing actions, improving robustness and interpretability in complex indoor and urban navigation.

Key results

4B
Slow reasoner size

Qwen-3.5-4B model used for deliberative visual-language reasoning.

2B
Fast action expert size

Qwen-3.5-2B model used for reactive waypoint control.

30M
Pretraining samples

Multi-task corpus spanning five navigation tasks.

77.3%
POI arrival success

Success within 2 m of the named point-of-interest entrance on ABotN-POIBench.

95.4%
Indoor point-goal success

Zero-collision success rate on the indoor ABotN-PointBench split.

92.9%
Outdoor point-goal success

Success rate under the outdoor fewer-than-three-collisions criterion.

What the paper found

Alibaba Group’s AMAP CV Lab introduces ABot-N1, a visual-language navigation foundation model that replaces monolithic observation-to-action policies with an asynchronous slow-fast design. Its 4B-parameter Qwen-3.5-4B reasoner interprets instructions, visual history, and scene semantics, producing explicit Chain-of-Thought traces plus affordance and target pixels; a 2B-parameter Qwen-3.5-2B action expert then converts those grounded signals into continuous SE(2) waypoints at control frequency. This structured interface unifies point-goal, object-goal, POI-goal, instruction-following, and person-following, reducing coordinate drift, separating semantic recognition from approach control, and making failures more interpretable than systems such as NVIDIA’s GR00T N1-style brain-body architectures. Trained on 30M samples with pixel-grounded supervision and GRPO reinforcement learning, ABot-N1 is evaluated on established benchmarks and the newly released ABotN-PointBench and ABotN-POIBench. It reaches 77.3% POI entrance-arrival success, a 35.0% gain over POINav, and achieves 95.4% zero-collision success indoors and 92.9% success outdoors. The model also improves open-vocabulary object navigation and person tracking, while real-robot deployment on Alibaba’s AMap TuTu quadruped demonstrates that visual re-grounding can support safer, long-horizon urban navigation with low-fidelity maps.

Original abstract

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that map observations directly to actions, yet they often suffer from coordinate drift and poor handling of long-tail semantics. Furthermore, these black-box mappings lack interpretability, hindering the simultaneous achievement of generality, robustness, and transparency. We present ABot-N1, a step toward a general Visual Language Navigation foundation model, that addresses these challenges by decoupling cognition from control via a slow-fast architecture guided by dual visual-language signals. More specifically, a slow vision-language reasoner performs explicit Chain-of-Thought reasoning while producing a pixel goal. This compact set of image-space anchor points serves as a universal interface for diverse tasks, including point-goal, object-goal, poi-goal, instruction-following, and person-following. Subsequently, a fast action expert leverages both the textual cues and the pixel guidance to generate continuous waypoints at the native control frequency. By bridging high-level intents and low-level control through pixel-grounded anchors paired with explicit linguistic traces, our approach ensures robust, generalizable, and interpretable navigation across simulation and real-world benchmarks. ABot-N1 establishes new state-of-the-art records, delivering massive gains specifically in urban-scale navigation: boosting POI arrival by 35.0% (to 77.3%) and achieving 95.4%/92.9% SR in complex indoor and outdoor scenes. It also maintains superior robustness across object-reaching, person-following, and instruction-following tasks. New Point-Goal/POI-Goal benchmarks are released as open source to advance the field of urban-scale navigation.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis