NTH

Vision Pretraining for Dense Spatial Perception

AuthorsZelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue

July 7, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches vision models to pay attention to boundaries so they can learn richer spatial structure and improve tasks like depth estimation.

Key results

160.75M
training corpus

curated images used to train LingBot-Vision

1B
model size

LingBot-Vision ViT-g/16 parameter scale

0.296
NYU-Depth v2 RMSE

best frozen-feature depth probing result for LingBot-Vision

70.0
DAVIS-2017 J&F

training-free video object segmentation score

73.5
YouTube-VOS J&F

training-free video object segmentation score

23x
student parameter reduction

0.3B student versus 7B DINOv3 on NYU-Depth v2

What the paper found

Vision Pretraining for Dense Spatial Perception proposes a boundary-centric self-supervised objective that treats boundaries as native learning signals rather than downstream outputs. Instead of random masking, the method uses masked boundary modeling: a teacher Vision Transformer discovers boundary-bearing tokens online, forces them into the student’s masked set, and supervises them with a categorical boundary-field loss over discretized distance and orientation bins. This reparameterization stabilizes dense self-distillation and enables parameter-free a-contrario validation of candidate line segments, so unsupported structure never becomes a training target. Scaled to LingBot-Vision, a 1B Vision Transformer trained on 160.75M curated images, the approach matches or surpasses foundation models up to 7B parameters on dense spatial perception. On NYU-Depth v2 linear probing, LingBot-Vision reaches 0.296 RMSE, ahead of DINOv3’s 0.309, while on DAVIS-2017 and YouTube-VOS it posts 70.0 and 73.5 J&F-Mean, respectively, near the strongest distilled baselines. The paper also shows that a 0.3B student matches the 7B DINOv3 on NYU-Depth v2 with roughly 23× fewer parameters. These pretrained encoders transfer directly to LingBot-Depth 2.0, where replacing the initialization and scaling RGB-D training data from 3M to 150M samples yields leading depth completion on 14 benchmarks, especially on hard transparent and reflective scenes.

Original abstract

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis