NTH

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

AuthorsJunpeng Jing, Ronglai Zuo, Zhelun Shen, Shangchen Zhou, Rolandos Alexandros Potamias, Stefanos Zafeiriou, Krystian Mikolajczyk, Jiankang Deng

June 30, 2026 2 min read
Watch on YouTube
The one-line take

LAS2 shows that stereo matching can stay fast enough for real devices while still improving zero-shot accuracy, thanks to a smarter architecture and a carefully staged training recipe.

Key results

1.8M
synthetic training pairs

labeled stereo pairs used in Stage 1

0.5M
real-world unlabeled pairs

stereo pairs used in Stage 3 pseudo-label adaptation

2.88
LAS2-M KITTI 2012 D1

zero-shot benchmark score

3.61
LAS2-M KITTI 2015 D1

zero-shot benchmark score

8.1
LAS2-M latency H200

inference latency in ms

101
LAS2-M latency Orin

inference latency in ms

What the paper found

Lite Any Stereo V2, or LAS2, from Imperial College London, reframes efficient zero-shot stereo matching around deployment reality rather than MAC counts alone. The paper replaces the heavier 3D aggregation used in prior LAS with a pure 2D cost-aggregation pipeline built on FasterNet, then trains a family of feed-forward models, LAS2-S/M/L, plus an iterative LAS2-H variant. Its three-stage training recipe combines supervised learning on 1.8M synthetic stereo pairs, self-distillation with feature alignment, and pseudo-labeled adaptation on 0.5M real-world unlabeled pairs, where pseudo supervision is cleaned by left-right consistency, edge masking, sky masking, and error clamping. On four zero-shot benchmarks—KITTI 2012, KITTI 2015, ETH3D, and Middlebury—LAS2-M reaches 2.88 D1 on KITTI 2012, 3.61 D1 on KITTI 2015, 2.59 Bad 1.0 on ETH3D, and 5.47 Bad 2.0 on Middlebury, while running in 8.1 ms on H200 and 101 ms on Orin. LAS2-H further improves to 2.64 D1 on KITTI 2012 and 3.31 D1 on KITTI 2015 with 15.1 ms on H200 and 344 ms on Orin, outperforming Fast-FoundationStereo in accuracy while remaining much faster. The ablations show that FasterNet is the best practical block choice and that the final training stage drives the largest gain, reducing LAS2-M’s error substantially over stage 1 alone. Overall, the paper argues that efficient stereo can now be both strong and deployable without relying on foundation-model priors.

Original abstract

Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation-model priors, making them difficult to deploy on resource-constrained platforms. In contrast, efficient stereo models offer faster inference but are commonly considered less capable of strong zero-shot generalization. In this paper, we challenge this assumption by introducing Lite Any Stereo V2 (LAS2), an ultra-fast model series designed for efficient zero-shot stereo matching. LAS2 is developed from both architecture and training perspectives. Architecturally, we revisit efficient stereo design under practical deployment settings and propose a 2D-only cost aggregation framework, optimized for real inference latency rather than theoretical MACs alone. For training, we develop a three-stage strategy that combines synthetic supervision, self-distillation, and real-world knowledge distillation. To improve the reliability of real-world pseudo supervision, we further introduce pseudo-label filtering and an error-clamping operation, enabling smoother synthetic-to-real transfer. We instantiate LAS2 as a family of models, including feed-forward variants for different efficiency budgets and an iterative variant for higher accuracy. Extensive experiments show that LAS2 achieves state-of-the-art accuracy among efficient stereo methods while maintaining significantly lower latency. Specifically, LAS2-H achieves stronger overall zero-shot performance than the iterative method Fast-FoundationStereo, with 1.8x and 2.7x faster inference on H200 and Orin, respectively. The project page, demos, and code are available at https://tomtomtommi.github.io/LiteAnyStereoV2/.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis