NTH

HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction

AuthorsChong Cheng, Peilin Tao, Nanjie Yao, Guanzhi Ding, Xianda Chen, Yuansen Du, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Zhengqing Chen, Hao Wang

June 1, 2026 2 min read
Watch on YouTube
The one-line take

HorizonStream is a new streaming Transformer that helps 3D reconstruction stay stable over extremely long video sequences by separating short-range matching from long-range geometric memory.

Key results

19.75
KITTI ATE avg

Mean trajectory error on KITTI, outperforming streaming baselines such as Lingbot-map and LongStream in Table 1.

25.30
VBR ATE avg

Mean trajectory error on VBR, better than Lingbot-map and TTT3R in Table 3.

10000
VBR sequence length

The method is reported to generalize stably to sequences exceeding 10,000 frames with constant memory and linear time.

48
training clip length

HorizonStream is trained on 48-frame clips and still generalizes to much longer sequences.

What the paper found

HorizonStream is a streaming 3D reconstruction model from HKUST(GZ) and Horizon Robotics that targets the core failure mode of online geometry: long-sequence drift under causal, bounded-memory constraints. The paper reframes streaming reconstruction as an evidence influence kernel and shows that common designs create pathological memory patterns, including sliding-window cutoffs, refresh discontinuities, attention sinks, and state saturation. HorizonStream factorizes this kernel into Geometric Local Attention and Geometric Linear Attention: the local module uses 3D-aware content matching with Spatiotemporal RoPE and head-wise reliability gating, while the linear module maintains an O(1) recurrent geometric state with learned channel-wise decay rates, giving different features distinct retention horizons. Metric Readout Tokens then recover stable scale and rigid pose from the persistent state, and an optional loop-closure stage refines revisits. Trained on only 48-frame clips, the model generalizes stably to sequences over 10,000 frames with constant memory and linear time. Across KITTI, Oxford Spires, ScanNet++, TUM RGB-D, Waymo, and VBR, it outperforms prior streaming methods and often approaches offline systems; for example, on KITTI the average ATE drops to 19.75 versus 25.29 for Lingbot-map and 51.90 for LongStream, and on VBR it reaches 25.30 versus 27.53 for Lingbot-map and 64.99 for TTT3R. Ablations confirm that both channel-wise retention and local 3D attention are necessary for long-horizon stability.

Original abstract

Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short-lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike-like attention sinks. To resolve this, we formalize geometric propagation as an \emph{evidence influence kernel} and propose HorizonStream, a long-horizon Transformer that explicitly factorizes this kernel. For the long-range temporal factor, Geometric Linear Attention learns channel-wise decay rates to enable bounded, multi-timescale propagation of geometric evidence. For the short-range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that HorizonStream, trained on only 48-frame clips, generalizes stably to sequences exceeding 10,000\ frames with constant memory and linear time, achieving state-of-the-art streaming 3D reconstruction performance. Project Page: https://3dagentworld.github.io/horizonstream/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis