HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
AuthorsChong Cheng, Peilin Tao, Nanjie Yao, Guanzhi Ding, Xianda Chen, Yuansen Du, Xiaoyang Guo, Wei Yin, Weiqiang Ren, Qian Zhang, Zhengqing Chen, Hao Wang
Resources
HorizonStream is a new streaming Transformer that helps 3D reconstruction stay stable over extremely long video sequences by separating short-range matching from long-range geometric memory.
Key results
Mean trajectory error on KITTI, outperforming streaming baselines such as Lingbot-map and LongStream in Table 1.
Mean trajectory error on VBR, better than Lingbot-map and TTT3R in Table 3.
The method is reported to generalize stably to sequences exceeding 10,000 frames with constant memory and linear time.
HorizonStream is trained on 48-frame clips and still generalizes to much longer sequences.
What the paper found
HorizonStream is a streaming 3D reconstruction model from HKUST(GZ) and Horizon Robotics that targets the core failure mode of online geometry: long-sequence drift under causal, bounded-memory constraints. The paper reframes streaming reconstruction as an evidence influence kernel and shows that common designs create pathological memory patterns, including sliding-window cutoffs, refresh discontinuities, attention sinks, and state saturation. HorizonStream factorizes this kernel into Geometric Local Attention and Geometric Linear Attention: the local module uses 3D-aware content matching with Spatiotemporal RoPE and head-wise reliability gating, while the linear module maintains an O(1) recurrent geometric state with learned channel-wise decay rates, giving different features distinct retention horizons. Metric Readout Tokens then recover stable scale and rigid pose from the persistent state, and an optional loop-closure stage refines revisits. Trained on only 48-frame clips, the model generalizes stably to sequences over 10,000 frames with constant memory and linear time. Across KITTI, Oxford Spires, ScanNet++, TUM RGB-D, Waymo, and VBR, it outperforms prior streaming methods and often approaches offline systems; for example, on KITTI the average ATE drops to 19.75 versus 25.29 for Lingbot-map and 51.90 for LongStream, and on VBR it reaches 25.30 versus 27.53 for Lingbot-map and 64.99 for TTT3R. Ablations confirm that both channel-wise retention and local 3D attention are necessary for long-horizon stability.
Original abstract
Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short-lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike-like attention sinks. To resolve this, we formalize geometric propagation as an \emph{evidence influence kernel} and propose HorizonStream, a long-horizon Transformer that explicitly factorizes this kernel. For the long-range temporal factor, Geometric Linear Attention learns channel-wise decay rates to enable bounded, multi-timescale propagation of geometric evidence. For the short-range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that HorizonStream, trained on only 48-frame clips, generalizes stably to sequences exceeding 10,000\ frames with constant memory and linear time, achieving state-of-the-art streaming 3D reconstruction performance. Project Page: https://3dagentworld.github.io/horizonstream/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.