RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
AuthorsYuchuan Ding, Linfei Li, Lin Zhang, Ying Shen
RaysUp is a fast, lightweight way to turn low-resolution vision foundation model features into sharp, high-resolution maps using geometry-aware ray representations.
Key results
RaysUp model size reported in the ultra-lightweight comparison.
RaysUp uses about 16% of AnyUp's parameters.
RaysUp is described as delivering approximately 7× faster inference than AnyUp.
What the paper found
RaysUp, from Tongji University, is an ultra-light, VFM-agnostic feature upsampling framework that reconstructs low-resolution outputs from foundation models such as DINOv2, DINOv3, SigLIP2, and PE Spatial at arbitrary target resolutions. Instead of plain 2D interpolation, it lifts reconstruction into a geometry-aware ray domain by combining a Spatially Decoupled Guidance Encoder, any-resolution cross-attention, Ray Positional Encoding based on 6D Plücker ray coordinates, and geometry-aware neighborhood attention. The design explicitly targets semantic fidelity and boundary sharpness while keeping computation small: the model uses only 0.14M parameters, about 16% of AnyUp, and reaches roughly 7× faster inference. On ImageNet-trained upsampling, it achieves state-of-the-art or near-state-of-the-state performance across COCO-Stuff, Pascal-VOC, ADE20K, Cityscapes, NYUv2, DAVIS, and ProxyCLIP-style open-vocabulary segmentation, with particularly strong gains on depth and surface-normal estimation. The ablation results show that RayPE is essential for geometric consistency, while the decoupled guidance encoder improves anisotropic structure modeling with only 0.14M parameters. Overall, RaysUp reframes universal feature upsampling as ray-aligned reconstruction, delivering a practical accuracy-efficiency trade-off that remains stable even at 2K × 2K resolution.
Original abstract
Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong generalization ability. However, their patchified or pooled outputs are inherently low-resolution, limiting their effectiveness in tasks requiring fine-grained, pixel-level reasoning. Existing feature upsampling approaches either degrade semantic fidelity or rely on VFM-specific retraining and heavy architectures, hindering efficiency and scalability. To address these challenges, we propose RaysUp, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions. Unlike conventional 2D interpolation or attention-based schemes, RaysUp lifts feature reconstruction into a geometry-aware ray domain. Specifically, we introduce a Spatially Decoupled Guidance Encoder for direction-aware guidance encoding, an Any-Resolution Cross-Attention mechanism for resolution-flexible reconstruction, and a novel Ray Positional Encoding (RayPE) that injects implicit 3D geometric priors via 6D Plucker ray coordinates. Finally, a Geometry-Aware Neighborhood Attention module further ensures content-adaptive bilateral aggregation while preserving geometric consistency. Extensive experiments across diverse dense prediction tasks demonstrate that RaysUp achieves state-of-the-art performance while using only 16% of the parameters of AnyUp and delivering approximately 7x faster inference. These results highlight a substantially improved accuracy-efficiency trade-off and establish RaysUp as a practical and scalable solution for universal feature upsampling. Code is available at https://github.com/MAP-RaysUp/RaysUp.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.