UniSHARP: Universal Sharp Monocular View Synthesis
AuthorsMeixi Song, Dizhe Zhang, Hao Ren, Ruiyang Zhang, Bo Du, Ming-Hsuan Yang, Lu Qi
Resources
UniSHARP lets a single monocular view-synthesis system render sharp scenes across perspective, fisheye, and panoramic cameras by aligning them in a shared omnidirectional latent space.
Key results
FoV-stratified benchmark count for perspective cameras
FoV-stratified benchmark count for wide-FoV cameras
FoV-stratified benchmark count for fisheye cameras
FoV-stratified benchmark count for panoramic cameras
UniSHARP panoramic benchmark result on HM3D
UniSHARP runtime in seconds per sample
What the paper found
UniSHARP extends the SHARP feedforward 3D Gaussian splatting framework into a universal monocular view synthesis system that works across perspective, wide-FoV, fisheye, and 360-degree panoramic cameras. The core technical move is a unified ray-distance space, where Gaussian primitives are anchored along predicted rays and radial distances rather than image-plane coordinates, allowing the model to transfer across heterogeneous projection models without camera-specific branches. It combines Geometry Anchored Gaussians with Feature Conditioned Gaussian residuals, fusing 2D semantic encodings and 3D ray-based features, and adds panoramic-specific regularization such as spherical Gaussian initialization and distortion-aware probabilistic dropout. To evaluate this setting, the authors build a FoV-stratified benchmark spanning 36,873 perspective, 10,692 wide-FoV, 14,163 fisheye, and 42,754 panoramic validation pairs. On this benchmark, UniSHARP reports strong gains over prior monocular baselines: on HM3D panoramas it reaches 29.244 PSNR, 0.895 SSIM, and 0.065 LPIPS, compared with Matrix3D at 23.398 PSNR; on OmniRooms-Wide it achieves 25.243 PSNR and 0.076 LPIPS; and on ScanNet++ Fisheye it reaches 20.660 PSNR and 0.184 LPIPS. The pose-free variant also remains competitive, and inference is efficient at 3.1 seconds per sample, much faster than PanoDreamer at 8.6 seconds and Matrix3D at 38.8 seconds.
Original abstract
In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of camera systems, from conventional perspective cameras to wide-field-of-view, fisheye and omnidirectional panoramic settings. To overcome the pinhole-specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, Gaussian primitives are arranged along rays and radial distances in a ray-based universal representation, while 2D semantic and 3D spatial features extracted from UniK3D-inspired encoders are jointly decoded to generate the complete Gaussian cloud. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems across various scenes. The benchmark is further stratified by field of view (FoV) to enable fine-grained assessment of the universal monocular rendering task. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. The project page can be found at: https://insta360-research-team.github.io/Unisharp-website/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.