NTH

UniSHARP: Universal Sharp Monocular View Synthesis

AuthorsMeixi Song, Dizhe Zhang, Hao Ren, Ruiyang Zhang, Bo Du, Ming-Hsuan Yang, Lu Qi

June 22, 2026 2 min read
Watch on YouTube
The one-line take

UniSHARP lets a single monocular view-synthesis system render sharp scenes across perspective, fisheye, and panoramic cameras by aligning them in a shared omnidirectional latent space.

Key results

36873
Perspective validation pairs

FoV-stratified benchmark count for perspective cameras

10692
Wide-FoV validation pairs

FoV-stratified benchmark count for wide-FoV cameras

14163
Fisheye validation pairs

FoV-stratified benchmark count for fisheye cameras

42754
Panorama validation pairs

FoV-stratified benchmark count for panoramic cameras

29.244
HM3D PSNR

UniSHARP panoramic benchmark result on HM3D

3.1
Inference time

UniSHARP runtime in seconds per sample

What the paper found

UniSHARP extends the SHARP feedforward 3D Gaussian splatting framework into a universal monocular view synthesis system that works across perspective, wide-FoV, fisheye, and 360-degree panoramic cameras. The core technical move is a unified ray-distance space, where Gaussian primitives are anchored along predicted rays and radial distances rather than image-plane coordinates, allowing the model to transfer across heterogeneous projection models without camera-specific branches. It combines Geometry Anchored Gaussians with Feature Conditioned Gaussian residuals, fusing 2D semantic encodings and 3D ray-based features, and adds panoramic-specific regularization such as spherical Gaussian initialization and distortion-aware probabilistic dropout. To evaluate this setting, the authors build a FoV-stratified benchmark spanning 36,873 perspective, 10,692 wide-FoV, 14,163 fisheye, and 42,754 panoramic validation pairs. On this benchmark, UniSHARP reports strong gains over prior monocular baselines: on HM3D panoramas it reaches 29.244 PSNR, 0.895 SSIM, and 0.065 LPIPS, compared with Matrix3D at 23.398 PSNR; on OmniRooms-Wide it achieves 25.243 PSNR and 0.076 LPIPS; and on ScanNet++ Fisheye it reaches 20.660 PSNR and 0.184 LPIPS. The pose-free variant also remains competitive, and inference is efficient at 3.1 seconds per sample, much faster than PanoDreamer at 8.6 seconds and Matrix3D at 38.8 seconds.

Original abstract

In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of camera systems, from conventional perspective cameras to wide-field-of-view, fisheye and omnidirectional panoramic settings. To overcome the pinhole-specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, Gaussian primitives are arranged along rays and radial distances in a ray-based universal representation, while 2D semantic and 3D spatial features extracted from UniK3D-inspired encoders are jointly decoded to generate the complete Gaussian cloud. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems across various scenes. The benchmark is further stratified by field of view (FoV) to enable fine-grained assessment of the universal monocular rendering task. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. The project page can be found at: https://insta360-research-team.github.io/Unisharp-website/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis