NTH

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

AuthorsSensen Gao, Zhaoqing Wang, Qihang Cao, Dongdong Yu, Changhu Wang, Jia-Wang Bian

July 12, 2026 2 min read
Watch on YouTube
The one-line take

PixWorld unifies 3D scene reconstruction and generation in pixel space, adding geometry-aware supervision to improve both visual quality and 3D fidelity.

Key results

1.044B
parameter_count

Total trainable parameters of PixWorld

67K
training_scenes

Multi-view scenes from RealEstate10K and DL3DV-10K used for training

18.88
single_image_realestate_psnr

RealEstate10K 1-view generation PSNR

0.614
single_image_realestate_auc5

RealEstate10K 1-view generation AUC@5

16.50
dl3dv_generation_psnr

DL3DV-10K 1-view generation PSNR

71.04
worldscore_average

PixWorld average score on WorldScore

What the paper found

PixWorld, from AISphere and Nanyang Technological University, unifies 3D scene reconstruction and generation by moving diffusion from latent space into pixel space and supervising a pixel-aligned 3D Gaussian Splatting representation directly through differentiable rendering. Unlike prior latent-space systems such as Gen3R, it removes the VAE/RAE bottleneck and adds a geometry perception loss built on a frozen 3D foundation model, π 3, so the model matches rendered views not only photometrically but also in geometry-aware feature space. The method uses a 24-layer DiT with 1.044B parameters, trains from scratch on about 67K multi-view scenes from RealEstate10K and DL3DV-10K plus 10M BLIP-3o images, and evaluates on RealEstate10K, DL3DV-10K, and WorldScore. It achieves the best reported single-image generation scores on RealEstate10K, reaching 18.88 PSNR and 0.614 AUC@5, improves PSNR to 16.50 and AUC@5 to 0.485 on DL3DV-10K, and wins the WorldScore average at 71.04. For reconstruction, it attains 26.21 PSNR on RealEstate10K and 23.18 on DL3DV-10K with 4-view inputs, while an ablation shows that removing geometry perception drops RealEstate10K 1-view PSNR from 19.12 to 17.99 and AUC@5 from 0.642 to 0.562, confirming that the new structural loss is central to long-horizon geometric fidelity.

Original abstract

3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generation. Recent works attempt to unify them in latent space, but with notable drawbacks: the diffusion objective is defined on latent features rather than the underlying 3D representation, and both branches suffer from information loss introduced by latent encoding, while requiring a pretrained Variational Autoencoder (VAE) or Representation Autoencoder (RAE). In this paper, we reformulate these two tasks under a unified pixel-space diffusion paradigm and introduce PixWorld, a single model that jointly addresses 3D reconstruction and generation. By supervising diffusion directly on rendered images, PixWorld removes the above limitations and aligns optimization with 3D scene fidelity. Beyond photometric and perceptual supervision that operates at the 2D image level and lacks 3D geometric awareness, we further introduce a geometry perception loss that aligns rendered views with their ground truth in the geometry-aware feature space of a pretrained 3D foundation model, providing 3D structural supervision. PixWorld consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis