NTH

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives

AuthorsChong Zeng, Yue Dong, Pieter Peers, Lvmin Zhang, Maneesh Agrawala

September 17, 2026 2 min read
Watch on YouTube
The one-line take

RenderFormer-V2 uses specialized transformer attention to render complex scenes with diverse materials, environments, and volumetric effects without retraining for each scene.

Key results

10M
Training scenes

Approximate procedurally generated scenes used for training.

70TB
Training data volume

Total size of the rendered training dataset.

207M
Model parameters

Approximate size of the complete RenderFormer-V2 model.

100K
Maximum demonstrated scene scale

The renderer handles scenes exceeding 100K primitives.

28.0810
RenderFormer-V2 PSNR

PSNR at 2048 × 2048 resolution before high-resolution fine-tuning.

26.2423
RenderFormer baseline PSNR

PSNR of RenderFormer at 2048 × 2048 resolution.

What the paper found

RenderFormer-V2 is a two-stage transformer neural renderer that converts heterogeneous scene primitives into images with learned global illumination, without per-scene training or specialized rendering code. The paper includes a Microsoft contribution and extends RenderFormer beyond triangle-only geometry by jointly encoding textured triangles, volumetric scattering elements, triangular lights, environment maps, and camera rays. Its main innovation is rendering-aware sparse attention: Hilbert-curve serialization enables local sliding-window attention, while attention sinks retain global registers, light-source tokens, and mean-pooled summarization tokens for long-range transport. The view-dependent stage replaces full self-attention with shifted-window Swin attention, improving resolution scalability. A 9D neural material-appearance embedding decouples reflectance from a fixed BRDF, while a pretrained VAE encodes 32 × 32 texture, normal, and displacement patches. Training uses approximately 10M procedurally generated Blender Cycles scenes spanning 1K–64K primitives and 256²–2048² resolutions, totaling 70TB. The resulting model contains 207M parameters and can handle scenes exceeding 100K primitives, including caustics, environment lighting, textures, displacement, and volumetric scattering. At 2048 × 2048 resolution, RenderFormer-V2 reaches PSNR 28.0810 versus RenderFormer’s 26.2423, while high-resolution fine-tuning raises it to 30.7304. The method’s limitations include fixed 32 × 32 texture patches, no explicit temporal coherence, and a maximum of 8 light sources inherited from training.

Original abstract

We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to primitive transport, and a view-dependent stage that transforms the internal neural scene representation into image pixels. Different from RenderFormer, our model employs a novel combined windowed-attention and rendering-informed attention sink in the view-independent stage to improve scalability while maintaining render accuracy. To further improve versatility, RenderFormerV2 supports heterogeneous scene primitives, including environment maps and participating media, and it employs a material encoding independent of the underlying surface reflectance model that encodes material appearance via a novel neural embedding. We demonstrate the versatility of RenderFormer-V2 on a variety of scenes and perform an extensive ablation of the improved attention mechanism.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis