RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
AuthorsChong Zeng, Yue Dong, Pieter Peers, Lvmin Zhang, Maneesh Agrawala
Resources
RenderFormer-V2 uses specialized transformer attention to render complex scenes with diverse materials, environments, and volumetric effects without retraining for each scene.
Key results
Approximate procedurally generated scenes used for training.
Total size of the rendered training dataset.
Approximate size of the complete RenderFormer-V2 model.
The renderer handles scenes exceeding 100K primitives.
PSNR at 2048 × 2048 resolution before high-resolution fine-tuning.
PSNR of RenderFormer at 2048 × 2048 resolution.
What the paper found
RenderFormer-V2 is a two-stage transformer neural renderer that converts heterogeneous scene primitives into images with learned global illumination, without per-scene training or specialized rendering code. The paper includes a Microsoft contribution and extends RenderFormer beyond triangle-only geometry by jointly encoding textured triangles, volumetric scattering elements, triangular lights, environment maps, and camera rays. Its main innovation is rendering-aware sparse attention: Hilbert-curve serialization enables local sliding-window attention, while attention sinks retain global registers, light-source tokens, and mean-pooled summarization tokens for long-range transport. The view-dependent stage replaces full self-attention with shifted-window Swin attention, improving resolution scalability. A 9D neural material-appearance embedding decouples reflectance from a fixed BRDF, while a pretrained VAE encodes 32 × 32 texture, normal, and displacement patches. Training uses approximately 10M procedurally generated Blender Cycles scenes spanning 1K–64K primitives and 256²–2048² resolutions, totaling 70TB. The resulting model contains 207M parameters and can handle scenes exceeding 100K primitives, including caustics, environment lighting, textures, displacement, and volumetric scattering. At 2048 × 2048 resolution, RenderFormer-V2 reaches PSNR 28.0810 versus RenderFormer’s 26.2423, while high-resolution fine-tuning raises it to 30.7304. The method’s limitations include fixed 32 × 32 texture patches, no explicit temporal coherence, and a maximum of 8 light sources inherited from training.
Original abstract
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-transport effects such as caustics, volumetric scattering, environment lighting, textured and displaced surfaces and out-of-distribution materials without per-scene training or specialized code. RenderFormer-V2 models global light transport as a sequence-to-sequence transformation. Following its predecessor, RenderFormer-V2 also employs a two stage process: a view-independent stage that resolves intra-scene primitive to primitive transport, and a view-dependent stage that transforms the internal neural scene representation into image pixels. Different from RenderFormer, our model employs a novel combined windowed-attention and rendering-informed attention sink in the view-independent stage to improve scalability while maintaining render accuracy. To further improve versatility, RenderFormerV2 supports heterogeneous scene primitives, including environment maps and participating media, and it employs a material encoding independent of the underlying surface reflectance model that encodes material appearance via a novel neural embedding. We demonstrate the versatility of RenderFormer-V2 on a variety of scenes and perform an extensive ablation of the improved attention mechanism.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.