SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
AuthorsPaul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi
Resources
SynCity 3000 turns image-to-3D generation into a scalable scene-level diffusion system that can build large, coherent 3D worlds from prompts and layouts.
Key results
Procedurally generated scene-like samples used to fine-tune the TRELLIS-based 3D generators.
Optimization steps for the TRELLIS sparse-structure transformer.
Optimization steps for the TRELLIS structured-latent transformer.
Full SynCity 3000 score for structural similarity to the input template.
User-study win rate against TRELLIS using the SynCity 3000 template.
What the paper found
SynCity 3000, from the University of Oxford’s Visual Geometry Group with support from Meta Research, generates globally coherent, layout-controlled 3D worlds from text. Its two-stage pipeline first uses Flux ControlNet and a MultiDiffusion-inspired overlapping-window process to create a dimetric 2D scene template, then fine-tunes the TRELLIS image-to-3D model for convolutional inference over sliding 3D windows. The model produces a coarse voxel structure, DINOv2-conditioned structured latents, and ultimately 3D Gaussian Splats that can scale beyond object-sized outputs. To overcome the shortage of scene-scale training data, the authors procedurally place Objaverse-XL assets on randomized terrains, producing 320k synthetic samples for fine-tuning TRELLIS’s sparse-structure and latent transformers for 260k and 660k steps on 2× NVIDIA RTX A6000 GPUs. On template faithfulness, the full method reaches 0.5247 SSIM and 14.4137 PSNR, outperforming off-the-shelf TRELLIS and alternatives including TripoSG and Hunyuan3D-2.1. In a 27-person user study, SynCity 3000 achieved a 78.6% overall scene-preference win rate against TRELLIS using its template, while users unanimously preferred its flexible layout control over the grid-based SynCity. Layout instructions in the experiments were generated with ChatGPT 5 Instant. Remaining limitations include dimetric-perspective requirements, reduced texture fidelity from window averaging, occasional duplicated structures, and a cartoonish appearance inherited partly from FLUX and Objaverse-XL.
Original abstract
We present SynCity 3000, a framework for generating 3D scenes that are globally coherent while enabling fine-grained layout control. Building on the ability of current image-to-3D generators to produce complex 3D assets from a single image, we extend this capability to the scale of entire scenes by adapting the generator to be applicable as a convolutional operator. We achieve this by fine-tuning the model on scene-like data generated by a new synthetic data engine, which we propose to address the scarcity of 3D scene data for training. The convolutional generator is then applied to a dimetric image of the entire scene, generated from the user prompt, resulting in 3D scenes of arbitrary size and complexity. Across diverse prompts and layouts, SynCity 3000 produces large, coherent, and detailed scenes, addressing the shortcomings of prior approaches to 3D scene generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.