GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
AuthorsGuangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang
Resources
GenFirst stabilizes end-to-end latent generative models by letting generation shape the latent space before progressively adding reconstruction pressure.
Key results
SiT reaches a gFID of 0.97 with classifier-free guidance on ImageNet 256 × 256.
SiT reaches 1.45 without classifier-free guidance on ImageNet 256 × 256.
The MMDiT-based EiT system reaches 0.90 on GenEval.
GenFirst reaches comparable quality with a 2-fold reduction in training cost.
GenFirst reaches competitive FID using at least 70-fold fewer steps than baseline SiT training.
What the paper found
A ByteDance Seed paper introduces GenFirst, a strategy for directly training a variational autoencoder and its generative prior together instead of freezing a reconstruction-optimized latent space. The authors identify latent collapse as a prior–entropy imbalance: generative prior fitting compresses posterior means and variances, while an explicitly weighted entropy term preserves uncertainty. They also show that reconstruction learns quickly, whereas generation needs longer optimization, so GenFirst applies strong generative pressure first and strengthens reconstruction later. The framework supports exact-likelihood continuous autoregressive priors from FARMER and flow-matching SiT priors, with improvements demonstrated on ImageNet 256 × 256 and text-to-image tasks. SiT reaches a gFID of 0.97 with classifier-free guidance and 1.45 without guidance on ImageNet 256 × 256, while an MMDiT-based system reaches 0.90 on GenEval, outperforming models including FLUX.2-dev and Qwen-Image in the reported comparison. GenFirst also reaches comparable quality with a 2-fold reduction in training cost and reaches competitive FID using at least 70-fold fewer steps than baseline SiT training. Extensions combine generation with representation learning using Qwen3-VL and SigLIP or VLM-NLL supervision, and jointly model continuous text and image latents.
Original abstract
Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.