NTH

GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling

AuthorsGuangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang

September 2, 2026 2 min read
Watch on YouTube
The one-line take

GenFirst stabilizes end-to-end latent generative models by letting generation shape the latent space before progressively adding reconstruction pressure.

Key results

0.97
SiT ImageNet gFID with CFG

SiT reaches a gFID of 0.97 with classifier-free guidance on ImageNet 256 × 256.

1.45
SiT ImageNet gFID without CFG

SiT reaches 1.45 without classifier-free guidance on ImageNet 256 × 256.

0.90
Text-to-image GenEval

The MMDiT-based EiT system reaches 0.90 on GenEval.

2
Training-cost reduction

GenFirst reaches comparable quality with a 2-fold reduction in training cost.

70
Fewer baseline training steps

GenFirst reaches competitive FID using at least 70-fold fewer steps than baseline SiT training.

What the paper found

A ByteDance Seed paper introduces GenFirst, a strategy for directly training a variational autoencoder and its generative prior together instead of freezing a reconstruction-optimized latent space. The authors identify latent collapse as a prior–entropy imbalance: generative prior fitting compresses posterior means and variances, while an explicitly weighted entropy term preserves uncertainty. They also show that reconstruction learns quickly, whereas generation needs longer optimization, so GenFirst applies strong generative pressure first and strengthens reconstruction later. The framework supports exact-likelihood continuous autoregressive priors from FARMER and flow-matching SiT priors, with improvements demonstrated on ImageNet 256 × 256 and text-to-image tasks. SiT reaches a gFID of 0.97 with classifier-free guidance and 1.45 without guidance on ImageNet 256 × 256, while an MMDiT-based system reaches 0.90 on GenEval, outperforming models including FLUX.2-dev and Qwen-Image in the reported comparison. GenFirst also reaches comparable quality with a 2-fold reduction in training cost and reaches competitive FID using at least 70-fold fewer steps than baseline SiT training. Extensions combine generation with representation learning using Qwen3-VL and SigLIP or VLM-NLL supervision, and jointly model continuous text and image latents.

Original abstract

Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis