i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
AuthorsBoya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
This paper delivers a fully open, highly competitive text-to-image diffusion model plus a carefully tested training recipe that could become a new starting point for open generative AI research.
Key results
size of the final i1 model
systematic ablations performed to study model and data design choices
TPU v6e hours used for the experiments
256-resolution pretraining steps for the final i1 model
high-resolution fine-tuning steps at 512 and 1024 resolution
absolute percentage-point gain on the five-benchmark average
What the paper found
i1, from Princeton University, is presented as a simple and fully open text-to-image diffusion recipe that closes much of the gap to proprietary systems by combining 300+ controlled experiments and 700K+ TPU v6e hours into a 3B-parameter model trained only on public data. The paper shows that several seemingly standard choices are not necessary or are suboptimal: a single strong text encoder, T5Gemma-2B, paired with a larger 2-transformer-block adapter is more effective than stacking multiple encoders, AdaLN noise/timestep conditioning adds little and is removed, and dual-stream MMDiT with long skip connections gives the best performance-parameter trade-off. On the data side, long synthetic captions from Qwen3-VL-30B-A3B outperform short captions, but require inference-time prompt rewriting to recover performance on short prompts; equal weighting across datasets is a strong default, and repeating a diverse dataset causes only marginal degradation. The final i1 model uses public datasets, 2M low-resolution training steps, 0.5M at 512 resolution, and 0.3M at 1024 resolution, then reaches 86.73 DPG-Bench, 70.1 PRISM, 0.8531 CVTG-2K, and 0.922 LongText-Bench, outperforming the best prior fully open model by 29.5 absolute points on average across five benchmarks.
Original abstract
Diffusion models have consistently driven progress in text-to-image generation. However, it is challenging to attribute recent progress to specific modeling and data choices: state-of-the-art open-weight models provide limited ablations, and do not disclose their training data and full training details. The research community needs fully open (weights, data, and code) models as a foundation for further research; yet existing fully open models still fall significantly short of leading models in performance. In this project, we conduct a systematic investigation of the modeling and data design choices in text-to-image diffusion training and inference with 300+ controlled experiments totaling 700K+ TPU v6e hours. Our experiments highlight several empirical findings (e.g., equal weighting is a strong default for mixing curated datasets) and simple design decisions (e.g., larger text encoder adapters improve performance with minimal added parameters) for training strong models. Guided by these insights, we train i1, a 3B-parameter text-to-image diffusion model using only publicly available datasets. i1 is competitive with leading models on five representative benchmarks (GenEval, DPG, PRISM, CVTG-2K, and LongText), and outperforms the best existing fully open model by 29.5 absolute percentage points on average. We provide the i1 checkpoints, training and inference code, and the data processing pipeline. Together, our findings and the i1 recipe establish a practical foundation for future open research in text-to-image diffusion models. Our code is available at https://github.com/zlab-princeton/i1.
Read the original paperMore in Diffusion Models
Browse all 58 papers →FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon
FoMo uses the moment when two images diverge during diffusion generation as an automated measure of how perceptually different they are.
LongLive-Plug: Once-for-All Distillation for Video Generation
Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen
LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.
Simplex Diffusion Models
Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli
Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.