NTH

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

AuthorsBoya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu

June 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper delivers a fully open, highly competitive text-to-image diffusion model plus a carefully tested training recipe that could become a new starting point for open generative AI research.

Key results

3B
parameter count

size of the final i1 model

300+
controlled experiments

systematic ablations performed to study model and data design choices

700K+
compute budget

TPU v6e hours used for the experiments

2M
low-res training steps

256-resolution pretraining steps for the final i1 model

0.5M/0.3M
512/1024 training steps

high-resolution fine-tuning steps at 512 and 1024 resolution

29.5
average improvement vs best fully open model

absolute percentage-point gain on the five-benchmark average

What the paper found

i1, from Princeton University, is presented as a simple and fully open text-to-image diffusion recipe that closes much of the gap to proprietary systems by combining 300+ controlled experiments and 700K+ TPU v6e hours into a 3B-parameter model trained only on public data. The paper shows that several seemingly standard choices are not necessary or are suboptimal: a single strong text encoder, T5Gemma-2B, paired with a larger 2-transformer-block adapter is more effective than stacking multiple encoders, AdaLN noise/timestep conditioning adds little and is removed, and dual-stream MMDiT with long skip connections gives the best performance-parameter trade-off. On the data side, long synthetic captions from Qwen3-VL-30B-A3B outperform short captions, but require inference-time prompt rewriting to recover performance on short prompts; equal weighting across datasets is a strong default, and repeating a diverse dataset causes only marginal degradation. The final i1 model uses public datasets, 2M low-resolution training steps, 0.5M at 512 resolution, and 0.3M at 1024 resolution, then reaches 86.73 DPG-Bench, 70.1 PRISM, 0.8531 CVTG-2K, and 0.922 LongText-Bench, outperforming the best prior fully open model by 29.5 absolute points on average across five benchmarks.

Original abstract

Diffusion models have consistently driven progress in text-to-image generation. However, it is challenging to attribute recent progress to specific modeling and data choices: state-of-the-art open-weight models provide limited ablations, and do not disclose their training data and full training details. The research community needs fully open (weights, data, and code) models as a foundation for further research; yet existing fully open models still fall significantly short of leading models in performance. In this project, we conduct a systematic investigation of the modeling and data design choices in text-to-image diffusion training and inference with 300+ controlled experiments totaling 700K+ TPU v6e hours. Our experiments highlight several empirical findings (e.g., equal weighting is a strong default for mixing curated datasets) and simple design decisions (e.g., larger text encoder adapters improve performance with minimal added parameters) for training strong models. Guided by these insights, we train i1, a 3B-parameter text-to-image diffusion model using only publicly available datasets. i1 is competitive with leading models on five representative benchmarks (GenEval, DPG, PRISM, CVTG-2K, and LongText), and outperforms the best existing fully open model by 29.5 absolute percentage points on average. We provide the i1 checkpoints, training and inference code, and the data processing pipeline. Together, our findings and the i1 recipe establish a practical foundation for future open research in text-to-image diffusion models. Our code is available at https://github.com/zlab-princeton/i1.

Read the original paper

More in Diffusion Models

Browse all 58 papers →
02Diffusion

LongLive-Plug: Once-for-All Distillation for Video Generation

Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen, Bohan Zhang, Xiao Fu, Qianli Ma, Chen-Hsuan Lin, Weian Mao, Bryan Chu, Song Han, Yukang Chen

LongLive-Plug distills key video-generation capabilities into reusable LoRA adapters that can accelerate and improve many downstream diffusion models without retraining each one.

Read analysis
03Diffusion

Simplex Diffusion Models

Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang, James Thornton, Arnaud Doucet, Valentin De Bortoli

Simplex Diffusion Models keep uncertainty alive during discrete denoising, enabling faster and stronger generation for text, code, and math tasks.

Read analysis