NTH

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

AuthorsDong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, Chong Luo, Jianmin Bao, Ji Li, Lei Shi, Qinhong Yang, Xiuyu Wu, Xuelu Feng, Yan Lu, Yanchen Dong, Yitong Wang, Yunuo Chen

June 6, 2026 3 min read
Watch on YouTube
The one-line take

Lens is a more compute-efficient text-to-image model that uses richer data, smarter batching, and post-training tricks to match or beat larger systems while generating images faster.

Key results

3.8B
Model size

Lens is described as a 3.8B-parameter foundational text-to-image model.

19.3%
Training compute vs Z-Image

Lens requires about 19.3% of the pretraining compute used by Z-Image.

800M
Pretraining dataset size

Lens is trained on Lens-800M, a dataset of 800M densely captioned image-text pairs.

109
Average caption length

The Lens-800M captions are generated by GPT-4.1 and average approximately 109 words.

8406
RL dataset size

Lens-RL-8K contains 8,406 prompts for reinforcement-learning post-training.

3.15
Inference time

Lens generates a 1024² image in 3.15 seconds on a single NVIDIA H100 GPU, while Lens-Turbo reaches 4-step generation in 0.84 seconds.

What the paper found

Lens, from the Microsoft Lens Team, is a 3.8B-parameter foundational text-to-image model built to cut training compute without sacrificing quality. Its core result is efficiency: Lens matches or exceeds larger open-source systems such as Z-Image, LongCat-Image, FLUX.2, and Qwen-Image on benchmarks including OneIG, GenEval, LongText, and CVTG, while using only about 19.3% of Z-Image’s pretraining compute. The gains come from three design choices: Lens-800M, an 800M-pair dataset captioned by GPT-4.1 with dense 109-word average descriptions; mixed-resolution and mixed-aspect-ratio training over 512², 768², and 1024² buckets, which gives the model generalization up to 1440² and aspect ratios from 1:2 to 2:1; and convergence-oriented architecture choices, especially the FLUX.2 semantic VAE and the GPT-OSS 20B-A3B language encoder, which also enables multilingual prompt following from English-only training. After pretraining, Lens uses DiffusionNFT reinforcement learning on Lens-RL-8K, an 8,406-prompt taxonomy-balanced set with GPT-4.1-mini rubric rewards, to reduce artifacts and improve alignment. A reasoner module with training-free system-prompt search refines user prompts, and a distilled Lens-Turbo variant reaches 4-step generation in 0.84 seconds on a single NVIDIA H100, versus 3.15 seconds for the 20-step model at 1024².

Original abstract

We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters across various benchmarks, while requiring significantly less training compute. For example, Lens requires only about 19.3% of the training compute used by Z-Image. The training efficiency of Lens stems from two key strategies beyond its compact model size. First, we maximize data information density per training batch by (i) training on Lens-800M, a dataset of 800M densely captioned image-text pairs whose captions are generated by GPT-4.1 and contain approximately 109 words on average, providing richer semantic supervision than conventional short captions, and (ii) constructing each batch from images with multiple resolutions and diverse aspect ratios, thereby enlarging the effective visual coverage of each optimization step. Second, we improve convergence speed through careful architectural choices, including adopting a semantic VAE that provides better latent representations and employing a strong language encoder that accelerates optimization while enabling multilingual generalization from English-only training data. After pre-training, we apply RL with taxonomy-driven prompts (Lens-RL-8K) and structured reward rubrics to suppress artifacts and improve visual quality, a reasoner module with training-free system prompt search to better align user requests with the model, and distillation-based acceleration for 4-step inference. Through efficient training and systematic optimization, Lens generalizes to arbitrary aspect ratios from 1:2 to 2:1 and resolutions up to 1440^2, and supports prompts in several commonly used languages. Thanks to its compact size, Lens generates a 1024^2 image in 3.15 seconds on a single NVIDIA H100 GPU, while its distilled turbo version performs 4-step generation in 0.84 seconds.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis