Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
AuthorsDong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, Jiawei Zhang, Jinjing Zhao, Sirui Zhang, Yang Yue, Zhiyang Liang, Baining Guo, Chong Luo, Jianmin Bao, Ji Li, Lei Shi, Qinhong Yang, Xiuyu Wu, Xuelu Feng, Yan Lu, Yanchen Dong, Yitong Wang, Yunuo Chen
Resources
Lens is a more compute-efficient text-to-image model that uses richer data, smarter batching, and post-training tricks to match or beat larger systems while generating images faster.
Key results
Lens is described as a 3.8B-parameter foundational text-to-image model.
Lens requires about 19.3% of the pretraining compute used by Z-Image.
Lens is trained on Lens-800M, a dataset of 800M densely captioned image-text pairs.
The Lens-800M captions are generated by GPT-4.1 and average approximately 109 words.
Lens-RL-8K contains 8,406 prompts for reinforcement-learning post-training.
Lens generates a 1024² image in 3.15 seconds on a single NVIDIA H100 GPU, while Lens-Turbo reaches 4-step generation in 0.84 seconds.
What the paper found
Lens, from the Microsoft Lens Team, is a 3.8B-parameter foundational text-to-image model built to cut training compute without sacrificing quality. Its core result is efficiency: Lens matches or exceeds larger open-source systems such as Z-Image, LongCat-Image, FLUX.2, and Qwen-Image on benchmarks including OneIG, GenEval, LongText, and CVTG, while using only about 19.3% of Z-Image’s pretraining compute. The gains come from three design choices: Lens-800M, an 800M-pair dataset captioned by GPT-4.1 with dense 109-word average descriptions; mixed-resolution and mixed-aspect-ratio training over 512², 768², and 1024² buckets, which gives the model generalization up to 1440² and aspect ratios from 1:2 to 2:1; and convergence-oriented architecture choices, especially the FLUX.2 semantic VAE and the GPT-OSS 20B-A3B language encoder, which also enables multilingual prompt following from English-only training. After pretraining, Lens uses DiffusionNFT reinforcement learning on Lens-RL-8K, an 8,406-prompt taxonomy-balanced set with GPT-4.1-mini rubric rewards, to reduce artifacts and improve alignment. A reasoner module with training-free system-prompt search refines user prompts, and a distilled Lens-Turbo variant reaches 4-step generation in 0.84 seconds on a single NVIDIA H100, versus 3.15 seconds for the 20-step model at 1024².
Original abstract
We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters across various benchmarks, while requiring significantly less training compute. For example, Lens requires only about 19.3% of the training compute used by Z-Image. The training efficiency of Lens stems from two key strategies beyond its compact model size. First, we maximize data information density per training batch by (i) training on Lens-800M, a dataset of 800M densely captioned image-text pairs whose captions are generated by GPT-4.1 and contain approximately 109 words on average, providing richer semantic supervision than conventional short captions, and (ii) constructing each batch from images with multiple resolutions and diverse aspect ratios, thereby enlarging the effective visual coverage of each optimization step. Second, we improve convergence speed through careful architectural choices, including adopting a semantic VAE that provides better latent representations and employing a strong language encoder that accelerates optimization while enabling multilingual generalization from English-only training data. After pre-training, we apply RL with taxonomy-driven prompts (Lens-RL-8K) and structured reward rubrics to suppress artifacts and improve visual quality, a reasoner module with training-free system prompt search to better align user requests with the model, and distillation-based acceleration for 4-step inference. Through efficient training and systematic optimization, Lens generalizes to arbitrary aspect ratios from 1:2 to 2:1 and resolutions up to 1440^2, and supports prompts in several commonly used languages. Thanks to its compact size, Lens generates a 1024^2 image in 3.15 seconds on a single NVIDIA H100 GPU, while its distilled turbo version performs 4-step generation in 0.84 seconds.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.