Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
AuthorsAmber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin
Resources
A controlled study using animated letters shows that balanced video data and accurate captions are crucial for training effective text-to-video models.
Key results
MMDiT flow-matching model used for controlled experiments
Videos used per baseline training experiment
FG PSNR on 3-letter evaluation with equal 1-, 2-, and 3-letter mixing
Maximum relative compute required to match validation loss with corrupted captions
Maximum reported recovery of lost FG PSNR after clean-data fine-tuning
What the paper found
Researchers from Meta Superintelligence Labs and Purdue University introduce Moving Alphabet, a procedural testbed that renders one to three moving letters with controllable appearance, motion, duration, and caption corruption, isolating training-data effects that are difficult to study in real video datasets used by systems such as OpenAI’s Sora. They train an 800M-parameter MMDiT flow-matching model with T5-XXL and CLIP ViT-L/14 text encoders, using 150K videos per baseline experiment. Balanced mixtures consistently outperform specialist datasets: equal mixing of 1-, 2-, and 3-letter scenes reaches 12.9 dB FG PSNR on 3-letter evaluation, versus 10.1 dB for pure 3-letter training, while equal mixing of 2-, 4-, and 8-second clips achieves 12.2 dB average FG PSNR with fewer tokens than long-video-only training. Caption correctness is especially damaging when degraded: FG PSNR falls from 12.3 to 4.3 dB as precision drops from 1.0 to 0.3, and corrupted captions can require up to 4 times the compute to reach the same validation loss. Classifier-free guidance improves categorical prompt adherence for moderate corruption but cannot restore pixel fidelity; fine-tuning on clean data recovers at most 55% of lost FG PSNR in the reported moderate-corruption case, and severe caption errors remain largely irreversible. The central recommendation is to prioritize balanced data distributions and accurate captions during pre-training rather than relying on inference-time or post-training fixes.
Original abstract
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw videos and captioning to create video-text pairs for learning text-to-video mappings. We study how data distribution and caption quality impact text-to-video models. To enable controlled experiments, we introduce Moving Alphabet, a procedural testbed that renders letters with varying fonts, colors, sizes, and positions, moving in different directions and speeds against a black background. This design allows precise control over data distribution and caption quality by corrupting ground-truth metadata. Our experiments yield three findings: a) a diverse and balanced distribution of video content and duration is critical for generalization; b) caption quality significantly affects both model performance and training efficiency, suggesting that text-to-video models are bounded by video understanding capabilities; and c) classifier-free guidance and fine-tuning on high-quality data provide partial recovery from models trained on corrupted captions, but cannot fully compensate for poor pre-training data. We believe these insights can inform the development of large-scale text-to-video models, and we advocate for greater attention to the science of pre-training data.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.