Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
AuthorsRunhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao
Resources
SpectraReward turns pretrained multimodal models into plug-and-play judges that can guide image generators without extra reward-model training.
Key results
SpectraReward improvement over the BAGEL baseline at 512 resolution.
GenEval score for BAGEL trained with Self-SpectraReward at 512 resolution.
WISE score when Self-SpectraReward is evaluated with self-chain-of-thought.
Largest reward MLLM evaluated, Qwen3-VL-235B-A22B; scaling did not produce monotonic gains.
What the paper found
Researchers from the University of Hong Kong, Peking University, and ByteDance Seed introduce SpectraReward, a training-free reward function that converts a frozen pretrained multimodal large language model into a text-to-image reinforcement-learning evaluator. Instead of asking an MLLM to assign a noisy scalar score or decompose a prompt into verification questions, the method performs one image-conditioned, teacher-forced forward pass and averages the log-likelihood of the original prompt tokens; higher likelihood means the image can more reliably be “read back” as the requested description. Self-SpectraReward applies the same principle inside a unified multimodal model, using BAGEL’s understanding branch to reward its own generation branch, with no external reward model or preference labels. Across two generators, three RL algorithms, nine MLLM backbones spanning 4B to 235B, and five out-of-distribution benchmarks, the method consistently improves prompt following. With BAGEL at 512 resolution, SpectraReward raises TIIF-Bench overall short-prompt performance by 10.0 points and reaches 89.5 on GenEval with Self-SpectraReward, while Self-SpectraReward reaches 0.76 on WISE with self-chain-of-thought evaluation. Notably, reward-model scale is non-monotonic: Qwen3-VL-235B-A22B underperforms smaller models, while policy-aligned self-reward matches or surpasses much larger external evaluators. The strongest result suggests that tokenizer, vision-encoder, and data-distribution alignment can matter more than raw MLLM size for image-generation RL.
Original abstract
In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.