NTH

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

AuthorsRunhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao

July 19, 2026 2 min read
Watch on YouTube
The one-line take

SpectraReward turns pretrained multimodal models into plug-and-play judges that can guide image generators without extra reward-model training.

Key results

10.0
TIIF-Bench short-prompt gain

SpectraReward improvement over the BAGEL baseline at 512 resolution.

89.5
Self-SpectraReward GenEval

GenEval score for BAGEL trained with Self-SpectraReward at 512 resolution.

0.76
Self-SpectraReward WISE

WISE score when Self-SpectraReward is evaluated with self-chain-of-thought.

235B
Reward MLLM scale range

Largest reward MLLM evaluated, Qwen3-VL-235B-A22B; scaling did not produce monotonic gains.

What the paper found

Researchers from the University of Hong Kong, Peking University, and ByteDance Seed introduce SpectraReward, a training-free reward function that converts a frozen pretrained multimodal large language model into a text-to-image reinforcement-learning evaluator. Instead of asking an MLLM to assign a noisy scalar score or decompose a prompt into verification questions, the method performs one image-conditioned, teacher-forced forward pass and averages the log-likelihood of the original prompt tokens; higher likelihood means the image can more reliably be “read back” as the requested description. Self-SpectraReward applies the same principle inside a unified multimodal model, using BAGEL’s understanding branch to reward its own generation branch, with no external reward model or preference labels. Across two generators, three RL algorithms, nine MLLM backbones spanning 4B to 235B, and five out-of-distribution benchmarks, the method consistently improves prompt following. With BAGEL at 512 resolution, SpectraReward raises TIIF-Bench overall short-prompt performance by 10.0 points and reaches 89.5 on GenEval with Self-SpectraReward, while Self-SpectraReward reaches 0.76 on WISE with self-chain-of-thought evaluation. Notably, reward-model scale is non-monotonic: Qwen3-VL-235B-A22B underperforms smaller models, while policy-aligned self-reward matches or surpasses much larger external evaluators. The strongest result suggests that tokenizer, vision-encoder, and data-distribution alignment can matter more than raw MLLM size for image-generation RL.

Original abstract

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL. Project Page: https://huangrh99.github.io/SpectraReward/

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis