Think Before You Score: Thinking Reward Model for Visual Generation
AuthorsXuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
Affiliations[
Resources
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
Key results
TRM's structured SFT data include approximately 20K image-generation cases.
TRM's structured SFT data include approximately 28K image-editing cases.
PD-GRPO is trained on approximately 4000 difficulty-aware preference pairs.
TRM reaches 71.2% preference accuracy on GenAI-T2I.
TRM reaches 67.9% preference accuracy on MMRB2-T2I.
TRM-guided reinforcement learning improves BAGEL's TIIF-Long score by 5.81 points.
What the paper found
Think Before You Score introduces the Thinking Reward Model, or TRM, for evaluating image generation and editing. Instead of mapping a prompt and image directly to a scalar, TRM first creates a case-adaptive rubric, checks atomic criteria with binary judgments, summarizes dimensions such as prompt alignment, aesthetics or source consistency, and then produces a fine-grained pointwise score. It is trained from approximately 20K image-generation cases and 28K image-editing cases, followed by 4000 difficulty-aware preference pairs. The second stage uses Pairwise Dual-Group Relative Policy Optimization, or PD-GRPO, which improves relative discrimination without the score polarization associated with Bradley–Terry optimization. Initialized from Qwen3.5-9B, TRM reaches 71.2% on GenAI-T2I and 67.9% on MMRB2-T2I, outperforming the reported 9B open-source reward models and GPT-4.1 on those benchmarks, while remaining competitive with Gemini 3 Pro. As a training signal, TRM-guided FlowGRPO improves BAGEL’s TIIF-Long score by 5.81 points, from 75.62 to 81.43, and also improves FLUX.1-dev, SD3.5-M, FLUX.2-Klein, and SenseNova-U1.5 across generation and editing benchmarks. The central result is that explicitly deciding what matters for each visual case produces more calibrated rewards and more useful reinforcement-learning gradients than direct or fixed-criterion scoring.
Original abstract
Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou
MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.