NTH

Think Before You Score: Thinking Reward Model for Visual Generation

AuthorsXuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

Affiliations[

September 30, 2026 2 min read
Watch on YouTube
The one-line take

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Key results

20K
Image-generation training cases

TRM's structured SFT data include approximately 20K image-generation cases.

28K
Image-editing training cases

TRM's structured SFT data include approximately 28K image-editing cases.

4000
Preference pairs

PD-GRPO is trained on approximately 4000 difficulty-aware preference pairs.

71.2%
TRM GenAI-T2I accuracy

TRM reaches 71.2% preference accuracy on GenAI-T2I.

67.9%
TRM MMRB2-T2I accuracy

TRM reaches 67.9% preference accuracy on MMRB2-T2I.

5.81
BAGEL TIIF-Long gain

TRM-guided reinforcement learning improves BAGEL's TIIF-Long score by 5.81 points.

What the paper found

Think Before You Score introduces the Thinking Reward Model, or TRM, for evaluating image generation and editing. Instead of mapping a prompt and image directly to a scalar, TRM first creates a case-adaptive rubric, checks atomic criteria with binary judgments, summarizes dimensions such as prompt alignment, aesthetics or source consistency, and then produces a fine-grained pointwise score. It is trained from approximately 20K image-generation cases and 28K image-editing cases, followed by 4000 difficulty-aware preference pairs. The second stage uses Pairwise Dual-Group Relative Policy Optimization, or PD-GRPO, which improves relative discrimination without the score polarization associated with Bradley–Terry optimization. Initialized from Qwen3.5-9B, TRM reaches 71.2% on GenAI-T2I and 67.9% on MMRB2-T2I, outperforming the reported 9B open-source reward models and GPT-4.1 on those benchmarks, while remaining competitive with Gemini 3 Pro. As a training signal, TRM-guided FlowGRPO improves BAGEL’s TIIF-Long score by 5.81 points, from 75.62 to 81.43, and also improves FLUX.1-dev, SD3.5-M, FLUX.2-Klein, and SenseNova-U1.5 across generation and editing benchmarks. The central result is that explicitly deciding what matters for each visual case produces more calibrated rewards and more useful reinforcement-learning gradients than direct or fixed-criterion scoring.

Original abstract

Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards. We further observe that conventional pairwise preference optimization can induce score polarization, and introduce Pairwise Dual-Group Relative Policy Optimization (PD-GRPO), which leverages pairwise supervision to improve reward discrimination while preserving fine-grained pointwise scoring. Extensive experiments on image generation and editing reward-modeling benchmarks demonstrate that TRM achieves state-of-the-art performance among open-source reward models while remaining highly competitive with proprietary alternatives. Moreover, using TRM as a reward for reinforcement learning consistently improves diverse visual generation models, demonstrating that its fine-grained, case-adaptive rewards translate into effective optimization signals for visual generation.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis
03Generative Model

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou

MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.

Read analysis