Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
AuthorsXin Jin, Huanqia Cai, Zhen Li, Zechao Zhan, Dengyang Jiang, Aiming Hao, Yuming Jiang, Chunle Guo, Peng Gao, Ming-Ming Cheng, Steven C. H. Hoi
Resources
This paper upgrades reward models for image generation by teaching them to reason about score distributions instead of predicting a single scalar, making them both more accurate and easier to deploy.
Key results
27B GDSO teacher on the internal test set
9B RISD student on the internal test set
27B GDSO teacher score calibration
9B RISD student score calibration
net GSB improvement over the SFT baseline
What the paper found
Z-Reward, from Alibaba Group’s Z-Image Team and Nankai University, reframes visual reward modeling as a reasoning-conditioned score distribution instead of a single scalar. The paper uses Qwen3.5-27B as a teacher and Qwen3.5-9B as a student, with Group-wise Direct Score Optimization (GDSO) to train the teacher on pointwise score calibration and same-prompt score-gap supervision, then Reasoning-Internalized Score Distillation (RISD) to transfer the teacher’s distributional judgment into a compact deployable model that no longer emits reasoning chains at inference time. On the internally annotated benchmark, the 27B GDSO teacher reaches 89.6% human preference accuracy and 0.7620 PLCC, outperforming SFT, RewardDance, and GRPO, while the 9B RISD student reaches 88.6% human preference accuracy and 0.7391 PLCC, closely matching the larger teacher and beating the OPD baseline. The annotation scheme itself is fine-grained, using four dimensions—text–image alignment, realism, aesthetics, and physical plausibility—scored on a nine-level half-point scale from 1.0 to 5.0. In text-to-image optimization, the deployed Z-Reward signal is differentiable through score expectations and yields a 41.3% net human-preference improvement over the SFT baseline, showing that internalized reasoning can serve as both a calibrated evaluator and an effective optimization objective.
Original abstract
Reward models are central to text-to-image post-training, but visual preference is subjective and better represented as a distribution over rubric scores than as a deterministic scalar. Existing scalar, score-token, and pairwise reward models over-compress uncertainty and fine-grained score differences, while reasoning-based generative rewards provide stronger judgments but are costly to deploy and difficult to use as direct optimization signals. We propose Z-Reward, a teacher-student reward modeling framework that decouples reasoning-heavy judgment from efficient reward deployment. The teacher is a large VLM that uses reasoning to infer rubric-aligned score distributions, and is trained with Group-wise Direct Score Optimization (GDSO), which combines policy-gradient rewards from distribution expectations with direct pointwise and pairwise supervision on score distributions and score gaps. The student is trained with Reasoning-Internalized Score Distillation (RISD), which transfers the teacher's reasoning-conditioned score distribution into a compact VLM without requiring explicit reasoning chains at inference time. On our internally annotated evaluation set, the 27B GDSO teacher reaches 89.6% human preference accuracy, outperforming SFT, RewardDance, and GRPO, while the 9B RISD student reaches 88.6%, outperforming the OPD baseline and closely matching the larger teacher. We further show that Z-Reward can serve as a differentiable reward signal for text-to-image optimization, yielding a 41.3% net human-preference improvement over the SFT baseline.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.