RULER: Instance-aware Rubric Rewards for SVG Generation
AuthorsHangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
AffiliationsAnt Group · The Hong Kong University of Science and Technology (Guangzhou) · Independent Researcher · University of Oxford
Resources
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Key results
Rendered SVG samples used to compare automated metrics with human judgments.
Spearman correlation between rubric scores and human quality ratings.
Goodman-Kruskal agreement with human pairwise preferences.
RULER score on the MMSVG-Illustration benchmark.
RULER score on the MMSVG-Icon benchmark.
What the paper found
RULER addresses a central problem in text-to-SVG generation: there is usually no single correct image, while scalar rewards such as CLIP and Aesthetic can mis-rank stylized vectors and invite reward hacking. The method uses Claude-Opus-4.6 to convert each instruction into a six-item, instance-aware rubric covering semantic fidelity, visual quality, and rendering style; Qwen3-VL-8B then scores rendered SVG rollouts item by item, and the weighted satisfactions train a Qwen3-8B policy with Group Relative Policy Optimization. This requires neither paired SVG references nor human preference labels. In an evaluation of 900 human-annotated samples, rubric scoring aligned with human judgments at Spearman correlation 0.7929 and Goodman-Kruskal ranking agreement 0.7574, outperforming CLIP and Aesthetic. On MMSVG-Illustration and MMSVG-Icon, RULER raised rubric scores to 0.693 and 0.683, compared with Qwen3-8B baselines of 0.432 and 0.395, respectively, matching the substantially larger DeepSeek-V3 and exceeding dedicated systems such as OmniSVG and JanusCoder. The study also shows why reward design matters: conventional CLIP, Aesthetic, and HPS reinforcement learning inflated aesthetic scores while producing longer, visually degraded SVGs. For evaluation, GPT-5-mini served as an independent rubric judge, while Claude and OpenAI’s GPT-5.5 generated alternative rubrics with similar gains, suggesting the approach is not tied to one frontier model.
Original abstract
Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.
Read the original paperMore in Generative Models
Browse all 63 papers →Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.
Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou
MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.