NTH

RULER: Instance-aware Rubric Rewards for SVG Generation

AuthorsHangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

AffiliationsAnt Group · The Hong Kong University of Science and Technology (Guangzhou) · Independent Researcher · University of Oxford

September 30, 2026 2 min read
Watch on YouTube
The one-line take

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Key results

900
Human sample count

Rendered SVG samples used to compare automated metrics with human judgments.

0.7929
Rubric human correlation

Spearman correlation between rubric scores and human quality ratings.

0.7574
Rubric ranking agreement

Goodman-Kruskal agreement with human pairwise preferences.

0.693
MMSVG-Illustration rubric score

RULER score on the MMSVG-Illustration benchmark.

0.683
MMSVG-Icon rubric score

RULER score on the MMSVG-Icon benchmark.

What the paper found

RULER addresses a central problem in text-to-SVG generation: there is usually no single correct image, while scalar rewards such as CLIP and Aesthetic can mis-rank stylized vectors and invite reward hacking. The method uses Claude-Opus-4.6 to convert each instruction into a six-item, instance-aware rubric covering semantic fidelity, visual quality, and rendering style; Qwen3-VL-8B then scores rendered SVG rollouts item by item, and the weighted satisfactions train a Qwen3-8B policy with Group Relative Policy Optimization. This requires neither paired SVG references nor human preference labels. In an evaluation of 900 human-annotated samples, rubric scoring aligned with human judgments at Spearman correlation 0.7929 and Goodman-Kruskal ranking agreement 0.7574, outperforming CLIP and Aesthetic. On MMSVG-Illustration and MMSVG-Icon, RULER raised rubric scores to 0.693 and 0.683, compared with Qwen3-8B baselines of 0.432 and 0.395, respectively, matching the substantially larger DeepSeek-V3 and exceeding dedicated systems such as OmniSVG and JanusCoder. The study also shows why reward design matters: conventional CLIP, Aesthetic, and HPS reinforcement learning inflated aesthetic scores while producing longer, visually degraded SVGs. For evaluation, GPT-5-mini served as an independent rubric judge, while Claude and OpenAI’s GPT-5.5 generated alternative rubrics with similar gains, suggesting the approach is not tied to one frontier model.

Original abstract

Generating Scalable Vector Graphics (SVG) code from natural-language instructions is an open-ended task without absolute visual ground truth, leaving both evaluation and policy optimization without a faithful signal. Scalar metrics (CLIP, Aesthetic) calibrated on natural images transfer poorly to stylized vector content, and reusing them as RL rewards triggers reward hacking. We address both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization. Because the rubric is derived from text alone, RULER requires neither paired SVG ground truth nor human preference labels. On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683, surpassing dedicated SVG specialists and matching the substantially larger DeepSeek-V3, with ablations identifying rubric design as the active lever for RL on open-ended SVG generation. The project page is available at https://hangyuran.github.io/RULER/.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
02Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis
03Generative Model

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou

MovieGrid turns long videos into jointly modeled spatial grids, helping generative models produce more coherent stories with many connected shots.

Read analysis