NTH

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

AuthorsTianyi Xiong, Zhengyuan Yang, Xiaofei Wang, Chung-Ching Lin, Ruichun Ma, Kevin Lin, Zhendong Wang, Linjie Li, Chenxi Liu, Ruibo Chen, Ramani Duraiswami, Heng Huang, Lijuan Wang

August 28, 2026 2 min read
Watch on YouTube
The one-line take

RubSE helps vision-language models iteratively repair generated interfaces by turning visual feedback into prioritized, structured rubrics instead of letting each code edit disrupt the whole page.

Key results

15
Final-round settings improved

RubSE outperformed naïve self-evolution in 15 of 18 final-round settings.

1.20
Average overall-score gain

Average final-round improvement over naïve self-evolution in overall score points.

12.8%
Trajectory collapse reduction

Average collapse rate for frontier models with RubSE, down from 18.9%.

32.8%
Recovery rate

Average recovery rate for frontier models using RubSE, up from 20.7%.

1.60
GPT-5.4 API cost multiplier

RubSE costs 1.60 times naïve self-evolution for GPT-5.4.

What the paper found

RubSE addresses a central failure in UI-to-code self-evolution: visual repair coupling, where a local HTML/CSS edit propagates through layout, styling, and component dependencies, fixing one mismatch while damaging faithful regions. The framework converts screenshot feedback into structured visual-repair context through an EVOLVE–SELECT–HISTORY loop. EVOLVE generates typed rubrics across five aspects—layout geometry, spacing density, typography, styling, and completeness—SELECT chooses one prioritized repair target, and HISTORY prevents repeated or over-broad edits. Across six vision-language models and three benchmarks, including Design2Code and UI2Code-Real, RubSE improves naïve self-evolution in 15 of 18 final-round settings, with average gains of 1.20 overall-score points and 0.11 aspect-mean points; it also wins in 14 of 18 best-round settings. The evaluation covers OpenAI’s GPT-5.4 and GPT-5.2, Anthropic’s Claude-Sonnet-4.5, and Qwen3-VL-32B-Instruct, Qwen3.5-9B, and Qwen3.6-35B-A3B. RubSE reduces trajectory collapse for frontier models from 18.9% to 12.8% on average and improves recovery from 20.7% to 32.8%. Stronger GPT-5.4-generated rubrics also transfer to weaker Qwen code improvers by emphasizing global structure and actionable visual relations rather than isolated CSS patches. The trade-off is computational: RubSE costs 1.60 times as much as naïve self-evolution for GPT-5.4, while adding only 2.5% latency for Qwen3.5-9B.

Original abstract

Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, termed visual repair coupling: a local code edit may propagate through layout, style, and component dependencies, correcting one visual mismatch while degrading regions that were previously faithful. To address this issue, we present RubSE, a Rubric-guided Self-Evolution framework that uses rubrics to represent visual feedback as a structured visual-repair context. At each refinement round, RubSE generates typed candidate rubrics, selects one prioritized repair target, and stores previously selected rubrics as history, thereby steering each revision toward a well-scoped visual repair while discouraging repeated or over-broad changes. Evaluations across six VLMs and three UI-to-code benchmarks demonstrate that RubSE substantially outperforms naïve self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling. Further analysis shows that RubSE mitigates trajectory collapse by improving recovery from severe visual regressions, and that stronger rubric generators can transfer effective visual-repair guidance to weaker code improvers.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis