Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
AuthorsTianyi Xiong, Zhengyuan Yang, Xiaofei Wang, Chung-Ching Lin, Ruichun Ma, Kevin Lin, Zhendong Wang, Linjie Li, Chenxi Liu, Ruibo Chen, Ramani Duraiswami, Heng Huang, Lijuan Wang
Resources
RubSE helps vision-language models iteratively repair generated interfaces by turning visual feedback into prioritized, structured rubrics instead of letting each code edit disrupt the whole page.
Key results
RubSE outperformed naïve self-evolution in 15 of 18 final-round settings.
Average final-round improvement over naïve self-evolution in overall score points.
Average collapse rate for frontier models with RubSE, down from 18.9%.
Average recovery rate for frontier models using RubSE, up from 20.7%.
RubSE costs 1.60 times naïve self-evolution for GPT-5.4.
What the paper found
RubSE addresses a central failure in UI-to-code self-evolution: visual repair coupling, where a local HTML/CSS edit propagates through layout, styling, and component dependencies, fixing one mismatch while damaging faithful regions. The framework converts screenshot feedback into structured visual-repair context through an EVOLVE–SELECT–HISTORY loop. EVOLVE generates typed rubrics across five aspects—layout geometry, spacing density, typography, styling, and completeness—SELECT chooses one prioritized repair target, and HISTORY prevents repeated or over-broad edits. Across six vision-language models and three benchmarks, including Design2Code and UI2Code-Real, RubSE improves naïve self-evolution in 15 of 18 final-round settings, with average gains of 1.20 overall-score points and 0.11 aspect-mean points; it also wins in 14 of 18 best-round settings. The evaluation covers OpenAI’s GPT-5.4 and GPT-5.2, Anthropic’s Claude-Sonnet-4.5, and Qwen3-VL-32B-Instruct, Qwen3.5-9B, and Qwen3.6-35B-A3B. RubSE reduces trajectory collapse for frontier models from 18.9% to 12.8% on average and improves recovery from 20.7% to 32.8%. Stronger GPT-5.4-generated rubrics also transfer to weaker Qwen code improvers by emphasizing global structure and actionable visual relations rather than isolated CSS patches. The trade-off is computational: RubSE costs 1.60 times as much as naïve self-evolution for GPT-5.4, while adding only 2.5% latency for Qwen3.5-9B.
Original abstract
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, termed visual repair coupling: a local code edit may propagate through layout, style, and component dependencies, correcting one visual mismatch while degrading regions that were previously faithful. To address this issue, we present RubSE, a Rubric-guided Self-Evolution framework that uses rubrics to represent visual feedback as a structured visual-repair context. At each refinement round, RubSE generates typed candidate rubrics, selects one prioritized repair target, and stores previously selected rubrics as history, thereby steering each revision toward a well-scoped visual repair while discouraging repeated or over-broad changes. Evaluations across six VLMs and three UI-to-code benchmarks demonstrate that RubSE substantially outperforms naïve self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling. Further analysis shows that RubSE mitigates trajectory collapse by improving recovery from severe visual regressions, and that stronger rubric generators can transfer effective visual-repair guidance to weaker code improvers.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.