NTH

Reinforcement Learning from Intermediate Renders for Image-to-Code Generation

AuthorsOmri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel

AffiliationsWeizmann Institute of Science · MITProject page: https://ir4rl.github.io

October 4, 2026 2 min read
Watch on YouTube
The one-line take

IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.

Key results

700
SVG training set

svg-stack examples used to train the OmniSVG-4B IR4RL model

97.48
MMSVGBench Illustrations DINO

DINO score achieved by OmniSVG-4B with combined process and outcome rewards

2.5k
SVG output length

Token count for IR4RL outputs on MMSVGBench Illustrations

86.9
DaTikZ-v3 DreamSim

DreamSim score achieved by IR4RL for Image-to-TikZ generation

0.6k
TikZ output length

Token count for IR4RL outputs on DaTikZ-v3

10
Best process-reward settings

Process-reward weight alpha identified in the SVG ablation

What the paper found

IR4RL, or Intermediate Renders for Reinforcement Learning, addresses a core weakness of image-to-code reinforcement learning: outcome-only rewards score the final rendering but cannot identify which tokens helped or harmed it. The method closes each executable SVG or TikZ prefix, renders it, computes the visual-score change between consecutive prefixes, and propagates these delta rewards backward through tokens with exponential discounting, while combining them with the final outcome reward under GRPO. Using OmniSVG-4B trained on 700 svg-stack examples, IR4RL reaches a DINO score of 97.48 on MMSVGBench Illustrations while reducing output length to 2.5k tokens, outperforming supervised fine-tuning, RAFT, and outcome-only RL. On Image-to-TikZ, it achieves 86.9 DreamSim and produces 0.6k-token programs on DaTikZ-v3, surpassing DeTikZify baselines. Ablations identify a process-reward weight of 10 and propagation factor of 0.9 as effective settings, with rendering after every completed drawing command performing best. The approach also compares favorably with larger general-purpose systems, including Qwen3-VL-235B, Gemini 3 Flash, Anthropic’s Sonnet 5, and OpenAI’s GPT-5.2. IR4RL changes training supervision without changing inference, but requires representations whose partial programs remain meaningfully renderable.

Original abstract

Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that is poorly aligned with the contribution of individual tokens. A generated program may contain operations that accurately reproduce some parts of the target image alongside others that introduce errors, yet all tokens are trained from the same final outcome. We observe that many intermediate code prefixes are not only executable, but already produce meaningful partial renders that reflect progress toward the target. This property provides a natural source of denser supervision during generation. Based on this observation, we introduce IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. We evaluate our approach on Image-to-SVG and Image-to-TikZ generation. Across both tasks, our method improves over supervised fine-tuning and standard GRPO, yielding new state-of-the-art open-source models. This shows that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code models.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis
03Code Generation

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo

CodeMidas turns existing codebases into scalable, automatically verified RL environments that train coding agents to perform better across diverse software tasks.

Read analysis