NTH

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

AuthorsJeonghye Kim, Minseon Kim, Young Jin Kim, Matheus Pereira, Marc-Alexandre Côté, Alessandro Sordoni, Xingdi Yuan, Zhengyan Shi

AffiliationsKAIST · Introduction Figure 1: ProgramDistill evaluates reference-guided software engineering for interactive web applications. A coding agent is given a working reference application whose source implementat · Microsoft Research Montréal · Microsoft AI https://aka.ms/froggy Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical

September 25, 2026 2 min read
Watch on YouTube
The one-line take

ProgramDistill tests whether coding agents can rebuild web-app features by exploring working reference software rather than merely following written specifications.

Key results

26
Interactive web applications

Applications used to generate the ProgramDistill benchmark

1975
Replay-verified behaviors

Behaviors mined and verified through deterministic browser replay

4063
Generated tasks

Atomic and cumulative SWE tasks created automatically

84.3%
GPT-6 Astra partial-repair score

Mean binary score on the 300-task partial-application suite

49.2%
GPT-6 Astra full-reconstruction score

Cumulative workflow recovery in full-application reconstruction

64.0%
Depth-8 GPT-6 Astra score

Partial-application binary success at restoration depth 8

What the paper found

ProgramDistill, developed with Microsoft, introduces a benchmark for reference-guided software engineering: an agent observes a working interactive web app, infers its behavior without source access, repairs an incomplete version, and validates the result through replayable browser traces. Its automated mine-craft-patch pipeline, built with GPT-5.6 Sol, factorizes 26 applications into 1975 replay-verified behaviors and 4063 tasks, using prerequisite lineages, source masks, deterministic execution, and gold patches as executable verifiers. The benchmark spans atomic repair, cumulative restoration, and full-application reconstruction, with difficulty controlled by restoration depth. On partial reconstruction, OpenAI’s GPT-6 Astra achieves an 84.3% mean binary score, ahead of Anthropic’s Claude Opus 5 at 68.7%, but Astra’s performance falls from 100% at depth 1 to 64.0% at depth 8. In full-application reconstruction, Astra reaches 49.2% cumulative workflow recovery, compared with 28.8% for Claude Opus 5. The study also evaluates GPT-5.3 Codex, Google DeepMind’s Gemini models, and xAI’s Grok 4.6, showing that deeper tasks expose compositional failures: agents observe and validate less per required behavior as repair burden grows. ProgramDistill therefore serves both as a diagnostic benchmark and as a potential curriculum and reinforcement-learning environment for coding agents.

Original abstract

Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. We build ProgramDistill by factorizing applications into features of different granularities, each associated with replayable behaviors executable via its gold patch. Our pipeline, mine-craft-patch, discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction. In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8. ProgramDistill thus provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis