ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
AuthorsJeonghye Kim, Minseon Kim, Young Jin Kim, Matheus Pereira, Marc-Alexandre Côté, Alessandro Sordoni, Xingdi Yuan, Zhengyan Shi
AffiliationsKAIST · Introduction Figure 1: ProgramDistill evaluates reference-guided software engineering for interactive web applications. A coding agent is given a working reference application whose source implementat · Microsoft Research Montréal · Microsoft AI https://aka.ms/froggy Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical
Resources
ProgramDistill tests whether coding agents can rebuild web-app features by exploring working reference software rather than merely following written specifications.
Key results
Applications used to generate the ProgramDistill benchmark
Behaviors mined and verified through deterministic browser replay
Atomic and cumulative SWE tasks created automatically
Mean binary score on the 300-task partial-application suite
Cumulative workflow recovery in full-application reconstruction
Partial-application binary success at restoration depth 8
What the paper found
ProgramDistill, developed with Microsoft, introduces a benchmark for reference-guided software engineering: an agent observes a working interactive web app, infers its behavior without source access, repairs an incomplete version, and validates the result through replayable browser traces. Its automated mine-craft-patch pipeline, built with GPT-5.6 Sol, factorizes 26 applications into 1975 replay-verified behaviors and 4063 tasks, using prerequisite lineages, source masks, deterministic execution, and gold patches as executable verifiers. The benchmark spans atomic repair, cumulative restoration, and full-application reconstruction, with difficulty controlled by restoration depth. On partial reconstruction, OpenAI’s GPT-6 Astra achieves an 84.3% mean binary score, ahead of Anthropic’s Claude Opus 5 at 68.7%, but Astra’s performance falls from 100% at depth 1 to 64.0% at depth 8. In full-application reconstruction, Astra reaches 49.2% cumulative workflow recovery, compared with 28.8% for Claude Opus 5. The study also evaluates GPT-5.3 Codex, Google DeepMind’s Gemini models, and xAI’s Grok 4.6, showing that deeper tasks expose compositional failures: agents observe and validate less per required behavior as repair burden grows. ProgramDistill therefore serves both as a diagnostic benchmark and as a potential curriculum and reinforcement-learning environment for coding agents.
Original abstract
Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. We build ProgramDistill by factorizing applications into features of different granularities, each associated with replayable behaviors executable via its gold patch. Our pipeline, mine-craft-patch, discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction. In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8. ProgramDistill thus provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.