NTH

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

AuthorsNiels Mündler-Sasahara, Hristo Venev, Dawn Song, Martin Vechev, Jingxuan He

July 19, 2026 3 min read
Watch on YouTube
The one-line take

This work turns the compiler into a real-time coding partner that catches AI-generated Rust mistakes while the code is still being written.

Key results

13.1%
Average compiler-error rate with generative compilation

Aggregate rate across the evaluated Rust model-task configurations, improving on 20.7% with post-generation feedback.

20.7%
Average compiler-error rate with post-generation feedback

Baseline rate after standard compiler-feedback repair, down from 65.9% without compiler feedback.

11
Best functional-correctness configurations

Generative compilation achieved the highest functional-correctness score in 11 of 14 model-task configurations.

33.3%
Early error-detection point

Mean relative file position where generative compilation detected unrecoverable errors, compared with 100% for post-generation checking.

3
Median diagnostic delay

Median number of lines between the eventual compiler-error span and generative compilation’s report.

20
CRUST-Bench translation instances

Challenging C-to-Rust repository-level instances used in the Translation evaluation.

What the paper found

Researchers from ETH Zurich and the University of California, Berkeley, including Martin Vechev and Dawn Song, introduce generative compilation, a method that gives large language models compiler feedback while they are still producing code. Instead of waiting for a complete file, a lightweight transformation called a sealor closes each partial Rust program with syntax and typed placeholders, then reuses rustc to detect genuine dead ends; diagnostics are mapped back to the original prefix, allowing black-box APIs such as Anthropic’s Claude Opus 4.8, OpenAI’s GPT 5.3 Codex, and Google DeepMind’s Gemini 3.5 Flash to revise earlier code without constrained decoding. The approach is formally developed for Featherweight Rust, with completeness and selective soundness mechanized in Lean, and extended to real Rust using control-flow placeholders holediv() and holeval(). Across seven models on 20 CRUST-Bench translation instances and 30 UpdatedAPI projects, post-generation feedback reduced the average compiler-error rate from 65.9% to 20.7%, while generative compilation lowered it further to 13.1% and achieved the best functional-correctness result in 11 of 14 model-task configurations. In a replay analysis, generative compilation detected errors at 33.3% of file generation on average, compared with 40.3% for function-level checking and 100% for post-generation feedback, while its median diagnostic appeared only 3 lines after the eventual error source, versus 89 lines for post-generation checking. The result is an on-the-fly compiler loop that reduces error cascades while preserving rich diagnostics and black-box model compatibility.

Original abstract

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not guide intermediate generation steps, such as those during autoregressive LLM decoding. Constrained decoding intervenes earlier by rejecting invalid tokens during sampling, but requires white-box model access and costly reimplementation for semantic constraints.We introduce generative compilation, the first approach to obtaining compiler feedback on partial programs during generation. The core technical device is a sealor: a lightweight, mostly syntax-guided transformation that converts partial programs into complete ones that standard compilers can diagnose. It is designed such that possible-to-complete partial programs are never rejected, while preserving enough code context to catch genuine dead ends early. We construct such a sealor on a core Rust-like calculus and prove that it satisfies these properties, all mechanized in Lean. We extend it to the first partial-program checker for real Rust. We evaluate our method on challenging repository-level Rust coding tasks, across both frontier black-box and open-weight models. We show that generative compilation reduces non-compiling outputs and improves functional correctness, relative to standard post-generation feedback. It does so by detecting a broad range of errors close to their source and early during generation, thereby reducing errors cascades and enabling focused diagnostics. More broadly, generative compilation is a step toward making compilers a first-class citizen of AI-assisted programming active during generation, rather than a separate post-generation check.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis