NTH

When LLM Decompilers Recompile More and Preserve Less

AuthorsChang Liu, Edward Raff, Kristopher Micinski

September 14, 2026 3 min read
Watch on YouTube
The one-line take

This work shows that LLM decompilers can produce code that compiles and passes known tests while silently changing program behavior or erasing vulnerabilities.

Key results

4.9%
Passers that diverge

Candidates passed every shipped test but diverged under Decompile-Diverge.

13%
Maximum single-system divergence

Highest divergence rate among systems on established LLM decompilation corpora.

300
GitHub benchmark functions

Real library functions used in the GitHub evaluation track.

287
CVE benchmark functions

Vulnerability-grounded functions used to measure behavioral and crash preservation.

90%
LLM4Decompile GitHub build rate

Raised from Ghidra’s 75% build rate, while Matched fell from 74% to 62%.

8.7%
CVE Crash Absence

Share of vulnerable functions whose recompiled output silently removed the reference crash.

What the paper found

This paper argues that LLM decompilers can recompile more code while preserving less of the original behavior. Its Decompile-Diverge framework synthesizes a driver for each function, uses AFL++ to fuzz the reference implementation, and compares bounded observable state—including outputs, writes, crashes, and hangs—without relying on shipped tests. Across eight systems in nine configurations evaluated on HumanEval-Decompile, ExeBench, AnghaBench, and MBPP, 4.9% of candidates that passed every shipped test still diverged, reaching 13% for one system. A focused example shows SK2Decompile matching only 70/100 generated inputs and LLM4Decompile 66/100, despite both passing 10/10 shipped tests. On 300 GitHub library functions and 287 CVE-grounded functions, LLM4Decompile raises Ghidra’s build rate from 75% to 90%, but its Matched rate drops from 74% to 62%. On the CVE track, 8.7% of functions exhibit Crash Absence: the recompiled code silently removes a vulnerability-triggering crash. Source analysis attributes these failures to LLMs inventing fields, types, callees, constants, and guards that replace visible unknowns such as undefined4 and DAT_* in Ghidra or Hex-Rays output. Conservative rewriting with DeGPT-Qwen more closely preserves front-end behavior, while aggressive systems such as SK2Decompile and LLM4Decompile improve compilation at the cost of semantic fidelity. The inclusion of GLM-5.2 and Qwen3.6-35B-A3B further shows that this is an evaluation problem spanning multiple contemporary model families, not a defect isolated to one decompiler.

Original abstract

Decompilation recovers high-level source from compiled machine code and serves as a foundation for security tasks such as vulnerability detection and malware analysis. Traditional decompilers like Ghidra and Hex-Rays expose whatever they cannot resolve as visible placeholders and often emit pseudocode that will not compile or execute; LLM-based decompilers produce clean, idiomatic C and are now judged almost entirely by recompilability and re-executability: whether the output builds and passes its shipped input/output tests. We show that these metrics can reward the wrong path: a function may recompile and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability may disappear from the recompiled code with no visible trace of the crash. Neither failure is caught by existing suites. To address this gap, we propose Decompile-Diverge, a behavioral comparison oracle not relying on fixed or hand-crafted tests: for each function it synthesizes a driver, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect changes in the function's behavior. Across eight systems in nine configurations on established LLM decompilation corpora, candidates that pass every shipped test still diverge from the original on our input corpus: 4.9% overall, and as many as 13% for a single system. On 300 real GitHub library functions and 287 CVE-grounded functions, recompilability and behavioral agreement can come apart: the strongest refinement LLM lifts Ghidra's build rate from 75% to 90%, while its Matched rate falls from 74% to 62%; on disclosed vulnerabilities, up to one tenth exhibit Crash Absence in its output. Source-level analysis traces this divergence to introduced fields, types, callees, and guards that replace the visible unknowns traditional tools leave behind.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis