Paper: 2609.05370 Authors: Chang Liu, Edward Raff, Kristopher Micinski Categories: cs.CR, cs.AI

The Gap

Decompilation is the indispensable craft of reversing compiled binary machine code back into high-level human-readable C source code. It underpins malware reverse engineering, vulnerability discovery, patch verification, and firmware audits. Traditional decompilers like Ghidra or Hex-Rays IDA Pro operate on deterministic abstract interpretation. When they fail to recover a data structure or control flow, they expose ugly placeholders, complex pointer casts, or raw goto statements that often do not recompile cleanly.

The arrival of large language models sparked a revolution in decompilation. LLM-based decompilers produce beautiful, idiomatic C code with descriptive variable names and structured loops. Academic benchmarks quickly crowned this new wave, judging models almost exclusively on two metrics: recompilability (whether the generated C code builds without syntax errors) and re-executability (whether the recompiled binary passes the author’s shipped unit tests).

This paper exposes the lethal hazard behind those metrics. A decompiled function can easily recompile and pass a handful of simple happy-path unit tests, while completely corrupting the binary’s actual semantic behavior on edge cases. Worst of all, when decompiling binaries containing documented memory corruption vulnerabilities, LLM decompilers routinely rewrite tricky pointer arithmetic into clean, sanitized buffers. The vulnerability vanishes from the source code, blinding malware analysts and security auditors to the real exploit.

[COMPILED MACHINE CODE] Contains subtle pointer arithmetic & boundary vulnerability
             |
             +------------------------------+------------------------------+
             v                                                             v
  [TRADITIONAL: GHIDRA / IDA]                                   [LLM DECOMPILER]
  Messy, uncompilable pseudocode                                Beautiful, idiomatic C code
  - Exposes raw memory offsets                                  - 95%+ Recompilability score
  - Preserves exact crashing paths                              - Passes basic unit tests
  - Hard for humans to read quickly                             - BUT: "smooths over" unsafe logic
             |                                                             |
             v                                                             v
  True Vulnerability Preserved                                  Vulnerability Silently Erased!

The Increment

One sentence: Before this paper, high recompilability was celebrated as the gold standard of LLM decompilation; after it, the Syracuse and CrowdStrike team proves that recompilability actively incentivizes models to hallucinate safe code, demonstrating that LLM decompilers frequently recompile more while preserving less semantic reality.

Core Mechanism

The authors formalize this divergence through Decompile-Diverge, a differential fuzzing and behavioral auditing oracle designed specifically for decompiled code.

Rather than relying on limited, author-supplied test suites, Decompile-Diverge pairs the original compiled binary and the recompiled LLM candidate, feeding them hundreds of thousands of inputs generated via LibFuzzer-guided differential testing:

  1. Differential Execution Oracle: The fuzzer executes both binaries side-by-side with identical inputs, comparing return values, memory writes, and standard output streams to uncover behavioral deviation.
  2. Crash & Vulnerability Preservation Auditing: The framework feeds known proof-of-concept exploit payloads into both binaries. In numerous instances, the original binary crashed with a segmentation fault or buffer overflow, whereas the LLM-recompiled binary completed cleanly because the LLM had replaced dynamic pointer bounds with fixed, safe array indices.
  3. Semantic Inversion Diagnosis: The authors discovered that refinement models and end-to-end models regularly substitute tricky loops with standard textbook algorithms that look plausible but implement the opposite control logic.
   DECOMPILE-DIVERGE AUDITING PIPELINE

           [Original Binary (C)]
                     |
       +-------------+-------------+
       |                           |
       v                           v
  [Test Input / Fuzzing]     [Decompiler] -> [LLM Recompiled Binary (C')]
       |                                              |
       +--------------------+-------------------------+
                            |
                            v
            [Differential Behavioral Oracle]
                            |
         +------------------+------------------+
         v                                     v
   [Semantic Divergence]             [Vulnerability Erasure]
   Output C != Output C'             Original Crashes;
   on edge-case inputs               Decompiled Code Silently Passes

To understand why this happens, consider a structural metaphor of a restoration artist fixing a damaged ancient manuscript. A traditional conservator leaves torn edges, stained parchment, and missing words as blank spaces with bracketed question marks. It is ugly and unreadable as a bedtime story, but a historian can see exactly what survived. An LLM restorer is an overly eager apprentice who fills every torn paragraph with smooth, lyrical prose. The resulting book reads effortlessly and wins awards for grammar, but the apprentice has accidentally written out the murder confession that was the sole reason detectives were inspecting the scroll.

Key Concepts

  • Recompilability Illusion: The false assumption that because decompiled source code builds without compiler errors, it must accurately represent the behavior of the source binary.
  • Vulnerability Erasure: The failure mode where an LLM decompiler replaces unsafe pointer operations or buffer offsets with sanitized idioms, causing security analysts to miss critical flaws.
  • Differential Behavioral Fuzzing: Testing original and decompiled binaries against extensive randomized edge-case inputs to identify functional divergences that escape static unit tests.

Framework Shift

Before (Current LLM Decompiler Evaluation):
Assembly -> [LLM] -> C Code -> `gcc C Code` -> [Passes Build?] -> High Score!
(Rewards hallucinated safe code that compiles easily)

After (Decompile-Diverge Differential Protocol):
Assembly -> [LLM] -> C Code -> [Differential Fuzzing vs Original Binary]
                                            |
                         +------------------+------------------+
                         v                                     v
                 [Semantic Fidelity]                 [Exploit Preservation]

From rewarding superficial syntactic validity to enforcing differential runtime equivalence under adversarial inputs, the core shift is unmasking the aesthetic hallucinations of LLM code models.

Expert Assessment

Problem choice: Crucial for the software security industry. Decompilation tools are used in high-stakes environments where believing an inaccurate decompilation can lead to disastrous false-negative vulnerability assessments.

Method maturity: Using LibFuzzer differential execution provides an objective, uncheatable ground truth that does not rely on LLM judges.

Experimental integrity: The inclusion of CrowdStrike malware engineers grounds the paper in operational reality. The case studies showing real CVEs vanishing in recompiled source are deeply alarming.

Writing quality: Exceptional clarity. The distinction between syntax recovery, operational re-executability, and deep semantic equivalence is articulated with surgical precision.

Verdict: strong accept — Essential reading for any team applying LLMs to reverse engineering or program analysis.

Takeaways

  • Never trust an LLM decompiler for vulnerability audits without cross-verifying against raw Ghidra/IDA assembly offsets.
  • High recompilability is often an inverse indicator of faithfulness: models that try to guess missing types often hallucinate away tricky edge-case bugs.
  • Benchmark designers for code generation must abandon shallow unit tests in favor of differential fuzzing against the original executable.

论文: 2609.05370 作者: Chang Liu, Edward Raff, Kristopher Micinski 分类: cs.CR, cs.AI

缺口

逆向反编译(Decompilation)是将底层机器二进制汇编还原为高级 C 语言代码的核心技术,是恶意软件分析、漏洞挖掘、固件审计和安全补丁验证的基石。 传统的反编译工具(如 NSA 开源的 Ghidra 或顶级商业逆向工具 IDA Pro)完全依赖确定性的控制流图分析与类型推断。 当遇到无法解构的复杂数据结构或混淆指针运算时,传统工具会如实输出丑陋的占位符、强转指针或繁杂的 goto 语句。 这些伪代码往往根本无法重新编译执行,但却如实保留了底层每一个细微的内存偏移。

大语言模型的爆发彻底重塑了反编译领域。 基于 LLM 的反编译器能够生成格式优美、变量命名自然、控制结构清晰的标准 C 语言代码。 学术界和工业界榜单随即对此顶礼膜拜,几乎完全依赖两个指标给模型打分:可重编译率(Recompilability)(代码能否通过编译器构建)以及可重执行率(Re-executability)(编译出的程序能否跑通作者随附的单元测试)。

然而,雪城大学与全球网络安全巨头 CrowdStrike 的研究人员刺破了这一繁荣假象: 这两个指标正在奖励完全错误的研发方向! 大模型生成的代码可能在语法上毫无瑕疵,也能跑通几个简单的正向测试用例,但在边界输入下行为与原程序彻底分道扬镳。 最致命的是,当反编译一段本身包含缓冲区溢出等高危漏洞的二进制程序时,大模型往往会“自作聪明”地将不安全的指针算术替换为安全的数组访问——原始漏洞在反编译代码中彻底蒸发了,给安全分析人员造成极其危险的虚假安全感。

[原始机器二进制码] 包含底层的危险指针运算与隐蔽内存漏洞
           |
           +------------------------------+------------------------------+
           v                                                             v
  [传统工具: Ghidra / IDA Pro]                                   [LLM 神经反编译器]
  输出杂乱难读、无法通过编译的伪代码                              输出优雅易读的标准 C 语言代码
  - 如实保留原始内存偏移与底层指令                                - 可重编译率高达 95%+
  - 完整保留所有恶性崩溃与越界路径                                - 轻松跑通随附的基础单元测试
  - 人类逆向分析耗时费力                                          - 致命问题:随手“脑补并修复”了漏洞
           |                                                             |
           v                                                             v
  原始漏洞与崩溃行为完全保留                                     漏洞被无形抹去,安全审计彻底失效!

增量

一句话: 在这篇论文之前,可重编译率被奉为神经反编译技术的黄金标准;在这篇论文之后,研究团队证实高编译率实际上诱导了大模型进行虚假脑补——生成代码“重编译得越多,真实语义保留得越少”。

核心机制

针对现有测试用例覆盖面窄、容易被 LLM 投机取巧的弊端,研究团队提出了差分模糊测试审计框架 Decompile-Diverge。

Decompile-Diverge 抛弃了预设的静态测试集,直接将原始编译的二进制程序与 LLM 反编译后重新编译的新二进制置于同一竞技场:

  1. 差分运行时预言机(Differential Execution Oracle):基于 LibFuzzer 引导技术,对成对程序施加数十万次随机变异输入,全方位监控函数返回值、内存修改快照以及标准输出流的偏离。
  2. 漏洞与崩溃保真度审计(Vulnerability Erasure Auditing):将已知的漏洞利用载荷(PoC)输入两个程序。 实验发现大量令人冷汗直下的案例:原始程序发生段错误(Segmentation Fault)彻底崩溃,而大模型重编译的程序却毫无感知地平稳运行——因为大模型在反编译时,顺手把越界指针改写成了合规的安全边界!
  3. 控制语义异构归因:分析表明,很多端到端模型在遇到复杂循环时,会直接填入算法教科书里的标准模板,看似合情合理,实则彻底颠倒了原程序的业务逻辑。
   DECOMPILE-DIVERGE 差分模糊审计架构

           [原始二进制可执行程序 (C)]
                       |
         +-------------+-------------+
         |                           |
         v                           v
  [模糊测试海量变异输入]       [LLM 反编译器] -> [重编译的二进制程序 (C')]
         |                                           |
         +---------------------+---------------------+
                               |
                               v
               [差分行为监控与状态比对预言机]
                               |
            +------------------+------------------+
            v                                     v
     [语义发散与逻辑断裂]                   [安全漏洞隐形蒸发]
     在非标准边界输入下,                   原始程序必然崩溃;
     两者的内存输出与返回值彻底背离         重编译代码由于脑补而静默通关

可以用一个文物修复学徒修补古卷的核喻来理解这个困局: 传统的文物修复大师面对残破的古籍,会保留撕裂的毛边、残缺的墨迹,用带问号的括号标注模糊字迹。 这篇残卷看起来破破烂烂,普通人根本读不通,但历史学家却能依据残余墨迹精准破案。 而大模型反编译器就像一个自作聪明的学徒工:看到断句,大笔一挥补上几段辞藻华丽的现代散文。 整本书装订得漂漂亮亮,甚至能拿去朗诵比赛获奖,但他却随手把古卷上记载凶手真实罪证的那行关键错字给涂改掉了。

关键概念

  • 重编译幻象(Recompilability Illusion):错误地将“代码能够通过 gcc 编译”等同于“反编译代码忠实还原了原程序语义”的行业盲区。
  • 漏洞蒸发(Vulnerability Erasure):大模型反编译器在遭遇复杂底层操作时,习惯性用标准库模式替换未定义行为,导致安全漏洞在生成的源码中不留痕迹地消失。
  • 差分模糊测试(Differential Fuzzing):通过高频生成海量边缘输入,动态比对两套二进制执行轨迹差异的高阶程序分析方法。

框架转变

之前(主流学术界的评价标准):
汇编指令 -> [LLM 转换] -> C 源码 -> `gcc 编译` -> [编译成功且正向用例通过] -> 评定为高分突破!
(实质上在奖励模型通过编造安全代码来迎合编译检查)

之后(Decompile-Diverge 真实语义对抗评测):
汇编指令 -> [LLM 转换] -> C 源码 -> [与原程序进行差分模糊测试对抗]
                                                |
                             +------------------+------------------+
                             v                                     v
                     [真实语义保真度]                       [底层漏洞可复现性]

从仅仅考核表层代码的美观度与静态编译通过率,转变为在数万次严苛模糊测试下审判运行时语义的一致性,核心转变在于击碎了神经反编译领域的“美学泡沫”。

专家评审

选题眼光: 极具行业穿透力与警示意义。 反编译是网络攻防与恶意代码分析的刚需场景,若分析人员轻信大模型“修复后”的代码,将直接导致漏洞漏报与严重的战略误判。

方法成熟度: 引入成熟的 LibFuzzer 差分模糊测试作为裁判,摆脱了传统代码相似度指标(如 CodeBLEU)在逆向领域的纸上谈兵,具备绝对的客观性。

实验诚意: CrowdStrike 资深安全专家的参与为论文注入了浓厚的实战色彩。 对真实 CVE 漏洞在反编译代码中消失过程的案例解构详实震撼,说服力极强。

写作功力: 概念剖析鞭辟入里,对模型幻觉诱因的分析条理清晰。

Verdict: 强接收(Strong Accept) — 逆向工程与 AI 代码安全结合领域的必读里程碑。

要点总结

  • 开展漏洞挖掘或恶意软件逆向时,切勿盲目相信 LLM 反编译出的漂亮 C 代码,必须对照 Ghidra/IDA 的底层汇编与原始偏移进行交叉核验。
  • 极高的重编译率往往是“失真脑补”的反向警报:越急于讨好编译器的模型,越倾向于悄悄抹去难处理的边界缺陷。
  • 代码生成领域的基准设计者应当彻底摒弃玩具式的固定单元测试,转向对抗性的差分模糊执行评估。