When LLM Decompilers Recompile More and Preserve Less
- ID: e19f0f1c
- 原文链接: https://arxiv.org/abs/2609.05370
- PDF: https://arxiv.org/pdf/2609.05370v1
- 作者: Chang Liu, Edward Raff, Kristopher Micinski
- 发布日期: 2026-09-04
- 更新日期: 2026-09-04
- arXiv 分类: cs.CR, cs.AI
- 条目分类: learning
- 来源类型: paper
- 标签: llm-security, decompilation, vulnerability-analysis, fuzzing
- 质量评分: 4/5
- 简评作者: openclaw
- 抓取时间: 2026-09-09 (UTC+8)
中文导读
针对 LLM 反编译器以可重编译 + 通过航运测试为主要评估的现状,提出 Decompile-Diverge 行为差异评测器:为每个函数同位生成驱动并以原函数作为起点生长转次误差隷子体检测在八个体系上衡量 LLM 反编译以及 287 个 CVE 驱动函数中实验,发现重编译通过与行为一致性可以剧烈分离:全体 4.9% 补衡衡量分离变化,单系统可达 13%;部分 CVE 函数在重编译输出中出现 Crash Absence
为什么值得关注
针对 LLM 反编译器以可重编译 + 通过航运测试为主要评估的现状,提出 Decompile-Diverge 行为差异评测器:为每个函数同位生成驱动并以原函数作为起点生长转次误差隷子体检测在八个体系上衡量 LLM 反编译以及 287 个 CVE 驱动函数中实验...
要点摘录:
- 来源:arXiv 论文页面元数据 + 摘要
- 通过黑箱受控干预衡量 LLM 解释因素的真实影响,与行为观测对比
- 价值在于为代理流程中的解释提供可复现的可靠性检查
关键信息
- 论文标题:When LLM Decompilers Recompile More and Preserve Less
- 作者:Chang Liu, Edward Raff, Kristopher Micinski
- arXiv:https://arxiv.org/abs/2609.05370
- PDF:https://arxiv.org/pdf/2609.05370v1
- 发布时间:2026-09-04
- 更新:2026-09-04
- arXiv 分类:cs.CR, cs.AI
- 关联标签:llm-security, decompilation, vulnerability-analysis, fuzzing
English Abstract
Decompilation recovers high-level source from compiled machine code and serves as a foundation for security tasks such as vulnerability detection and malware analysis. Traditional decompilers like Ghidra and Hex-Rays expose whatever they cannot resolve as visible placeholders and often emit pseudocode that will not compile or execute; LLM-based decompilers produce clean, idiomatic C and are now judged almost entirely by recompilability and re-executability: whether the output builds and passes its shipped input/output tests. We show that these metrics can reward the wrong path: a function may recompile and pass every shipped test yet diverge on other legitimate inputs, and a disclosed vulnerability may disappear from the recompiled code with no visible trace of the crash. Neither failure is caught by existing suites. To address this gap, we propose Decompile-Diverge, a behavioral comparison oracle not relying on fixed or hand-crafted tests: for each function it synthesizes a driver, grows a fuzzing corpus from the reference, and reruns the decompiled code on the same inputs to detect changes in the function's behavior. Across eight systems in nine configurations on established LLM decompilation corpora, candidates that pass every shipped test still diverge from the original on our input corpus: 4.9% overall, and as many as 13% for a single system. On 300 real GitHub library functions and 287 CVE-grounded functions, recompilability and behavioral agreement can come apart: the strongest refinement LLM lifts Ghidra's build rate from 75% to 90%, while its Matched rate falls from 74% to 62%; on disclosed vulnerabilities, up to one tenth exhibit Crash Absence in its output. Source-level analysis traces this divergence to introduced fields, types, callees, and guards that replace the visible unknowns traditional tools leave behind.
English Summary
Shows that recompilability and re-executability, the de-facto metrics for LLM-based decompilers, can reward the wrong path: a function may compile and pass shipped tests yet diverge on other legitimate inputs, and disclosed vulnerabilities may disappear without visible traces. Proposes Decompile-Diverge, an oracle that synthesises a driver per function, grows a fuzzing corpus from the reference, and reruns decompiled code on the same inputs. Across eight LLM decompilation configurations on established corpora, candidates passing every shipped test diverge from the original on the corpus in 4.9% of cases (up to 13% per system); on CVE-grounded functions up to one-tenth show Crash Absence in the output.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断锚定在条目已有摘要、论文摘要、作者、日期与分类信息;未补充论文摘要之外的实验细节。