AI 编程 5.0 · 必读 2026-08-13 · 论文

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the...

给 LLM 生成 GPU kernel 造了一台契约级验证器:十二道对抗性契约门,每道都是正确 kernel 必须满足的性质,其中数道零容差任何阈值选择都无法把失败解释掉对外审计某公开系统已验收的 2,638 个机器生成 kernel:39.5% 在任何容差下已损坏62.1% 至少违反一道契约;领域标准松检验放过的 kernel 里,有 1,487 个被验证器拒收,反向只有 14 个结论以四种独立方式防御(7/7 阳性对照阈值校准扫描与参考基准正确性代码 98.5% 一致分层人工审计)对内附带首个 Blackwell tcgen05 原生 GDN 族训练 backward,对双精度 oracle 独立验证

打开原文回到归档

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

  • ID: e2589713
  • 原文链接: https://arxiv.org/abs/2608.12700
  • PDF: https://arxiv.org/pdf/2608.12700v1
  • 作者: Rishi Shah, Rishav Shrestha
  • 日期: 2026-08-13
  • 更新: N/A
  • 分类: coding
  • 来源类型: paper
  • 标签: llm-generated-code, gpu-kernels, verification, contracts, blackwell, benchmark
  • 质量评分: 5/5
  • 抓取时间: 2026-08-17T23:51:34+08:00

中文导读

给 LLM 生成 GPU kernel 造了一台“契约级验证器”:十二道对抗性契约门,每道都是正确 kernel 必须满足的性质,其中数道零容差——任何阈值选择都无法把失败解释掉。对外审计某公开系统已验收的 2,638 个机器生成 kernel:39.5% 在任何容差下已损坏、62.1% 至少违反一道契约;领域标准松检验放过的 kernel 里,有 1,487 个被验证器拒收,反向只有 14 个。结论以四种独立方式防御(7/7 阳性对照、阈值校准扫描、与参考基准正确性代码 98.5% 一致、分层人工审计)。对内附带首个 Blackwell tcgen05 原生 GDN 族训练 backward,对双精度 oracle 独立验证。

为什么值得关注

LLM 生成 kernel 的照妖镜:已验收 kernel 中 39.5% 任何容差下都错,十二道零容差契约门可直接抄进验收层。

收录理由:本周最硬的测量工作:对“LLM 生成性能代码”报告正确率的系统性证伪,验收门可直接复用

Abstract

Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run, break when the shape changes, or accumulate in fp16 where the reference keeps an fp32 total. We build the instrument that checks correctness properly: a contract-grade verifier of twelve adversarial gates, each a property a correct kernel must satisfy, several of them tolerance-free, so no choice of threshold can explain a failure away. Aimed outward, the verifier audits 2,638 machine-generated kernels that a public system's own harness had already accepted as correct. It finds 39.5% broken beyond any tolerance argument and 62.1% carrying at least one violation. The field's standard test accepts 1,487 kernels the verifier rejects, against only 14 the other way. We defend the finding four independent ways: a 7/7 positive control, a threshold-calibration sweep, 98.5% agreement with the reference benchmark's own correctness code, and a stratified hand-audit. Aimed inward, the verifier judges a kernel of our own: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family, including the reverse-state stage the field still runs on a fallback. We establish its correctness independently, against a double-precision oracle, and train five family members through it. The correctness signal behind reported progress in kernel generation is far weaker than the numbers suggest, and a set of tolerance-free contracts would close most of the gap.

元数据

  • arXiv ID: 2608.12700
  • 主分类: cs.LG
  • 分类: cs.LG, cs.AR, cs.DC
  • 评论: 17 pages, 3 figures. Also archived at doi:10.5281/zenodo.21563213
Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.12700(2026-08-17)。