基础设施 5.0 · 必读 2026-09-23 · 论文

Xtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing

Xtrace(arXiv 2609.28769)把 probe 直接插进编译后的 GPU kernel 二进制,而不是插在编译前,从而避开现有工具的trace 的是另一个二进制问题方法上复用插入点的 dead 寄存器,按 compiler hazard table 消解所有冒险,再调度指令顺序对 LLM 生产 kernel:保留 94-98% 指令(现有工具 8-48%),开销 0.9-2.8%(现有 3.8-75.6%);用于引导 coding agent 优化 FlashAttention-3 时少 3.9x 迭代,并 trace 闭源 cuDNN kernel 把 FA4 吞吐抬 5.2-13.3%支持 19 种 NVIDIA/AMD 架构,代码公开

打开原文回到归档

Xtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing

原文链接: https://arxiv.org/abs/2609.28769
作者: Zhuobin Huang et al.
发布时间: 2026-09-23
源: arxiv

摘要

Xtrace(arXiv 2609.28769)把 probe 直接插进编译后的 GPU kernel 二进制,而不是插在编译前,从而避开现有工具的「trace 的是另一个二进制」问题。方法上复用插入点的 dead 寄存器,按 compiler hazard table 消解所有冒险,再调度指令顺序。对 LLM 生产 kernel:保留 94-98% 指令(现有工具 8-48%),开销 0.9-2.8%(现有 3.8-75.6%);用于引导 coding agent 优化 FlashAttention-3 时少 3.9x 迭代,并 trace 闭源 cuDNN kernel 把 FA4 吞吐抬 5.2-13.3%。支持 19 种 NVIDIA/AMD 架构,代码公开。

English Summary

Xtrace (arXiv 2609.28769) inserts probes directly into the compiled GPU kernel binary, sidestepping the long-standing 'the trace is of a different binary than the one that ran' problem that afflicts tools that instrument before compilation. The technique reuses only registers that hold dead values at the insertion address, resolves all hazards through the compiler's hazard table, and schedules the instruction order accordingly. On LLM production kernels it retains 94-98% of instructions (existing tools 8-48%) at 0.9-2.8% overhead (existing tools 3.8-75.6%); a coding agent using Xtrace to optimize FlashAttention-3 needed 3.9x fewer iterations, and tracing closed-source cuDNN kernels lifted FA4 throughput 5.2-13.3%. It supports 19 NVIDIA/AMD architectures and ships with public code.

为什么值得关注

主题线扩展 AAIF gpu/profiling/intra-kernel/binary-instrumentation 等主题。

信息源