AI 编程 4.0 · 优秀 2026-08-21 · 论文

HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization

LLM 写 GPU kernel 常见做法是同一思路在 Triton/CUDA/汇编三种实现空间里各自碰运气从不互通HIERA 把性能瓶颈特征(memory-bound / compute-bound 等负载画像)建模为与实现空间无关的中间规划层:先在一个空间里学到瓶颈归因,再把这套归因迁移到其他空间指导搜索在覆盖三空间的内核基准上超过前沿模型自带优化 agent 与多种自适应采样基线,试错预算更少跨实现空间的性能归因能迁移,是对LLM 只会背模板质疑的机制化回应

打开原文回到归档

HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization

  • ID: 4794868e
  • 原文链接: https://arxiv.org/abs/2608.21157
  • PDF: https://arxiv.org/pdf/2608.21157v1
  • 作者: Jinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li
  • 日期: 2026-08-21
  • 更新: 2026-08-21
  • 分类: coding
  • 来源类型: paper
  • 标签: llm-coding, gpu-kernel, performance-engineering, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-25T15:40:00+00:00

中文导读

LLM 写 GPU kernel 常见做法是同一思路在 Triton/CUDA/汇编三种实现空间里各自碰运气、从不互通。HIERA 把性能瓶颈特征(memory-bound / compute-bound 等负载画像)建模为与实现空间无关的中间规划层:先在一个空间里学到瓶颈归因,再把这套归因迁移到其他空间指导搜索。在覆盖三空间的内核基准上超过前沿模型自带优化 agent 与多种自适应采样基线,试错预算更少。跨实现空间的性能归因能迁移,是对「LLM 只会背模板」质疑的机制化回应。

为什么值得关注

LLM 写 GPU kernel 常见做法是同一思路在 Triton/CUDA/汇编三种实现空间里各自碰运气、从不互通。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
  • 作者:Jinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li
  • arXiv:https://arxiv.org/abs/2608.21157
  • 发布时间:2026-08-21
  • arXiv 分类:cs.DC, cs.AI

English Abstract

High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient methods for automated GPU kernel generation and optimization has become increasingly important. Existing LLM-based methods typically optimize within a fixed implementation space, limiting either optimization flexibility or search efficiency. We propose \textsc{HIERA}, a hierarchical search-space planning framework for GPU kernel optimization. \textsc{HIERA} constructs contract-augmented task specifications, selects an appropriate implementation space across PyTorch operators, CUDA libraries, and custom CUDA kernels, and uses profiling feedback and expert knowledge to guide structured iterative refinement. Experiments on KernelBench across multiple various workload levels and base LLMs show that \textsc{HIERA} delivers stronger overall implementation validity, sample efficiency, and optimization performance than existing training-free methods, while remaining competitive with the training-based CUDA-L1 without additional model training. A case study on a specialized stencil operator from scientific computing further achieves a \(1.53\times\) speedup over cuDNN, demonstrating the potentiality of the general framework beyond standard machine-learning workloads.

English Summary

HIERA models workload bottleneck features as an implementation-space-agnostic intermediate planning layer for LLM GPU kernel optimization, transferring bottleneck attribution learned in one space (e.g. Triton) to guide search in others (CUDA, PTX), beating frontier models' built-in optimization agents and adaptive sampling baselines across a three-space kernel benchmark with fewer trial budgets.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。