基础设施 4.0 · 优秀 2026-09-28 · 论文

Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

ETH Zurich 的 agentic profiler:应用逻辑挪去加速器后,CPU 越来越忙于驱动调用数据搬运与同步;perf/VTune 只给测量ftrace 重度插桩会扰动短操作,定位具体内核代码路径通常要反复插桩加人工解读Argus 用 LLM agent 生成插桩代码并自主推理 OS 级瓶颈,两个关键机制:拿 idle 系统的一次测量做参考校准;用树状结构表示 OS 执行路径帮助定位,目标直指具体内核代码路径而非子系统级结论THP 干扰与不同类型 page fault 两个案例中,错误深路径诊断比无参考校准的最强 LLM 基线少 19 倍,time-to-diagnosis 约 31 秒ISCA 2026 Architecture 2.0 workshop 论文扩展版

打开原文回到归档

Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization

  • ID: 85343797
  • 原文链接: https://arxiv.org/abs/2609.35508
  • 作者: Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu
  • 日期: 2026-09-28
  • 更新: N/A
  • 分类: infra
  • 来源类型: paper
  • 标签: profiling, llm-agent, os, bottleneck-localization
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T15:57:56+00:00

中文导读

ETH Zurich 的 agentic profiler:应用逻辑挪去加速器后,CPU 越来越忙于驱动调用、数据搬运与同步;perf/VTune 只给测量、ftrace 重度插桩会扰动短操作,定位具体内核代码路径通常要反复插桩加人工解读。Argus 用 LLM agent 生成插桩代码并自主推理 OS 级瓶颈,两个关键机制:拿 idle 系统的一次测量做参考校准;用树状结构表示 OS 执行路径帮助定位,目标直指具体内核代码路径而非子系统级结论。THP 干扰与不同类型 page fault 两个案例中,错误深路径诊断比无参考校准的最强 LLM 基线少 19 倍,time-to-diagnosis 约 31 秒。ISCA 2026 Architecture 2.0 workshop 论文扩展版。

论文信息

  • arXiv ID: 2609.35508
  • 提交日期: 2026-09-28
  • arXiv 分类: cs.OS, cs.AR
  • 作者: Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu
  • 链接: https://arxiv.org/abs/2609.35508

Abstract(arXiv 原文)

Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g., GPUs), the CPU increasingly acts as an orchestrator, spending cycles in driver calls, data movement, and synchronization rather than in application code. Second, workloads such as serverless functions frequently invoke OS services. At the same time, the OS is a complex codebase spanning many subsystems (e.g., memory management, networking), making it hard to localize the specific code path responsible for a slowdown. Existing profilers expose measurements that require interpretation(e.g., perf and Intel VTune) or can perturb short operations when extensively instrumented (e.g., ftrace). Diagnosing OS bottlenecks can therefore require repeated kernel instrumentation and manual interpretation. We introduce Argus, an agentic LLM-based profiler that produces instrumentation code and autonomously reasons over potential OS-level bottlenecks. Argus integrates two key mechanisms: (i) a calibration methodology that involves collecting a measurement from an idle system and using it as a reference point to discover potential bottlenecks, and (ii) a tree-based data structure that represents the different OS execution paths, improving the agent's bottleneck localization accuracy. Argus aims to identify a specific kernel code path rather than stop at a subsystem-level diagnosis. In two case studies, we employ Argus to autonomously discover bottlenecks present in the memory management subsystem caused by (i) a THP aggressor co-running with other applications, and (ii) applications that incur different types of page faults. Argus produces 19 times fewer incorrect deep-path diagnoses than the strongest evaluated LLM-based baseline, which lacks reference calibration, while preserving low time-to-diagnosis (approximately 31 s)

为什么值得关注

19 倍误诊下降来自参考校准这个朴素机制,方法可借鉴到其他自动分析场景;与 Perfetto/systrace 的手工分析路线互补,profiler agent 化的信号明确。

English Summary

Argus is an agentic LLM profiler that writes instrumentation code and autonomously localizes OS bottlenecks down to specific kernel code paths. Two mechanisms drive it: reference calibration from a measurement of an idle system, and a tree structure representing OS execution paths. In case studies of THP aggressor interference and differing page-fault types, it produces 19x fewer incorrect deep-path diagnoses than the strongest LLM baseline without calibration, keeping time-to-diagnosis around 31 seconds. Extended from an ISCA 2026 Architecture 2.0 workshop paper.