Agent 与自动化 4.0 · 优秀 2026-10-01 · 论文

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free V...

KaliBench(NeurIPS 2026 评测与数据集 track 接收)衡量 LLM 把自然语言意图翻译成 Kali Linux 真实 CLI 命令的能力:8,504 条查询-命令对1,642 个工具23 个能力维度5 个安全阶段,经由手册锚定的构建管线加确定性规范化与别名感知评测安全运营依赖严格 CLI,小的语法错误flag-value 绑定错误或参数次序错误都会让执行失效,而这正是既有知识问答型评测测不到的部分验证管线组合 LLM 校验沙箱终端执行与人工修正;基于这些细粒度确定性信号可以构造免运行时的可验证奖励用于训练

打开原文回到归档

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

  • ID: 27ff405c
  • 原文链接: https://arxiv.org/abs/2610.02206
  • PDF: https://arxiv.org/pdf/2610.02206
  • 作者: Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
  • 日期: 2026-10-01
  • 更新: 2026-10-01
  • 分类: agents
  • 来源类型: paper
  • 标签: cybersecurity, tool-use, benchmark, cli, verifiable-rewards, neurips-2026
  • 质量评分: 4/5
  • Fetch: 2026-10-03T04:21:07Z

中文导读

KaliBench(NeurIPS 2026 评测与数据集 track 接收)衡量 LLM 把自然语言意图翻译成 Kali Linux 真实 CLI 命令的能力:8,504 条查询-命令对1,642 个工具23 个能力维度5 个安全阶段,经由手册锚定的构建管线加确定性规范化与别名感知评测安全运营依赖严格 CLI,小的语法错误flag-value 绑定错误或参数次序错误都会让执行失效,而这正是既有知识问答型评测测不到的部分验证管线组合 LLM 校验沙箱终端执行与人工修正;基于这些细粒度确定性信号可以构造免运行时的可验证奖励用于训练

为什么值得关注

KaliBench 用 8,504 条真实 CLI 翻译对测安全 agent 的手:语法flag 绑定参数次序,错了就执行失败

以上导读与价值判断锚定论文摘要与元数据,完整英文摘要见下文。

关键信息

  • 论文标题:KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
  • 作者:Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
  • arXiv:https://arxiv.org/abs/2610.02206
  • 发布时间:2026-10-01
  • arXiv 分类:cs.CL, cs.AI, cs.CR
  • 关联标签:cybersecurity, tool-use, benchmark, cli, verifiable-rewards, neurips-2026

English Abstract

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

English Summary

KaliBench (accepted at NeurIPS 2026 Evaluations and Datasets Track) measures LLMs' natural-language-to-CLI translation on Kali Linux: 8,504 query-command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases, built via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation. Security operations rely on strict CLIs where minor syntax errors, incorrect flag-value bindings, or argument misordering invalidate execution - a gap knowledge-based evaluations miss. Verification combines LLM validation, sandboxed terminal execution, and human-in-the-loop refinement; the fine-grained deterministic signals further enable runtime-free verifiable rewards for training.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。