Agent 与自动化 4.0 · 优秀 2026-09-03 · 论文

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

SOC 中的 LLM 智能体架构:把拓扑推理从语义推理中解耦企业规模下两个硬限制有限上下文装不下数千主机的认证图,自由文本生成无法保证遏制动作与拓扑一致Sentinel-RL 用异构图注意力编码器把实时认证子图压缩为定长状态,PPO 策略把状态映射到受限调查动作集,LLM 循环只负责消费策略建议并在 critic 门控下产出分析师可读叙述;在 LANL 数据集上实例化

打开原文回到归档

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

  • ID: c59a5c9b
  • 原文链接: https://arxiv.org/abs/2609.04159
  • PDF: https://arxiv.org/pdf/2609.04159v1
  • 作者: Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild
  • 日期: 2026-09-03
  • 更新: 2026-09-03
  • 分类: cs.CR, cs.AI
  • 来源类型: arxiv
  • 标签: soc, security-operations, graph-attention, ppo, hybrid-architecture, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-09-06T04:24:00Z

中文导读

LLM 智能体被越来越多地提议为自主 SOC 分析师,但企业规模下有两个硬限制:有限上下文窗口装不下数千主机的认证图,自由文本生成无法保证推荐的遏制动作与拓扑一致。Sentinel-RL 把拓扑推理从语义推理中解耦:异构图注意力编码器把实时认证子图压缩为定长状态,PPO 策略把状态映射到受限调查动作集,LLM 智能体循环只消费策略建议,并在 critic 门控下产出分析师可读叙述。系统在 LANL 综合、多源网络安全事件数据集与 Indiana University Quartz HPC 集群上实例化,四项结果:两阶段 CREATE 摄取模式在单台 32 核节点上 14.2 分钟装入 2400 万边认证子图,约为 MERGE 管线的 24 倍快;滑动窗口告警引擎 50 次试验均在 2.5 秒内触发 25 事件/10 秒阈值;PPO 200 轮训练收敛到平均回合回报 8.74±0.31,红队标注事件上 held-out 精确率 0.91、召回 0.87;完整检测-调查-建议-人工批准循环中位耗时 6.3 秒。论文还贡献了热节点死锁绕过、锚节点共置 HPC 部署模式,以及覆盖误报经济学、可逆性保证、审计合规与人工批准边界的企业级就绪分析。

为什么值得关注

别让 LLM 硬扛认证图:图注意力编码拓扑PPO 出受限动作LLM 只写有 critic 把关的叙述SOC 智能体的分工范式

English Abstract

Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates on. We present Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning: a heterogeneous graph attention encoder summarizes the live authentication subgraph into a fixed-dimensional state, a Proximal Policy Optimization (PPO) policy maps this state to a constrained set of investigative actions, and an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives gated by a critic.

Obsidian 证据

  • 元数据与摘要经 opencli arxiv paper 2609.04159 核对(2026-09-06T04:24:00Z);中文导读锚定摘要陈述的事实与数字。