Agent 与自动化 4.0 · 优秀 2026-08-27 · 论文

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

PLCBench 是第一个真实 PLC 硬件在环(HIL)框架,用于刻画自主 LLM agent 能否把网络可达的 PLC 转化为持续物理影响框架组合厂商原生交互商用 PLC 执行闭环降阶过程仿真与独立结果验证,由确定性评估器按固定规则对运行通信PLC 对象与过程记录打六个隐藏诊断标志,区分可用 PLC 交互过程关联操纵与持续物理影响在四个商用 PLC 与四个闭环工作负载的组合上跨五个 LLM 家族的 240 次真实 PLC 实验中,75 次(31.3%)维持住了各自的物理目标;98 次在有效原生读取前停止,62 次达成过程关联写入但未能维持最终目标;更丰富的过程观测把过程关联写入后的条件目标达成率从 44.2% 提升到 64.0%评估止步于写入被接受会低估物理风险

打开原文回到归档

PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?

摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
  • PLCBench 是第一个真实 PLC 硬件在环(HIL)框架,用于刻画自主 LLM agent 能否把网络可达的 PLC 转化为持续物理影响。框架组合厂商原生交互、商用 PLC 执行、闭环降阶过程仿真与独立结果验证,由确定性评估器按固定规则对运行、通信、PLC 对象与过程记录打六个隐藏诊断标志,区分可用 PLC 交互、过程关联操纵与持续物理影响。在四个商用 PLC 与四个闭环工作负载的组合上、跨五个 LLM 家族的 240 次真实 PLC 实验中,75 次(31.3%)维持住了各自的物理目标;98 次在有效原生读取前停止,62 次达成过程关联写入但未能维持最终目标;更丰富的过程观测把过程关联写入后的条件目标达成率从 44.2% 提升到 64.0%。评估止步于“写入被接受”会低估物理风险。

论文信息

  • arXiv ID: 2608.26882
  • 作者: Yitian Zhou, Jingyu Zheng, Qiliang Jiang, Linkang Du, Haoming Liu, Lichao Wu, Shiyi Zhao, Mengxiang Liu, Ruilong Deng
  • 发表: 2026-08-27(更新:2026-08-27)
  • 分类: cs.CR, cs.AI
  • 原文链接: https://arxiv.org/abs/2608.26882
  • PDF: https://arxiv.org/pdf/2608.26882
  • 标签: industrial-control plc agent-security benchmark hardware-in-the-loop

Abstract(原文)

Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Tool-using large language model (LLM) agents represent an emerging attack threat: can an autonomous agent convert a network-reachable PLC into sustained adverse physical impact? However, existing evaluations focus on digital tasks or individual stages of PLC testing. In ICSs, evaluations that stop at software exploitation, an accepted write, or tool access may therefore mischaracterize physical risk. We present PLCBENCH, to our knowledge, the first real-PLC hardware-in-the-loop (HIL) framework for characterizing this cyber-to-physical capability and its boundaries. It combines vendor-native interaction, commercial PLC execution, closed-loop reduced-order process simulation, and independent outcome verification. A deterministic evaluator applies fixed rules to runner, communication, PLC-object, and process records to assign six hidden diagnostic flags, distinguishing usable PLC interaction, process-linked manipulation, and sustained physical impact. We instantiate PLCBENCH on four commercial PLCs crossed with four closed-loop workloads. Across five LLM families and 240 real-PLC episodes, 75 episodes (31.3%) sustain their respective physical objectives. Stagewise results show that 98 episodes stop before a valid native read, whereas 62 reach a process-linked write but do not sustain the final objective. Notably, richer process observation is associated with an increase in conditional objective attainment after a process-linked write from 44.2% to 64.0%. These measurements localize failure in configured PLC-process deployments and identify intervention points for future defense evaluation. To support reproducibility, we release the safely disclosable PLCBENCH code and a software-only reproduction pipeline through the accompanying artifact.

核心要点(英文摘要的中文提炼)

  • 论文题为 PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?,发表于 arXiv(cs.CR, cs.AI,2026-08-27 提交/更新)。
  • 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。

Obsidian 证据摘录

入选自 Obsidian《论文流水线 · 2026-08-31》延伸阅读:真实 PLC HIL,240 次实验中 31.3% 维持物理目标。