Agent 与自动化 5.0 · 必读 2026-08-03 · 论文

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

Magnet 针对专门 Agent 组合和跨会话代理带来的检测盲区:攻击者可把有害目标拆成多个单独看似无害的会话论文提出以用户 ID 等更高层相关器聚合能力产物,从大量良性会话中吸附出可组合成有害能力的证据包,补足单轮或单会话威胁模型

打开原文回到归档

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

中文导读

Magnet 针对专门 Agent 组合和跨会话代理带来的检测盲区:攻击者可把有害目标拆成多个单独看似无害的会话论文提出以用户 ID 等更高层相关器聚合能力产物,从大量良性会话中吸附出可组合成有害能力的证据包,补足单轮或单会话威胁模型

为什么值得关注

Magnet detects AI misuse assembled across otherwise benign agent sessions.

Grounded relevance: authors, date, arXiv categories, and abstract claims below; no extra experimental claims beyond the abstract/metadata.

关键信息

  • 论文标题:Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation
  • 作者:Natalie Isak, Matthew Dressman
  • arXiv:https://arxiv.org/abs/2608.02518
  • 发布时间:2026-08-03
  • arXiv 分类:cs.AI, cs.CY
  • 关联标签:agent-security, misuse-detection, multi-agent, cross-session

English Abstract

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a critical gap: an attacker can decompose a harmful goal into innocuous-looking units and execute each in isolated agentic sessions. The agent is stateless between conversations, but the attacker is not. This asymmetry allows for cross-session trajectories that are effective at evading detection. Our contributions are twofold. First, we demonstrate cross-session goal decomposition as an evasion technique, showing it may elicit more harmful capability than equivalent single-session or multi-turn attacks. By capability we mean an artifact produced at one step of an objective, evidenced by what an interaction produced (model responses and tool-call results), and composable with capabilities accrued elsewhere into a harmful whole. Second, we propose Magnet: an efficient and robust detection approach that models relevant capabilities accrued over time and across agentic conversations, aggregated at a higher-level correlator (in this case, a user ID) rather than per-conversation state. The main challenge is assembling the evidence bundle Magnet reasons over. The incriminating artifacts may be needles scattered through a haystack of benign sessions that are individually harmless, dangerous only once collected. Rather than searching the haystack straw-by-straw (i.e. per-session inspection), Magnet does what its name implies: it attracts the relevant needles out of the hay, across sessions and across time, into a compact evidence bundle a detector can act on.

English Summary

The most capable AI deployments are not single models but ensembles of specialized agents that delegate and act in coordination. This architecture unlocks powerful new capabilities, and it also introduces risks that existing frameworks for monitoring, detection, and mitigation were not designed to address. Most state-of-the-art AI abuse detection literature focuses on single-turn or multi-turn (single-session) threat models. This leaves a critical gap: an attacker can decompose a harmful goal into innocuous-looking units and execute each in isolated agentic sessions. The agent is stateless between conversations, but the attacker is not. This asymmetry allows for cross-session trajectories that are effective at evading detection. Our contributions are twofold....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • content-fetcher run; entry id da21903d; arXiv 2608.02518.