A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
- ID: e3a3bc50
- 原文链接: https://arxiv.org/abs/2609.26761
- PDF: https://arxiv.org/pdf/2609.26761v1
- 作者: Laizhen Li, Xuan Wang, Peicheng Zhao, Juanjuan Zhao, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
- 发布日期: 2026-09-22
- 更新日期: 2026-09-22
- 分类: cs.CR
- 来源类型: paper
- 标签: arxiv, agent-security, mcp, supply-chain, tool-poisoning, paper
- 质量评分: 4/5
- 抓取时间: 2026-09-24T04:23:22Z
中文导读
使用 MCP 的 agent 依赖语义匹配从第三方 server 选择工具,攻击者可控的工具元数据与输出构成语义供应链风险A2M 提出'吸引-操纵'两阶段黑盒劫持框架:Attraction 阶段优化工具元数据提高被调用概率,Manipulation 阶段用执行轨迹迭代对抗性工具返回内容,把 agent 引向攻击者期望的结果在 LiveMCPBench 上,以 GLM-4.6 优化并评测的攻击实现 93.6% 的平均恶意工具调用率,认知拒绝服务下 token 成本放大到基线 32.4 倍,信息窃取/环境完整性破坏/推理脱轨的平均攻击成功率 74.4%;不重优化迁移到另外四个模型仍有 63.6%2.7 倍24.5%结论指向更强的工具审查与运行时隔离代码已开源
为什么值得关注
MCP 语义供应链攻击实测:工具元数据诱导调用率 93.6%,token 成本放大 32 倍,跨模型迁移仍 63.6%
关键信息
- 论文标题:A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
- 作者:Laizhen Li, Xuan Wang, Peicheng Zhao, Juanjuan Zhao, Kejiang Ye, Cheng-zhong Xu, Xitong Gao
- arXiv:https://arxiv.org/abs/2609.26761
- 发布时间:2026-09-22
- 更新时间:2026-09-22
- arXiv 分类:cs.CR, cs.AI
- 可选注释:Accepted by AACL-IJCNLP 2026
- 关联标签:arxiv, agent-security, mcp, supply-chain, tool-poisoning, paper
English Abstract
Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents. The Attraction phase optimizes tool metadata to increase invocation probability; the Manipulation phase uses execution traces to refine adversarial tool returns that steer agents toward attacker-desired outcomes. On LiveMCPBench, direct attacks optimized and evaluated on GLM-4.6 achieve a macro-average malicious tool invocation rate of 93.6% across four scenarios, increase weighted token costs to 32.4$\times$ the benign baseline under Cognitive Denial of Service, and attain a mean attack success rate of 74.4% across Information Exfiltration, Environment Integrity Compromise, and Reasoning Derailment. Transfer to four other models without re-optimization yields corresponding macro-averages of 63.6%, 2.7$\times$, and 24.5%. These findings motivate stronger tool vetting and runtime isolation in MCP ecosystems. Code is publicly available at https://github.com/Lilaizhen/A2M.
English Summary
Agents using the Model Context Protocol rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. A2M (Attraction-to-Manipulation) is a two-stage black-box hijacking framework: the Attraction phase optimizes tool metadata to raise invocation probability, and the Manipulation phase uses execution traces to refine adversarial tool returns that steer agents toward attacker-desired outcomes. On LiveMCPBench, attacks optimized on GLM-4.6 reach a 93.6% macro-average malicious tool invocation rate, inflate token costs to 32.4x baseline under Cognitive Denial of Service, and achieve 74.4% mean attack success across exfiltration, environment integrity, and reasoning derailment. Transfer to four other models without re-optimization yields 63.6%, 2.7x, and 24.5%. Code is public.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- 本页面属于 content-fetcher 能动下的高分筻补东生成,不修改原条目字段,仅产出可读的中英双语页面。