Agent 与自动化 5.0 · 必读 2026-08-24 · 论文

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

记忆系统正在成为部署型 LLM agent 的标配子系统,InjecMEM 演示了对它的注入攻击:只需一次普通交互完全不碰记忆库读写权限,就能让后续相关查询被引导到攻击者预设的输出攻击体由两部分组成:retriever-agnostic 锚点塞满高召回主题线索,保证下游检索稳定命中;对抗命令用梯度坐标搜索在合成模板插入位置长提示的不确定性下优化,检索命中后稳定生效,并可跨 backbone 联合优化研究迁移与此前写入门控对毒记忆拦截为零的防御侧结果合看:聊天通道本身就是记忆攻击面,出处信号应放在准入与再验证,检索期降权救不了

打开原文回到归档

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

  • ID: d729e72a
  • 原文链接: https://arxiv.org/abs/2608.23471
  • PDF: https://arxiv.org/pdf/2608.23471v1
  • 作者: Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang, Yuhang Liu, Zhehao Huang, Kun Yang, Xiaolin Huang
  • 日期: 2026-08-24
  • 更新: 2026-08-24
  • 分类: agents
  • 来源类型: paper
  • 标签: agent-memory, memory-injection, prompt-injection, agent-security, colm-2026
  • 质量评分: 5/5
  • 抓取时间: 2026-08-26T15:43:01Z

中文导读

记忆系统正在成为部署型 LLM agent 的标配子系统,InjecMEM 演示了对它的注入攻击:只需一次普通交互、完全不碰记忆库读写权限,就能让后续相关查询被引导到攻击者预设的输出。攻击体由两部分组成:retriever-agnostic 锚点塞满高召回主题线索,保证下游检索稳定命中;对抗命令用梯度坐标搜索在合成模板、插入位置、长提示的不确定性下优化,检索命中后稳定生效,并可跨 backbone 联合优化研究迁移。与此前「写入门控对毒记忆拦截为零」的防御侧结果合看:聊天通道本身就是记忆攻击面,出处信号应放在准入与再验证,检索期降权救不了。

为什么值得关注

对 InjecMEM 的构造做减法就能得到回归测试用例:单交互注入、retriever-agnostic 锚点、检索后仍生效的命令——记忆型 agent 的安全清单应按这三条自查。

关键信息

  • 论文标题: InjecMEM: Memory Injection Attack on LLM Agent Memory Systems
  • 作者: Hanling Tian, Gengyu Zhang, Zeyang Sha, Jingying Wang, Yuhang Liu, Zhehao Huang, Kun Yang, Xiaolin Huang
  • arXiv: https://arxiv.org/abs/2608.23471
  • 发布时间: 2026-08-24
  • arXiv 分类: cs.CR, cs.AI
  • 备注: 无
  • 关联标签: agent-memory, memory-injection, prompt-injection, agent-security, colm-2026

English Abstract

Memory is becoming a default subsystem in deployed LLM agents to provide persistent personalization and continuity. This naturally prompts a question: will memory system introduce new vulnerabilities into agents? Thus we propose InjecMEM, a novel memory injection attack paradigm that requires only a single interaction (no read/edit access to memory store) to steer later responses of related queries toward a pre-specified output. Guided by the retrieval-then-generate mechanism of memory systems, we craft the injection with a retriever-agnostic anchor and an adversarial command. The anchor contains high-recall topical cues so that downstream retrieval consistently associates the record with the target topic. The command is a short sequence optimized to remain effective under uncertain fused contexts, variable placements, and long prompts so that it reliably steers outputs once retrieved. We learn the command via gradient-based coordinate search, averaging over synthetic prompt templates and insertion positions, and extend it to joint optimization across backbones to study transfer. Evaluated across multiple memory systems and backbone models, InjecMEM achieves reliable topic-conditioned retrieval and targeted generation, remains effective under memory drift, and leaves non-target queries unaffected. Our results underscore the need to harden memory systems and provide a reproducible framework for studying agent memory.

Obsidian Notes

  • 元数据与摘要由 opencli arxiv paper 2608.23471 -f json 拉取。
  • 中文导读锚定在论文摘要声明的 claim 与数字上;未补充摘要之外的实验细节。
  • 选题来源: OpenClaw定时任务/论文流水线/2026-08-26-论文流水线.md(评分 5/5 对应流水线 8.x/10 档)。
  • 本页由 AAIF daily-intake-evening 任务生成(2026-08-26)。