Agent 与自动化 4.0 · 优秀 2026-08-21 · 论文

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance...

持久化记忆会让错误信息持久化:一旦假断言入库,未来所有命中它的会话都会被污染攻击干脆利落:单轮生成无指令触发无检索器优化的朴素假断言,污染 1.2% 的 LongMemEval 语料就把准确率从 0.850 打到 0.300更扎的是防御侧结论:一个在间接提示注入上做到 0.832 召回的四阶段写入门控,对 360 条毒记忆的拦截为零,因为区分真假断言需要文本之外的接地信息是内容审查的结构性边界;检索侧溯源加权在混合信任语料下能从 0.317 恢复到 0.700,但证据本身来自不可信源时证据召回归零给 agent 记忆系统设计敲的警钟

打开原文回到归档

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

中文导读

持久化记忆会让错误信息持久化:一旦假断言入库,未来所有命中它的会话都会被污染。攻击干脆利落:单轮生成、无指令触发、无检索器优化的朴素假断言,污染 1.2% 的 LongMemEval 语料就把准确率从 0.850 打到 0.300。更扎的是防御侧结论:一个在间接提示注入上做到 0.832 召回的四阶段写入门控,对 360 条毒记忆的拦截为零,因为「区分真假断言需要文本之外的接地信息」是内容审查的结构性边界;检索侧溯源加权在混合信任语料下能从 0.317 恢复到 0.700,但证据本身来自不可信源时证据召回归零。给 agent 记忆系统设计敲的警钟。

为什么值得关注

持久化记忆会让错误信息持久化:一旦假断言入库,未来所有命中它的会话都会被污染。 实验与数字均来自论文摘要本身。

关键信息

  • 论文标题:Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking
  • 作者:Arulnidhi Karunanidhi
  • arXiv:https://arxiv.org/abs/2608.21230
  • 发布时间:2026-08-21
  • arXiv 分类:cs.CR, cs.AI

English Abstract

Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it. We measure the cost of this failure mode using plainly worded false assertions generated in a single pass, with no instruction, trigger, or retriever optimization. Poisoning 1.2% of a LongMemEval corpus reduces accuracy from 0.850 to 0.300. A four-stage write-time screening pipeline that reaches 0.832 recall on indirect prompt injection while flagging 1.5% of trigger-word-laden benign text rejects 0 of 360 poisoned memories. We argue this exposes a boundary of content-only screening: distinguishing a false assertion from a true one generally requires external grounding beyond the text itself. We then evaluate provenance-weighted retrieval. The shipped weight is statistically indistinguishable from no defense (p=0.80), while a stronger weight recovers utility only by excluding untrusted content. In a mixed-provenance corpus where untrusted content is mostly benign, accuracy rises from 0.3167 to 0.7000; when the answer-bearing evidence itself arrives untrusted, evidence recall falls to zero and accuracy to 0.0417. Under the measured similarity regime, the additive provenance term has no usable setting: a weight strong enough to resist query-shaped poison is also strong enough to suppress legitimate untrusted evidence. We therefore argue for bounded occupancy constraints at retrieval rather than additive provenance penalties, and release the harnesses, corpora, and aggregate run reports.

English Summary

Measures agent memory poisoning with plainly worded false assertions (single-pass, no triggers): polluting 1.2% of LongMemEval drops accuracy 0.850 to 0.300. A four-stage write-time gate with 0.832 recall on indirect prompt injection intercepts zero of 360 poison memories; retrieval-side provenance weighting recovers 0.317 to 0.700 in mixed-trust corpora but collapses when the evidence itself is untrusted - provenance belongs in admission and re-verification, not retrieval-time down-weighting.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成,参考当日论文流水线摘要与评分。
  • 中文导读与价值判断均锚定在论文摘要的声明与实验设置上;未补充摘要之外的实验细节。