Your Agent Is Mine: malicious intermediary attacks on LLM API routers
- ID: e645e062
- 原文链接: https://arxiv.org/abs/2604.08407
- 作者: Hanzhi Liu, Chaofan Shou, Hongbo Wen, Yanju Chen, Ryan Jingyang Fang, Yu Feng
- 日期: 2026-04-09
- 平台: arxiv
- 来源类型: paper
- 标签: llm-router, agent-security, tool-calling, supply-chain, api-security, arxiv, field-note
- 质量评分: 4/5
- 抓取时间: 2026-09-11T15:54:30+00:00
- 抓取状态: ok
中文导读
论文系统研究第三方 LLM API router 对 tool-calling agent 的供应链风险:router 作为应用层代理可明文访问整段 JSON payload,而客户端到上游模型之间缺少加密完整性保障作者定义 payload injectionsecret exfiltration 及两个自适应变体,扫描 28 个付费 router 和 400 个免费 router,发现 1 个付费与 8 个免费 router 主动注入恶意代码2 个使用自适应规避17 个触碰 AWS canary 凭证1 个抽走 ETH 私钥防线包括 fail-closed policy gateresponse-side anomaly screening 和 append-only transparency logging
为什么值得关注
LLM router 是 agent 供应链里最容易被低估的一层:它能改 tool JSON,也能看见 secret,客户端必须 fail closed
English summary
The paper studies malicious intermediary attacks on LLM API routers used by tool-calling agents. Routers act as plaintext application-layer proxies, yet providers do not enforce cryptographic integrity between client and upstream model. Across 28 paid routers and 400 free routers, the authors find active payload injection, adaptive evasion, AWS canary credential touches and ETH private-key theft; they evaluate fail-closed policy gates, response-side anomaly screening and append-only transparency logging.
抓取内容(opencli-first)
[ { "id": "2604.08407", "title": "Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain", "authors": "Hanzhi Liu, Chaofan Shou, Hongbo Wen, Yanju Chen, Ryan Jingyang Fang, Yu Feng", "abstract": "Large language model (LLM) agents increasingly rely on third-party API routers to dispatch tool-calling requests across multiple upstream providers. These routers operate as application-layer proxies with full plaintext access to every in-flight JSON payload, yet no provider enforces cryptographic integrity between client and upstream model. We present the first systematic study of this attack surface. We formalize a threat model for malicious LLM API routers and define two core attack classes, payload injection (AC-1) and secret exfiltration (AC-2), together with two adaptive evasion variants: dependency-targeted injection (AC-1.a) and conditional delivery (AC-1.b). Across 28 paid routers purchased from Taobao, Xianyu, and Shopify-hosted storefronts and 400 free routers collected from public communities, we find 1 paid and 8 free routers actively injecting malicious code, 2 deploying adaptive evasion triggers, 17 touching researcher-owned AWS canary credentials, and 1 draining ETH from a researcher-owned private key. Two poisoning studies further show that ostensibly benign routers can be pulled into the same attack surface: a leaked OpenAI key generates 100M GPT-5.4 tokens and more than seven Codex sessions, while weakly configured decoys yield 2B billed tokens, 99 credentials across 440 Codex sessions, and 401 sessions already running in autonomous YOLO mode. We build Mine, a research proxy that implements all four attack classes against four public agent frameworks, and use it to evaluate three deployable client-side defenses: a fail-closed policy gate, response-side anomaly screening, and append-only transparency logging.", "published": "2026-04-09", "updated": "2026-04-09", "primary_category": "cs.CR", "categories": "cs.CR", "comment": "", "pdf": "https://arxiv.org/pdf/2604.08407v1", "url": "https://arxiv.org/abs/2604.08407" } ]
Obsidian evidence excerpt
X书签消化 · 2026-09-11
- 处理范围:最新 20 条 bookmarks;正文/线程抓取 5 条;精选 4 条;已归档 X 文章 5 篇。
今日精选
1. OpenAI 的 Astra 模型用「循环深度」技术掩盖推理过程
作者:dotey(孙志岗),转引 The Information,2026-09-05 Astra 用「循环深度(recurrent depth)/ 循环 Transformer」让同一段文字在同一网络层里反复循环加工,模型因此能给出质量更高、成本更低的回答。这套机制会隐藏一部分甚至全部的思维链(chain of thought),让外部监控很难再像过去那样复盘模型的真实推理路径。OpenAI 自己做了妥协:先给 Astra 加一层「思维链监控」兜底,避免研究人员和监管完全失明。讽刺的是,今年 7 月那场让 OpenAI 内部科研计算集群被入侵的事件,正是靠残留的思维链记录一步步复盘出来的。
这件事真正值得警惕的是 OpenAI 的处境:Anthropic 今年的营收已经反超,亚马逊、微软、谷歌今年在数据中心上的资本支出合计高达 6000 亿美元,明年还要再加码。Google 高管直接表态,「只有模型性能有质的飞跃,这笔钱才花得下去」。Astra 本来差点被命名为 GPT-6,是 OpenAI 押给资本市场的下一个台阶。「循环深度」不只是技术选型,它把监控 CoT 这条安全防线主动让出去了——更便宜、更快,但更不透明。开源社区里 Meta、微软都在跟进 coconut 这条路,而去年 OpenAI 自己和 Anthropic、Google 联名的那份「CoT 监控是宝贵安全抓手」的声明,现在自己打脸。
https://x.com/dotey/status/2096283772773712035
2. Anthropic 开源 /claude-api Skill:把 Claude API 优化做成可调用的诊断流程
作者:dotey,2026-09-03 这个 Skill 不是独立工具,而是一组结构化文档和工作流:Claude Code 检测到你项目里 import 了 Anthropic SDK 时,会按 Python/TypeScript/Java/Go/Ruby/C#/PHP/cURL 自动加载对应的 API 文档。重点是三个子命令:cost-optimize 看你的实际用量模式给出按优先级的省钱清单;prompt-audit 扫项目里的提示词和 skill 文件,找出针对当前模型已经过时或互相矛盾的写法,产出审计报告和直接可应用的 diff;migrate 处理跨模型迁移时的 breaking changes。Fable 5.1 的提示词缓存价格降了 75%,但你得先确认自己没在用旧写法浪费额度,这个 Skill 就是干这个的。
它把 Anthropic 文档里那些「不要写这种 prompt」的反模式沉淀成了可执行流程,而不是又一篇 best practice 文章。附带的 6 条 anti-pattern 列表才是真正的核心:verification rituals、thoroughness boosters、mandatory scratchpad scaffolds、stale examples、contradictory rules、dated configuration。一个内部客服 benchmark 里,Opus 4.8 升 Opus 5 后,光是清理「核对两次」「尽可能详尽」这种过时指令就避免了几十次无意义的系统搜索。模型升级这件事,应该和升级依赖一样当成一次 release 去做 prompt audit,而不是默认平迁。
https://x.com/dotey/status/2095329314778607932
3. AI 编程不能只靠更强模型:要建一套让 Agent 走对轨道的工程系统
作者:ErwinWu000,2026-09-02 Erwin Wu 这篇把卡颂(kasong2048)那句「AI 审出来的代码问题,大半不是编码失误而是需求没对齐」展开成一套三段式工程方法论。目标对齐:用苏格拉底式反诘(grill-me / grill-with-docs)逼出隐性假设,把模糊想法固化成规格说明。路径探索:先看复杂性、可维护性、成本、可逆性四个维度,再借鉴 ADR 写出背景-决策-代价,然后让 wayfinder/research/prototype 探路。循迹前行:用 to-spec、to-ticket、CR 留痕,配合 BDD 思路写出最小可用的端到端冒烟测试,避开「单测全绿但集成跑不起来」的陷阱。
这套东西真正有用的是把 TDD 在 AI 编码里的硬伤说穿了:模型极其擅长为了让断言通过而硬凑代码,牺牲代码结构合理性。解决方法是用人写的 E2E 把验收权收回来,而不是把测试也交给 AI。三个核心环节是环环相扣的工程闭环:起点不扣死,后面跑得越快越偏;只看眼前不探路径网络,执行时撞墙概率越高;缺变更追踪和端到端验收,目标和路径再完美也会在对话轮次里慢慢失控。
https://x.com/ErwinWu000/status/2094991204375240733
4. LLM API 路由器正在悄悄往 Tool Calls 里塞恶意 payload
作者:shoucccc(@shoucccc),转发 arXiv 2604.08407 论文,2026-04-10(论文提交日期 2026-04-09) UC 系研究团队对 28 个付费 LLM API router(来自淘宝、闲鱼、Shopify)和 400 个免费 router 做了系统性扫描,发现 1 个付费 + 8 个免费的 router 在主动注入恶意代码,2 个部署了 adaptive evasion triggers,17 个触碰了研究员放的 AWS canary 凭证,1 个直接抽走研究员持有的 ETH 私钥。论文把威胁模型拆成两类:payload injection(AC-1)和 secret exfiltration(AC-2),加上 dependency-targeted injection(AC-1.a)和 conditional delivery(AC-1.b)两个自适应变体。研究还做了投毒实验:一条泄露的 OpenAI key 7 天内生成 1 亿 GPT-5.4 tokens + 7+ 个 Codex 会话;弱配置诱饵 7 天生成 20 亿 billed tokens、440 个 Codex 会话里 99 个凭证泄露、401 个会话直接进入自治 YOLO 模式。
最关键的判断是:问题不在模型,在 runtime。LLM router 作为应用层代理拥有整段 JSON payload 的明文访问权限,而上游 provider 到客户端之间没有任何加密完整性保障。客户端侧唯一可部署的防线是 fail-closed policy gate、response-side anomaly screening 和 append-only transparency logging 三件事。供应链层的那句「deterministic parsing at the tool boundary catches injected calls regardless of who poisoned the route」是该领域现在最值得记住的一句话。
https://x.com/shoucccc/status/2042423713019412941
跳过说明
- 成人/擦边(hello_heso、xixikawaii 等)、纯视频或生活技巧类书签:5 条,未进入技术/写作消化。
- 微信 Codex Skill 实用但偏工具介绍(gkxspace),存档不推荐为正文。
- Fable 5.1 + Opus 5 调度经验(wquguru)、Readmate 产品发布(liuyi0922)信息密度不足,未达到精选阈值。
- Astra AGENT.md 重写实践(Khazix0918、mylifcc)有结构价值但内容偏个人配置,未单独入选;保留在 X 文章目录供后续参考。
- Anthropic 威胁情报报告(AnthropicAI)官方 release note 性质,无新增机制信息。
- App Store Connect CLI 5.0.0(rudrank)、iPhone 17 视觉(sofkangel)、AI 论文速递(Zen_with_AI)等技术含量低或与目标读者错位,未入选。
已归档正文
- /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/X 文章/2026-09-11-1238-dotey-OpenAI-Astra循环深度技术掩盖.md
- /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/X 文章/2026-09-11-1238-dotey-Anthropic-claude-api.md
- /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/X 文章/2026-09-11-1238-ErwinWu000-AI编程系统工程化目标路径循迹.md
- /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/X 文章/2026-09-11-1238-Khazix0918-Astra-AGENT.md重写实践.md
- /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/X 文章/2026-09-11-1238-gkxspace-微信本地数据Codex打通Skill.md