模型与实验室 5.0 · 必读 2026-08-10 · 论文

Stealing Reasoning Traces from Proprietary LLM APIs

多家 frontier 提供商把 step-by-step reasoning 以加密文本块返回客户端,下一轮请求再带回,而不是纯服务端保存;这些加密块在同一提供商生态内可跨 session / 用户 / 模型移植利用兼容性可做可扩展的 decryption jailbreak:用较弱模型 + jailbreak,把较强模型的加密 thoughts 转成明文论文还指出 reasoning token 计数与 API 计费 thinking tokens 在多数 prompt 上 1:1...

打开原文回到归档

Stealing Reasoning Traces from Proprietary LLM APIs

  • ID: b49aad65
  • Original: https://arxiv.org/abs/2608.09867
  • Added: 2026-08-12
  • Source: arxiv / kotekjedi et al.
  • Original Date: 2026-08-10
  • AAIF Category: models
  • Quality Score: 5
  • Status: active
  • Source Type: paper
  • Language: en
  • Tags: reasoning-traces, encrypted-cot, llm-api, security, jailbreak

中文摘要

多家 frontier 提供商把 step-by-step reasoning 以加密文本块返回客户端,下一轮请求再带回,而不是纯服务端保存;这些加密块在同一提供商生态内可跨 session / 用户 / 模型移植利用兼容性可做可扩展的 decryption jailbreak:用较弱模型 + jailbreak,把较强模型的加密 thoughts 转成明文论文还指出 reasoning token 计数与 API 计费 thinking tokens 在多数 prompt 上 1:1;公开分享的 Claude Code / Codex session 中若带加密 reasoning,可能被解码并泄露个人数据,初步扫描约 7000 条公开 trace,报出 API key / 邮箱 / 密码等敏感字段团队已走 responsible disclosure;实验室已修补部分问题,仍在继续处理

English Summary

Several frontier providers return step-by-step reasoning as encrypted blobs that the client must echo on the next request rather than keeping server-side. The paper shows these blobs are portable across sessions, users, and models within the same provider, enabling scalable decryption jailbreaks - a weaker model plus a jailbreak can lift the encrypted thoughts of a stronger model into plaintext. Authors report that reasoning-token counts map 1:1 to billed thinking tokens on most prompts and that public Claude Code / Codex sessions containing encrypted reasoning can leak personal data; a scan of ~7,000 public traces surfaced API keys, emails, and passwords. The team has gone through responsible disclosure; some issues are patched, others remain in handling.

来源 / Obsidian 引用

  • 本机 Obsidian 同日 digest 落地(ClawFeed 24h / AK-RSS-Digest 89 源 / X-Hot-Brief / 内容选题编排)
  • 评分依据:原文正文(不靠标题/摘要/源声誉)