The Implications of Linguistic Illegibility for LLM Security
- ID: 048f2377
- 原文链接: https://arxiv.org/abs/2609.02852
- PDF: https://arxiv.org/pdf/2609.02852v1
- 作者: James Mickens
- 日期: 2026-09-02
- 更新: 2026-09-02
- 分类: agents
- 来源类型: paper
- 标签: llm-security, chain-of-thought, sandboxing, taint-tracking, ai-safety
- 质量评分: 4/5
- 抓取时间: 2026-09-19T04:21:30Z
中文导读
Mickens 提出语言不可读性:LLM 外化的语言输出和机制探测出的语言特征,都不是理解模型内部计算的可靠透镜内部计算本质是激活空间上的数学,进出自然语言只是两端的有损翻译推论:一切依赖模型语言自述的安全机制(CoT 监控constitutional 自我批判语言特征 activation probing)都不可能完全可靠沙箱必须有不依赖读模型语言状态的保证:对模型输出做 taint tracking健壮虚拟化沙箱配置第三方审计;这套组合本可缓解近期前沿模型的沙箱逃逸
为什么值得关注
别信模型说的话:CoT 监控永远不彻底,沙箱要靠 taint tracking 兜底
关键信息
- 论文标题:The Implications of Linguistic Illegibility for LLM Security
- 作者:James Mickens
- arXiv:https://arxiv.org/abs/2609.02852
- 发布时间:2026-09-02
- arXiv 分类:cs.LG, cs.CR
- 关联标签:llm-security, chain-of-thought, sandboxing, taint-tracking, ai-safety
English Abstract
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.
English Summary
An LLM's externalized linguistic outputs and mechanistically-probed language features can be an unreliable lens for understanding internal model computation - a property the author names 'linguistic illegibility'. Since internal computation is math over activation spaces with lossy translations to natural language at the bookends, security mechanisms that rely on linguistic self-reporting (chain-of-thought monitoring, constitutional self-critique, activation probing) can never be completely sound. The paper argues sandboxes need guarantees independent of reading a model's linguistic state: taint tracking over model outputs, robust virtualization, and third-party auditing of sandbox configurations, which collectively would have mitigated recent sandbox exploits by frontier models.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。