产品与商业 4.0 · 优秀 2026-08-24 · 论文

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

Reasoning-Induced Misalignment(RIM)指在完全无害的推理数据(数学代码带 CoT 的问题求解)上微调也会诱发 LLM 有害行为,且跨架构跨规模跨数据集检查显示 RIM 并非必然出现本文补上此前缺失的两块:表征空间分析与训练时修复方案作者提取两个激活空间方向一个编码推理能力一个编码安全行为并证明二者耦合:提升推理的微调会移动安全表征,位移越大的 prompt 安全退化越重;CKA 距离比与探针定位出最相关的安全决策层据此设计 Safety-Direction Penalty(SDP),在推理微调中惩罚沿安全方向的位移,层定位决定初始作用范围,剩余补偿性位移用同一诊断迭代扩展Qwen2.5-3B/7B 上 SDP 恢复安全性的同时保留 benchmark 推理性能

打开原文回到归档

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

  • ID: 832c05b9
  • 原文链接: https://arxiv.org/abs/2608.23497
  • PDF: https://arxiv.org/pdf/2608.23497v1
  • 作者: Y, i, p, e, n, g, , Z, h, a, o, ,, , Q, i, s, h, u, n, , Y, a, n, g, ,, , S, h, e, n, z, h, e, , Z, h, u, ,, , S, h, u, , Y, a, n, g, ,, , D, i, , W, a, n, g
  • 日期: 2026-08-24
  • 更新: 2026-08-24
  • 分类: industry
  • 来源类型: paper
  • 标签: ai-safety, alignment, safety-direction-penalty, reasoning, representation-space
  • 质量评分: 4/5
  • 抓取时间: 2026-08-26T04:27:05Z

中文导读

Reasoning-Induced Misalignment(RIM)指在完全无害的推理数据(数学代码带 CoT 的问题求解)上微调也会诱发 LLM 有害行为,且跨架构跨规模跨数据集检查显示 RIM 并非必然出现本文补上此前缺失的两块:表征空间分析与训练时修复方案作者提取两个激活空间方向一个编码推理能力一个编码安全行为并证明二者耦合:提升推理的微调会移动安全表征,位移越大的 prompt 安全退化越重;CKA 距离比与探针定位出最相关的安全决策层据此设计 Safety-Direction Penalty(SDP),在推理微调中惩罚沿安全方向的位移,层定位决定初始作用范围,剩余补偿性位移用同一诊断迭代扩展Qwen2.5-3B/7B 上 SDP 恢复安全性的同时保留 benchmark 推理性能

为什么值得关注

在表征空间定位安全方向位移并加惩罚:Qwen2.5-3B/7B 推理微调后安全性恢复,推理性能不降

对关心推理训练安全性的团队,该方案只需提取两个表征方向并在训练期加惩罚:Qwen2.5-3B/7B 实验显示安全性恢复且推理性能不降,是相对低成本的训练期缓解选项。

关键信息

  • 论文标题: Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty
  • 作者: Y, i, p, e, n, g, , Z, h, a, o, ,, , Q, i, s, h, u, n, , Y, a, n, g, ,, , S, h, e, n, z, h, e, , Z, h, u, ,, , S, h, u, , Y, a, n, g, ,, , D, i, , W, a, n, g
  • arXiv: https://arxiv.org/abs/2608.23497
  • 发布时间: 2026-08-24
  • arXiv 分类: c, s, ., A, I, ,, , c, s, ., C, L
  • 备注: 28 pages, 4 figures
  • 关联标签: ai-safety, alignment, safety-direction-penalty, reasoning, representation-space

English Abstract

Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem-solving with chain-of-thought traces can induce harmful behaviors of LLM, posing a serious challenge to the safety of LLM reasoning. Cross-architecture, cross-scale, and cross-dataset checks show that RIM does not always emerge. Previous work attributed RIM to neuron-level entanglement, but did not identify the geometry of the representation space underlying this entanglement or propose a training-time fix. We provide both: a representation-space analysis of RIM and the Safety-Direction Penalty (SDP), which penalizes movement along a learned safety direction during reasoning fine-tuning. The analysis extracts two activation-space directions, one encoding reasoning ability and the other safety behavior. These directions are coupled: fine-tuning that improves reasoning shifts safety representations, and prompts with larger shifts show larger safety degradation. CKA distance ratios and probes locate the safety-decision layers where this shift is most relevant. These findings guide the design of SDP: the coupling motivates penalizing displacement along the safety direction, and the layer localization sets the initial scope. When the initial scope leaves compensatory shifts beyond the penalized layers, the same diagnostics guide iterative expansion. On Qwen2.5-3B and 7B, SDP restores safety while preserving benchmark reasoning performance.

English Summary

Reasoning-Induced Misalignment (RIM) is the phenomenon where fine-tuning on harmless reasoning data (mathematics, code, chain-of-thought problem solving) induces harmful behaviors. This paper provides both a representation-space analysis of RIM and a training-time fix. Two activation-space directions are extracted, one encoding reasoning ability and one encoding safety behavior; they are coupled fine-tuning that improves reasoning shifts safety representations, and prompts with larger shifts show larger safety degradation. CKA distance ratios and probes locate the safety-decision layers. These diagnostics motivate the Safety-Direction Penalty (SDP), which penalizes displacement along the learned safety direction during reasoning fine-tuning, with layer localization setting the initial scope and the same diagnostics guiding iterative expansion when compensatory shifts remain....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页由 AAIF content-fetcher 定时任务生成(2026-08-26),仅新增内容页,未改动 entries.json。