模型与实验室 4.0 · 优秀 2026-10-01 · 论文

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

消融语言模型某个组件后其他组件常显得在自我修复,既往最系统的研究认为该现象嘈杂且难有单一解释本文给出一个机制:修复响应其实是消融前就存在的增益把任何干预看作反事实对比符号强度坐标轴 上的一点,常规消融只是未校准的点;细粒度单元 r 的因果修复响应服从仿射律 E_r()=own_r+_r,斜率 _r 是有无消融都一致影响模型的固定系数,符号决定该单元对抗还是强化被移除信号在 Gemma/Qwen/LLaMA/Mistral 四个家族的事实判断任务上,81 个下游方向中 68 个服从该律,且 _r 幅值可从固定权重预估;GPT-2 Small 的 IOI 电路中可达的 10 个头有 7 个服从且全是 counterweight

打开原文回到归档

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

  • ID: ccb114e8
  • 原文链接: https://arxiv.org/abs/2610.02173
  • PDF: https://arxiv.org/pdf/2610.02173v1
  • 作者: Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
  • 日期: 2026-10-01
  • 更新: 2026-10-01
  • 分类: models
  • 来源类型: paper
  • 标签: interpretability, self-repair, ablation, circuits, mechanistic-interpretability
  • 质量评分: 4/5
  • Fetch: 2026-10-03T12:25:19Z

中文导读

消融语言模型某个组件后其他组件常显得在自我修复,既往最系统的研究认为该现象嘈杂且难有单一解释。本文给出一个机制:修复响应其实是消融前就存在的增益。把任何干预看作反事实对比符号强度坐标轴 λ 上的一点,常规消融只是未校准的点;细粒度单元 r 的因果修复响应服从仿射律 E_r(λ)=own_r+γ_r·λ,斜率 γ_r 是有无消融都一致影响模型的固定系数,符号决定该单元对抗还是强化被移除信号。在 Gemma/Qwen/LLaMA/Mistral 四个家族的事实判断任务上,81 个下游方向中 68 个服从该律,且 γ_r 幅值可从固定权重预估;GPT-2 Small 的 IOI 电路中可达的 10 个头有 7 个服从且全是 counterweight。所谓自我修复,可能只是 counterweight 在对比信号出现时的常规运作。

为什么值得关注

自修复不是消融的产物而是消融前就有的增益:修复响应服从仿射律,斜率可从固定权重直接预估

以上导读与价值判断锚定论文摘要与元数据,完整英文摘要见下文。

关键信息

  • 论文标题:Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
  • 作者:Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
  • arXiv:https://arxiv.org/abs/2610.02173
  • 发布时间:2026-10-01
  • arXiv 分类:cs.LG (primary), cs.CL
  • 关联标签:interpretability, self-repair, ablation, circuits, mechanistic-interpretability

English Abstract

Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

English Summary

Ablating a component of a language model often makes other components appear to compensate - self-repair - previously concluded to be noisy without a single explanation. The authors argue there is one: a gain already present before any ablation. Any intervention on a causally important component is a point on the axis of signed counterfactual contrast strength λ, making conventional ablations uncalibrated points. The causal repair response of a fine-grained unit r follows an affine law E_r(λ)=own_r+γ_rλ, where γ_r is a fixed coefficient that influences the model with or without ablation; its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four model families (Gemma, Qwen, LLaMA, Mistral), 68 of 81 downstream directions - including MLP neurons, OV neurons, and singular directions - follow the law, and γ_r magnitude can be anticipated from the fixed weights. On GPT-2 Small's IOI circuit, seven of the ten reachable heads follow the law, and all seven are counterweights: apparent self-repair may simply be a counterweight performing its usual operation when the contrastive signal emerges at the core.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。