Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs
- ID: af3c939e
- 原文链接: https://arxiv.org/abs/2609.10439
- PDF: https://arxiv.org/pdf/2609.10439v1
- 作者: Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou
- 日期: 2026-09-09
- 更新: 2026-09-09
- 分类: models
- 来源类型: paper
- 标签: machine-unlearning, privacy, quantization, llm, field-note
- 质量评分: 4/5
- 抓取时间: 2026-09-11T04:23:41
中文导读
层级机器遗忘:forget-to-retain 显著性分数挑层只动最有效的少数层;量化后遗忘知识回潮的问题有了实证缓解路径弥散小更新最容易被低位舍入抹掉
为什么值得关注
层级机器遗忘:forget-to-retain 显著性分数挑层只动最有效的少数层;量化后遗忘知识回潮的问题有了实证缓解路径弥散小更新最容易被低位舍入抹掉
关键信息
- 论文标题:Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs
- 作者:Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou
- arXiv:https://arxiv.org/abs/2609.10439
- 发布时间:2026-09-09
- arXiv 分类:cs.LG, cs.AI
- 关联标签:machine-unlearning, privacy, quantization, llm, field-note
English Abstract
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can degrade utility and remain brittle under deployment changes such as post-training quantization, where forgotten knowledge may partially re-emerge. We propose Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score. This score identifies layers with high influence on the forget set and low sensitivity to the retain set, allowing FOM-UL to concentrate updates where they are most effective while leaving most of the model unchanged. This targeted update strategy improves the forgetting-utility trade-off and provides an empirical path toward quantization-resilient unlearning by reducing the chance that small, diffuse updates are erased by low-bit rounding. Across TOFU, KnowUnDo, and MUSE-style evaluations, FOM-UL reduces residual memorization compared with strong GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines while preserving retain-set utility close to the vanilla model. Under 8-bit and 4-bit post-training quantization, FOM-UL maintains stronger memorization suppression and utility preservation than competing methods, and adversarial prompt evaluations show lower recovery of forgotten content. Overall, FOM-UL provides an efficient unlearning strategy that improves targeted forgetting, utility preservation, and deployment robustness without claiming formal guarantees of erasure.
English Summary
FOM-UL is a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score, identifying layers with high influence on the forget set and low sensitivity to the retain set, concentrating updates where they are most effective while leaving most of the model unchanged. Across TOFU, KnowUnDo, and MUSE-style evaluations, it reduces residual memorization versus GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines while preserving retain-set utility close to the vanilla model, maintains stronger suppression under 8-bit and 4-bit post-training quantization, and shows lower recovery of forgotten content under adversarial prompts, without claiming formal guarantees of erasure.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。