Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
- ID: 879fffc9
- 原文链接: https://arxiv.org/abs/2609.40286
- PDF: https://arxiv.org/pdf/2609.40286v1
- 作者: Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
- 日期: 2026-09-30
- 更新: 2026-09-30
- 分类: models
- 来源类型: paper
- 标签: unlearning, multilingual, safety, benchmark
- 质量评分: 4/5
- 抓取时间: 2026-10-02T12:30:30.134235+00:00
中文导读
在一种语言里遗忘一个事实并不保证其他语言中也被移除改变查询语言甚至只改要求答案的语言都能重新打开看似已遗忘的知识(跨语言漏洞)对所有语言都做遗忘既不可扩展也会放大对无关能力的损伤论文形式化 language-budgeted multilingual unlearning:在语言预算内选一组源语言最大化跨语言擦除,并发布覆盖 174 个语言-文字对25 类原子改写的 Cross-Lingual Unlearning Tensor 基准提出的 COVER 只需良性校准数据与冻结模型即可部署,挑选源语言最大化无遗忘监督语言的覆盖;单独强的源语言并不可靠地组合成强集合三个模型家族两个遗忘集上,COVER 相比均匀选择把平均 held-out 残余访问降低 7.8-27.3%,并在 LORELEI 低资源语言真实新闻文档上验证迁移
为什么值得关注
单语言遗忘挡不住换语言重问:174 语言基准 + COVER 在语言预算内选源语言,把残余访问压低 7.8-27.3%
关键信息
- 论文标题: Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
- 作者: Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
- arXiv: https://arxiv.org/abs/2609.40286
- arXiv 分类: cs.CL, cs.AI
- 发布时间: 2026-09-30
- 关联标签: unlearning, multilingual, safety, benchmark
English Abstract
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
English Summary
Unlearning a fact in one language does not guarantee removal in others; changing the query or even the requested answer language can reopen forgotten knowledge. The paper formalizes language-budgeted multilingual unlearning and introduces the Cross-Lingual Unlearning Tensor benchmark (174 language-script pairs, 25 paraphrase types). COVER selects source languages to maximize coverage of languages receiving no forget supervision, requiring only benign calibration data and the frozen model; naively strong individual sources do not compose into strong sets. Across three model families and two forget sets, COVER cuts mean held-out residual access by 7.8-27.3% versus uniform selection, with gains transferring to real LORELEI low-resource news documents.
Obsidian Notes
- Content grounded in an opencli fetch during content-fetcher backfill (papers via
opencli arxiv paper -f json; articles viaopencli web read). - 中文导读 block is the entry's summary_zh (verbatim); the value note is the entry's one_liner; English blocks are verbatim excerpts from the fetched source. No claims beyond the fetched metadata/source text were added.