ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
- ID: d3e4f352
- 原文链接: https://arxiv.org/abs/2608.20338
- PDF: https://arxiv.org/pdf/2608.20338
- 作者: Sahil Kale, Ian Harris
- 日期: 2026-08-20
- 更新: 2026-08-20
- 分类: models
- 来源类型: arxiv
- 标签: unlearning, benchmark, llm, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-23T05:25:57Z
中文导读
ConceptGuard 提出概念级情境敏感的遗忘评测基准:现有方法用互不相干的事实集与直接回忆来评测 unlearning,无法检验删除有害概念时能否保留相关良性知识
为什么值得关注
ConceptGuard 提出概念级情境敏感的遗忘评测基准:现有方法用互不相干的事实集与直接回忆来评测 unlearning,无法检验删除有害概念时能否保留相关良性知识
Grounding: introduces the notion of dual-use concepts and builds forget/retain sets that are explicitly complementary in concept usage, so unlearning can be gauged at the concept level rather than via disjoint fact sets.
关键信息
- 论文标题: ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
- 作者: Sahil Kale, Ian Harris
- arXiv: https://arxiv.org/abs/2608.20338
- 发布时间: 2026-08-20
- arXiv 分类: cs.CL
- 关联标签: unlearning, benchmark, llm, arxiv
English Abstract
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge. We argue that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning. To better evaluate unlearning techniques from such a practical viewpoint, we introduce the notion of dual-use concepts: concepts that can be used in both harmful and benign contexts. Building on these concepts, we construct a benchmark called ConceptGuard where forget and retain sets are explicitly complementary in concept usage. Our benchmark uniquely enables unlearning to be explored and gauged at the level of concepts, instead of sparse facts, and evaluation is intent-sensitive with the goal of maximizing contextual separation to promote safer behavior. We demonstrate that current unlearning techniques perform poorly under this setting, showing weak contextual separation alongside poor performance in ROUGE and concept-level metrics. Our results reveal strong forgetting-utility trade-offs, limited gains in contextual sensitivity, and poor consistency in concept-level control across methods, and provide ideas for unlearning approaches that better align with real-world safety requirements. Our dataset is publicly available.
English Summary
ConceptGuard is a benchmark for context-sensitive, concept-level unlearning in LLMs: existing evaluations use disjoint forget/retain fact sets with direct recall, and so fail to test whether harmful concepts are removed while related benign knowledge survives.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。