On the Threat Model of Weird Generalization and Emergent Misalignment
- ID: 58f86ed5
- 原文链接: https://arxiv.org/abs/2608.23476
- PDF: https://arxiv.org/pdf/2608.23476v1
- 作者: M, i, r, i, a, m, , W, a, n, n, e, r, ,, , M, a, r, k, , D, r, e, d, z, e, ,, , W, i, l, l, i, a, m, , W, a, l, d, e, n
- 日期: 2026-08-24
- 更新: 2026-08-24
- 分类: industry
- 来源类型: paper
- 标签: ai-safety, weird-generalization, fine-tuning, threat-model, emergent-misalignment
- 质量评分: 4/5
- 抓取时间: 2026-08-26T04:27:05Z
中文导读
Weird generalization(WG)指在小的领域数据集上做窄域微调,却引发模型行为的广泛意外变化本文系统检验哪些微调数据特征是 WG 出现的必要条件:数据集规模构成语言呈现风格相对模型参数知识的新颖度;同时分析 WG 的度量对评测问题集有多敏感三个开源权重模型四个数据集的实验显示:WG 程度(1)更多取决于数据构成和语言而非规模;(2)对预训练中见过的熟悉数据反而更强;(3)对所用评测问题集高度敏感结论:WG 是训练与评测数据两侧都相当脆弱的性质,更合理的定位是需要精心数据工程的对抗性威胁,而非日常微调固有的重大风险
为什么值得关注
三个开源模型四个数据集:weird generalization 高度依赖数据构成与语言,更像对抗性威胁而非日常微调的固有风险
对开源模型微调与模型供应链安全,这篇把 weird generalization 风险定位到数据构成、语言与相对模型参数知识的新颖度上:小规模、构成特殊的第三方数据集应按对抗性威胁对待,而非日常微调的固有风险。
关键信息
- 论文标题: On the Threat Model of Weird Generalization and Emergent Misalignment
- 作者: M, i, r, i, a, m, , W, a, n, n, e, r, ,, , M, a, r, k, , D, r, e, d, z, e, ,, , W, i, l, l, i, a, m, , W, a, l, d, e, n
- arXiv: https://arxiv.org/abs/2608.23476
- 发布时间: 2026-08-24
- arXiv 分类: c, s, ., C, L
- 备注: N/A
- 关联标签: ai-safety, weird-generalization, fine-tuning, threat-model, emergent-misalignment
English Abstract
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
English Summary
Narrow fine-tuning on small, domain-specific datasets can produce broad, surprising behavior changes weird generalization (WG). This paper investigates which features of fine-tuning data are necessary for WG to arise: dataset size, composition, language, presentation style, and novelty relative to the model's parametric knowledge, plus how sensitive WG measurement is to the evaluation question set. Experiments with three open-weight models on four datasets show the degree of WG (1) depends heavily on dataset composition and language, more than on size; (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- 本页由 AAIF content-fetcher 定时任务生成(2026-08-26),仅新增内容页,未改动 entries.json。