模型与实验室 5.0 · 必读 2026-08-06 · 论文

Learning When to Trust via Selective Context Preference Optimization

论文把模型抗误导上下文的问题改写为 selective trust:只训练模型抵抗外部信号,可能得到一个忽略所有上下文的鲁棒但无用模型MIST 为同一 reasoning item 构造 cleanmisleadingcorrect-contextirrelevant-context 四种条件,并用 SC2W 衡量误导信号把原本答对的问题翻错的频率SCOPE 从 clean-correct/misleading-wrong 失败中挖偏好对,用均衡 DPO 降低误导易感性,同时保留正确上下文收益

打开原文回到归档

Learning When to Trust via Selective Context Preference Optimization

Source: https://arxiv.org/abs/2608.06377
PDF: https://arxiv.org/pdf/2608.06377v1
Content fetched: 2026-08-09T15:34:20.317851+00:00
Grounding: OpenCLI arXiv metadata and abstract; Obsidian evidence: OpenClaw定时任务/论文流水线/2026-08-09-论文流水线.md

Metadata

  • Author(s): Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
  • Original date: 2026-08-06
  • Platform: arxiv
  • AAIF quality score: 5

中文摘要

论文把模型抗误导上下文的问题改写为 selective trust:只训练模型抵抗外部信号,可能得到一个忽略所有上下文的“鲁棒但无用”模型。MIST 为同一 reasoning item 构造 clean、misleading、correct-context、irrelevant-context 四种条件,并用 SC2W 衡量误导信号把原本答对的问题翻错的频率。SCOPE 从 clean-correct/misleading-wrong 失败中挖偏好对,用均衡 DPO 降低误导易感性,同时保留正确上下文收益。

English Summary

SCOPE recasts robustness to misleading context as selective trust. MIST renders each reasoning item under clean, misleading, correct-context, and irrelevant-context conditions, while SC2W measures context-induced flips from clean-correct to wrong. The paper uses balanced DPO preference pairs to reduce susceptibility without discarding useful context.

Abstract excerpt

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong.

Why it matters for AAIF

上下文安全的目标不是一概不信外部信号,而是学会选择性信任。