Learning When to Trust via Selective Context Preference Optimization
Source: https://arxiv.org/abs/2608.06377
PDF: https://arxiv.org/pdf/2608.06377v1
Content fetched: 2026-08-09T15:34:20.317851+00:00
Grounding: OpenCLI arXiv metadata and abstract; Obsidian evidence: OpenClaw定时任务/论文流水线/2026-08-09-论文流水线.md
Metadata
- Author(s): Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong
- Original date: 2026-08-06
- Platform: arxiv
- AAIF quality score: 5
中文摘要
论文把模型抗误导上下文的问题改写为 selective trust:只训练模型抵抗外部信号,可能得到一个忽略所有上下文的“鲁棒但无用”模型。MIST 为同一 reasoning item 构造 clean、misleading、correct-context、irrelevant-context 四种条件,并用 SC2W 衡量误导信号把原本答对的问题翻错的频率。SCOPE 从 clean-correct/misleading-wrong 失败中挖偏好对,用均衡 DPO 降低误导易感性,同时保留正确上下文收益。
English Summary
SCOPE recasts robustness to misleading context as selective trust. MIST renders each reasoning item under clean, misleading, correct-context, and irrelevant-context conditions, while SC2W measures context-induced flips from clean-correct to wrong. The paper uses balanced DPO preference pairs to reduce susceptibility without discarding useful context.
Abstract excerpt
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong.
Why it matters for AAIF
上下文安全的目标不是一概不信外部信号,而是学会选择性信任。