Agent 与自动化 4.0 · 优秀 2026-09-18 · 论文

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstenti...

多跳检索失败在结构可预测的查询子群中聚集两条形式化结果:置信错误约简可行当且仅当检索特征携带关于成功与否的互信息(LLM-judge 管线满足dense-only 明显更弱,解释了两者的 AUC-AC 差距);没有任何单一 ANN 分数特征在所有失败域都最优(MuSiQue 上是查询长度HoVer 上是 hop-1 集中度)据此构建 RegimeAbstain:用最多 9 个无需额外 LLM 调用的查询-ANN 结构特征计算检索置信分 RCS,实现校准弃答策略,在三个多跳基准两种检索架构五个失败域(置信错误率 14.5%-62.1%)上全部取得最优或并列最优 AUC-AC

打开原文回到归档

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

Source: <https://arxiv.org/abs/2609.22056&gt;
Authors: Andre Bacellar
Published: 2026-09-18
Categories: cs.IR, cs.CL, cs.LG
arXiv: 2609.22056

Abstract

Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer). We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.

Summary

Argues multi-hop retrieval failures cluster in structurally predictable subpopulations. Proves (1) CWAR reduction requires retrieval features to carry mutual information about success (true for LLM-judge, weak in dense-only); (2) no single ANN-score feature dominates across failure regimes (query length on MuSiQue, hop-1 concentration on HoVer). Realizes RegimeAbstain, a logistic Retrieval Confidence Score over up to nine query-ANN structural features, no extra LLM call. Across MuSiQue / 2WikiMultiHopQA / HoVer and five failure regimes, RCS is best-or-co-best in AUC-AC against eight baselines; on MuSiQue LLM-judge it cuts CWAR from 39.5% to 20.6% at 50% coverage (ECE=0.035) and transfers to 2WikiMultiHopQA with -0.5pp AUC loss.

摘要

证明多跳检索失败集中于可预测的结构化子人群。两项定理:(1)CWAR 可降仅当检索特征与成功互信息,在 LLM-judge 架构下成立、于纯密集检索中不成立;(2)无单一 ANN 分数特征能跨失败范式领先,查询长度在 MuSiQue 主导,hop-1 集中度在 HoVer 主导。产物 RegimeAbstain:对多达 9 个查询-ANN 结构特征作 logistic RCS,不需额外 LLM 调用。五个范式下 RCS 对八个基线均最佳或并佳 AUC-AC;MuSiQue LLM-judge 上 50% 覆盖率 CWAR 从 39.5% 降至 20.6%(ECE=0.035),跨集迁移到 2WikiMultiHopQA 仅 -0.5pp AUC 损失。