Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Source: <https://arxiv.org/abs/2609.22056>
Authors: Andre Bacellar
Published: 2026-09-18
Categories: cs.IR, cs.CL, cs.LG
arXiv: 2609.22056
Abstract
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings. Second (Feature Regime Complementarity): no single ANN score feature achieves best predictive performance across all failure regimes; the dominant feature differs between datasets (query length on MuSiQue, hop-1 concentration on HoVer). We instantiate these principles in RegimeAbstain, which computes a Retrieval Confidence Score (RCS), a logistic function of up to nine query-ANN structural features, all available without any additional LLM call, and uses it to implement a calibrated abstention policy. We define the Confident-Wrong-Answer Rate (CWAR) metric and evaluate across three multi-hop benchmarks (MuSiQue, 2WikiMultiHopQA, HoVer) and two retrieval architectures (LLM-judge and dense-only), covering five failure regimes with CWAR from 14.5% to 62.1%. RCS achieves best or co-best AUC-AC in all five conditions against eight confidence baselines. On MuSiQue (LLM-judge), RCS reduces CWAR from 39.5% to 20.6% at 50% coverage (47.8% relative reduction), with ECE=0.035. A model trained on MuSiQue transfers to 2WikiMultiHopQA with only -0.5pp AUC loss, confirming the domain-agnostic structure of regime features.
Summary
Argues multi-hop retrieval failures cluster in structurally predictable subpopulations. Proves (1) CWAR reduction requires retrieval features to carry mutual information about success (true for LLM-judge, weak in dense-only); (2) no single ANN-score feature dominates across failure regimes (query length on MuSiQue, hop-1 concentration on HoVer). Realizes RegimeAbstain, a logistic Retrieval Confidence Score over up to nine query-ANN structural features, no extra LLM call. Across MuSiQue / 2WikiMultiHopQA / HoVer and five failure regimes, RCS is best-or-co-best in AUC-AC against eight baselines; on MuSiQue LLM-judge it cuts CWAR from 39.5% to 20.6% at 50% coverage (ECE=0.035) and transfers to 2WikiMultiHopQA with -0.5pp AUC loss.
摘要
证明多跳检索失败集中于可预测的结构化子人群。两项定理:(1)CWAR 可降仅当检索特征与成功互信息,在 LLM-judge 架构下成立、于纯密集检索中不成立;(2)无单一 ANN 分数特征能跨失败范式领先,查询长度在 MuSiQue 主导,hop-1 集中度在 HoVer 主导。产物 RegimeAbstain:对多达 9 个查询-ANN 结构特征作 logistic RCS,不需额外 LLM 调用。五个范式下 RCS 对八个基线均最佳或并佳 AUC-AC;MuSiQue LLM-judge 上 50% 覆盖率 CWAR 从 39.5% 降至 20.6%(ECE=0.035),跨集迁移到 2WikiMultiHopQA 仅 -0.5pp AUC 损失。