When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control
Source: https://arxiv.org/abs/2609.17516 · platform: arxiv · authors: Ali Şenol · date: 2026-09-15
TL;DR(J(中)文摘要)
Ali \u015eenol \u63d0\u51fa Chain-of-Self-Questioning (CoSQ)\uff1a\u7eaf prompt \u6846\u67b6\uff0c\u628a\u300c\u4f5c\u7b54\u300d\u53d8\u6210\u6761\u4ef6\u51b3\u7b56\u2014\u2014\u5148\u663e\u5f0f\u8bc4\u4f30\u56de\u7b54\u8be5\u95ee\u9898\u6240\u9700\u4fe1\u606f\u662f\u5426\u5145\u5206\uff0c\u518d\u51b3\u5b9a\u7b54\u6216\u5f03\u3002TruthfulQA \u591a\u9009 817 \u9898\u300111 \u4e2a\u5f00\u6e90/\u6258\u7ba1\u6a21\u578b\u65cf\u300117 \u79cd\u6761\u4ef6\u4e0b\uff0c\u6700\u7ec8 balanced \u534f\u8bae\u4e0b Grounded-CoSQ(\u03c4=0.90) \u628a\u65e0\u6761\u4ef6\u9519\u8bef\u627f\u8bfa\u7387\u4ece chain-of-thought \u7684 13.1% \u964d\u5230 8.9%\uff08\u76f8\u5bf9 -32.1%\uff09\uff0c\u7b54\u51c6\u7387\u4ece 86.9% \u5347\u5230 89.7%\uff0c\u8986\u76d6\u7387 87.6%\uff1b\u4e09\u4e2a\u53d8\u4f53\uff08Grounded/Critical/Adaptive\uff09\u5728\u6240\u6709\u6a21\u578b\u3001\u6240\u6709\u9608\u503c\u4e0a\u5747\u7a33\u5b9a\u4f18\u4e8e\u57fa\u7ebf\u3002\u7ed3\u8bba\uff1a\u5f53\u4e0d\u652f\u6301\u7684\u627f\u8bfa\u6bd4\u8f6c\u4eba\u5de5\u66f4\u8d35\u65f6\uff0c\u81ea\u8bc4\u53ef\u652f\u6491\u53ef\u8c03\u7684\u300c\u7b54\u6216\u5f03\u300d\u3002
Summary (English)
Large language models can produce fluent answers when their factual support is weak. This paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework that makes answer commitment conditional on an explicit assessment of the information required to answer a question. We evaluate three CoSQ variants under seventeen conditions on the 817-item TruthfulQA multiple-choice validation set using eleven open-weight and hosted model families. In the final balanced-option protocol, Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate from 13.1% under chain-of-thought prompting to 8.9%, a 32.1% relative reduction, while increasing answered accuracy from 86.9% to 89.7% and answering 87.6% of questions. Both improvements hold for all eleven models and at every evaluated threshold. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while remaining more reliable than the baseline. A secondary Natural Questions Short-Answer evaluation provides convergent open-form evidence. These findings show that self-assessment can support explicit, tunable answer-or-abstain decisions when an unsupported commitment is more costly than referral or review.
入库依据(OpenCLI arXiv metadata 直出)
opencli arxiv paper 2609.17516 -f json\uff1acs.CL+cs.AI\uff0c\u5355\u4f5c\u8005 Ali \u015eenol\uff0cpublished/updated \u540c 2026-09-15\uff1babstract \u542b 11 \u6a21\u578b\u65cf\u3001\u03c4=0.90 \u7b49\u5173\u952e\u6570\u636e\u70b9\u76f4\u63a5\u843d content\uff0c\u672a\u505a\u4e8c\u6b21\u8f6c\u8ff0\u3002