Uncertainty Decomposition for Clarification Seeking in LLM Agents
- ID: 23a3ceb7
- 原文链接: https://arxiv.org/abs/2606.19559
- PDF: https://arxiv.org/pdf/2606.19559v1
- 作者: Gregory Matsnev
- 日期: 2026-06-17
- 更新: N/A
- 分类: agents
- 来源类型: paper
- 标签: uncertainty, clarification, webshop, alfworld, agent-capability, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-17T04:23:51+00:00Z
中文导读
面向交互式 LLM agent 的提示级不确定性分解:把动作置信度与请求不确定性 u 拆开,任务欠规约时主动向用户请求澄清部署约束(黑盒 API交互延迟预算无标注轨迹)排除了 logprob多采样与训练式方法,提示式估计是唯一可行族作者发布 WebShop-Clarification 与 ALFWorld-Clarification 两个基准(50% 任务故意欠规约),在 GPT-5.1DeepSeek-v3.2-expGLM-4.7Qwen3.5-35BGPT-OSS-120B 五个骨干上对比 ReAct+UE 与 UAM:ALFWorld-Clarification 澄清 F1 平均提升 73%(对 ReAct+UE)与 36%(对 UAM),WebShop-Clarification 上每个骨干均领先
为什么值得关注
面向交互式 LLM agent 的提示级不确定性分解:把动作置信度与请求不确定性 u 拆开,任务欠规约时主动向用户请求澄清部署约束(黑盒 API交互延迟预算无标注轨迹)排除了 logprob多采样与训练式方法。
收录理由:把不确定性量化落到主动澄清这一具体 agent 能力上,基准可直接复用
关键信息
- 论文标题:Uncertainty Decomposition for Clarification Seeking in LLM Agents
- 作者:Gregory Matsnev
- arXiv:https://arxiv.org/abs/2606.19559
- 发布时间:2026-06-17
- arXiv 分类:cs.AI, cs.CL
- 关联标签:uncertainty, clarification, webshop, alfworld, agent-capability, arxiv
English Abstract
Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and call for underspecification-aware, decomposed, and communicable uncertainty representations that can unlock new agent capabilities such as proactive clarification seeking and shared mental-model building. Practical deployment constraints -- black-box APIs, interactive latency budgets, and the absence of labeled trajectories -- rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the most viable family for surfacing such signals at deployment time. We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when the task specification is ambiguous. To evaluate it, we introduce two clarification-augmented benchmarks (WebShop-Clarification and ALFWorld-Clarification) in which 50% of tasks are deliberately underspecified, and systematically compare the proposed decomposition against ReAct+UE and Uncertainty-Aware Memory (UAM) across five LLM backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B) on these variants together with the standard WebShop, ALFWorld, and REAL benchmarks for fault detection. Averaged across the five backbones, the proposed decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and by 36% over UAM, and leads clarification F1 on every backbone on WebShop-Clarification and on four of five backbones on ALFWorld-Clarification, indicating that the gains generalize beyond a single LLM.
English Summary
A prompt-based uncertainty decomposition for interactive LLM agents that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when a task specification is ambiguous. Deployment constraints (black-box APIs, latency budgets, no labeled trajectories) rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the viable family. Two clarification-augmented benchmarks are introduced (WebShop-Clarification, ALFWorld-Clarification) with 50% of tasks deliberately underspecified. Across five backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B), the decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and 36% over UAM on average, and leads on every backbone on WebShop-Clarification.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断锚定于条目与论文摘要,未补充摘要之外的内容。