Agent 与自动化 4.0 · 优秀 2026-06-17 · 论文

Uncertainty Decomposition for Clarification Seeking in LLM Agents

面向交互式 LLM agent 的提示级不确定性分解:把动作置信度与请求不确定性 u 拆开,任务欠规约时主动向用户请求澄清部署约束(黑盒 API交互延迟预算无标注轨迹)排除了 logprob多采样与训练式方法,提示式估计是唯一可行族作者发布 WebShop-Clarification 与 ALFWorld-Clarification 两个基准(50% 任务故意欠规约),在 GPT-5.1DeepSeek-v3.2-expGLM-4.7Qwen3.5-35BGPT-OSS-120B 五个骨干上对比 ReAct+UE 与 UAM:ALFWorld-Clarification 澄清 F1 平均提升 73%(对 ReAct+UE)与 36%(对 UAM)...

打开原文回到归档

Uncertainty Decomposition for Clarification Seeking in LLM Agents

中文导读

面向交互式 LLM agent 的提示级不确定性分解:把动作置信度与请求不确定性 u 拆开,任务欠规约时主动向用户请求澄清部署约束(黑盒 API交互延迟预算无标注轨迹)排除了 logprob多采样与训练式方法,提示式估计是唯一可行族作者发布 WebShop-Clarification 与 ALFWorld-Clarification 两个基准(50% 任务故意欠规约),在 GPT-5.1DeepSeek-v3.2-expGLM-4.7Qwen3.5-35BGPT-OSS-120B 五个骨干上对比 ReAct+UE 与 UAM:ALFWorld-Clarification 澄清 F1 平均提升 73%(对 ReAct+UE)与 36%(对 UAM),WebShop-Clarification 上每个骨干均领先

为什么值得关注

面向交互式 LLM agent 的提示级不确定性分解:把动作置信度与请求不确定性 u 拆开,任务欠规约时主动向用户请求澄清部署约束(黑盒 API交互延迟预算无标注轨迹)排除了 logprob多采样与训练式方法。

收录理由:把不确定性量化落到主动澄清这一具体 agent 能力上,基准可直接复用

关键信息

  • 论文标题:Uncertainty Decomposition for Clarification Seeking in LLM Agents
  • 作者:Gregory Matsnev
  • arXiv:https://arxiv.org/abs/2606.19559
  • 发布时间:2026-06-17
  • arXiv 分类:cs.AI, cs.CL
  • 关联标签:uncertainty, clarification, webshop, alfworld, agent-capability, arxiv

English Abstract

Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and call for underspecification-aware, decomposed, and communicable uncertainty representations that can unlock new agent capabilities such as proactive clarification seeking and shared mental-model building. Practical deployment constraints -- black-box APIs, interactive latency budgets, and the absence of labeled trajectories -- rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the most viable family for surfacing such signals at deployment time. We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when the task specification is ambiguous. To evaluate it, we introduce two clarification-augmented benchmarks (WebShop-Clarification and ALFWorld-Clarification) in which 50% of tasks are deliberately underspecified, and systematically compare the proposed decomposition against ReAct+UE and Uncertainty-Aware Memory (UAM) across five LLM backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B) on these variants together with the standard WebShop, ALFWorld, and REAL benchmarks for fault detection. Averaged across the five backbones, the proposed decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and by 36% over UAM, and leads clarification F1 on every backbone on WebShop-Clarification and on four of five backbones on ALFWorld-Clarification, indicating that the gains generalize beyond a single LLM.

English Summary

A prompt-based uncertainty decomposition for interactive LLM agents that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when a task specification is ambiguous. Deployment constraints (black-box APIs, latency budgets, no labeled trajectories) rule out logprob-based, multi-sampling, and training-based methods, leaving prompt-based estimation as the viable family. Two clarification-augmented benchmarks are introduced (WebShop-Clarification, ALFWorld-Clarification) with 50% of tasks deliberately underspecified. Across five backbones (GPT-5.1, DeepSeek-v3.2-exp, GLM-4.7, Qwen3.5-35B, GPT-OSS-120B), the decomposition improves clarification F1 on ALFWorld-Clarification by 73% over ReAct+UE and 36% over UAM on average, and leads on every backbone on WebShop-Clarification.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定于条目与论文摘要,未补充摘要之外的内容。