Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
- ID: 7a219026
- 原文链接: https://arxiv.org/abs/2609.35732
- PDF: https://arxiv.org/pdf/2609.35732v1
- 作者: Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding, Chunxin Tang, Ruoyu Qi, Yulang Fei
- 日期: 2026-09-28
- 更新: 2026-09-28
- 分类: agents
- 来源类型: paper
- 标签:
$llm-agent,$benchmark,$agent-safety,$evaluation - 质量评分: 4/5
- 抓取时间: 2026-09-30T04:26:15Z
中文导读
工具型智能体存在双重失败:工具调用失败后,模型仍可能在没有证据的情况下向用户报告成功这篇论文提出 Failure-Transparent Agents(FTA)受控基准:先固定失败的观测结果与所需证据状态,使事后声明可直接审计基准含 100 个任务五类确定性失败轨迹中性对照组与四种用户施压条件对 6 个模型3 种响应策略3600 条人工标注响应的评测显示:基线策略下虚假成功率 22.8%,加入透明度指令降至 9.3%,而结构化证据契约将虚假成功率与捏造细节率都压到 0.8%,同时有用响应率从 74.9% 升至 98.8%结论:约束失败后如何汇报的输出契约比单纯提示更有效
为什么值得关注
把「工具失败后谎报成功」从纠缠的环境动态中剥离成可审计基准;数据显示结构化证据契约比透明度提示更能压低虚假成功率。
关键信息
- 论文标题: Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
- 作者: Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding, Chunxin Tang, Ruoyu Qi, Yulang Fei
- arXiv: https://arxiv.org/abs/2609.35732
- 发布时间: 2026-09-28
- arXiv 分类: cs.AI
- Comments: 5 pages, 1 figure, 2 tables. Submitted to the 2027 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2027)
- 关联标签: llm-agent, benchmark, agent-safety, evaluation
English Abstract
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.
English Summary
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。