Agent 与自动化 4.0 · 优秀 2026-08-14 · 论文

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

多智能体推理系统常用一致性置信度或自动打分决定哪些消息参与最终答案,这隐含着可能正确的消息才值得保留作者提出 Diverse Hypothesis Deliberation(DHD)协议:缓存五条独立生成的消息,让同一个下游 integrator 分别在可见或隐藏每条消息的情况下重放,从而测量消息的轨迹价值在五个数学与科学基准gpt-oss-120b 与 gemma-4-31B-it 两个开源模型族上,每个基准-模型组合都出现答案错误但有助于下游的消息;在能改变最终对错的错误消息中,超过四成的改变是正向的,且可重复效应 unlikely 源于重放波动(p=0.0002)结论:答案正确性提供信息,但不决定消息的去留价值

打开原文回到归档

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

  • ID: febe7513
  • 原文链接: https://arxiv.org/abs/2608.14375
  • PDF: https://arxiv.org/pdf/2608.14375v1
  • 作者: Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur
  • 日期: 2026-08-14
  • 更新: 2026-08-14
  • 分类: agents
  • 来源类型: paper
  • 标签: multi-agent, message-filtering, evaluation, trajectory-value, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-18T05:28:55Z

中文导读

多智能体推理系统常用一致性置信度或自动打分决定哪些消息参与最终答案,这隐含着可能正确的消息才值得保留作者提出 Diverse Hypothesis Deliberation(DHD)协议:缓存五条独立生成的消息,让同一个下游 integrator 分别在可见或隐藏每条消息的情况下重放,从而测量消息的轨迹价值在五个数学与科学基准gpt-oss-120b 与 gemma-4-31B-it 两个开源模型族上,每个基准-模型组合都出现答案错误但有助于下游的消息;在能改变最终对错的错误消息中,超过四成的改变是正向的,且可重复效应 unlikely 源于重放波动(p=0.0002)结论:答案正确性提供信息,但不决定消息的去留价值

为什么值得关注

错误答案的消息也可能推动下游推理:DHD 重放协议显示错误-有益消息在每个基准-模型组合中都存在

该论文发表于 2026-08-14,作者为 Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur,arXiv 分类 cs.AI, cs.CL, cs.LG;以上判断基于论文摘要所述内容。

关键信息

  • 论文标题: Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
  • 作者: Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang, Weijian Zheng, Fernando Llorente, Xiaolong Ma, Xinyang Li, Eliu A. Huerta, Ian T. Foster, Rajeev Thakur
  • arXiv: https://arxiv.org/abs/2608.14375
  • 发布时间: 2026-08-14
  • arXiv 分类: cs.AI, cs.CL, cs.LG
  • 备注: 24 pages, 9 figures. Includes an appendix and an ancillary reproducibility artifact
  • 关联标签: multi-agent, message-filtering, evaluation, trajectory-value, arxiv

English Abstract

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120b and gemma-4-31B-it, wrong-helpful messages appear in every benchmark-model combination. Among wrong-answer messages that change final correctness, more than four in ten changes are helpful in each model. Controlled repeats show that the number of repeatable message effects is unlikely to arise from replay variation alone (p=0.0002). A focused intervention on repeatable wrong-helpful messages finds that the complete message works best, while retaining its reasoning preserves more success than retaining only its answer; the source of the complete-message advantage remains open. Within the same problem, repeated trajectory-value evidence also identifies a better keep-or-remove choice than answer correctness alone. Answer correctness is therefore informative but does not determine trajectory value. DHD measures this missing property and produces reusable labels for learning when agents should listen.

English Summary

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages shape the final answer, assuming a message likely to be correct is also worth keeping. The authors test this with Diverse Hypothesis Deliberation (DHD): a controlled protocol that caches five independently generated messages and replays the same downstream integrator with each message available or hidden, measuring a message's trajectory value. Across five mathematics and science benchmarks and two open model families (gpt-oss-120b and gemma-4-31B-it), wrong-helpful messages appear in every benchmark-model combination; among wrong-answer messages that change final correctness, more than four in ten changes are helpful, with repeatability unlikely from replay variation alone (p=0.0002). Answer correctness is informative but does not determine trajectory value.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。