Agent 与自动化 4.0 · 优秀 2026-08-26 · 论文

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World S...

主张答案准确率不足以作为 LLM 数据 agent 的可靠性信号:结构化数据任务里基准正确的答案可能出自无效计算轨迹提出 Trace Integrity 判据(显式可执行schema 合法算子忠实可重放答案一致可审计)与执行契约(把用户意图绑定到 schema算子计划查询与验证状态),并给出 CAIT(正确答案/无效轨迹)率BIRD Mini-Dev 上 Direct SQLOperation Summary+SQLContract-First SQL 答案准确率 20%/22%/24%,轨迹完整通过率 39%/43%/40%,CAIT 率高达 55%/59.1%/45.8%准确率轨迹有效性与静默失败风险是三种不同信号

打开原文回到归档

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

中文导读

主张答案准确率不足以作为 LLM 数据 agent 的可靠性信号:结构化数据任务里基准正确的答案可能出自无效计算轨迹提出 Trace Integrity 判据(显式可执行schema 合法算子忠实可重放答案一致可审计)与执行契约(把用户意图绑定到 schema算子计划查询与验证状态),并给出 CAIT(正确答案/无效轨迹)率BIRD Mini-Dev 上 Direct SQLOperation Summary+SQLContract-First SQL 答案准确率 20%/22%/24%,轨迹完整通过率 39%/43%/40%,CAIT 率高达 55%/59.1%/45.8%准确率轨迹有效性与静默失败风险是三种不同信号

为什么值得关注

BIRD Mini-Dev 上 CAIT 率高达 45.8%–59.1%:近半到六成基准答对背后是无效计算轨迹。对做 text-to-SQL / 数据 agent 评测的人,这组指标把答案准确率、轨迹有效性、静默失败风险拆成三个必须分开报告的信号。

关键信息

  • 论文标题:Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
  • 作者:Srimonti Dutta, Akshata Kishore Moharir
  • arXiv:https://arxiv.org/abs/2608.26036
  • 发布时间:2026-08-26
  • arXiv 分类:cs.AI, cs.CL
  • 关联标签:evaluation、agents、text-to-sql、reliability、arxiv

English Abstract

Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. We identify the Structure Gap as the deployment failure mode that makes Trace Integrity necessary: natural-language reasoning and free-form rationales do not reliably specify the operator-level programs required by real-world systems. We operationalize Trace Integrity with execution contracts, structured artifacts that bind user intent to schema elements, operator plans, assumptions, executable queries, verification status, and final-answer linkage. We also introduce CAIT (Correct Answer / Invalid Trace) Rate, which measures how often answer-only evaluation counts computationally unsupported outputs as successes. In an empirical demonstration on BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL achieve answer accuracies of 20%, 22%, and 24%, while their Trace Integrity Pass Rates are 39%, 43%, and 40% and their CAIT Rates remain high at 55%, 59.1%, and 45.8%, showing that answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals. Real-world LLM data agents should, therefore, be evaluated not only by whether their outputs match a reference answer, but by whether those outputs are backed by auditable computation.

English Summary

Argues answer accuracy is an insufficient reliability signal for LLM data agents: on structured-data tasks a benchmark-correct answer can come from an invalid trace. Defines Trace Integrity (explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, auditable computation), operationalized via execution contracts binding user intent to schema elements, operator plans, queries, and verification status, plus CAIT (Correct Answer / Invalid Trace) Rate. On BIRD Mini-Dev, Direct SQL, Operation Summary+SQL, and Contract-First SQL reach 20%/22%/24% answer accuracy but 39%/43%/40% trace-integrity pass rates and 55%/59.1%/45.8% CAIT rates: answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。