Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
- ID: 193d45c0
- 原文链接: https://arxiv.org/abs/2608.26036
- PDF: https://arxiv.org/pdf/2608.26036v1
- 作者: Srimonti Dutta, Akshata Kishore Moharir
- 日期: 2026-08-26
- 更新: 2026-08-26
- 分类: agents
- 来源类型: paper
- 标签: evaluation、agents、text-to-sql、reliability、arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-28T05:15:03Z
中文导读
主张答案准确率不足以作为 LLM 数据 agent 的可靠性信号:结构化数据任务里基准正确的答案可能出自无效计算轨迹提出 Trace Integrity 判据(显式可执行schema 合法算子忠实可重放答案一致可审计)与执行契约(把用户意图绑定到 schema算子计划查询与验证状态),并给出 CAIT(正确答案/无效轨迹)率BIRD Mini-Dev 上 Direct SQLOperation Summary+SQLContract-First SQL 答案准确率 20%/22%/24%,轨迹完整通过率 39%/43%/40%,CAIT 率高达 55%/59.1%/45.8%准确率轨迹有效性与静默失败风险是三种不同信号
为什么值得关注
BIRD Mini-Dev 上 CAIT 率高达 45.8%–59.1%:近半到六成基准答对背后是无效计算轨迹。对做 text-to-SQL / 数据 agent 评测的人,这组指标把答案准确率、轨迹有效性、静默失败风险拆成三个必须分开报告的信号。
关键信息
- 论文标题:Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems
- 作者:Srimonti Dutta, Akshata Kishore Moharir
- arXiv:https://arxiv.org/abs/2608.26036
- 发布时间:2026-08-26
- arXiv 分类:cs.AI, cs.CL
- 关联标签:evaluation、agents、text-to-sql、reliability、arxiv
English Abstract
Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. We identify the Structure Gap as the deployment failure mode that makes Trace Integrity necessary: natural-language reasoning and free-form rationales do not reliably specify the operator-level programs required by real-world systems. We operationalize Trace Integrity with execution contracts, structured artifacts that bind user intent to schema elements, operator plans, assumptions, executable queries, verification status, and final-answer linkage. We also introduce CAIT (Correct Answer / Invalid Trace) Rate, which measures how often answer-only evaluation counts computationally unsupported outputs as successes. In an empirical demonstration on BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL achieve answer accuracies of 20%, 22%, and 24%, while their Trace Integrity Pass Rates are 39%, 43%, and 40% and their CAIT Rates remain high at 55%, 59.1%, and 45.8%, showing that answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals. Real-world LLM data agents should, therefore, be evaluated not only by whether their outputs match a reference answer, but by whether those outputs are backed by auditable computation.
English Summary
Argues answer accuracy is an insufficient reliability signal for LLM data agents: on structured-data tasks a benchmark-correct answer can come from an invalid trace. Defines Trace Integrity (explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, auditable computation), operationalized via execution contracts binding user intent to schema elements, operator plans, queries, and verification status, plus CAIT (Correct Answer / Invalid Trace) Rate. On BIRD Mini-Dev, Direct SQL, Operation Summary+SQL, and Contract-First SQL reach 20%/22%/24% answer accuracy but 39%/43%/40% trace-integrity pass rates and 55%/59.1%/45.8% CAIT rates: answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。