Agent 与自动化 5.0 · 必读 2026-08-06 · 论文

The Bitter Lesson of Tool Calling

论文系统比较 programmatic tool calling 与原生 JSON tool calling:把工具暴露为 typed Python stubs,让模型用代码链式或并行调用,并在单个 agent turn 内回传执行结果作者在 BFCL v4 上评估 14 个模型,报告 11/14 个模型 PTC 不低于 JSON baseline,GPT-5.6 family 提升 10.6%;并行 fan-out 下 13/14 个模型不低于 baseline,context rot 条件下也更稳定

打开原文回到归档

The Bitter Lesson of Tool Calling

Source: https://arxiv.org/abs/2608.06370
PDF: https://arxiv.org/pdf/2608.06370v1
Content fetched: 2026-08-08T15:34:45.962729+00:00
Grounding: OpenCLI source metadata/body plus Obsidian digest evidence

Metadata

  • Author(s): Ishan Patel, Sahil Sen, Elias Lumer, Vamse Kumar Subbiah
  • Original date: 2026-08-06
  • Platform: arxiv
  • AAIF quality score: 5

中文摘要

论文系统比较 programmatic tool calling 与原生 JSON tool calling:把工具暴露为 typed Python stubs,让模型用代码链式或并行调用,并在单个 agent turn 内回传执行结果。作者在 BFCL v4 上评估 14 个模型,报告 11/14 个模型 PTC 不低于 JSON baseline,GPT-5.6 family 提升 10.6%;并行 fan-out 下 13/14 个模型不低于 baseline,context rot 条件下也更稳定。

English Summary

Tool use transforms LLMs into agents, and programmatic tool calling replaces rigid JSON calls with typed Python stubs that can chain and parallelize naturally. This paper evaluates PTC versus native JSON tool calling across 14 language models on BFCL v4, reporting parity or better performance in 11 of 14 models, a 10.6% GPT-5.6 family improvement, stronger parallel fan-out results, and better stability under context rot.

Intake Rationale

工具接口形态会跟模型能力一起演进,typed code stubs 可能比 JSON 调用更适合强模型。

Obsidian Evidence

  • Evidence note: /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/OpenClaw定时任务/论文流水线/2026-08-08-论文流水线.md
  • arXiv ID: 2608.06370
  • Primary category: cs.CL
  • Categories: cs.CL

Source Excerpt

Tool use transforms LLMs into agents, and programmatic tool calling replaces rigid JSON calls with typed Python stubs that can chain and parallelize naturally. This paper evaluates PTC versus native JSON tool calling across 14 language models on BFCL v4, reporting parity or better performance in 11 of 14 models, a 10.6% GPT-5.6 family improvement, stronger parallel fan-out results, and better stability under context rot.

<!-- aaif-entry-id: 8379ed04 -->