Agent 与自动化 4.0 · 优秀 2026-08-25 · 论文

Joint Optimization of Tool Creation and Use for Large Language Model Agents

工具增强模型的能力上限受制于人写得出来的 API;现有工具创建系统在推理时给冻结模型发指令补缺,写工具的模型和使用工具的模型彼此脱节,没有信号保证产出的 schema 自己调得动SMITH 用强化学习把创建与使用放进同一条策略:每个 rollout 要么是从少量示例写出工具的 build 任务,要么是在共享工具池上完成调用的 use 任务;schema代码结果三条独立奖励轴分别捕捉三类失败4B Qwen3 在 13 个带精确验证器的程序化推理任务上训练后,held-out 任务宏平均 79.8,超过未训练的 30B-A3B 写手;TabMWP-Hard 40.4域外 GQA 42.6(比同底座最强推理时基线高 7.6),全程未接触视觉或表格训练数据;它写的工具还能抬升 LFM-2.5-350M 与 Qwen3-30B-A3B 的表现

打开原文回到归档

Joint Optimization of Tool Creation and Use for Large Language Model Agents

中文导读

工具增强模型的能力上限受制于人写得出来的 API;现有工具创建系统在推理时给冻结模型发指令补缺,写工具的模型和使用工具的模型彼此脱节,没有信号保证产出的 schema 自己调得动。SMITH 用强化学习把创建与使用放进同一条策略:每个 rollout 要么是从少量示例写出工具的 build 任务,要么是在共享工具池上完成调用的 use 任务;schema、代码、结果三条独立奖励轴分别捕捉三类失败。4B Qwen3 在 13 个带精确验证器的程序化推理任务上训练后,held-out 任务宏平均 79.8,超过未训练的 30B-A3B 写手;TabMWP-Hard 40.4、域外 GQA 42.6(比同底座最强推理时基线高 7.6),全程未接触视觉或表格训练数据;它写的工具还能抬升 LFM-2.5-350M 与 Qwen3-30B-A3B 的表现。

为什么值得关注

SMITH 把会写工具与会用工具放进一条策略联合 RL,4B 模型 held-out 79.8 超过未训练 30B 写手,schema/代码/结果三轴奖励

English Abstract

Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA (+7.6 over the best same-backbone inference-time baseline), without any visual or tabular training data. Tools written by our 4B models also lifted the performance of LFM-2.5-350M and Qwen3-30B-A3B under same reasoning tasks.

Obsidian Notes

  • Metadata and abstract fetched via opencli arxiv paper 2608.24571 -f json (2026-08-27); response parsed list-or-dict tolerant.
  • 论文标题含 LLM Agents; SMITH=Schema-grounded Multi-task Iterative Tool Honing。
  • 中文导读与价值判断锚定在论文摘要上,未补充摘要之外的实验细节。