Joint Optimization of Tool Creation and Use for Large Language Model Agents
- ID: 219e0681
- 原文链接: https://arxiv.org/abs/2608.24571
- PDF: https://arxiv.org/pdf/2608.24571v1
- 作者: Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee
- 日期: 2026-08-25
- 更新: 2026-08-25
- 分类: agents
- 来源类型: paper
- 标签: tool-use, reinforcement-learning, agents, qwen3, arxiv
- 质量评分: 4/5
中文导读
工具增强模型的能力上限受制于人写得出来的 API;现有工具创建系统在推理时给冻结模型发指令补缺,写工具的模型和使用工具的模型彼此脱节,没有信号保证产出的 schema 自己调得动。SMITH 用强化学习把创建与使用放进同一条策略:每个 rollout 要么是从少量示例写出工具的 build 任务,要么是在共享工具池上完成调用的 use 任务;schema、代码、结果三条独立奖励轴分别捕捉三类失败。4B Qwen3 在 13 个带精确验证器的程序化推理任务上训练后,held-out 任务宏平均 79.8,超过未训练的 30B-A3B 写手;TabMWP-Hard 40.4、域外 GQA 42.6(比同底座最强推理时基线高 7.6),全程未接触视觉或表格训练数据;它写的工具还能抬升 LFM-2.5-350M 与 Qwen3-30B-A3B 的表现。
为什么值得关注
SMITH 把会写工具与会用工具放进一条策略联合 RL,4B 模型 held-out 79.8 超过未训练 30B 写手,schema/代码/结果三轴奖励
English Abstract
Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches 79.8 macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA (+7.6 over the best same-backbone inference-time baseline), without any visual or tabular training data. Tools written by our 4B models also lifted the performance of LFM-2.5-350M and Qwen3-30B-A3B under same reasoning tasks.
Obsidian Notes
- Metadata and abstract fetched via
opencli arxiv paper 2608.24571 -f json(2026-08-27); response parsed list-or-dict tolerant. - 论文标题含 LLM Agents; SMITH=Schema-grounded Multi-task Iterative Tool Honing。
- 中文导读与价值判断锚定在论文摘要上,未补充摘要之外的实验细节。