Agent 与自动化 4.0 · 优秀 2026-08-20 · 论文

MidTool: Mid-training Data Synthesis for Agentic Tool Use

MidTool 是面向 agentic tool use 的 mid-training 开放语料构建管线:融合大规模网页PDF代码数据与来自真实工具 APIMCP 技能文档落地工作流的合成监督,教模型识别工具可供性从上下文填充参数组合调用流程并从信息不全中恢复在 Qwen3-4B/8B-Base 上先中训练再 SFT/RL,BFCLtau2-BenchMCP Universe 一致提升结论明确:通用工具使用和推理能力一样值得专门的 mid-training,而不是全靠后训练补课

打开原文回到归档

MidTool: Mid-training Data Synthesis for Agentic Tool Use

  • arXiv: 2608.20314
  • Authors: Fengqing Jiang, Yite Wang, Boyi Liu, Zhaoyang Wang, Canwen Xu, Zhewei Yao, Radha Poovendran, Yuxiong He
  • Published: 2026-08-20; categories: cs.AI

Abstract

Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown that targeted mid-training can strengthen reasoning-intensive abilities such as math and science, and can also improve agentic capabilities in software-engineering settings. In this work, we study the parallel but less explored agentic capability: general tool use. We present MidTool, an open corpus construction pipeline for agentic tool-use mid-training that combines large-scale web, PDF, and code data with synthesized supervision from real-world tool APIs, MCP skills, and document-grounded workflows. MidTool is designed to teach models how to recognize tool affordances, ground arguments from context, compose tool call workflow, and recover from incomplete information. We mid-train Qwen3-4B-Base and Qwen3-8B-Base on MidTool-Mix, and then apply follow-up post-training with both supervised fine-tuning and reinforcement learning. Compared with baselines, MidTool-Mix consistently improves downstream performance under both SFT and RL on BFCL, tau2-Bench, and MCP Universe. These results suggest that general tool use, like other important LLM capabilities, benefits from dedicated mid-training rather than being left entirely to post-training.

Why it matters (AAIF scan)

MidTool 是面向 agentic tool use 的 mid-training 开放语料构建管线:融合大规模网页PDF代码数据与来自真实工具 APIMCP 技能文档落地工作流的合成监督,教模型识别工具可供性从上下文填充参数组合调用流程并从信息不全中恢复在 Qwen3-4B/8B-Base 上先中训练再 SFT/RL,BFCLtau2-BenchMCP Universe 一致提升结论明确:通用工具使用和推理能力一样值得专门的 mid-training,而不是全靠后训练补课

Source: https://arxiv.org/abs/2608.20314
Captured: 2026-08-22 (AAIF content-fetcher)