EDGE / KOPA-Bench: Multi-Step Tool-Calling over Korean Open Public APIs
Source: https://arxiv.org/abs/2609.05395
Authors: Dain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh, Kyuseong Lim
Published: 2026-09-04
Categories: cs.AI, cs.CL
PDF: https://arxiv.org/pdf/2609.05395v1
Abstract
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.
Key Findings
- Benchmark: KOPA-Bench — 145 real-world multi-step tool-calling tasks over live Korean government APIs, filling the gap that open-source on-premise agents underperform in this regime.
- Synthesis method: EDGE (Execution-grounded Dynamic Graph for tool-calling data synthEsis) builds a graph of how each tool's output can feed another's input; keeps only the links that succeed when actually called against the live APIs; traverses these verified links to synthesize executable multi-step trajectories.
- Result: A 9B model fine-tuned via GRPO on EDGE data nearly matches the untuned 27B model from the same family, with substantial gains on KOPA-Bench and BFCL.
- Note: Accepted to EMNLP 2026 Industry Track. 30 pages, 7 figures, 26 tables.
中文概要
本文提出 KOPA-Bench(145 个面向韩国公开政府 API 的多步工具调用任务)以及数据合成方法 EDGE。设计要点:数据主权要求下、机构需部署开源本地 LLM 智能体与现实政府 API 连接,但现有开源模型在多步场景下表现不足。EDGE 在生成过程中以现实执行结果为准:构建工具输出输入图,仅保留实际调用能成功的边,以这些被验证的路径生成可执行多步轨迹。细调后以 GRPO 训练的 9B 模型几乎贴近同系 27B 未调模型,同时在 KOPA-Bench 与 BFCL 上都有较大提升。