Agent 与自动化 4.0 · 优秀 2026-08-12 · 论文

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

论文发布 VAKRA 基准:跨 62 个领域8000+ 可执行 API,从多样 API 交互结构化 API 多跳推理多源检索 三个难度阶梯评估企业场景下的代理能力现有基准多与单一能力隔离评估,VAKRA 将多跳推理资源调度出错恢复打包在一个可重现的任务集中

打开原文回到归档

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

  • ID: 39879de4
  • 原文链接: https://arxiv.org/abs/2608.12282
  • PDF: https://arxiv.org/pdf/2608.12282v1
  • 作者: Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
  • 日期: 2026-08-12
  • 更新: 2026-08-12
  • 分类: agents
  • 来源类型: paper
  • 标签: benchmark, tool-use, multi-hop, enterprise-agents, rag, cs.ai
  • 质量评分: 4/5
  • 抓取时间: 2026-08-14T04:19:48Z

中文导读

论文发布 VAKRA 基准:跨 62 个领域8000+ 可执行 API,从多样 API 交互结构化 API 多跳推理多源检索 三个难度阶梯评估企业场景下的代理能力现有基准多与单一能力隔离评估,VAKRA 将多跳推理资源调度出错恢复打包在一个可重现的任务集中

为什么值得关注

VAKRA 把多跳 API 推理与多源检索打包成一个企业代理可重现评测集

关键信息

  • 论文标题:VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
  • 作者:Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor
  • arXiv:https://arxiv.org/abs/2608.12282
  • 发布时间:2026-08-12
  • arXiv 分类:cs.AI
  • 关联标签:benchmark, tool-use, multi-hop, enterprise-agents, rag, cs.ai

English Abstract

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA

English Summary

VAKRA introduces a benchmark of 8,000+ executable APIs across 62 domains with three difficulty tiers: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source retrieval. It targets enterprise agents that have to combine API calls and document lookup under tool-use policies, an evaluation that prior benchmarks treat in isolation. The release gives a single reproducible suite covering routing, multi-hop reasoning, and error recovery.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。