研究与学习 4.0 · 优秀 2026-09-22 · 论文

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

coding agent 必须先产出可解析的工具调用,harness 才能执行本文证明本地推理服务栈本身会混入工具使用评测结果:Ollama 默认 tools= 请求按模型被静态模板开关拦截,Phi-3 与 Gemma-3 在推理前即被拒绝,且 rejection/重试耗尽若不作为结构化失败元数据保留,会被下游误判为模型不发调用报出 0% 保真率对被放行的模型,在原生通道之外补充文本工具列表能恢复大部分测得保真率;而统一走文本协议反而降低原生支持工具调用的 Llama-3.2 的保真率Ollama/llama.cpp/vLLM/SGLang 四栈对同一请求处理各不相同;约束解码消除解析失败但可能不终止;按轮合并与按实例统计的估计差最高约 55 个点作者给出把 serving 行为纳入评测协议的检查清单

打开原文回到归档

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

中文导读

coding agent 必须先产出可解析的工具调用,harness 才能执行本文证明本地推理服务栈本身会混入工具使用评测结果:Ollama 默认 tools= 请求按模型被静态模板开关拦截,Phi-3 与 Gemma-3 在推理前即被拒绝,且 rejection/重试耗尽若不作为结构化失败元数据保留,会被下游误判为模型不发调用报出 0% 保真率对被放行的模型,在原生通道之外补充文本工具列表能恢复大部分测得保真率;而统一走文本协议反而降低原生支持工具调用的 Llama-3.2 的保真率Ollama/llama.cpp/vLLM/SGLang 四栈对同一请求处理各不相同;约束解码消除解析失败但可能不终止;按轮合并与按实例统计的估计差最高约 55 个点作者给出把 serving 行为纳入评测协议的检查清单

为什么值得关注

本地工具使用评测的隐藏混淆变量:Ollama 静态开关四栈行为不一统计口径差 55 个点,serving 栈应纳入评测协议

关键信息

  • 论文标题:Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation
  • 作者:Lijuan Tang, Yuemeng Zheng
  • arXiv:https://arxiv.org/abs/2609.26693
  • 发布时间:2026-09-22
  • 更新时间:2026-09-22
  • arXiv 分类:cs.CL, cs.AI, cs.SE
  • 可选注释:9 pages, 4 figures, 3 tables. Accepted at the 2nd Workshop for Research on Agent Language Models (REALM) @ EMNLP 2026
  • 关联标签:arxiv, tool-use, evaluation, serving, ollama, vllm, paper

English Abstract

A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: some models are accepted and return calls as text, some return native tool_calls, while Phi-3 and Gemma-3 are rejected before inference. In our harness, rejection and retry exhaustion are not preserved as structured failure metadata, so downstream analysis can misclassify them as model non-calls and naively report 0% fidelity. Adding a text tool list while retaining the native channel recovers much of the measured fidelity for accepted models, whereas a uniform text protocol reduces fidelity for Llama-3.2, which has native tool-call support. Cross-stack probes on Ollama, llama.cpp, vLLM, and SGLang show different handling of the same request. Constrained decoding removes parse failures but can induce non-termination, and turn-pooled versus per-instance estimates differ by up to about 55 points. We conclude with a checklist for treating serving behavior as part of the evaluation protocol.

English Summary

A coding agent must emit a valid tool call before the harness can execute its chosen action. This paper shows measured tool-use outcomes can depend on the local serving stack rather than model behavior alone. In Ollama, the default tools= request is gated per model by a static template flag: Phi-3 and Gemma-3 are rejected before inference, and harness rejection/retry exhaustion may be misclassified as model non-calls, naively reporting 0% fidelity. Adding a text tool list alongside the native channel recovers much of the measured fidelity for accepted models, while a uniform text protocol reduces fidelity for Llama-3.2, which has native tool-call support....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • 本页面属于 content-fetcher 能动下的高分筻补东生成,不修改原条目字段,仅产出可读的中英双语页面。