基础设施 4.0 · 优秀 2026-09-09 · 论文

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

这篇来自企业一线的论文主张:企业部署的是"系统"而不是"权重",可用能力由权重服务路径精度输出契约和 harness 共同决定,但作者审计的 18 个基准全都在给宣传用的模型 ID 打分,这本身就是测量误差 论文提出 IB2 协议让这一误差变得可报告:gold-blind 的能力绑定预检先验证服务路径能否执行评测契约;首过计分把失败纳入分数把不支持的能力挡在外面;裁定环节结构性盲分 在 11 个系统的实测中:同一权重不同服务路径的两次完整运行分别未通过绑定门的不同谓词,而第三次通过了宣传的模型 ID 完全暴露不了这种差异;七个套件中四个在六系统区间内饱和,区分度几乎全部来自受治理的数据库操作和多表 join;服务侧选择让同一声明版本+精度从 77.38 变到 82.54;把失败响应从分母剔除会直接改变排名结论

打开原文回到归档

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

  • ID: ffa418a5
  • 原文链接: https://arxiv.org/abs/2609.10494
  • PDF: https://arxiv.org/pdf/2609.10494v1
  • 作者: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
  • 日期: 2026-09-09
  • 更新: 2026-09-09
  • 分类: infra
  • 来源类型: paper
  • 标签: evaluation, enterprise-ai, benchmark, serving-route, field-note
  • 质量评分: 4/5
  • 抓取时间: 2026-09-11T04:23:41

中文导读

18 个被审计基准都在给模型 ID 打分而企业实际部署的是系统;IB2 用能力绑定预检+失败计入分数+盲裁定让服务路径差异可测量同一权重换个服务路径就能从 77.38 走到 82.54

为什么值得关注

18 个被审计基准都在给模型 ID 打分而企业实际部署的是系统;IB2 用能力绑定预检+失败计入分数+盲裁定让服务路径差异可测量同一权重换个服务路径就能从 77.38 走到 82.54

关键信息

  • 论文标题:IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
  • 作者:Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
  • arXiv:https://arxiv.org/abs/2609.10494
  • 发布时间:2026-09-09
  • arXiv 分类:cs.CL, cs.AI, cs.LG
  • 关联标签:evaluation, enterprise-ai, benchmark, serving-route, field-note

English Abstract

Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.

English Summary

Enterprises deploy systems, not checkpoints: usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. The IB2 protocol makes this measurement error reportable via a gold-blind capability-binding preflight, reliability-inclusive first-pass scoring, and structurally score-blind adjudication. Across eleven systems: two complete single-route runs on identical weights later failed distinct predicates of the binding gate while a third passed; four of seven suites saturate within a six-system band with spread almost entirely from governed database work and multi-tab joins; serving-arm choice moved one declared revision and precision from 77.38 to 82.54; excluding failed responses from denominators changes the point ordering.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。