IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
- ID: ffa418a5
- 原文链接: https://arxiv.org/abs/2609.10494
- PDF: https://arxiv.org/pdf/2609.10494v1
- 作者: Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
- 日期: 2026-09-09
- 更新: 2026-09-09
- 分类: infra
- 来源类型: paper
- 标签: evaluation, enterprise-ai, benchmark, serving-route, field-note
- 质量评分: 4/5
- 抓取时间: 2026-09-11T04:23:41
中文导读
18 个被审计基准都在给模型 ID 打分而企业实际部署的是系统;IB2 用能力绑定预检+失败计入分数+盲裁定让服务路径差异可测量同一权重换个服务路径就能从 77.38 走到 82.54
为什么值得关注
18 个被审计基准都在给模型 ID 打分而企业实际部署的是系统;IB2 用能力绑定预检+失败计入分数+盲裁定让服务路径差异可测量同一权重换个服务路径就能从 77.38 走到 82.54
关键信息
- 论文标题:IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier
- 作者:Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
- arXiv:https://arxiv.org/abs/2609.10494
- 发布时间:2026-09-09
- arXiv 分类:cs.CL, cs.AI, cs.LG
- 关联标签:evaluation, enterprise-ai, benchmark, serving-route, field-note
English Abstract
Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.
English Summary
Enterprises deploy systems, not checkpoints: usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. The IB2 protocol makes this measurement error reportable via a gold-blind capability-binding preflight, reliability-inclusive first-pass scoring, and structurally score-blind adjudication. Across eleven systems: two complete single-route runs on identical weights later failed distinct predicates of the binding gate while a third passed; four of seven suites saturate within a six-system band with spread almost entirely from governed database work and multi-tab joins; serving-arm choice moved one declared revision and precision from 77.38 to 82.54; excluding failed responses from denominators changes the point ordering.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。