研究与学习 5.0 · 必读 2026-09-03 · 论文

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Ob...

两项预注册审计挑战 LLM judge 作为测量工具的前提假设:同名模型今天与明天给出同一读数52,988 次受审请求中,同窗口重复排名 Spearman 仅 0.400(要求 0.90),次日字节级重放也只有 0.78(要求 0.99),而执行记录显示请求本身无异常三个机制:标签到含义的映射本身引入与信号同强的偏差候选差距低于仪器噪声底七个数量级字节相同输入返回不同排名且精确置换读数放大该噪声预注册后续实验进一步排除出路:等待无改善(0.805 vs 0.800)四家供应商共享同一噪声底(中位数 0.74-0.88)批不变内核自托管仅在服务器空闲时有效作者给出三层快照同一性阶梯八条设计规则与报告清单:共享端点上的模型名不是冻结仪器,预注册评估必须先测量仪器本身2026-09-03 提交 cs.AI

打开原文回到归档

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

Source: https://arxiv.org/abs/2609.04198
Authors: Haoyaun Zhu, Jie Zhang
Published: 2026-09-03
Categories: cs.AI, cs.LG
PDF: https://arxiv.org/pdf/2609.04198

Abstract

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.

Key Points

  • LLM judges are measurement instruments resting on a rarely stated assumption: the same request sent to the same model name reads the same tomorrow. Two preregistered campaigns audited this with every threshold fixed in advance.
  • Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99 - each time with the execution record at ceiling.
  • Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound.
  • Preregistered follow-ups bounded the problem: waiting did not help on the days sampled (0.805 vs 0.800, replicated over five further days); four providers share the noise floor (medians 0.74-0.88, predicted by none of the exposed metadata fields); self-hosting on batch-invariant kernels helped only while the server was quiet.
  • The paper distills the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance.
  • Bottom line: on a shared endpoint, a model name is not a frozen instrument - a preregistered evaluation must measure its instrument before freezing any gate on it.

中文概要

两项预注册审计挑战 LLM judge 作为测量工具的前提假设:同名模型今天与明天给出同一读数52,988 次受审请求中,同窗口重复排名 Spearman 仅 0.400(要求 0.90),次日字节级重放也只有 0.78(要求 0.99),而执行记录显示请求本身无异常三个机制:标签到含义的映射本身引入与信号同强的偏差候选差距低于仪器噪声底七个数量级字节相同输入返回不同排名且精确置换读数放大该噪声预注册后续实验进一步排除出路:等待无改善(0.805 vs 0.800)四家供应商共享同一噪声底(中位数 0.74-0.88)批不变内核自托管仅在服务器空闲时有效作者给出三层快照同一性阶梯八条设计规则与报告清单:共享端点上的模型名不是冻结仪器,预注册评估必须先测量仪器本身2026-09-03 提交 cs.AI