研究与学习 5.0 · 必读 2026-08-06 · 论文

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Inform...

论文把 AIVAT 方差缩减与 continuously monitored confidence sequences 结合,提出 AV-AIVAT,用于在不破坏统计保证的前提下提前停止两智能体强弱评估作者在 15 个 LLM agent 配置71439 手 HUNL paired hands 上报告:原始 outcomes 相比 AIVAT-corrected outcomes 需要约 74 倍中位样本量才能达到同等停止条件,并区分 asymptotic screening 与 finite-sample certification

打开原文回到归档

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

Source: https://arxiv.org/abs/2608.06362
PDF: https://arxiv.org/pdf/2608.06362v1
Content fetched: 2026-08-09T04:20:50.800307+00:00
Grounding: OpenCLI arXiv metadata and abstract

Metadata

  • Author(s): Boning Li, Yu Chen, Longbo Huang
  • Original date: 2026-08-06
  • Platform: arxiv
  • AAIF quality score: 5
  • Primary category: cs.GT
  • Categories: cs.GT, cs.AI, cs.CL, cs.LG, cs.MA
  • Comment: 34 pages, 5 figures

中文摘要

论文把 AIVAT 方差缩减与 continuously monitored confidence sequences 结合,提出 AV-AIVAT,用于在不破坏统计保证的前提下提前停止两智能体强弱评估作者在 15 个 LLM agent 配置71439 手 HUNL paired hands 上报告:原始 outcomes 相比 AIVAT-corrected outcomes 需要约 74 倍中位样本量才能达到同等停止条件,并区分 asymptotic screening 与 finite-sample certification

English Summary

AV-AIVAT combines AIVAT variance reduction with continuously monitored confidence sequences so agent evaluations can stop as soon as evidence is sufficient without invalidating the stated confidence level. Across 15 LLM-agent configurations and 71,439 paired HUNL hands, raw outcomes required a median 74x as many hands as AIVAT-corrected outcomes under the AsympCS target, with separate discussion of exact finite-sample certification via EB-CS.

Intake Rationale

Agent evaluation should optimize for auditable early stopping, not only fixed-budget aggregate win rates.

Source Metadata

Source Abstract

Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-information games through conditional mean-zero corrections, by a median $54\times$ across 15 LLM agent configurations spanning 71,439 paired Heads-Up No-Limit Hold'em (HUNL) hands, but does not say when to stop. We combine AIVAT with continuously monitored Confidence Sequences (CSs) into anytime-valid AIVAT (AV-AIVAT), whose online value model learns only from past games so that no game scores its own correction. At the nominal 95\% level and a target precision of $\pm1$ Big Blind, raw outcomes need a median $74\times$ as many hands as AIVAT-corrected outcomes to stop under the Asymptotic CS (AsympCS). Exact finite-sample certification uses the Empirical-Bernstein CS (EB-CS), which needs an independently justified bound on corrected payoffs. We establish such a bound structurally for Leduc hold'em and characterize a width floor set by the CS's bet cap and that bound, which governs how much of a variance gain becomes earlier stopping; the descriptive HUNL EB-CS runs show a median $1.37\times$ stopping-time ratio. AV-AIVAT turns variance reduction into efficient, auditable early stopping while separating asymptotic screening from exact certification, so an evaluation can stop the moment its evidence suffices and hand a third party everything needed to recheck the verdict at that very stopping time.

<!-- aaif-entry-id: 94bb097c -->