Agent 与自动化 4.0 · 优秀 2026-08-19 · 论文

Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning

指出验证文献里 "level" 至少混用了五种含义(粒度/抽象/风险层/系统栈层/ground truth 来源),提出统一轴线的元标准 VAL:验证规范从哪来判定能保证什么L0 是模型自我声明,L2 是有客观 ground truth 的正确性,L3/L4 是可判定系统的单性质/领域完备性,L5 在无限制场景下不可达核心是 completeness blind spot:替换式和采样式验证器能确认候选成立但无法证明没漏掉候选;完备性只在形式可刻画性质上可达,开放世界验证封顶在 L2适合做验证体系设计的前置阅读

打开原文回到归档

Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning

Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: verification granularity, concept abstraction, risk tier, system-stack layer, and the epistemic source of the ground truth. We propose Verification Autonomy Levels (VAL), a meta-standard that classifies any verification scheme along a single axis: where does the verification spec come from, and what does the verdict guarantee? VAL ranges from L0 (LLM self-declaration; no deterministic anchor) through L2 (objective ground truth; correctness only) to L3/L4 (decidable systems with single-property or domain-level completeness), with L5 impossible in the unrestricted case. Central to VAL is the completeness blind spot: substitution- and sampling-based verifiers can confirm that proposed candidates hold, but cannot prove that no candidate was missed. We further identify a dichotomy the literature has not stated: completeness is reachable only for formally specifiable properties, whereas empirical open-world verification (fact-checking, diagnosis) caps at anchored correctness (L2). We document this gap empirically across four domains (symbolic mathematics, behavior monitoring, medical diagnosis, and code generation, the last a reverse validation with predictions stated before evidence) and in the strongest formal-verification baseline in our survey, whose authors note the verifier focuses on the correctness of each step. We show the levels of granularity, concept hierarchy, risk, and system stack are orthogonal to VAL, resolving a systematic conflation across 17 surveyed papers. Code and full assessment are released as supplementary material.

Authors: Yajie Yin Published: 2026-08-19 Categories: cs.CL arXiv: 2608.19009

Source: https://arxiv.org/abs/2608.19009
Captured: 2026-08-21 (AAIF daily-intake-evening)