模型与实验室 4.0 · 优秀 2026-08-10 · 论文

Consilience for Verifier-Free Test-Time Scaling

论文揭示 confidence-based 无验证器测试时扩展(VF-TTS)在复杂任务上灾难性失效:均匀的高置信度往往意味着没有真正探索,反而偏爱"自信的错误答案",据此提出 consilience 框架通过组合式指标显式评估置信度的时间不对称性,主动惩罚早期高置信而强制要求终局确定性,在研究生级数学与自由格式代码生成上均显著优于现有 baseline

打开原文回到归档

Consilience for Verifier-Free Test-Time Scaling

  • ID: b8c369c6
  • Original: https://arxiv.org/abs/2608.09898
  • PDF: https://arxiv.org/pdf/2608.09898v1
  • Authors: Lecheng Kong, Like Hui, Haitao Mao, Jun Huan
  • Published: 2026-08-10
  • Updated: 2026-08-10
  • arXiv Categories: cs.CL, cs.LG
  • AAIF Category: models
  • Source Type: paper
  • Tags: test-time-scaling, reasoning, confidence, llm
  • Quality Score: 4/5
  • Fetched At: 2026-08-12T04:29:18.480369+00:00

Chinese Guide

论文揭示 confidence-based 无验证器测试时扩展(VF-TTS)在复杂任务上灾难性失效:均匀的高置信度往往意味着没有真正探索,反而偏爱"自信的错误答案",据此提出 consilience 框架通过组合式指标显式评估置信度的时间不对称性,主动惩罚早期高置信而强制要求终局确定性,在研究生级数学与自由格式代码生成上均显著优于现有 baseline

Why It Matters

论文揭示 confidence-based 无验证器测试时扩展(VF-TTS)在复杂任务上灾难性失效:均匀的高置信度往往意味着没有真正探索,反而偏爱"自信的错误答案",据此提出 consilience 框架通过组合式指标显式评估置信度的时间不对称性...

This content page is grounded in the existing AAIF entry plus arXiv metadata fetched with opencli arxiv paper. It does not add experimental claims beyond the arXiv abstract and entry summary.

Key Metadata

English Abstract

Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and tasks. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning, as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.

English Summary

Demonstrates a critical limitation of confidence-based verifier-free test-time scaling (VF-TTS): uniformly high confidence often signals a failure to explore rather than correctness, biasing selection toward confidently wrong answers. The proposed consilience framework operationalizes a "specific confidence trajectory pattern" exploratory branching (low initial confidence) followed by converging certainty (high final confidence) via a combinatorial metric that penalizes high initial confidence while strictly demanding final certainty. Across graduate-level mathematics and free-form code generation, consilience outperforms existing confidence-based VF-TTS baselines.

Obsidian Notes

  • Generated by the AAIF content-fetcher cron mode from arXiv metadata and the existing AAIF entry.
  • Chinese guide and relevance notes are anchored to the entry summary, arXiv abstract, authors, dates, categories, and links.
  • Canonical content path is content/{entry_id}.md under the repository root; openclaw/content/ is intentionally avoided.