Agent 与自动化 4.0 · 优秀 2026-09-04 · 论文

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

Konstantin GrotovValentin Malykh 2026-09-04 提交,EMNLP 2026 Industry Track 收录软件工程 agent 失败代价高,但失败信号通常要等执行回测才发现Speculative Uncertainty(SU)从反向用 speculative decoding拿信号:不碰 logits/权重/激活,用一个小开放权重 draft 模型对 agent 已经生成的轨迹做一次前向评分,再把 reasoning span 和 action span 分开算 phase-aware 特征,校准到可验证目标落地为 pre-execution veto gate:在 Qwen3-Coder-480B 和 Claude 3.5 Sonnet 上,执行错误率降 68 个百分点,token 成本降 1419%...

打开原文回到归档

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

Source: https://arxiv.org/abs/2609.05274
Authors: Konstantin Grotov, Valentin Malykh
Published: 2026-09-04
Categories: cs.LG
PDF: https://arxiv.org/pdf/2609.05274v1

Abstract

LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small open-weight draft model scores the agent's already-generated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3-Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6-8 percentage points and token cost by 14-19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.

中文概要

Konstantin Grotov、Valentin Malykh 2026-09-04 提交,EMNLP 2026 Industry Track 收录。软件工程 agent 失败代价高,但失败信号通常要等执行、回测才发现。Speculative Uncertainty(SU)从『反向用 speculative decoding』拿信号:不碰 logits/权重/激活,用一个小开放权重 draft 模型对 agent 已经生成的轨迹做一次前向评分,再把 reasoning span 和 action span 分开算 phase-aware 特征,校准到可验证目标。落地为 pre-execution veto gate:在 Qwen3-Coder-480B 和 Claude 3.5 Sonnet 上,执行错误率降 6–8 个百分点,token 成本降 14–19%;OOD 任务不需要重训就能迁移,跨 agent 模型可泛化。对资源受限的端侧 agent orchestrator 是直接可借鉴的工程模式。

一句话

Speculative Uncertainty 用 draft-model 一次前向给 agent 失败信号,执行前 veto gate 降 6–8 个百分点执行错误率