基础设施 4.0 · 优秀 2026-09-28 · 论文

TokenCast: Forecasting Token Consumption During LLM Agent Execution

同一任务在 LLM 智能体多次执行间 token 消耗可差一个数量级,且总消耗在执行前难以预测TokenCast 为每个执行段学习可组合的成本表示(自身消耗+带来的上下文增长),相邻段复合成累计估计,把前段上下文被后继每次调用重读的隐性成本纳入预测;执行中随新证据刷新预测且不增加额外 LLM 调用,SWE-bench Verified 上平均累计预测耗时 32.8ms/次跨 4 个任务套件6 个智能体模型的 96 种组合中,对最强对比方法的 MAE 平均降低 14.5%;离线预算控制重放中,在相同轨迹完成率下比固定预算策略平均省 21.3% token代码已开源(github.com/DEFENSE-SEU/TokenCast)

打开原文回到归档

TokenCast: Forecasting Token Consumption During LLM Agent Execution

  • ID: 9ce4e19e
  • 原文链接: https://arxiv.org/abs/2609.35760
  • PDF: https://arxiv.org/pdf/2609.35760v2
  • 作者: Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
  • 日期: 2026-09-28
  • 更新: 2026-09-29
  • 分类: infra
  • 来源类型: paper
  • 标签: $token-cost, $agent-observability, $budget-control, $swe-bench
  • 质量评分: 4/5
  • 抓取时间: 2026-09-30T04:26:15Z

中文导读

同一任务在 LLM 智能体多次执行间 token 消耗可差一个数量级,且总消耗在执行前难以预测TokenCast 为每个执行段学习可组合的成本表示(自身消耗+带来的上下文增长),相邻段复合成累计估计,把前段上下文被后继每次调用重读的隐性成本纳入预测;执行中随新证据刷新预测且不增加额外 LLM 调用,SWE-bench Verified 上平均累计预测耗时 32.8ms/次跨 4 个任务套件6 个智能体模型的 96 种组合中,对最强对比方法的 MAE 平均降低 14.5%;离线预算控制重放中,在相同轨迹完成率下比固定预算策略平均省 21.3% token代码已开源(github.com/DEFENSE-SEU/TokenCast)

为什么值得关注

面向智能体基础设施的成本可观测性:执行中即可预测总 token 消耗,预算控制重放平均省 21.3% token,对配额管理与成本归因有直接工程价值。

关键信息

  • 论文标题: TokenCast: Forecasting Token Consumption During LLM Agent Execution
  • 作者: Chaoqian Ouyang, Ling Yue, Libin Zheng, Hanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, Shimin Di
  • arXiv: https://arxiv.org/abs/2609.35760
  • 发布时间: 2026-09-28
  • arXiv 分类: cs.LG, cs.AI, cs.SE
  • 关联标签: token-cost, agent-observability, budget-control, swe-bench

English Abstract

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.

English Summary

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。