Agent 与自动化 4.0 · 优秀 2026-08-12 · 论文

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

回答LLM agent 的控制流何时值得搬上 GPU:模型调用之间的小型确定性迁移(路由结果状态更新发出副作用)在何种条件下有足够并发度,以及 GPU 算出的路由决策留在设备端会改变什么用固定划分份额 F精确离线份额 P*局部上界 U在线达成份额 A 形式化 ready-cohort 边界,零服务时间情形下可用专用动态规划精确算 P*实测 36 种配置全部快于 host 往返(行中位比 1.19x-2.39x),14,557,440 次批量调用与独立实现的 host oracle 全一致对 SoC 内 CPUNPU 往返的端侧 agent 编排有直接平移价值

打开原文回到归档

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

中文导读

回答“LLM agent 的控制流何时值得搬上 GPU”:模型调用之间的小型确定性迁移(路由结果、状态更新、发出副作用)在何种条件下有足够并发度,以及 GPU 算出的路由决策留在设备端会改变什么。用固定划分份额 F、精确离线份额 P*、局部上界 U、在线达成份额 A 形式化 ready-cohort 边界,零服务时间情形下可用专用动态规划精确算 P*。实测 36 种配置全部快于 host 往返(行中位比 1.19x-2.39x),14,557,440 次批量调用与独立实现的 host oracle 全一致。对 SoC 内 CPU↔NPU 往返的端侧 agent 编排有直接平移价值。

为什么值得关注

agent 控制流 device-resident 的收益边界:36 种配置全快于 host 往返,千万级调用与 oracle 全一致。

收录理由:给端侧 agent 编排的“控制流上不上 NPU”问题提供了形式化边界与全配置实测

Abstract

LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enough concurrent work for GPU execution, and what changes when a GPU-computed route decision remains on device. We formalize the ready-cohort boundary using fixed-partition share F, exact offline share P*, local upper bound U, and online achieved share A. Under zero service time, unlimited capacity, and equal relative launch deadlines, a specialized dynamic program computes P* exactly. In a stationary Poisson replay of one pinned 851-session public trace panel, the primary condition at 100,000 target active sessions, K=256, and a 50 ms launch deadline gives F=30.19%, P*=43.00%, and U=45.85%. Exact packing recovers 81.83% of the opportunity lost at fixed window boundaries. The outcome-derived route key is a conditioning proxy, not proof of executable identity. A separate mechanism study keeps a GPU-computed binary decision on device instead of returning four bytes to the host and redispatching. Across four named GPU placements, the device-resident path is faster in all 36 configurations; within-placement row-median ratios range from 1.19x to 2.39x. Across both admissible mechanisms, all 14,557,440 tested batched invocations match a separately implemented host oracle. A fixed nested device graph that removes no host decision is slower in all 60 configurations across five placements. Together, the studies establish two measurable gates for GPU agent control: deadline-feasible cohort supply and observation placement. A joined finite online runtime is required to measure A, CPU displacement, and service-level benefit.

元数据

Obsidian 证据:OpenClaw定时任务/论文流水线/2026-08-17-论文流水线.md(2026-08-17 周度回顾);元数据抓取自 opencli arxiv paper 2608.12123(2026-08-17)。