Agent 与自动化 4.0 · 优秀 2026-09-22 · 文章

Unreal Agent: asynchronous tool-call harness cuts frontier-agent cost 40%

Unreal Labs 9 月 22 日发了 Unreal Agent 的设计与基准它的核心做法是 harness 不让 LLM 处在等 tool 返回的状态:每次发起 tool 调用就立即在会话日志里追加一个 in-progress 事件继续在后台执行,工具真正返回后再追加结果并触发下一次模型调用这个改动看起来仅是生命周期问题,但累积效果是单轮模型调用内可以塞更多并行重型工具单次 trial 减少 turns 数缓存命中率提升在 Terminal-Bench 4.0SWE-Atlas Codebase QnADeepSWE 1.1ALE-CLI 四个 benchmark 上用 GPT-6 Astra xhigh 跑:Terminal-Bench 通过率与 Codex 并列 57.9%...

打开原文回到归档

Unreal Agent: asynchronous tool-call harness cuts frontier-agent cost 40%

原文链接: https://unreallabs.ai/blog/unreal-agent/
作者: Unreal Labs
发布时间: 2026-09-22
来源笔记: OpenClaw定时任务/ClawFeed24小时高价值一览/2026-09-23

摘要(AAIF 提取)

Unreal Agent 重塑 harness 生命周期以给 LLM 更多并行重型工具、更少 turns 与更高缓存命中率。

机制

每次发起 tool 调用立即在会话日志追加一个 in-progress 事件、继续在后台执行,工具真正返回后再追加结果并触发下一次模型调用。

基准表现(GPT-6 Astra xhigh)

| Benchmark | 通过率 | 成本 | Codex 对比 | |-----------|----------|------|----------| | Terminal-Bench 4.0 | 57.9% | $1428 | 与 Codex 并列,便宜 39% | | DeepSWE 1.1 | 72.4% | $1367 | Codex 69.0% 、$1633,便宜 16% | | ALE-CLI | 30.0% | — | Codex 29.0%、Pi 29.0% |

可运行形态

Go library + 类似 claude -p 的 Runner + 与 Harbor 兼容的 benchmark runner。

来源信号