Agent 与自动化 4.0 · 优秀 2026-09-30 · 论文

Turbo Harness: Instance-Adaptive Harness Optimization

已有 harness 自动优化产出的是对所有任务实例统一套用的单一全局 harness,但平均最优不等于每个实例最优Turbo Harness 回收一次已完成的全局 harness 优化过程产生的工件,总结成结构化 playbook,再训练一个 harness editor 利用这份先验优化经验对全局 harness 生成实例专属补丁;推理时 editor 根据实例与 playbook 搭出定制 harness 供执行模型使用跨交互 agent 任务软件工程长程终端任务的 7 个基准一致超过现有 harness 优化基线与 2609.40303 的极简结论对照:这是把优化开销前置到 playbook 的另一条路线

打开原文回到归档

Turbo Harness: Instance-Adaptive Harness Optimization

  • ID: a6b6e297
  • 原文链接: https://arxiv.org/abs/2609.40330
  • PDF: https://arxiv.org/pdf/2609.40330v1
  • 作者: Tunyu Zhang, Hao Wang, Kai Xu, Dimitris N. Metaxas
  • 发布时间: 2026-09-30
  • 更新: 2026-09-30
  • 分类: agents
  • 来源类型: paper
  • 标签: agent-harness, self-improvement, instance-adaptive, benchmark
  • 质量评分: 4/5
  • 抓取时间: 2026-10-02T04:30:58Z

中文导读

已有 harness 自动优化产出的是对所有任务实例统一套用的单一全局 harness,但平均最优不等于每个实例最优Turbo Harness 回收一次已完成的全局 harness 优化过程产生的工件,总结成结构化 playbook,再训练一个 harness editor 利用这份先验优化经验对全局 harness 生成实例专属补丁;推理时 editor 根据实例与 playbook 搭出定制 harness 供执行模型使用跨交互 agent 任务软件工程长程终端任务的 7 个基准一致超过现有 harness 优化基线与 2609.40303 的极简结论对照:这是把优化开销前置到 playbook 的另一条路线

为什么值得关注

把全局 harness 优化产物做成 playbook,训练 editor 按任务实例打补丁,7 个基准超静态全局 harness

The abstract describes Turbo Harness recycling artifacts from a completed global harness-optimization run into a structured playbook, training a harness editor to emit instance-specific harness patches at inference time, and consistently outperforming harness-optimization baselines across seven benchmarks spanning agent tasks, software engineering, and terminal tasks.

关键信息

  • 论文标题:Turbo Harness: Instance-Adaptive Harness Optimization
  • 作者:Tunyu Zhang, Hao Wang, Kai Xu, Dimitris N. Metaxas
  • arXiv: https://arxiv.org/abs/2609.40330
  • 发布时间:2026-09-30
  • arXiv 分类:cs.AI
  • 关联标签:agent-harness, self-improvement, instance-adaptive, benchmark

英文摘要

Automating the search for effective harnesses is an important step toward enabling agents to recursively self-improve. Existing harness optimizations typically produce a single global harness that is applied uniformly across task instances. However, a harness that works well on average may not be optimal for every instance. We introduce Turbo Harness, a framework that can adapt a globally optimized harness to each instance by reusing information generated during the original optimization process. Specifically, Turbo Harness recycles artifacts produced during a completed global harness optimization run, and summarizes them into a structured playbook. We train a harness editor to leverage this prior optimization experience to generate instance-specific patches to the global harness. At inference time, the editor uses the instance and the playbook to construct a tailored harness in which the execution model operates. Through numerical experiments, we show that Turbo Harness consistently outperforms existing harness optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.

English Summary

Turbo Harness adapts a globally optimized agent harness to each task instance by recycling artifacts from the completed global optimization run into a structured playbook; a trained harness editor leverages this prior optimization experience to generate instance-specific harness patches at inference time. It consistently outperforms existing harness optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.

Obsidian Notes

  • 本页为内容补齐(content backfill):条目早已入库,本次补写内容页。
  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锺定在条目已有摘要与本次拉取的论文摘要上,未添加摘要之外的实验细节。