Agent 与自动化 4.0 · 优秀 2026-09-29 · 论文

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

研究测试时 AI4AI:在 Builder 与 Target 模型权重均冻结的前提下,让 Builder 学会为 Target 构建更好的执行环境(agent harness)核心是 Meta-Skill:从 Target 在开发集上的执行反馈中提炼出何时需要支持提供什么资源的原则库,再用冻结的技能库为未见任务构建 harness在 Harness-Bench 与 NewtonBench 上,全库 meta-skills 将 macro-average 性能较无技能构建提升 8.95 个百分点,较直接把技能库塞给 Target 提升 12.02 个百分点,说明把经验转成可执行的环境支持比直接给经验更有效

打开原文回到归档

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

  • ID: 008268b1
  • 原文链接: https://arxiv.org/abs/2609.38143
  • PDF: https://arxiv.org/pdf/2609.38143
  • 作者: C, h, e, n, g, , Q, i, a, n, ,, , K, u, n, l, u, n, , Z, h, u, ,, , B, e, i, b, i, n, , L, i, ,, , Z, h, e, n, h, a, i, l, o, n, g, , W, a, n, g, ,, , H, e, n, g, , J, i
  • 日期: 2026-09-29
  • 更新: 2026-09-29
  • 分类: agents
  • 来源类型: paper
  • 标签: arxiv, agent-harness, meta-learning, test-time-compute, ai4ai
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T04:22:00+00:00

中文导读

研究测试时 AI4AI:在 Builder 与 Target 模型权重均冻结的前提下,让 Builder 学会为 Target 构建更好的执行环境(agent harness)核心是 Meta-Skill:从 Target 在开发集上的执行反馈中提炼出何时需要支持提供什么资源的原则库,再用冻结的技能库为未见任务构建 harness在 Harness-Bench 与 NewtonBench 上,全库 meta-skills 将 macro-average 性能较无技能构建提升 8.95 个百分点,较直接把技能库塞给 Target 提升 12.02 个百分点,说明把经验转成可执行的环境支持比直接给经验更有效

为什么值得关注

Builder从执行反馈中提炼meta-skills为Target构建agent harness,双冻结下比直接给经验多拿12个百分点

Grounded in the abstract: full meta-skill library lifts macro-average performance by 8.95 points over no-skill harness construction and 12.02 points over handing the Target the skill library directly, on Harness-Bench and NewtonBench, with both model weights frozen.

关键信息

  • 论文标题:Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
  • 作者:C, h, e, n, g, , Q, i, a, n, ,, , K, u, n, l, u, n, , Z, h, u, ,, , B, e, i, b, i, n, , L, i, ,, , Z, h, e, n, h, a, i, l, o, n, g, , W, a, n, g, ,, , H, e, n, g, , J, i
  • arXiv: https://arxiv.org/abs/2609.38143
  • 发布时间:2026-09-29
  • arXiv 分类:c, s, ., A, I, ,, , c, s, ., C, L, ,, , c, s, ., L, G
  • 关联标签:arxiv, agent-harness, meta-learning, test-time-compute, ai4ai

English Abstract

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction, and 12.02 points over direct delivery of the same bank to the Target. These results highlight the value of translating experience into executable support. Gains when the same model serves both roles further suggest a path to system level self-improvement through learning to build better environments.

English Summary

Studies test-time AI4AI: with both Builder and Target model weights frozen, a Builder learns to construct better execution environments (agent harnesses) for a Target. Meta-Skills are principles distilled from the Target's execution feedback on a dev set specifying when support is needed and what resources to provide; the frozen skill bank then builds harnesses for unseen tasks. On Harness-Bench and NewtonBench, full-bank meta-skills improve macro-average performance by 8.95 points over no-skill construction and 12.02 points over directly handing the bank to the Target.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。