AI 编程 5.0 · 必读 2026-08-06 · 论文

Learning Globally Reusable Skills for Coding Agents

论文提出 GSE,把 coding agent 的技能演化从局部追加升级为全局技能关系图聚类合并和 replay 验证它显式维护 Skill Relation Graph 以保持技能库一致性,用 cluster-based consolidation 抽象可复用能力,并用 replay-driven verification 防止过拟合和行为回退;在 bug-revealing test generationfalse-positive bug report filtering 和内部工业 agent 上报告明显 F1 提升

打开原文回到归档

Learning Globally Reusable Skills for Coding Agents

Source: https://arxiv.org/abs/2608.06153
PDF: https://arxiv.org/pdf/2608.06153v1
Content fetched: 2026-08-08T15:34:45.962729+00:00
Grounding: OpenCLI source metadata/body plus Obsidian digest evidence

Metadata

  • Author(s): Chen Yang, Jiashuo Tian, Ziqi Wang, Xinyin Liu, Meiru Ye, Junjie Chen
  • Original date: 2026-08-06
  • Platform: arxiv
  • AAIF quality score: 5

中文摘要

论文提出 GSE,把 coding agent 的技能演化从局部追加升级为全局技能关系图、聚类合并和 replay 验证。它显式维护 Skill Relation Graph 以保持技能库一致性,用 cluster-based consolidation 抽象可复用能力,并用 replay-driven verification 防止过拟合和行为回退;在 bug-revealing test generation、false-positive bug report filtering 和内部工业 agent 上报告明显 F1 提升。

English Summary

GSE is a globalized skill-evolution framework for coding agents. It maintains a Skill Relation Graph, consolidates local updates into reusable capabilities, and uses replay-driven verification to avoid overfitting and regressions. Evaluations on OpenHands, mini-SWE-agent, bug-revealing test generation, false-positive filtering, and an internal industrial agent report broad precision/recall/F1 improvements.

Intake Rationale

技能库不能只局部追加,coding agent 需要关系图、合并和 replay 验证来保持可复用性。

Obsidian Evidence

  • Evidence note: /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/OpenClaw定时任务/论文流水线/2026-08-08-论文流水线.md
  • arXiv ID: 2608.06153
  • Primary category: cs.SE
  • Categories: cs.SE, cs.AI

Source Excerpt

GSE is a globalized skill-evolution framework for coding agents. It maintains a Skill Relation Graph, consolidates local updates into reusable capabilities, and uses replay-driven verification to avoid overfitting and regressions. Evaluations on OpenHands, mini-SWE-agent, bug-revealing test generation, false-positive filtering, and an internal industrial agent report broad precision/recall/F1 improvements.

<!-- aaif-entry-id: df568042 -->