Agent 与自动化 4.0 · 优秀 2026-08-20 · 论文

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

对 agent 技能归纳做了受控对照研究,比较两个维度:任务级 vs 子任务级归纳文本 vs 代码技能格式核心发现:任务级技能多数情况让 agent 跌破无记忆基线,子任务级则平均高于基线;文本技能比代码技能迁移更稳进一步提出特异性 抽象度合成的 skill utility 分数仅凭技能与任务描述无需任何任务执行即可预测迁移成功率,可在新任务运行前诊断技能记忆质量做 agent 记忆/技能库系统前的必读实证

打开原文回到归档

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

  • arXiv: 2608.20274
  • Authors: Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou
  • Published: 2026-08-20; categories: cs.AI, cs.CL

Abstract

Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and abstractness, which measures how evenly its relevance spreads across tasks. Neither property alone predicts task success, but their combined effect does, which we propose as a skill utility score. The score correlates consistently with task success when skills are transferred, and subtask-level and text skills score higher. Computing skill utility only needs the skills and task descriptions but not any task execution, so our score serves as a practical diagnostic of a skill memory before any new task runs.

Why it matters (AAIF scan)

对 agent 技能归纳做了受控对照研究,比较两个维度:任务级 vs 子任务级归纳文本 vs 代码技能格式核心发现:任务级技能多数情况让 agent 跌破无记忆基线,子任务级则平均高于基线;文本技能比代码技能迁移更稳进一步提出特异性 抽象度合成的 skill utility 分数仅凭技能与任务描述无需任何任务执行即可预测迁移成功率,可在新任务运行前诊断技能记忆质量做 agent 记忆/技能库系统前的必读实证

Source: https://arxiv.org/abs/2608.20274
Captured: 2026-08-22 (AAIF content-fetcher)