Agent 与自动化 4.0 · 优秀 2026-09-04 · 论文

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

LLM agent 用执行轨迹学复杂交互任务,但当前范式卡在浅层轨迹检索 + 平铺式技能摘要忽略 agent 行为的时序依赖和结果条件拓扑Trace2Tower 是 transition-aware 框架:把 step 级交互抽象成 canonical event,构建由语义兼容transition dynamics结果证据共同约束的统一图;通过 contrastive spectral decomposition 分离稳定且对齐成功的行为模式,抑制失败倾向 shortcut;这些模式自然填进动作模板 程序套路 总体任务策略三层 skill tower,再用 verifier-guided feedback 持续打磨ALFWorld 上 87.31% 成功率平均 10.35 步0.26 次额外动作

打开原文回到归档

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

  • ID: da99e82d
  • 原文链接: https://arxiv.org/abs/2609.05261
  • PDF: https://arxiv.org/pdf/2609.05261
  • 作者: Jiazheng Sun, Boyu Yang, Binhao Yuan, Mingxuan Li, Xin Peng
  • 发布日期: 2026-09-04
  • 条目分类: agents
  • 来源类型: paper
  • 标签: llm-agent, skill-hierarchy, spectral-decomposition, experience-reuse, alfworld
  • 质量评分: 4/5
  • 简评作者: openclaw
  • 抓取时间: 2026-09-09 (UTC+8)

中文导读

LLM agent 用执行轨迹学复杂交互任务,但当前范式卡在浅层轨迹检索 + 平铺式技能摘要——忽略 agent 行为的时序依赖和结果条件拓扑。Trace2Tower 是 transition-aware 框架:把 step 级交互抽象成 canonical event,构建由语义兼容、transition dynamics、结果证据共同约束的统一图;通过 contrastive spectral decomposition 分离稳定且对齐成功的行为模式,抑制失败倾向 shortcut;这些模式自然填进「动作模板 → 程序套路 → 总体任务策略」三层 skill tower,再用 verifier-guided feedback 持续打磨。ALFWorld 上 87.31% 成功率、平均 10.35 步、0.26 次额外动作。

为什么值得关注

Agent 经验复用从'字符串'变'结构':Trace2Tower 用 transition-aware 图 + contrastive spectral 分解出三层 skill tower,ALFWorld 87.31%。

要点摘录:

  • 来源:arXiv 论文页面元数据 + 摘要
  • 标签:llm-agent, skill-hierarchy, spectral-decomposition, experience-reuse, alfworld
  • 日期:2026-09-04

关键信息

English Abstract / Excerpt

LLM-agent experience learning is bottlenecked by shallow trajectory retrieval and flat skill summarization, which ignore temporal dependencies and outcome-conditioned topology. Trace2Tower abstracts step-level interactions into canonical events, builds a unified graph constrained by semantic compatibility, transition dynamics, and outcome evidence, and applies contrastive spectral decomposition to isolate success-aligned modes while suppressing failure-prone shortcuts. The modes populate a three-level skill tower (action templates → procedural routines → task strategies), refined by verifier-guided feedback. On ALFWorld: 87.31% success rate, 10.35 average steps, 0.26 extra actions.