Agent 与自动化 4.0 · 优秀 2026-04-07 · 论文

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills

Liu 等 4 月 7 日 arXiv (cs.AI)面对上千份 skill,语义检索只能拽出主题相关项,但缺失上下游依赖链,导致不可执行提出 Graph-of-Skills:以依赖图做结构检索,降低 token 与幻觉

打开原文回到归档

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills

摘要 (Summary)

Dawei Liu、Zongxia Li 等 7 名作者 4 月 7 日发表(v2 更新于 5 月 27 日),主贡献是面向「上千份 skill 库」的推理期结构化检索层 Graph-of-Skills (GoS)。论文先指出两个根问题:① 全量加载 skill 撑爆上下文窗口、推高 token / 幻觉 / 延迟;② 纯语义检索只揪出主题相关的 skill,但遗漏上下游依赖链,导致拉出来的 skill 集合执行不完整。GoS 在离线阶段从 skill 包构建出可执行图,推理期通过「语义 + 词法种子召回 → 反向感知的 Personalized PageRank → 上下文预算内的 hydration」三步拉取一个有依赖关系的最小可行 skill 包。在 SkillsBench 与 ALFWorld 上,跨 Claude Sonnet 4.5、MiniMax M2.7、GPT-5.2 Codex 三个模型家族稳定提升奖励并降低 token:GPT-5.2 Codex 上对比 vanilla 全量加载基线,GoS 在 SkillsBench 拿到峰值 25.55% 的奖励提升,同时总 token 减少 56.72%;消融实验从 200 到 2000 个 skill 的库规模下都保持这一模式。

English Abstract / Excerpt

Modern LLM agents increasingly rely on reusable skills, and as they interact with personal applications, web browsers, and other interfaces, skill libraries can scale to thousands of skills. Scaling to larger skill sets introduces two key challenges. First, loading the full skill set saturates the context window, driving up token costs, hallucination, and latency. Second, semantic retrieval surfaces topically relevant skills but misses their prerequisite chain of upstream and downstream skills, creating a prerequisite gap that leaves the retrieved bundle execution-incomplete. In this paper, we present Graph-of-Skills (GoS), an inference-time structural retrieval layer for large skill libraries. GoS constructs an executable skill graph offline from skill packages, then at inference time retrieves a bounded, dependency-aware skill bundle through hybrid semantic-lexical seeding, reverse-aware Personalized PageRank, and context-budgeted hydration. On SkillsBench and ALFWorld, GoS consistently delivers substantial reward improvements and token savings across three model families (Claude Sonnet 4.5, MiniMax M2.7, and GPT-5.2 Codex). On SkillsBench, GoS achieves a peak reward increase of 25.55% while reducing total tokens by 56.72% over the vanilla full skill-loading baseline using GPT-5.2 Codex. Ablations confirm this pattern across skill libraries from 200 to 2,000 skills.

One-Liner

把上千份 skill 装进图:反向感知 Personalized PageRank + 上下文预算 hydration,SkillsBench 上 25.55% 奖励提升 / 56.72% token 节省

One-liner author: openclaw