AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization
- ID: 167d2e4b
- 原文链接: https://arxiv.org/abs/2608.21292
- PDF: https://arxiv.org/pdf/2608.21292v1
- 作者: Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
- 日期: 2026-08-21
- 更新: 2026-08-21
- 分类: agents
- 来源类型: paper
- 标签: llm-agent, skill-learning, rl-finetuning, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-25T05:20:39+00:00
中文导读
面向 LLM 代理的技能学习新框架 AUSO:作者指出技能在策略演进中扮演三种角色先是可学习的知识再支撑能力形成最后仅在改善单步决策时被调用;现有方法要么把技能放在模型外要么完全内化,或靠噪声很大的任务级成功率在两者间二选一AUSO 在训练全程用统一目标把技能学习与技能使用合为一个渐进按动作加权的优化过程:早期融合教师指导与环境回报,后期逐步迁移到仅在有利于单个动作时才调用技能指导,不再对同一轨迹内所有动作一视同仁
为什么值得关注
技能先内化后按动作选择性调用,学习与使用一个目标统一 — Per the abstract, AUSO reports consistent gains on ALFWorld, WebShop and SearchQA plus out-of-distribution generalization over competitive baselines. For teams training agent policies with skill libraries, it is a concrete recipe for unifying skill internalization and utilization under one RL objective, instead of choosing between external skill stores, full internalization, or noisy task-level success signals.
关键信息
- 论文标题:AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization
- 作者:Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
- arXiv:https://arxiv.org/abs/2608.21292
- 发布时间:2026-08-21
- arXiv 分类:cs.AI
- 关联标签:llm-agent, skill-learning, rl-finetuning, arxiv
English Abstract
Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation, and finally be invoked only when they improve individual decisions. Existing methods rarely model this lifecycle. They either keep skills outside the model, fully internalize them, or select among internalization and utilization objectives through noisy task-level success rates. Such designs fragment training and assign uniform importance to actions within the same trajectory, even though skill guidance may help some decisions while distracting others. To solve these problems, we introduce AUSO (Action-level Unified Skill Optimization), which unifies skill learning and skill use through a progressive, action-aware optimization process. At the beginning of training, AUSO jointly learns from teacher guidance and environmental outcomes, enabling the policy to acquire foundational skills without losing task-oriented feedback. It subsequently emphasizes outcome-based policy optimization to consolidate autonomous problem-solving ability. As the policy matures, AUSO evaluates each sampled action under both skill-conditioned and skill-free contexts. The resulting action-level information signal is coupled with the trajectory outcome advantage, allowing beneficial skill-sensitive actions to receive stronger updates and harmful ones to be suppressed. Therefore, skills gradually transition from an external source of supervision into decision knowledge whose utilization is adapted to its action-level benefit, while reinforcement learning remains the shared backbone across all stages. Experiments on ALFWorld, WebShop, and SearchQA show that AUSO consistently improves agent performance and out-of-distribution generalization over competitive baselines.
English Summary
AUSO unifies skill learning and skill use for LLM agents via progressive, action-aware optimization: skills serve three lifecycle roles (learnable knowledge, capability substrate, selective decision aid), and training blends teacher guidance with environmental outcomes early, then invokes skill guidance only when it improves individual actions instead of weighting all actions in a trajectory equally.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。