Agent Lightning v1.0: Towards Harnessed Agentic RL
- ID: 958a9b26
- 原文链接: https://arxiv.org/abs/2608.17528
- PDF: https://arxiv.org/pdf/2608.17528
- 作者: Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo
- 日期: 2026-08-18
- 更新: 2026-08-18
- 分类: cs.AI, cs.SE
- 来源类型: arxiv
- 标签: agents, rl, harness, arxiv, paper
- 质量评分: 5/5
- 抓取时间: 2026-08-20T15:44:49Z
中文导读
Agent Lightning v1.0 用约 3500 行代码把部署时的 agent harness 直接接入 RL 后训练,提出 harnessed agentic RL 范式:环境交互循环由 harness 持有,训练器只观察 LLM 请求-响应对序列。论文系统处理了重分词、样本合并、优势计算、损失归一化与后端调度五个工程问题,覆盖指令跟随、搜索、编码三类 agent,并提供完整可复现的编码 agent RL 管线。只用 6K 训练样本与中等算力,Qwen3.5-9B 在 SWE-bench Verified 上从 41.8% 提升到 56.4%,绝对提升 14.6 个点。该架构思路已被 verl Uni-Agent、AReaL 2.0、slime、Polar 等框架采用。
为什么值得关注
把部署 harness 接进 RL 训练循环的轻量开源框架,6K 样本让 9B 模型 SWE-bench 提升 14.6 个点
English Abstract
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.
Obsidian 证据
- 来源 digest: 论文流水线 2026-08-20(评分 9.2)。
- 元数据与摘要经 opencli arxiv paper 核对;中文导读锚定摘要陈述的事实与数字。