AI 编程 4.0 · 优秀 2026-09-18 · 论文

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

用 RL 训练编码 agent 需要带可靠验证器的多样任务;现有方法多依赖 issuecommit 等开发工件,任务覆盖面受限CodeMidas 是一条 agentic 流水线,仅以源码为任务特定输入,把既有代码库中已实现的功能转成可执行 RL 环境:agent 探索已实现功能并形成行为规格基于原代码执行构造测试再通过执行检查与重复解题 rollout 验证过滤候选任务产出 5545 个训练任务,覆盖 3185 个开源代码库23 种语言15 个技术领域用 GRPO 在这些任务上训练 MiMo-V2.5,在覆盖 issue 解决等五个不同基准上全面提分

打开原文回到归档

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

  • ID: 39508e48
  • 原文链接: https://arxiv.org/abs/2609.22068
  • PDF: https://arxiv.org/pdf/2609.22068v1
  • 作者: Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo
  • 日期: 2026-09-18
  • 更新: 2026-09-18
  • 分类: coding
  • arXiv 分类: c, s, ., A, I
  • 来源类型: paper
  • 标签: rl-environments, agentic-coding, grpo, task-synthesis, verifiers
  • 质量评分: 4/5
  • 抓取时间: 2026-09-22T04:22:47Z

中文导读

用 RL 训练编码 agent 需要带可靠验证器的多样任务;现有方法多依赖 issuecommit 等开发工件,任务覆盖面受限CodeMidas 是一条 agentic 流水线,仅以源码为任务特定输入,把既有代码库中已实现的功能转成可执行 RL 环境:agent 探索已实现功能并形成行为规格基于原代码执行构造测试再通过执行检查与重复解题 rollout 验证过滤候选任务产出 5545 个训练任务,覆盖 3185 个开源代码库23 种语言15 个技术领域用 GRPO 在这些任务上训练 MiMo-V2.5,在覆盖 issue 解决等五个不同基准上全面提分

为什么值得关注

CodeMidas 只用源码就把既有代码库功能变成可执行 RL 环境,5545 任务/3185 仓库/23 语言,GRPO 训练全面提升

关键信息

  • 论文标题:CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
  • 作者:Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo
  • arXiv:https://arxiv.org/abs/2609.22068
  • PDF:https://arxiv.org/pdf/2609.22068v1
  • 发布时间:2026-09-18
  • 最近更新:2026-09-18
  • arXiv 主分类:cs.AI
  • arXiv 全部分类:c, s, ., A, I
  • 评论:N/A
  • 关联标签:rl-environments, agentic-coding, grpo, task-synthesis, verifiers

English Abstract

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.

English Summary

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。