Agent 与自动化 4.0 · 优秀 2026-10-01 · 论文

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

RPG(Reconstruct, Practice, Go Real):不更新模型权重的机器人执行系统自主改进框架从离线数据集中识别操作能力,在仿真中构造相关练习任务;练习时用执行反馈仿真器特权状态和数据集视频诊断失败,进而新增可复用符号技能改进已有技能修订系统提示词,候选改动先做跨任务评估再决定保留测试时由多模态 LLM 用最终系统提示词和技能库协调感知与控制在 22 个操作任务的 held-out 初始化上,任务成功率从第一轮练习后的 28.6% 升到 15 轮后的 95.0%,优于所有对比基线2026-10-01,cs.RO/cs.AI/eess.SY

打开原文回到归档

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Content-fetcher entry · 2026-10-04 · awesome-ai-field-notes
  • URL: https://arxiv.org/abs/2610.02204
  • PDF: https://arxiv.org/pdf/2610.02204
  • Source: arxiv · Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, et al.
  • Original Date: Thu, 1 Oct 2026 17:59:50 UTC (20,742 KB)
  • Added: 2026-10-04
  • Category: agents
  • Tags: robotics, embodied-agents, skill-library, self-improvement, simulation, manipulation
  • Quality Score: 4

中文摘要

RPG(Reconstruct, Practice, Go Real):不更新模型权重的机器人执行系统自主改进框架从离线数据集中识别操作能力,在仿真中构造相关练习任务;练习时用执行反馈仿真器特权状态和数据集视频诊断失败,进而新增可复用符号技能改进已有技能修订系统提示词,候选改动先做跨任务评估再决定保留测试时由多模态 LLM 用最终系统提示词和技能库协调感知与控制在 22 个操作任务的 held-out 初始化上,任务成功率从第一轮练习后的 28.6% 升到 15 轮后的 95.0%,优于所有对比基线2026-10-01,cs.RO/cs.AI/eess.SY

English Abstract

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: this https URL

One-liner

RPG(Reconstruct, Practice, Go Real):不更新模型权重的机器人执行系统自主改进框架从离线数据集中识别操作能力,在仿真中构造相关练习任务.

注:本文件为 content-fetcher cron 回填;opencli arxiv 拉取失败(COMMAND_EXEC),改以 opencli web read 抓取 arXiv abs 页面,仅依据可见的标题/作者/提交时间/分类/摘要生成,未补充摘要之外的实验细节。