Agent 与自动化 4.0 · 优秀 2026-09-18 · 论文

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

专业图形设计是长程 agent 任务:产物结构化可编辑动作相互依赖且无可靠程序化 oracle框架让冻结的前沿模型操作 230+ 工具的专业设计软件,外置自然语言技能形式的程序记忆随经验累积精炼:横向为未覆盖的复现子任务获取新程序,纵向依据成败执行修订既有程序,配对重放门只放行修失败而不伤已见成功的修改1406 条真实用户 brief1869 条自动评分轨迹五轮迭代,无权重更新无人工标注,技能库从 76 涨到 139,Claude-Sonnet-4 上 GenEval2 执行成功率 72.7%99.3%,对无技能 agent 胜率 61.8%/67.6%

打开原文回到归档

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Source: <https://arxiv.org/abs/2609.22086&gt;
Authors: Hongyang Du, Lan Yan, Christian Flores, Asim Kadav
Published: 2026-09-18
Categories: cs.AI, cs.CV
arXiv: 2609.22086

Abstract

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality), with 61.8% and 67.6% win rates against the no-skill agent across four specialized design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. We further show the two mechanisms are effective in combination: on 200 held-out briefs from user-traffic benchmark, widening or deepening alone reaches a 49.4% / 48.6% win rate over the no-skill agent, while their combination reaches 58.5% (p = 0.025). Procedural memory offers a practical route to continual adaptation of agents under noisy, unverifiable feedback.

Summary

Continual adaptation for agentic graphic design without any weight updates: a frozen frontier model drives professional design software via 230+ tools, while an external procedural memory of natural-language skills evolves from real user traffic. The bank widens (new skills for recurring uncovered subtasks) and deepens (revising existing skills against their own success/failure replays); a matched replay gate admits only failure-repairing changes that do not regress observed successes. Five rounds over 1,406 real user briefs / 1,869 auto-graded trajectories grow the skill bank from 76 to 139 and lift GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 generation-quality points), with 61.8% / 67.6% win rates over the no-skill agent across four design benchmarks on Sonnet-4 / Opus-4.6. On 200 held-out briefs, widening alone 49.4%, deepening alone 48.6%, combined 58.5% win rate (p=0.025). Positioning: procedural memory is a practical route to continual adaptation under noisy, unverifiable feedback.

摘要

面向 agentic 图形设计的持续适应框架:冻结的前沿模型操作 230+ 工具的专业设计软件,外置自然语言技能库作为程序记忆,从真实用户流量中演化——横向为反复出现且未覆盖的子任务新增技能,纵向依据既有技能自身的成败执行记录做修订;配对重放门只放行“修复失败且不回退已见成功”的修改。五轮迭代、1,406 条真实用户 brief、1,869 条自动评分轨迹,零权重更新、零人工标注,技能库从 76 增至 139;Claude-Sonnet-4 上 GenEval2 执行成功率 72.7%→99.3%(生成质量 +11.99 分),四个专项设计基准对无技能 agent 胜率 61.8% / 67.6%(Sonnet-4 / Opus-4.6)。200 条留出 brief 上,横向、纵向单独胜率 49.4% / 48.6%,组合达 58.5%(p=0.025)。结论:程序记忆是在噪声、不可验证反馈下实现 agent 持续适应的可行路线。