Source: Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments · platform: arxiv · authors: João Meneses dos Santos, Arlindo L. Oliveira · date: 2026-09-16
TL;DR(中文摘要)
语言 agent 在交互环境里仍然脆弱:成功依赖长程状态跟踪、合法动作执行与失败步骤恢复。本文扩展双过程 agent SwiftSage(快动作提议器 + 慢规划器),加入两个模块化认知扩展:AMM(Adaptive Memory Module,显著性门控的情景存储与触发式检索)和 SRM(Self-Reflection Module,有界的执行时校验与纠正干预)。两者都以 feature-flag 形式挂在同一执行底座上,可在 ScienceWorld 上做受控消融。四种配置(baseline、+AMM、+SRM、完整系统)中,完整系统取得最佳平均最终得分 64.62、成功率 43.17%、成功步效率 19.33 步,其中 SRM 是最强的独立贡献项。作者据此判断:该设定下的主要瓶颈是执行时控制;情景记忆要在运行循环稳定之后才最有用。
Summary (English)
Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. The authors extend SwiftSage, a dual-process agent combining a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptive Memory Module (AMM) for salience-gated episodic storage and trigger-driven retrieval, and a Self-Reflection Module (SRM) for bounded execution-time validation and corrective intervention. Both modules are feature-flagged extensions over the same execution substrate, enabling controlled ablations on ScienceWorld. Across four configurations — baseline, baseline+AMM, baseline+SRM, and the full system — the full system achieves the best mean final score (64.62), success rate (43.17%), and successful-step efficiency (19.33 steps), while SRM is the strongest standalone contributor. The results suggest execution-time control is the dominant bottleneck in this setting, while episodic memory becomes most useful once the runtime loop is stabilized.
入库依据
opencli arxiv paper 2609.19128 -f json 抓取 arXiv 元数据(标题/作者/摘要/分类/日期),内容由摘要直接支撑。
补充
arXiv 2609.19128,primary cs.AI,categories: cs.AI, cs.LG, cs.MA;comment: 13 pages, 1 figure;submitted 2026-09-16。