模型与实验室 4.0 · 优秀 2026-08-18 · 论文

Chain-of-Experience for Continual LLM Improvement

人类能从经验中持续学习,而常规 LLM 评测忽略了模型在推理阶段通过迭代交互获得改进的能力论文提出 Chain-of-Experience(CoE)设定:模型通过与自身或环境反馈的迭代交互积累经验轨迹,形成超越零样本推理的持续改进回路作者用模型自反馈与环境信号等多种反馈机制实例化 CoE,并系统评估其持续改进效果

打开原文回到归档

Chain-of-Experience for Continual LLM Improvement

  • ID: e53f96ec
  • 原文链接: https://arxiv.org/abs/2608.18027
  • PDF: https://arxiv.org/pdf/2608.18027v1
  • 作者: Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
  • 日期: 2026-08-18
  • 更新: 2026-08-18
  • 分类: cs.CL
  • 来源类型: arxiv
  • 标签: llm, experience, test-time, continual-improvement, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-20T05:27:37Z

中文导读

人类能从经验中持续学习,而常规 LLM 评测忽略了模型在推理阶段通过迭代交互获得改进的能力论文提出 Chain-of-Experience(CoE)设定:模型通过与自身或环境反馈的迭代交互积累经验轨迹,形成超越零样本推理的持续改进回路作者用模型自反馈与环境信号等多种反馈机制实例化 CoE,并系统评估其持续改进效果

为什么值得关注

Chain-of-Experience:让 LLM 在测试时通过自反馈与环境信号积累经验轨迹,形成持续改进回路

Grounded note: evaluated across math, coding, and knowledge domains with 8 LLMs including GPT-5, Gemini-2.5 Pro, and Claude-4.5 Sonnet; iterative experience beat feedback-free baselines with a 5.6% overall improvement and 19% lower API cost, and most gains emerged early in the iterations.

关键信息

  • 论文标题: Chain-of-Experience for Continual LLM Improvement
  • 作者: Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
  • arXiv: https://arxiv.org/abs/2608.18027
  • 发布时间: 2026-08-18
  • arXiv 分类: cs.CL
  • Comment: H.T. and Y.F. contributed to this work equally
  • 关联标签: llm, experience, test-time, continual-improvement, arxiv

English Abstract

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.

English Summary

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。