研究与学习 5.0 · 必读 2026-09-01 · 论文

The Rise of Verbal Reinforcement Learning

论文首次统一提出语言强化学习 (VRL)范式,把以自然语言作为反馈信号的学习体系汇总为三个支柱:语言作为任务地面信号(定义目标/状态/奖励结构)语言作为推理过程反馈(测试时引导推理,不更新参数)语言作为学习信号(训练阶段改变参数)作为统一架构的 survey,为语言代理反馈领域提供了一个清晰的分类谱系

打开原文回到归档

The Rise of Verbal Reinforcement Learning

  • ID: 95ece884
  • 原文链接: https://arxiv.org/abs/2609.01597
  • PDF: https://arxiv.org/pdf/2609.01597v1
  • 作者: Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu
  • 日期: 2026-09-01
  • 更新: 2026-09-01
  • 分类: learning
  • 来源类型: paper
  • 标签: llm-agent, reinforcement-learning, verbal-feedback, survey
  • 质量评分: 5/5
  • 抓取时间: 2026-09-03T04:29:22Z

中文导读

论文首次统一提出语言强化学习 (VRL)范式,把以自然语言作为反馈信号的学习体系汇总为三个支柱:语言作为任务地面信号(定义目标/状态/奖励结构)语言作为推理过程反馈(测试时引导推理,不更新参数)语言作为学习信号(训练阶段改变参数)作为统一架构的 survey,为语言代理反馈领域提供了一个清晰的分类谱系

为什么值得关注

论文首次统一提出语言强化学习 (VRL)范式,把以自然语言作为反馈信号的学习体系汇总为三个支柱:语言作为任务地面信号(定义目标/状态/奖励结构)语言作为推理过程反馈(测试时引导推理,不更新参数)语言作为学习信号(训练阶段改变参数)作为统一架构的 survey...

关键信息

  • 论文标题:The Rise of Verbal Reinforcement Learning
  • 作者:Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu
  • arXiv:https://arxiv.org/abs/2609.01597
  • 发布时间:2026-09-01
  • arXiv 分类:cs.CL, cs.AI
  • 关联标签:llm-agent, reinforcement-learning, verbal-feedback, survey

English Abstract

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and reward structures; (2) \textbf{Language as Deliberative Feedback}, where natural language guides reasoning at test time without the need to update model parameters; (3) \textbf{Language as Learning Signal}, where language-based feedback shapes model parameters through training. Within each pillar, we synthesize representative work, distinguish key subcategories of approaches, and outline the distinct role language plays in shaping agent behavior. Together, this taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.

English Summary

Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, when verbal feedback takes effect in an agent's lifecycle and what it modifies, yielding three pillars: (1) Language as Grounding Signal, where language defines the task itself by specifying goals, states, and reward structures; (2) Language as Deliberative Feedback, where natural language guides reasoning at test time without the need to update model parameters; (3) Language as Learning Signal, where language-based feedback shapes model parameters through training.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。