模型与实验室 4.0 · 优秀 2026-07-20 · 论文

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

开放任务上的 RL 常把 rubric 评价压成标量 reward,丢掉细粒度文本反馈提出 Experiential Learning (EL):把反馈模型从 LLM-as-a-Judge 改造成 LLM-as-a-Coach,把对 on-policy 响应的评估蒸馏为可迁移经验知识,经 teacher 条件化并用 on-policy context distillation 内化到 policy摘要称在 held-out 与未见开放任务上优于 rubric-based RL,泛化更好并缓解 reward hacking适合非自动验真任务的后训练信号设计

打开原文回到归档

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

  • source_url: https://arxiv.org/abs/2607.18110
  • source_type: paper
  • platform: arxiv
  • author: Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei
  • original_date: 2026-07-20
  • added_date: 2026-07-22
  • category: models
  • tags: post-training, experiential-learning, llm-as-judge, open-ended, reward-hacking, arxiv
  • quality_score: 4
  • arxiv_id: 2607.18110
  • arxiv_categories: cs.LG, cs.CL
  • pdf_url: https://arxiv.org/pdf/2607.18110v1

摘要(中文)

开放任务上的 RL 常把 rubric 评价压成标量 reward,丢掉细粒度文本反馈。提出 Experiential Learning (EL):把反馈模型从 LLM-as-a-Judge 改造成 LLM-as-a-Coach,把对 on-policy 响应的评估蒸馏为可迁移经验知识,经 teacher 条件化并用 on-policy context distillation 内化到 policy。摘要称在 held-out 与未见开放任务上优于 rubric-based RL,泛化更好并缓解 reward hacking。适合非自动验真任务的后训练信号设计。

Summary (English)

RL on open-ended tasks often compresses rubric evaluation into a scalar reward, discarding rich textual feedback. Experiential Learning (EL) repurposes the feedback model from LLM-as-a-Judge into LLM-as-a-Coach: the coach distills assessments of on-policy responses into transferable experiential knowledge, which conditions a teacher and is internalized by the policy via on-policy context distillation. Across two policy families, with self or proprietary feedback, EL outperforms rubric-based RL on held-out and unseen open-ended tasks, generalizes better, and mitigates reward hacking.

One-liner

开放任务后训练:用 Coach 经验知识替代标量 Judge reward,减轻刷分。

Source body / metadata

arXiv abstract grounded intake for 2607.18110. PDF: https://arxiv.org/pdf/2607.18110v1

RL on open-ended tasks often compresses rubric evaluation into a scalar reward, discarding rich textual feedback. Experiential Learning (EL) repurposes the feedback model from LLM-as-a-Judge into LLM-as-a-Coach: the coach distills assessments of on-policy responses into transferable experiential knowledge, which conditions a teacher and is internalized by the policy via on-policy context distillation. Across two policy families, with self or proprietary feedback, EL outperforms rubric-based RL on held-out and unseen open-ended tasks, generalizes better, and mitigates reward hacking.

Obsidian evidence

  • local_note: OpenClaw定时任务/论文流水线/2026-07-22-论文流水线.md
  • intake_run: daily-intake-evening 2026-07-22