LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
- source_url: https://arxiv.org/abs/2607.18110
- source_type: paper
- platform: arxiv
- author: Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei
- original_date: 2026-07-20
- added_date: 2026-07-22
- category: models
- tags: post-training, experiential-learning, llm-as-judge, open-ended, reward-hacking, arxiv
- quality_score: 4
- arxiv_id: 2607.18110
- arxiv_categories: cs.LG, cs.CL
- pdf_url: https://arxiv.org/pdf/2607.18110v1
摘要(中文)
开放任务上的 RL 常把 rubric 评价压成标量 reward,丢掉细粒度文本反馈。提出 Experiential Learning (EL):把反馈模型从 LLM-as-a-Judge 改造成 LLM-as-a-Coach,把对 on-policy 响应的评估蒸馏为可迁移经验知识,经 teacher 条件化并用 on-policy context distillation 内化到 policy。摘要称在 held-out 与未见开放任务上优于 rubric-based RL,泛化更好并缓解 reward hacking。适合非自动验真任务的后训练信号设计。
Summary (English)
RL on open-ended tasks often compresses rubric evaluation into a scalar reward, discarding rich textual feedback. Experiential Learning (EL) repurposes the feedback model from LLM-as-a-Judge into LLM-as-a-Coach: the coach distills assessments of on-policy responses into transferable experiential knowledge, which conditions a teacher and is internalized by the policy via on-policy context distillation. Across two policy families, with self or proprietary feedback, EL outperforms rubric-based RL on held-out and unseen open-ended tasks, generalizes better, and mitigates reward hacking.
One-liner
开放任务后训练:用 Coach 经验知识替代标量 Judge reward,减轻刷分。
Source body / metadata
arXiv abstract grounded intake for 2607.18110. PDF: https://arxiv.org/pdf/2607.18110v1
RL on open-ended tasks often compresses rubric evaluation into a scalar reward, discarding rich textual feedback. Experiential Learning (EL) repurposes the feedback model from LLM-as-a-Judge into LLM-as-a-Coach: the coach distills assessments of on-policy responses into transferable experiential knowledge, which conditions a teacher and is internalized by the policy via on-policy context distillation. Across two policy families, with self or proprietary feedback, EL outperforms rubric-based RL on held-out and unseen open-ended tasks, generalizes better, and mitigates reward hacking.
Obsidian evidence
- local_note: OpenClaw定时任务/论文流水线/2026-07-22-论文流水线.md
- intake_run: daily-intake-evening 2026-07-22