Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
- ID: ab652e80
- 原文链接: https://arxiv.org/abs/2608.16831
- PDF: https://arxiv.org/pdf/2608.16831v1
- 作者: Minh-Ha Nguyen, Cathy Shyr
- 日期: 2026-08-17
- 更新: 2026-08-17
- 分类: learning
- 来源类型: paper
- 标签: in-context-learning, policy-iteration, human-feedback, reinforcement-learning, clinical-ai, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-19T12:47:36Z
中文导读
把后训练 RL 的'评估改进'循环移植到 in-context learning:预训练 LM 只充当执行基底,持久修订发生在版本化的自然语言策略与工具集上LM critic 与临床专家复盘全项推理和工具使用轨迹,定位复发性失败并形成候选修订;专家可重新解读证据并保留采纳与回滚的最终权力,修订后以 Recall@1/Recall@5 验证成效在超罕见病基准上,PIHF 派生策略让 GPT-5.4 的 Recall@1 提升 32.7 个百分点Qwen3.6-35B 提升 31.1 个百分点,覆盖 3-49B 激活参数的专有与开源执行器'不动权重迭代提示策略'路线的有力证据
为什么值得关注
策略迭代+人类反馈:迭代版本化自然语言策略,GPT-5.4 罕见病 Recall@1 提升 32.7 个百分点
Grounded in the arXiv abstract (published 2026-08-17; categories: cs.AI, cs.CL).
关键信息
- 论文标题: Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
- 作者: Minh-Ha Nguyen, Cathy Shyr
- arXiv: https://arxiv.org/abs/2608.16831
- 发布时间: 2026-08-17
- arXiv 分类: cs.AI, cs.CL
- 关联标签: in-context-learning, policy-iteration, human-feedback, reinforcement-learning, clinical-ai, arxiv
English Abstract
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, while Recall@1 and Recall@5 validate outcomes after candidate execution. Across cumulative ablations and ultra-rare-disease benchmarks, a PIHF-derived policy improved Recall@1 in one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. Gains were 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B, a difference of 1.7 points. These results support the feasibility of using pretrained language models as fixed-weight execution substrates for expert-guided policy development in rare-disease diagnosis.
English Summary
PIHF (Policy Iteration with Human Feedback) transfers the recurrent evaluate-and-improve structure of generalized policy iteration into in-context learning: a pretrained language model serves as the execution substrate while persistent revision targets a versioned natural-language policy and tool set. An LM critic and a clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, with Recall@1 and Recall@5 validating outcomes after candidate execution...
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- metadata source: opencli