AI 编程 4.0 · 优秀 2026-08-31 · 论文

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

论文把工业级 LLM 后训练视作一个棕地维护场景:团队接手已部署 checkpoint, 在固定算力与数据配比预算下做定向改进,且不能破坏既有能力维护对象正变成 dataware 用一份受控的后训练 mixture 来决定模型行为,通过 bounded mixture patch 而非全量重训来更新作者从工业代码生成改进实践中归纳出三大挑战:零和式 mixture 设计yield 作为硬约束指标端到端集成的不确定性;并报告 yield-engineered patch 让 CodeForces pass@1 提升 2.59LiveCodeBench v6 pass@1 提升 6.11,均来自同一基座 checkpoint 的 16 次随机评测

打开原文回到归档

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

  • ID: 5313a3b4
  • 原文链接: https://arxiv.org/abs/2608.31102
  • PDF: https://arxiv.org/pdf/2608.31102v1
  • 作者: Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen, Ahmed E. Hassan
  • 日期: 2026-08-31
  • 更新: 2026-08-31
  • 分类: coding
  • 来源类型: paper
  • 标签: llm-agent, post-training, data-engineering, industry
  • 质量评分: 4/5
  • 抓取时间: 2026-09-02T04:23:36Z

中文导读

这篇来自软件工程(SE)社区工业界的论文(作者含 Ahmed E. Hassan 组)把 LLM 后训练重新表述为一个"棕地维护"问题:团队接手的是已部署 checkpoint,必须在固定算力与数据配比预算下落地定向改进,且不能让其他能力回退。被维护的工件被称作 dataware——模型行为由一份精选的后训练数据配比(mixture)治理,更新方式是有界的 mixture patch,而不是推倒重训。

论文提炼了三个反复出现的实践难点:零和的 mixture 设计(一份配比内此消彼长)、yield(蒸馏数据转化率)成为约束性指标、以及不确定条件下的端到端集成。案例来自一次工业级代码生成改进:把教师蒸馏更多转化为可用训练数据的干预,使被采纳的监督数据量提升 2.84 倍(同一 solution teacher、每个候选问题 4 次求解)。主评估中,yield 工程化后的 patch 使 CodeForces pass@1 +2.59(pass@3 +3.11),held-out LiveCodeBench v6 pass@1 +6.11(pass@3 +8.05),每个条件从同一固定 checkpoint 做 16 次随机评估、结果统计显著,内部 AIME/MATH 回归套件保持在容忍范围内。

为什么值得关注

这把后训练从"配方收藏"拉回到"工程学科"视角:真正可复用的不是某个 one-off recipe,而是治理 dataware 的维护纪律(预算约束、回归护栏、yield 指标)。对在真实业务里做模型迭代的团队,这是少见的把工业约束写清楚的一手材料。

关键信息

  • 论文标题:LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
  • 作者:Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen, Ahmed E. Hassan
  • arXiv:https://arxiv.org/abs/2608.31102
  • 发布时间:2026-08-31
  • arXiv 分类:cs.SE, cs.AI, cs.LG
  • 关联标签:llm-agent, post-training, data-engineering, industry

English Abstract

Industrial post-training is a brownfield regime. Teams inherit a deployed checkpoint and must land targeted improvements under fixed compute and mixture budgets without regressing the rest. The maintained artifact is increasingly dataware: behavior governed by a curated post-training mixture, updated via bounded mixture patches rather than clean-slate retraining. From an industrial code-generation improvement effort, we offer a maintainer's perspective on why this work is hard in practice, distilling three recurring challenges, zero-sum mixture design, yield as the binding metric, and end-to-end integration under uncertainty, and arguing that progress depends less on one-off recipes than on an engineering discipline for programming dataware. In our case study, interventions that raised the conversion of teacher distillation into usable training data increased accepted supervision by 2.84 times while using the same solution teacher and four solution attempts per candidate problem. In our primary evaluation, the yield-engineered patch improved CodeForces pass@1 by +2.59 points (+3.11 pass@3) and held-out LiveCodeBench v6 pass@1 by +6.11 (+8.05 pass@3), all statistically significant across 16 stochastic evaluations of each benchmark from one fixed checkpoint per condition, with internal AIME and MATH regression suites within tolerance.

English Summary

Industrial post-training is treated as a brownfield regime: teams inherit a deployed checkpoint and land targeted improvements under fixed compute and mixture budgets via bounded mixture patches. The paper distills three recurring challenges (zero-sum mixture design, yield as the binding metric, end-to-end integration under uncertainty), reports a case where yield-engineered interventions raised accepted distilled supervision 2.84x, and improved CodeForces pass@1 by +2.59 and held-out LiveCodeBench v6 pass@1 by +6.11 with statistical significance across 16 stochastic evaluations per condition.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。