TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
- ID: 03afb25a
- 原文链接: https://arxiv.org/abs/2608.26086
- PDF: https://arxiv.org/pdf/2608.26086v1
- 作者: Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang
- 日期: 2026-08-26
- 更新: 2026-08-26
- 分类: agents
- 来源类型: paper
- 标签: agents、benchmark、human-ai-collaboration、kaggle、arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-28T05:15:03Z
中文导读
把人和 agent 放进同一套 Kaggle 竞赛轨迹做版本级对比:134 个竞赛 4,465 条人类轨迹,其中 7 个竞赛同时由两类 agent 脚手架完成,得到 430 条配对人类轨迹与 207 条 agent 轨迹;每个代码版本都带分数时间戳和动作意图改动规模分数效果标签差距由此变得具体:专家在数据验证模型集成之间来回切换并重开放弃过的方案,而 Codex 反复调集成权重MLEvolve 原地改模型,都不按人类频率换方向从人类实践蒸馏的规划提示词能把被点名的行为拉向人类画像并提分,但整体努力画像仍是 agent 形状指令只关掉可指令化的那部分差距数据集schema标注器与抽取管线已开源
为什么值得关注
人机差距的归因数据集:把 benchmark 分数差拆成版本级的动作/意图/改动规模/分数效果标签,能直接定位 agent scaffold 卡在哪种低级行为循环(Codex 反复调集成权重、MLEvolve 原地改模型)。"指令只能关掉可指令化的那部分差距"这条结论,给提示词工程的收益边界提供了实证上界。
关键信息
- 论文标题:TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
- 作者:Jiarui Yan, Weiwei Sun, Sijie Li, Wenhan Li, Yiming Yang
- arXiv:https://arxiv.org/abs/2608.26086
- 发布时间:2026-08-26
- arXiv 分类:cs.LG, cs.AI
- 关联标签:agents、benchmark、human-ai-collaboration、kaggle、arxiv
English Abstract
Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4,465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.
English Summary
Pairs human and agent work on the same Kaggle competitions under one version-level schema: 4,465 human trajectories across 134 competitions, plus 430 paired human and 207 agent trajectories on seven competitions also worked by two agent scaffolds. Every code version carries score, timestamp, action, intent, edit size, and score effect. Experts alternate data work, validation, model changes, and ensembling, and reopen set-aside approaches; agent scaffolds collapse into narrow loops (Codex re-weights ensembles, MLEvolve mutates models in place) and rarely pivot. A planning prompt distilled from human practice shifts named behaviors toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。