研究与学习 3.0 · 值得看 2026-09-02 · X

Zen_with_AI 第 17 期 论文速览:mid-training 蒸馏长时程 agentAI4Science 基准自主编码 harness

Zen_with_AI 第 17 期小路和你读 AI 论文挑了 4 篇: Meta+UW+Princeton 的 Knowledge Distillation During Mid-Training Favor Reasoning over Factual Recall发现标准 KL 蒸馏偏推理压事实,提出 Switch Distillation 相对 NTP 拿到 1.61-1.71 推理增益且几乎不丢事实; Salesforce 的 Parsing the Streamappend-only 事件账本折叠成 typed run state,观察者侧 token 降到 1/14-1/15成本降到 1/5-1/7准确率 0.48 0.85-0.87...

打开原文回到归档

Zen_with_AI 第 17 期 论文速览:mid-training 蒸馏、长时程 agent、AI4Science 基准、自主编码 harness

Source: https://x.com/Zen_with_AI/status/2095188034899832954
Author: Zen_with_AI
Published: 2026-09-02

中文摘要

Zen_with_AI 第 17 期『小路和你读 AI 论文』挑了 4 篇:① Meta+UW+Princeton 的 Knowledge Distillation During Mid-Training Favor Reasoning over Factual Recall——Switch Distillation 相对 NTP 拿到 1.61-1.71× 推理增益且几乎不丢事实;② Salesforce 的 Parsing the Stream——append-only 事件账本折叠成 typed run state,观察者侧 token 降到 1/14-1/15、成本降到 1/5-1/7、准确率 0.48 → 0.85-0.87;③ UCSD Jingbo Shang 实验室的 SCILAWS-BENCH——用 118 个真实论文任务、约 800 万数据点区分『拟合好』和『科学站得住』;④ Harness-of-Harness——coding-agent 外面再套规划-编码-测试迭代循环,GameCraft/FrontierSWE/ProgramBench 上相对独立 harness 平均 +52.25%、最高 +82.86%。

Obsidian 证据摘要

来源:OpenClaw定时任务/X书签消化/2026-09-12-X书签消化.md 第 6 条
"9 月 2 日 Zen_with_AI 第 17 期'小路和你读 AI 论文'挑了 4 篇值得读的:① Meta+UW+Princeton 的 Knowledge Distillation During Mid-Training Favor Reasoning over Factual Recall……"