Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
- ID: 00e22c9d
- 原文链接: https://arxiv.org/abs/2608.31076
- PDF: https://arxiv.org/pdf/2608.31076v1
- 作者: Xuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, Shumin Deng
- 日期: 2026-08-31
- 更新: 2026-08-31
- 分类: agents
- 来源类型: paper
- 标签: llm-agent, scientific-research, evaluation, rubric
- 质量评分: 4/5
- 抓取时间: 2026-09-02T04:23:36Z
中文导读
AutoSciRub(zjunlp)是一个"先评估、后改进"框架,面向自主科研 agent。论文指出:开放式科研任务往往没有写清楚要做什么分析、用什么方法、什么算成功,导致 agent 漏掉关键分析、用错方法、或得出证据不足的结论。AutoSciRub 的做法是在执行研究之前,先把欠规范的指令分解成原子科学目标,结合相关文献和任务可见数据,合成一套具体、可执行、可验证的评估 rubric;执行阶段用 rubric 引导实验与分析,修订阶段按 rubric 逐条核查未满足的标准并做针对性修改。
实验数字(均来自论文摘要):在 ResearchClawBench 上,固定 Codex harness 时三个 backbone LLM 平均提升 2.08 分;固定 DeepSeek-V4-Flash backbone 时三个 agent harness 平均提升 2.95 分。在 AstaBench E2E Discovery 随机抽样的 20 个任务子集上,三个 agent harness 平均提升 16.8 分,同时完成任务数保持或增加。代码开源于 GitHub(zjunlp/AutoSciRub)。
为什么值得关注
它把"评估"前置成 agent 的控制机制:rubric 不是事后打分,而是执行前的可验证规格。对做科研/深度调研 agent 的人,这给出了一条不依赖更多算力、直接提高任务完成质量的路径——先让 agent 自己把"做好"的定义写清楚。
关键信息
- 论文标题:Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
- 作者:Xuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, Shumin Deng
- arXiv:https://arxiv.org/abs/2608.31076
- 发布时间:2026-08-31
- arXiv 分类:cs.CL, cs.AI, cs.IR, cs.LG, cs.MA, cs.SE
- 备注:Work in progress(Work in progress,数字可能随版本变化)
- 关联标签:llm-agent, scientific-research, evaluation, rubric
English Abstract
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).
English Summary
AutoSciRub is an evaluation-first framework for autonomous scientific research agents: it induces a task-specific executable rubric before execution, then uses it to guide execution, criterion-level verification, and iterative revision. On ResearchClawBench it yields average gains of 2.08 points (three backbones, fixed Codex harness) and 2.95 points (three harnesses, fixed DeepSeek-V4-Flash backbone); on a 20-task AstaBench E2E Discovery subset it averages +16.8 points across three harnesses while maintaining or increasing completed tasks.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。