IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
- ID: 374498bb
- 原文链接: https://arxiv.org/abs/2609.10539
- PDF: https://arxiv.org/pdf/2609.10539v1
- 作者: Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan
- 日期: 2026-09-09
- 更新: 2026-09-09
- 分类: cs.CL
- 来源类型: paper
- 标签: benchmark, research-specification, coding-agents, reproducibility, field-note
- 质量评分: 4/5
- 抓取时间: 2026-09-11T12:23:09+00:00
中文导读
一个研究 idea 可能新颖连贯科学上可信,但它的方法部分仍然不足以被忠实实现IdeaAMBIG 研究的就是这种"可编码化就绪度":规范是否给足了方法论信息,让胜任的实现者或 coding agent 不引入无依据假设就能建出目标方法 基准有 660 个证据锚定实例:163 个真实缺口(来自可复现性报告与 GitHub issues)+ 497 个注入 codification-ready 参考实现的受控合成缺口;评测三种能力:就绪度评估缺陷定位澄清动作生成 13 个 LLM 的结果很难看:最佳模型在真实实例上 Macro Defect Recovery Rate 只有 9.6%,但给它标注好的缺陷后澄清动作成功率 80.6%缺陷定位才是瓶颈 oracle 研究给出量化结论:给金标决议,下游可编码化率从 14% 涨到 98%;对"AI 科学家/coding agent 落地研究 idea"这条线,这是直接的需求清单
为什么值得关注
研究 idea 到实现之间的"规范缺口"基准:最佳模型真实缺口定位率仅 9.6%,给缺陷后澄清成功率 80.6%瓶颈在定位;金标决议能把可编码化率从 14% 拉到 98%
条目锚定 arXiv 2609.10539(2026-09-09 提交,分类 cs.CL),摘要自述贡献为上述机制与结论;详细信息以论文原文为准。
关键信息
- 论文标题: IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
- 作者: Yiling Ma, Yilun Zhao, Sihong Wu, Manasi Patwardhan, Arman Cohan
- arXiv: https://arxiv.org/abs/2609.10539
- 发布时间: 2026-09-09
- arXiv 分类: cs.CL
- 关联标签: benchmark, research-specification, coding-agents, reproducibility, field-note
English Abstract
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.
English Summary
IdeaAMBIG studies codification readiness of implementation-facing research-method specifications: whether they carry enough methodological information for a competent implementer or coding agent to build the intended method without unsupported assumptions. The benchmark has 660 evidence-grounded instances (163 real-world gaps from reproducibility reports and GitHub issues, 497 controlled synthetic gaps) and evaluates codification-readiness assessment, defect localization, and clarification action generation. Across 13 LLMs, the best model reaches only 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the annotated defect; an oracle study shows supplying the gold resolution raises downstream codification-ready rate from 14% to 98%.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。