研究与学习 4.0 · 优秀 2026-09-24 · 论文

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

arXiv 2609.30199 提出 ExplorationBench,把"科学探索能力评估"这一 wicked problem 落到可执行的"Alien Worlds"上:规则可执行 每条答案可被精确验证;规则与先验知识冲突 单凭预训练召回无法解题基准包含两个沙盒 AlienCode(31 个发现目标 / 70 任务)和 AlienLogic(24 / 70),各自附带一份有缺陷的手册任务级环境反馈和专用 tool-call schema,要求系统用它们去探索沙盒并解决 held-out 任务评测 10 个 AI 系统后,最强系统确实能习得并应用陌生规则,但不同轨迹间表现差异很大,且持续探索会让早期收益停滞甚至回退,是朝向"能真正获得新知识"的 AI 系统迈出的一步

打开原文回到归档

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

原文链接: https://arxiv.org/abs/2609.30199
作者: Ming Zhang et al.
发布时间: 2026-09-24
源: arXiv外部扫描 (2026-09-28)

摘要

arXiv 2609.30199 提出 ExplorationBench,把"科学探索能力评估"这一 wicked problem 落到可执行的"Alien Worlds"上:规则可执行 每条答案可被精确验证;规则与先验知识冲突 单凭预训练召回无法解题基准包含两个沙盒 AlienCode(31 个发现目标 / 70 任务)和 AlienLogic(24 / 70),各自附带一份有缺陷的手册任务级环境反馈和专用 tool-call schema,要求系统用它们去探索沙盒并解决 held-out 任务评测 10 个 AI 系统后,最强系统确实能习得并应用陌生规则,但不同轨迹间表现差异很大,且持续探索会让早期收益停滞甚至回退,是朝向"能真正获得新知识"的 AI 系统迈出的一步

English Summary

arXiv 2609.30199 turns scientific-exploration evaluation into a tractable benchmark built on verifiable Alien Worlds (rules are executable, conflicting with prior knowledge). It ships two sandboxes AlienCode (31 discovery targets / 70 tasks) and AlienLogic (24 / 70) each with a flawed manual and tool-call schema, evaluates 10 AI systems, and shows that the strongest can acquire and apply unfamiliar rules but performance varies across trajectories and continued exploration can stall or reverse earlier gains.

为什么值得关注

用 Alien Worlds 把"AI 能不能做科学发现"做成可执行可验证的基准

信息源