OpenAI o3 在 ARC-AGI 拿到 75.7%
- ID: digest_2026-06-16_07
- 原文链接: https://x.com/fchollet/status/1870169764762710376
- 作者: @fchollet
- 日期: 2024-12-20
- 分类: models
- 标签: openai, o3, arc-agi, reasoning, benchmark
- 质量评分: 5/5
- 抓取时间: 2026-06-17T12:22:50
o3 在 ARC-AGI 拿下 75.7%
中文翻译
OpenAI 公布下一代推理模型 o3,François Chollet(Keras 和 ARC-AGI 的作者)宣布他的团队与 OpenAI 合作,在 ARC-AGI 基准上对 o3 做了独立评估,结果被 Chollet 称为"AI 适应新任务能力的重大突破"。
关键数字:o3 在半私有评测集的低算力模式(每任务约 20 美元算力成本)下拿到 75.7%;高算力模式(每任务数千美元)拿到 87.5%。作为对照,ARC-AGI 此前的人类基准约 76%,GPT-4 时代最好的成绩也只在 30% 上下——o3 是首个让这个基准不再"看起来无解"的模型。
Chollet 特别强调这"不是单纯的暴力堆算力"——ARC-AGI 的设计目的就是测"没见过的推理模式",o3 在低算力模式就已经超过人类,说明它在新任务上的适应能力有了质的飞跃。这一结果直接推动了 ARC-AGI v2 基准的提上日程,因为 v1 在 o3 面前已经失去区分度。
推文附带一张示意图,展示了 o3 在不同算力预算下的分数曲线。
English Original
@fchollet
Today OpenAI announced o3, its next-gen reasoning model. We've worked with OpenAI to test it on ARC-AGI, and we believe it represents a significant breakthrough in getting AI to adapt to novel tasks.
It scores 75.7% on the semi-private eval in low-compute mode (for $20 per task in compute) and 87.5% in high-compute mode (thousands of $ per task). It's very expensive, but it's not just brute — these capabilities are new territory and they demand serious scientific attention.