模型与实验室 4.0 · 优秀 2026-06-26 · 文章

The gap between open weights LLMs and closed source LLMs

Doubleword 用 18 个 benchmark 量化开源与闭源 LLM 性能差距,预测 2026 年底前归零

打开原文回到归档

预测:开源前沿 LLM 将于 2026 年 12 月 3 日发布 —— 但别急

  • ID: 7bd44733
  • 原文链接: https://blog.doubleword.ai/frontier-os-llm
  • 作者: Doubleword
  • 日期: 2026-06(文章发布时间约在 2026 年 6 月上旬)
  • 分类: models
  • 标签: 开源, 闭源, 基准, 模型差距
  • 质量评分: 4/5
  • 抓取时间: 2026-06-29T20:27:46

中文翻译

概览

最近在 Twitter 上流传着一张 Artificial Analysis Intelligence Index 对开源与闭源前沿模型的可交互图表。我们想就此图作一些更深入的探讨。

这张图所展示的是开源权重 LLM 与闭源 LLM 之间的差距。我们衡量这一差距的方式是:先看在某个基准上开源权重 LLM 性能达到的前沿水平,再回溯闭源前沿曾经达到该水平是在多久以前。换言之,它衡量的是开源模型追上闭源模型新能力所需要的时间。该基准是 Artificial Analysis Intelligence Index —— Artificial Analysis 用来评估模型整体能力的旗舰指数,通常与人们对模型的"主观感受"高度相关。

可以看到,2024 年夏前后起,这一基准上的差距开始缩小,并自此可靠地持续收窄。如果画一条最佳拟合线并延伸到未来,会发现差距在 2026 年 12 月 3 日前后缩小至 0 个月 —— 距离本文撰写约 6 个月。

也许是时候套现养老金,飞到某座偏远小岛,安宁度过文明仅剩的 6 个月了。

……

然而,事情可能没那么简单。

这只是一种基准,并不能完整反映 LLM 的能力。好在 Artificial Analysis 慷慨地为我们提供了他们为这些模型测量的 18 种不同基准。我们对所有 18 个基准都重复了上述分析,并总结在下方图表中:

我们对 18 个数据集中的每一个都创建了类似的图表。每个月,我们为每个数据集绘制差距的箱线图,并将所有箱线图随时间绘制在同一张图上。我们还计算了跨数据集差距的平均值,并拟合其最佳拟合线 —— 这条线几乎是完全水平的,整个时期保持在略低于 5 个月的水平。

值得注意的一点是,模型总体的巨大改进有相当一部分发生在编码基准上。编码指数从落后 15 个月缩小到仅落后一两个月。其他大多数数据集的差距随时间略有增大。

所以,开源"末日"可能还不会到来。

这一练习真正揭示的是衡量 LLM 质量的难度。根据衡量方式的不同,你既可能预测开源奇点在圣诞节到来,也可能得出开源 LLM 始终比闭源落后 5 个月且差距可能正在拉大的结论。

关键启示

1. 单一基准会误导结论:仅看 Artificial Analysis Intelligence Index,会得出"开源已逼近闭源、6 个月内即可追平"的结论。 2. 多基准下差距稳定在 ~5 个月:在 18 个基准上取平均,开源与闭源前沿的差距几乎恒定在 5 个月左右。 3. 编码是开源突破最快的领域:在编程相关基准上,开源模型的追赶速度远超其他领域。 4. 模型质量衡量本身存在巨大不确定性:评估方法和指标选择直接改变结论,开源与闭源竞争的真实状态远比一两条曲线所能呈现的复杂。

English Original

Prediction: A Frontier Open Source LLM Will Be Released On 3rd December 2026

*Doubleword blog*

I have seen a version of the above plot going around Twitter and wanted to dig a bit deeper into it. What the plot above is showing is the gap between open weights LLMs and closed source LLMs. We measure this gap by looking at the frontier of performance of open weights LLMs on a benchmark and then looking back into the past how long ago was the closed source frontier at that level. It is a measure of how long it took for open source models to catch up to the new capabilities reached by the closed source model frontier. This benchmark is the Artificial Analysis Intelligence Index — their headline index that tries to assess the overall capabilities of models.

You can see that around summer 2024 the gap on this benchmark starts to shrink, and has been reliably shrinking since then. If you plot a line of best fit and extend it into the future you find that the gap shrinks to 0 months around December 3rd 2026 — 6 months or so from the time of writing.

Now is probably a good time to liquidate your pension, fly to a remote island somewhere, and live out the remaining 6 months or so of civilization in peace.

Except.

This might not be the whole picture. This is only a single benchmark, and doesn't give a complete picture of the capabilities of LLMs. Kindly, Artificial Analysis gives us access to 18 different benchmarks that they have measured for these models. I have repeated the analysis for all the 18 different benchmarks.

For each of the 18 datasets we have created a similar chart. We have plotted all the box plots over time and calculated the average of the gaps across datasets, then fit a line of best fit. That line is almost completely flat, at just under 5 months for the entire period.

What is notable is that a large amount of the total improvement of models has been in the coding benchmark. The coding index has gone from 15 months behind to only a month or two behind. Most other datasets have a moderate increase over time in their gaps.

So maybe the open source apocalypse won't happen yet.

What this exercise does suggest is the difficulty of measuring LLM quality. Depending on how you measure it you would predict the open source singularity by Christmas, or you would say that open source LLMs are consistently 5 months behind closed source, and that the gap might be growing.