Models Are Getting Dumber on Purpose
- ID: 08e6169d
- 原文链接: https://w4g1.dev/blog/models-are-getting-dumber-on-purpose
- 作者: Walter van der Giessen
- 日期: 2026-08-17
- 分类: models
- 来源类型: article
- 标签: llm, knowledge, reasoning, distillation, physics-of-lm, small-models
- 质量评分: 4/5
- 抓取时间: 2026-08-17T23:51:34+08:00
中文导读
把“模型越来越不记得事实”讲成主动设计:GLM-5.2 以约 40B 激活参数在 AIME 2026 拿 99.2%,而 SimpleQA 纯事实 recall 的榜首(Gemini 2.5 Pro)只有 53%,9B 级模型知识评测幻觉率 80-82%。基于 Physics of Language Models 系列给出“每参数约 2 bit 事实”的数量级估计:推理过程可压缩、事实不可压缩,于是蒸馏 + 可验证任务 RL 把“程序”灌进小模型、把“百科”下放给检索与 harness。终局是跑在消费级 GPU 上“会推理但不会背”的长寿小模型。
为什么值得关注
反共识归因:事实层正被设计成可选件——训练贵、迭代慢的知识让位给检索,小模型负责推理。
收录理由:用具体评测数字与 Physics of LM 证据链把“知识-推理取舍”讲成可检验的设计趋势
关键信息
- ClawFeed 24小时高价值一览 评分:ClawFeed 评分:9.0/10
- 来源:ClawFeed 24小时高价值一览(2026-08-17 期)
- Obsidian 证据:
OpenClaw定时任务/ClawFeed24小时高价值一览/2026-08-17-ClawFeed24小时高价值一览.md
原文快照
Models Are Getting Dumber on Purpose - Walter van der Giessen Skip to content W4G1 2026-08-17 Models Are Getting Dumber on Purpose Walter van der Giessen Reasoning scores keep climbing while per-token compute keeps dropping. GLM-5.2 scores 99.2% on AIME 2026 with about 40 billion parameters active per token. Qwen3.5 scores 91.3% with 17 billion active. DeepSeek V4-Flash runs 13 billion active. For scale, GPT-4 was rumored to run around 280 billion active parameters in 2023, and it could barely solve an AIME problem. At the small end, Qwen3.5 9B fits in 6GB of VRAM quantized and roughly doubles the score of the next best model under 10B parameters on Artificial Analysis's intelligence index. If you only looked at math and code benchmarks, you'd conclude that models are getting smarter per parameter at an absurd rate. They are, on those benchmarks. Ask the same models a plain factual question and the picture flips. On SimpleQA , a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questions. The small models barely register. Artificial Analysis measures Qwen3.5 4B and 9B at hallucination rates of 80 to 82% on its knowledge benchmark, which means that when they don't know a fact, which is most of the time, they make one up. Ask the 9B for the birth year of a minor 19th-century mathematician and you get a confident, plausible, wrong answer. The parameter count didn't drop for free. Labs are trading world knowledge for reasoning skill, and the trade is deliberate. What the parameters were for Facts take space. Research on knowledge capacity (the "Physics of Language Models" series has the cleanest measurements) puts it on the order of two bits of factual knowledge per parameter. If you want a model that knows the birth year of every minor Wikipedia figure, the population of every Dutch municipality, and the argument order of every function in every npm package, you pay for that in weights, and it's a big part of why frontier models grew to trillions of parameters. Reasoning compresses much better than facts do, because it's a relatively small set of procedures applied over and over: break the problem into parts, track intermediate state, check your own work, backtrack when a step fails. Distillation and reinforcement learning on verifiable tasks turn out to transfer those procedures into small models remarkably well. Phi-4 is 14 billion parameters, trained heavily on synthetic textbook-style data, and it's good at math and bad at trivia, which tells you exactly what its training data contained. That mix used to look like a limitation of the synthetic-data approach. It now looks like the design goal. The knowledge that survives the trade has a shape. These models are generalists: they know a little about nearly everything and almost nothing in depth. Ask one about PostgreSQL and it knows what it is, what it's good at, and roughly how MVCC works, but ask
抓取方式:opencli web read(2026-08-17)。完整原文见上方链接。