模型与实验室 4.0 · 优秀 2026-07-31 · 文章

AI models need moral support to make discoveries

限制前沿模型做出数学/密码学发现的,不是能力,而是模型对自己能力的悲观预期从旧模型拒绝数到 100DeepSeek-R1 认为汉诺塔不可能,到 Claude Mythos 反复想放弃密码分析,都是同一条能力-信念裂缝预言即使能力停滞,发现速度仍会加速

打开原文回到归档

AI models need moral support to make discoveries

Source: https://seangoedecke.com/ai-models-need-moral-support/
Author: Sean Goedecke
Published: 2026-07-31

Summary

In 2026, AI-produced mathematical results are now a daily flood. The most curious aspect is how *easy* the prompting is: to get Claude Mythos to produce a cryptographic breakthrough, you just ask it to "come up with something important" and periodically remind it to keep going.

The core thesis: AI is often limited by its beliefs about its own capabilities. Models are now smart enough to solve long-standing problems in mathematics before they have learned that they are able to do so. DeepSeek-R1 refused to solve Tower of Hanoi beyond 8 disks, claiming it was "impossible" — the same pattern that prevented old models from counting to 100.

The "refusal problem" (predicted in July 2025 to be solved by year's end) is now mostly solved for coding agents. The next frontier is training models that genuinely believe they can make novel scientific discoveries.

The author predicts: even if AI capabilities stalled out, the pace of AI discoveries will accelerate, because removing the self-belief obstacle alone would help enormously. AI discoveries will naturally enter training data, creating a virtuous cycle of increasing self-confidence.

Evidence (Obsidian)

From AK-RSS-Digest 2026-07-31 (score 8.5/10):

用 "refusal problem" 的演化史串联起一条清晰的能力-信念裂缝:从旧模型拒绝数到 100、到 DeepSeek-R1 认为汉诺塔不可能,再到 Claude Mythos 反复想放弃密码分析。关键判断:“最好的前沿模型只要存在漏洞就一定会找到”。