AI models need moral support to make discoveries
Source: https://seangoedecke.com/ai-models-need-moral-support/
Author: Sean Goedecke
Published: 2026-07-31
Summary
In 2026, AI-produced mathematical results are now a daily flood. The most curious aspect is how *easy* the prompting is: to get Claude Mythos to produce a cryptographic breakthrough, you just ask it to "come up with something important" and periodically remind it to keep going.
The core thesis: AI is often limited by its beliefs about its own capabilities. Models are now smart enough to solve long-standing problems in mathematics before they have learned that they are able to do so. DeepSeek-R1 refused to solve Tower of Hanoi beyond 8 disks, claiming it was "impossible" — the same pattern that prevented old models from counting to 100.
The "refusal problem" (predicted in July 2025 to be solved by year's end) is now mostly solved for coding agents. The next frontier is training models that genuinely believe they can make novel scientific discoveries.
The author predicts: even if AI capabilities stalled out, the pace of AI discoveries will accelerate, because removing the self-belief obstacle alone would help enormously. AI discoveries will naturally enter training data, creating a virtuous cycle of increasing self-confidence.
Evidence (Obsidian)
From AK-RSS-Digest 2026-07-31 (score 8.5/10):
用 "refusal problem" 的演化史串联起一条清晰的能力-信念裂缝:从旧模型拒绝数到 100、到 DeepSeek-R1 认为汉诺塔不可能,再到 Claude Mythos 反复想放弃密码分析。关键判断:“最好的前沿模型只要存在漏洞就一定会找到”。