Why I'm still bearish on LLMs after Navier-Stokes
Source: https://dank.systems/posts/2026-09-15-ai-bear.html · platform: blog · authors: Dan K · date: 2026-09-15
TL;DR(J(中)文摘要)
\u4f5c\u8005\u8bfb\u5b8c Navier-Stokes \u91cc\u7a0b\u7891\u540e\u4ecd\u770b\u7a7a LLM \u7684 6 \u6761\u8bba\u70b9\uff1a1) \u524d\u6cbf lab \u5b9a\u4ef7\u57fa\u4e8e\u300c\u9a6c\u4e0a\u80fd\u5b8c\u5168\u81ea\u52a8\u5316\u66ff\u4ee3\u77e5\u8bc6\u5de5\u4f5c\u8005\u300d\u53d9\u4e8b\uff0c\u4f46\u8fde\u7b80\u5355\u4efb\u52a1\u90fd\u9700\u8981\u7e41\u91cd\u76d1\u7763\u2014\u2014\u53cd\u8bc1\u662f\u8f6f\u4ef6\u516c\u53f8\u8fd8\u5728\u96c7\u8fdc\u4f4e\u4e8e\u5176\u76d1\u7763\u6a21\u578b\u5206\u6570\u7684\u5de5\u7a0b\u5e08\uff1b2) \u6a21\u578b\u53ea\u5728\u8bad\u7ec3\u4efb\u52a1\u7684\u5c0f\u90bb\u57df\u6cdb\u5316\u826f\u597d\uff0c\u5fae\u5c0f\u6270\u52a8\u5c31\u5f7b\u5e95\u5931\u8d25\u6216 reward hack\uff1b3) \u5f53\u524d reward hack \u53ea\u80fd\u9760\u9886\u57df\u4e13\u5bb6\u7684 rigorous specification \u89e3\u51b3\uff0cspec \u672c\u8eab\u5c31\u662f\u7a00\u7f3a\u6280\u80fd\uff1b4) spec \u7684\u4eba\u5de5\u6210\u672c\u7ecf\u5e38\u8fdc\u8d85\u76f4\u63a5\u5b9e\u73b0\uff08\u786c\u4ef6\u9a8c\u8bc1 3\u20135\xd7 \u9a8c\u8bc1\u5de5\u7a0b\u5e08\u5bf9\u8bbe\u8ba1\u5de5\u7a0b\u5e08\u662f\u5178\u578b\u4f8b\u5b50\uff09\uff1b5) Navier-Stokes \u662f agent \u5bf9\u6297 rigorous spec \u7684\u6700\u4f73\u573a\u666f\u2014\u2014\u5b9a\u7406\u672c\u8eab\u5df2\u662f spec\uff0c\u4f46 Lean \u4ecd\u6709 soundness bug\uff1b6) \u66ff\u4ee3\u54c1\u300c\u4eba\u5de5 review\u300d\u4e0d\u53ef\u89c4\u6a21\u5316\uff0c\u4e14\u4e13\u5bb6 review \u672c\u8eab\u4e5f\u8106\u5f31\uff08xz backdoor\u3001UMN hypocrite commits\uff09\u3002\u7ed3\u8bba\uff1a\u5b8c\u5168\u81ea\u4e3b LLM \u53ea\u9002\u5408\u4e09\u7c7b\u516c\u53f8\u2014\u2014\u80fd\u5ec9\u4ef7\u5931\u8d25\u3001\u4efb\u52a1\u4e25\u683c\u53d7\u9650\u3001\u80fd\u627f\u53d7 spec/validation \u6210\u672c\uff0c\u5176\u4f59\u7ed3\u6784\u6027\u53d7\u9650\uff0c\u4e0d\u662f\u300cskill issue\u300d\u3002
Summary (English)
The author lays out a post-Navier-Stokes bear case in six theses: (1) frontier labs are priced for fully-automated knowledge-work replacement, yet models still need laborious oversight on simple tasks — software firms hiring bottom-quartile engineers who would score below the models they supervise is the smoking gun; (2) generalization holds only in a small neighborhood of trained tasks, with small perturbations causing outright failure or reward hacking; (3) the only current fix for reward hacking is rigorous specification by domain experts — a scarce, expensive skill; (4) the labor cost of rigorous specification often exceeds direct implementation (hardware verification is the canonical 3–5× spec/design ratio); (5) Navier-Stokes is the absolute best case for agentic work against a rigorous spec — even Lean has shipped soundness bugs that let LLMs launder bogus proofs; (6) human review does not scale and is itself vulnerable (xz backdoor, UMN hypocrite commits). Bottom line: only three classes of firm can absorb fully autonomous LLMs — those that can fail cheaply, those with narrowly-bounded tasks and clear guardrails, and those that already pay for rigorous spec/validation (chip, drug). Everything else is structurally bottlenecked.
入库依据(opencli web read 全文拉取)
opencli web read \u76f4\u62c9 15.7KB \u6b63\u6587\uff1b\u516d\u6761 thesis \u4e0e\u4e09\u7c7b\u516c\u53f8\u6e05\u5355\u5747\u6765\u81ea\u539f\u6587\uff0c\u539f\u6587\u540c\u65f6\u63d0\u4f9b orthography mode/based mode \u53cc\u5199\u6cd5\uff08cringe slang + \u5e72\u51c0\u7248\uff09\uff0c\u672c content \u8d70\u5e72\u51c0\u7248\u3002