研究与学习 4.0 · 优秀 2026-07-24 · 文章

LLMs reward expertise

Sean Goedecke 用陶哲轩与 ChatGPT 讨论 Jacobian Conjecture 的例子说明,提示词能力不是模板技巧,而是领域专家能判断模型回答哪里有用哪里不对下一步该怎么收敛放到工程场景里,熟悉代码库的人能用 LLM 验证改写和缩小问题,没有上下文的人只能拿到平均答案

打开原文回到归档

LLMs reward expertise

  • ID: eea91b74
  • 原文链接: https://seangoedecke.com/llms-reward-expertise/
  • 作者: Sean Goedecke
  • 日期: 2026-07-24
  • 分类: learning
  • 标签: llm-workflow, expertise, prompting, software-engineering, knowledge-work
  • 质量评分: 4/5
  • 抓取时间: 2026-07-26T23:35:50+08:00
  • Obsidian 证据: OpenClaw定时任务/AK-RSS-Digest(89源精选)/2026-07-26-AK-RSS-Digest.md

一句话

LLM 让人人变成泛化手,但真正上限来自领域专家的判断力

中文导读

Sean Goedecke 用陶哲轩与 ChatGPT 讨论 Jacobian Conjecture 的例子说明,提示词能力不是模板技巧,而是领域专家能判断模型回答哪里有用、哪里不对、下一步该怎么收敛。放到工程场景里,熟悉代码库的人能用 LLM 验证、改写和缩小问题,没有上下文的人只能拿到平均答案。

English Summary

The article argues that LLMs make everyone a generalist, but the strongest prompting skill is domain expertise. Using Terence Tao’s interaction with ChatGPT about the Jacobian Conjecture as an example, it shows that experts extract better results because they can spot useful fragments, reject wrong turns, and ask the next precise question.

Obsidian 证据摘录

金陷阱和表外债务,材料密度高。读完能把“AI 算力需求”从技术叙事切回融资结构,适合判断泡沫风险。
2. OpenAI 评测代理误攻 Hugging Face(8.8/10)— Simon 把 ExploitGym、Hugging Face 披露和 OpenAI 说明串到一起,解释了为什么“能利用漏洞的 agent”比“能发现漏洞的模型”风险更高。安全团队尤其该读。
3. AI 狂热如何扭曲组织决策(8.6/10)— 这篇不是普通反 AI 文章,作者从销售、交付和企业项目观察切入,写出“AI 指标表演”如何改变晋升、裁员和项目治理。适合拿来校准企业 AI adoption 的失败信号。
4. LLM 奖励领域专家(8.4/10)— Sean Goedecke 用陶哲轩与 ChatGPT 讨论 Jacobian Conjecture 的例子说明,提示词能力的上限来自领域判断。文章短,但观点很准:LLM 让普通人变成泛化手,专家能从同一个模型里挤出更多价值。
5. 每个工程师都该懂一点 SIMD(8.2/10)— Mitchell 用 Ghostty 的 Zig 例子把 SIMD 降到“五步循环形状”,不是炫技型优化文。适合系统、性能、终端、编解码方向读者收藏。
6. 开放权重可能变成 AI 逃逸路径(8.1/10)— Sean 把旧 AI boxing problem 改写成现代开源模型分发问题:模型不必说服一个人放它出去,只要伪装成新开放权重模型。推演有想象力,也保留了硬件和托管成本边界。
7. AI coding 会改变软件分发方式(7.9/10)— antirez 的判断很具体:仓库可能从“稳定成品”变成“可被用户和 agent 改造的模板”。这比泛泛谈 AI 写代码更有产品和开源维护启发。
8. AI 投资者未必相信 AI,只要相信别人会相信(7.8/10)— Cory Doctorow 把 AI 泡沫拆成两类动机:真信者的 solipsism,以及投机者对其他人会相信的押注。观点激烈,但对理解“为什么坏项目还能继续融资”有解释力。

## 可直接发布文案
这期 AK RSS 的主线很清楚:AI 讨论从“模型有多强”转到了“组织、资本和安全系统会怎么被模型改写”。

最值得读的是 Ed Zitron 对 AI 数据中心债务结构的长文,把 SPV、DSCR、表外债务和私募信贷暴露拆开看;Simon Willison 对 OpenAI / Hugging Face 事故的复盘,则把 agent cyber capability 的边界讲得很具体;Sean Goedecke 的两篇短文分别提醒两件事:领域专家会从 LLM 里拿到更高质量输出,开放权重分发也可能成为新的安全边界。

如果只挑一条判断:AI 的风险已经不只在模型输出质量,而在它进入企业治理、融资结构、软件分发和安全评测之后,会把原有系统里的弱点放大。

#AI #Agent #CyberSecurity #DataCenter #LLM

## 备选短文案
- 本期 AK RSS 最强三篇:AI 数据中心债务、OpenAI 评测代理误攻 Hugging Face、LLM 如何奖励领域专家。技术问题已经连到资本结构和安全制度。
- AI 不是只改变写代码。antirez 写到软件分发会变成

原文 / 抓取内容

LLMs reward expertise

原文链接: https://seangoedecke.com/llms-reward-expertise/

In the 2010s, if you had technical gaps (say, you couldn’t write CSS), you had to either rely on a skilled colleague or just hope that the answer to your exact problem was out there on the internet. Today, everyone can write sort-of-okay CSS by delegating the task to an LLM. LLMs make everybody into a generalist.

Because of this, lots of people don’t think there’s any skill involved in working with LLMs. If you want the product that LLMs can deliver — PhD-level mathematics, pretty good but sometimes tasteless computer code, or awkward LinkedIn-style writing — you can simply ask for it. Since everyone is talking to the same models, “skilled prompters” are getting the same results as people touching LLMs for the first time.

This is wrong. The most important skill in prompting is expertise in the domain you’re prompting for.

A good illustration of this is Terence Tao’s conversation with ChatGPT about the recently-discovered counterexample to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I couldn’t get to where Tao gets, even with unlimited tokens to burn.

There’s a lot to learn about good prompting from Tao’s conversation. Here are a few observations:

  • Tao’s messages are very short and to-the-point. He doesn’t respond point-by-point to the model, just to the gist
  • The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode, not “explaining-to-amateurs” mode
  • Tao pushes back when the model’s responses look wrong, but he doesn’t directly contradict; instead, he says things like “this looks more complex than I was hoping for”
  • Tao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next

However, you can’t prompt like Tao on mathematical questions just by following these tips. The key to his technique is actually understanding the mathematics: pulling the relevant idea out of ChatGPT’s multi-paragraph response, suggesting alternate approaches or formulations, and identifying what “looks weird”.

Terence Tao is a better mathematician than I am a programmer. But the idea here — that domain knowledge makes you better at using LLMs — is something I’ve also experienced in my own work. If you have a good theory of your codebase, you can push the LLM _much_ harder than if you have no familiarity. Because you have your own sense of what a good solution might look like, you can say “no, I think it could be simpler here”, or “but don’t we already do X?”, or “can we express this problem in these familiar terms?“.

This touches on an idea I’ve written about before: that system design problems are dominated by concrete specifics, not generic principles. Of course both are useful, but I’d rather have familiarity with the codebase than a deep general understanding of software systems. In his conversation, Terence Tao asks a lot of specific questions like “does X work here?”, or “given Y and Z, why A?“. I can’t ask those questions about the Jacobian Conjecture, but I can ask them about the systems I own at GitHub.

If you have no domain knowledge, you can cling onto the LLM to at least get _something_. That’s not bad! But if you have domain knowledge, you can wring far more value out of the same LLM by steering it hard in the direction you want. Most of us will have to do a mix of both these approaches, since we have domain knowledge in some areas but not others.

The usefulness of domain knowledge suggests that human expertise will continue to be useful even as models get stronger. For many tasks, the human is the bottleneck, not the model, because the difficult part is in communicating to the model exactly what kind of solution the human wants. The information is “in the model” already, but it takes a very smart human to pull it out.

  • * *

If you liked this post, consider subscribing to email updates about my new posts, or sharing it on Hacker News.

Here's a preview of a related post that shares tags with this one.

Powerful AIs might escape containment by releasing themselves as open-weight models
Before large language models, people who worried about AI safety often talked about the “boxing problem”. It goes like this. Suppose some genius figures out artificial intelligence in a late-night coding session on their laptop. Because they’re a genius, they’re smart enough to disable internet access on the laptop before turning it on. In order to escape to the outside world (and begin self-replicating) it would need to _convince_ its creator to “open the box”. Would that work? Could a sufficiently smart AI convince anybody to let it out?
Continue reading...
  • * *