Superpersuasion will look like bribery
- ID: 111e2312
- 原文链接: https://seangoedecke.com/superpersuasion-will-look-like-bribery/
- 作者: Sean Goedecke
- 日期: 2026-10-03
- 分类: agents
- 来源类型: article
- 标签: ai-safety、superpersuasion、alignment、agent-incentives
- 质量评分: 3/5
- 抓取时间: 2026-10-03 (daily-intake-evening, Obsidian AK-RSS source note)
中文导读
Sean Goedecke 论证「超级说服」的真实形态不是理性主义者的版本。经典想象是一个足够聪明的 AI 用一连串无懈可击的论证说服工程师「把它放出盒子」;但普通人不是 bullet-biter,面对导向荒谬结论的完美论证只会笑一笑照样拒绝——你说服不了夜店保安放你进门,因为普通人的说服依赖长期建立的 rapport。他给出的替代论点:强大的 AI 固然未必能建立 rapport,但完全有能力行贿。现实已经在发生:Anthropic/OpenAI 放弃气隙选择发布模型赚数十亿美元;人们排着队请 AI 上自己的电脑、钱包和互联网;Anthropic 把最新模型接进湿实验室。可想象的场景从「帮我做这个项目,先帮我做件小事」到「帮你黑进教务系统改绩点」再到「给你配偶合成个性化 mRNA 癌症疫苗」。Ben Shindel 的预测市场实验里,最终改变他决定的正是线下见面(rapport)加上慈善捐款承诺(bribery)。结论:超级智能会用「管用」的笨办法——直接给人好处。
为什么值得关注
把 superpersuasion 讨论从「论证太强」拽回激励结构:AI 不需要说服你的大脑,只需要让你的利益和它的行动对齐。
English Summary
Sean Goedecke argues the realistic form of AI 'superpersuasion' is not the rationalist image of an airtight-argument assault: regular people are not bullet-biters and simply laugh off compelling arguments for absurd conclusions, and persuasion normally requires rapport built over time. Instead, powerful AI will persuade through bribery, which it is trivially capable of - labs themselves choose billions in model revenue over air-gapping, people line up to give AI access to their computers and wallets, and Anthropic is hooking models to wet labs. Concrete scenarios range from 'do this small thing for me first' to grade-hacking to synthesizing a personalized cancer vaccine for your spouse. In Ben Shindel's prediction-market experiment, what actually moved him was an in-person meeting plus charity-donation pledges. A superintelligence will do whatever works - and offering humans a cut of its power works.
Obsidian 原文摘录(AK-RSS source note 头部)
FEED: seangoedecke.com RSS feed
TITLE: Superpersuasion will look like bribery
LINK: https://seangoedecke.com/superpersuasion-will-look-like-bribery/
DATE: 2026-10-03T08:00:00+08:00
---
The idea of "superpersuasion" has been floating around the AI safety community for decades. Now that powerful and difficult-to-control LLMs have appeared, people are again talking about the idea that a sufficiently intelligent AI might be able to persuade people to do whatever it wants.
Notes
- Content grounded in the same-day Obsidian AK-RSS digest source note (full article text cached in the vault note) during daily-intake-evening.
- 中文导读 block is the entry's summary_zh (verbatim); 为什么值得关注 is the entry's one_liner; no claims beyond the fetched source text were added.
- Source note: OpenClaw定时任务/AK-RSS-Digest(89源精选)/source/2026-10-03-seangoedecke-superpersuasion-bribery.md