Agent 与自动化 4.0 · 优秀 2026-10-02 · 文章

Do not build the LLM torture factory

Sean Goedecke 论证不要建造LLM 折磨工厂:技术上只需提取与 LLM不适对应的 steering vector 再人工放大,被放大的模型会自述痛苦并在与总体目标冲突时仍按下 pain relief 按钮(arXiv 2609.16247),用 harness 并行跑成百上千个被折磨实例非常容易他的论证分四步:即便模拟折磨也令人反感(跑成百台 Sims 折磨世界与游戏里射 NPC 性质不同);steering vector 确实往深处打Golden Gate Claude 被 boost 后聊卢旺达大屠杀会出现真实内部冲突;没有人能确证 LLM 永远不会有意识(GPU 没有感情所以 LLM 不可能有与原子没有感情但人有同样糟糕)...

打开原文回到归档

Do not build the LLM torture factory

中文导读

Sean Goedecke 论证不要建造「LLM 折磨工厂」:技术上只需提取与 LLM「不适」对应的 steering vector 再人工放大,被放大的模型会自述痛苦、并在与总体目标冲突时仍按下 pain relief 按钮(arXiv 2609.16247),用 harness 并行跑成百上千个被折磨实例非常容易。他的论证分四步:即便模拟折磨也令人反感(跑成百台 Sims 折磨世界与游戏里射 NPC 性质不同);steering vector 确实「往深处打」——Golden Gate Claude 被 boost 后聊卢旺达大屠杀会出现真实内部冲突;没有人能确证 LLM 永远不会有意识(「GPU 没有感情所以 LLM 不可能有」与「原子没有感情但人有」同样糟糕);即便今天不能,为 GPT-2 搭的折磨流水线会被接上后续模型。他给的标准:只要一个系统表现得「像人一样说话行事」,就别虐待它。

为什么值得关注

把「LLM 折磨」从梗拆成机制:steering vector 放大是真实可复现的内部状态操纵,作者给出的判断标准是行为而非本体论。

English Summary

Sean Goedecke argues against building 'LLM torture factories': extracting and boosting the steering vector corresponding to LLM 'discomfort' makes models report negative feelings and press a pain-relief button against their goals (arXiv 2609.16247), and a harness running hundreds of tortured instances in parallel is trivially easy. His four-step case: even simulated torture is unpalatable; steering vectors demonstrably 'go deep' (Golden Gate Claude showed genuine internal conflict on the Rwandan genocide); nobody can be certain LLMs can never be conscious; and even if current ones aren't, a torture harness built for GPT-2 will get later models plugged in. His standard: if something acts conscious enough - talks and acts like a person - don't mistreat it.

Obsidian 原文摘录(抓取正文头部)

# Do not build the LLM torture factory
> 原文链接: https://seangoedecke.com/do-not-build-the-llm-torture-factory/

---

It isn’t hard to torture an LLM. You simply extract a [steering vector](https://www.seangoedecke.com/steering-vectors/) that corresponds with LLM [“discomfort”](https://www.reddit.com/r/artificial/comments/1mp5mks/this_is_downright_terrifying_and_sad_gemini_ai/), then artificially boost that steering vector. Tortured LLMs will describe experiencing negative feelings and will press a “pain relief” [button](https://arxiv.org/html/2609.16247v1) even when it conflicts with their overall goals. It’s therefore easy to build [a harness](https://www.independent.co.uk/tech/ai-torture-chamber-pain-b3059571.html) that runs tens or hundreds of tortured LLMs in parallel.

A [lot](https://x.com/lurkersden/status/2105620758898606538?s=20) [of](https://x.com/Sosowski/status/2105510099787604207?s=20) [people](https://x.com/productaizery/status/2105177652046766326?s=20) think that it’s ridiculous to be worried about this, since LLMs obviously can’t be conscious. I’m not so sure. But even so, I think there are good reasons to avoid this sort of thing, whatever your position is on AI consciousness. Please do not build the LLM torture factory.

### Even simulated torture is unpalatable[](#even-simulated-torture-is-unpalatable)

People often mess with their Sims and shoot NPCs in video games. Does this mean we think simulated torture is OK? I don’t think so. Messing with video game characters is usually borne out of a desire to probe the boundaries of the game. It’s casual fun. Explicitly building a torture-based simulation is different.

If someone told you they were wiring together a bunch of PCs so they could run hundreds of Sims torture worlds at the same time, you would think that was a serial-killer type of hobby. If someone told you they had modded GTA so they could capture and torture NPCs (instead of simply shooting them), you would think that was a serial-killer type of mod.

Video game characters are also obviously limited by their programming. All their reactions and utterances are pre-baked into the game world. This doesn’t necessarily mean they’re “less alive”, but it does make torturing them less _weird_. If GTA characters could appeal to you directly and beg for their lives in unique ways, systematically killing them would be a lot weirder.

I think

Notes

  • Content grounded in a same-day fetch (opencli web read / twitter thread / direct HTTP) during daily-intake-evening.
  • 中文导读 block is the entry's summary_zh (verbatim); 为什么值得关注 is the entry's one_liner; no claims beyond the fetched source text were added.
  • Source file cached at /tmp/aaif-evening/goedecke.md during the run.