模型与实验室 4.0 · 优秀 2026-08-19 · 文章

What Is Reasoning

Armin Ronacher 把 reasoning 的工程现实拆开:reasoning trace 不是模型内部神秘思考模块,而是模型被训练成把草稿写进特殊 token 框(如 GPT-OSS Harmony 格式的 analysis channel),再由 Responses API 解析器路由到独立 streamreasoning effort 也不是采样属性,而是 system prompt 里一行 Reasoning: low,改动会废掉 KV cacheDwarfStar 用 prefill 关闭 thinking用 Absolute maximum 提示词拉满自定义 think tool 能把推理撬到不该出现的位置,本质是 prompt injection 而非新能力对做 agent评估与安全的人是立刻可用的心智模型

打开原文回到归档

What Is Reasoning

中文导读

Armin Ronacher 把 reasoning 的工程现实拆开:reasoning trace 不是模型内部神秘思考模块,而是模型被训练成把草稿写进特殊 token 框(如 GPT-OSS Harmony 格式的 analysis channel),再由 Responses API 解析器路由到独立 stream。reasoning effort 也不是采样属性,而是 system prompt 里一行 Reasoning: low,改动会废掉 KV cache。DwarfStar 用 prefill 关闭 thinking、用 Absolute maximum 提示词拉满。自定义 think tool 能把推理撬到不该出现的位置,本质是 prompt injection 而非新能力。对做 agent、评估与安全的人是立刻可用的心智模型。

为什么值得关注

把 reasoning 从营销概念降级为 channel-routing convention:训练出的输出习惯而非神秘能力

原文摘录 (English Excerpt)

probably a good reason to separate them from what is normally shown to users.

At minimum, UIs need to detect them. The industry has done a good job at making reasoning traces sound special and exotic, but they really are just text: the model is trained to emit its thinking into a scratchpad as part of its response, before its final answer.

GPT-OSS’s Harmony response format makes this easy to see:

<|channel|>analysis<|message|> I need to work this out ... <|end|><|start|>assistant<|channel|>final<|message|> The answer is ... <|return|>

The markers are special tokens, but the reasoning between them uses “the same text” as the final answer (just that GPT chain-of-thought text sounds really funny). When the model samples the analysis channel token, a parser routes the following text into a separate stream exposed through the Responses API. For closed models, presumably a simple model redacts and summarizes it.

Reasoning Effort

How much budget goes to reasoning? Earlier APIs exposed reasoning token budgets, making it seem like a property of the sampling process. In reality, reasoning effort is baked into the system prompt. GPT-OSS puts this into the system prompt:

Reasoning: low

That’s it. Training produces the resulting behavior, such as emitting the token sequence that switches to the `ana

Obsidian 证据

  • 来源 digest: AK RSS Digest 2026-08-20(2026-08-20,评分 8.4)
  • 原文经 opencli web read / opencli arxiv paper 抓取核对,关键数字与摘要均锚定抓取内容。