Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
- ID: 1a072b28
- 原文链接: https://arxiv.org/abs/2609.04194
- PDF: https://arxiv.org/pdf/2609.04194v1
- 作者: Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
- 日期: 2026-09-03
- 分类: learning
- 来源类型: paper
- 标签: chain-of-thought, faithfulness, process-reward-models, llm-as-judge, colm-2026, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-09-05T12:33:07Z
- 会议/备注: Published at COLM 2026
中文导读
追问 CoT 文本是否真的编码了步级重要性:作者把推理步骤的重要性操作化为其优势(advantage)——该步骤对期望奖励的改变,用其与 LLM judge 从文本中判断的重要性对比。结论以标题点题:可读性不等于可解释性,CoT 看起来透明,但步骤文本携带的关于其功能作用的信息有限,直接动搬用 LLM judge 做错误诊断与步级监督的可靠性。
为什么值得关注
与 04198 同一天形成组合拳:judge 读数不可靠 + CoT 步级读数不编码功能作用,步级监督两大前提同时被检验。
关键信息
- 论文标题:Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
- 作者:Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
- arXiv:https://arxiv.org/abs/2609.04194
- 发布时间:2026-09-03
- arXiv 分类:cs.CL, cs.LG(primary: cs.CL)
- 关联标签:chain-of-thought, faithfulness, process-reward-models, llm-as-judge, colm-2026, arxiv
English Abstract
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.
English Summary
The paper asks whether chain-of-thought text actually encodes which reasoning steps matter: step importance is operationalized as its advantage (the change in expected reward from including that step, estimated via Monte Carlo rollouts), then LLM judges are evaluated against that ground truth. Capable judges beat a prevalence baseline but stay well below a noise ceiling; step importance is only partially recoverable from the reasoning trace.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。