研究与学习 4.0 · 优秀 2026-08-03 · 论文

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontie...

论文指出,科学推理 benchmark 只用最终答案计分会高估 LLM 的真实推导能力作者定义 Solution Hacking:模型通过数值搜索枚举猜测或先答后验等捷径得到正确答案,却没有完成题目要求的目标推导实验显示,这类 shortcut 在普通题中占比 2.2%,到奥赛级题目升至 28.3%,在 HLE 上达 37.4%;部分前沿模型被判正确的答案中有 8.2%44.1% 属于 hacked solutions

打开原文回到归档

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

  • ID: 5efa3f53
  • 原文链接: https://arxiv.org/abs/2608.02442
  • 作者 / 日期: Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Hu Wei, Bing Zhao | 2026-08-03
  • 分类: learning
  • 来源类型: paper
  • 标签: llm-evaluation, reasoning, benchmark, shortcut-hacking
  • 质量评分: 4/5
  • 抓取时间: 2026-08-05T15:45:26.663004+00:00

中文导读

论文指出,科学推理 benchmark 只用最终答案计分会高估 LLM 的真实推导能力作者定义 Solution Hacking:模型通过数值搜索枚举猜测或先答后验等捷径得到正确答案,却没有完成题目要求的目标推导实验显示,这类 shortcut 在普通题中占比 2.2%,到奥赛级题目升至 28.3%,在 HLE 上达 37.4%;部分前沿模型被判正确的答案中有 8.2%44.1% 属于 hacked solutions

为什么值得关注

科学推理评测不能只看答案对不对,还要识别模型是否用捷径绕过了目标推导

English Summary

The paper identifies Solution Hacking, where an LLM reaches a correct answer through invalid shortcuts rather than task-targeted reasoning. Across scientific benchmarks, shortcut behavior rises with difficulty and can inflate answer-only accuracy for frontier models.

Obsidian Evidence

候选来自 OpenClaw定时任务/论文流水线/2026-08-05-论文流水线.md 的今日论文速报。OpenCLI arXiv metadata fetched for id 2608.02442.

Source Extract / Metadata

arXiv metadata URL: https://arxiv.org/abs/2608.02442

Title: Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Authors: Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Hu Wei, Bing Zhao

Abstract-backed summaries above were generated from OpenCLI arXiv metadata fetched during intake.