研究与学习 4.0 · 优秀 2026-09-29 · 论文

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

指出现有长上下文评测已无法区分现代 harness(各 harness 准确率饱和评估成本相近),推出同时衡量有效性与效率的基准:任务需要词汇检索与语义匹配等多样检索策略,以及对全局/局部上下文的策略性自适应推理大量上下文语义相关但每步只有少量真正有用,既构成搜索难题也造成不同处理策略的准确率-成本权衡差异,例如用证据散点找出满足多条件的所有人物的任务

打开原文回到归档

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

  • ID: 890e947f
  • 原文链接: https://arxiv.org/abs/2609.38137
  • PDF: https://arxiv.org/pdf/2609.38137v1
  • 作者: Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye
  • 日期: 2026-09-29
  • 更新: 2026-09-29
  • 分类: learning
  • 来源类型: paper
  • 标签: arxiv, benchmark, long-context, harness-evaluation, retrieval
  • 质量评分: 4/5
  • 抓取时间: 2026-10-01T12:21:44+00:00

中文导读

论文指出现有长上下文评测已经饱和:各 harness 准确率接近、评估成本也大体相同,无法区分现代 harness。作者提出 LongHarness Bench,同时衡量长上下文 harness 的有效性与效率。任务要求多样检索策略(词法搜索、语义匹配)加上对全局与局部上下文的策略性自适应推理;大量上下文在语义上相关,但每一步只有一小部分真正有用——这既构成困难的搜索问题,也让不同处理策略呈现不同的准确率-成本权衡。典型例子:从散落在多个文档中的证据里找出满足若干条件的所有人物,先检查最具区分度的条件即可先行缩小搜索范围,再验证其余条件。

作者用多族前沿语言模型搭配 4 个 SOTA harness 做评测:最强组合在 4 个评测套件上的 macro-average 准确率也只有 68%,基准仍有区分力。更关键的发现是,同一个底层模型在不同 harness 下效率差异显著——效率因此被确立为长上下文评测的重要轴,该基准也为开发"策略性处理上下文而非穷举式处理"的 harness 提供了试验场。

为什么值得关注

长上下文评测已饱和?LongHarness Bench 用高干扰检索+推理任务同时量化 harness 的准确率与成本。对做长上下文/harness 工程的人来说,这是第一个把"效率"作为一等公民的长上下文基准:模型相同、harness 不同,成本可能差出一个档位,选型时只看准确率会漏掉一半信息。

关键信息

  • 论文标题:LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
  • 作者:Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye
  • arXiv:https://arxiv.org/abs/2609.38137
  • 发布时间:2026-09-29
  • arXiv 分类:cs.CL
  • 关联标签:arxiv, benchmark, long-context, harness-evaluation, retrieval

English Abstract

Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.

English Summary

Argues existing long-context evaluations are saturated and cannot distinguish modern harnesses, and introduces a benchmark scoring both effectiveness and efficiency. Tasks require diverse retrieval strategies (lexical search, semantic matching) plus strategic adaptive reasoning over global and local context; much of the context is semantically relevant but only a small subset useful at each step, creating a hard search problem and distinct accuracy-cost tradeoffs across processing strategies.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。