模型与实验室 4.0 · 优秀 2026-08-27 · 论文

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

企业文档问答评测长期受制于公司不愿公开内部通讯合成数据集又过于简单CorporateBench 提供经人工校验的多任务问答基准:语料超过 23 万份文档,模拟 12 到 10000 名员工的四家合成企业,语料采样自一个时间演化跨文档逻辑一致的知识库,即使数十万文档也保证逻辑自洽基准从信息抽取与知识库查询两个维度评测对五个 LLM 的评测显示:输入规模越接近真实企业尺度,表现越差CorporateBench 为企业通讯推理补上了评测生态的关键空白

打开原文回到归档

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

  • ID: dd859479
  • 原文链接: https://arxiv.org/abs/2608.27391
  • PDF: https://arxiv.org/pdf/2608.27391v1
  • 作者: Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
  • 日期: 2026-08-27
  • 更新: 2026-08-27
  • 分类: models
  • 来源类型: paper
  • 标签: benchmark, enterprise, temporal-knowledge-base, qa, evaluation
  • 质量评分: 4/5
  • 抓取时间: 2026-08-31T12:36:54Z
  • 备注: Accepted to EMNLP Findings

中文导读

企业文档问答评测长期受制于公司不愿公开内部通讯合成数据集又过于简单CorporateBench 提供经人工校验的多任务问答基准:语料超过 23 万份文档,模拟 12 到 10000 名员工的四家合成企业,语料采样自一个时间演化跨文档逻辑一致的知识库,即使数十万文档也保证逻辑自洽基准从信息抽取与知识库查询两个维度评测对五个 LLM 的评测显示:输入规模越接近真实企业尺度,表现越差CorporateBench 为企业通讯推理补上了评测生态的关键空白

为什么值得关注

企业文档问答评测长期受制于公司不愿公开内部通讯合成数据集又过于简单CorporateBench 提供经人工校验的多任务问答基准:语料超过 23 万份文档,模拟 12 到 10000 名员工的四家合成企业,语料采样自一个时间演化跨文档逻辑一致的知识库.

论文于 2026-08-27 提交至 arXiv(分类:c, s, ., A, I, ,, , c, s, ., C, L, ,, , c, s, ., I, R, ,, , c, s, ., L, G),arXiv 摘要页面:https://arxiv.org/abs/2608.27391。

关键信息

  • 论文标题:CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
  • 作者:Sil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking, David Mimno, Bishan Yang, Igor Labutov
  • arXiv:https://arxiv.org/abs/2608.27391
  • 发布时间:2026-08-27
  • arXiv 分类:c, s, ., A, I, ,, , c, s, ., C, L, ,, , c, s, ., I, R, ,, , c, s, ., L, G
  • 关联标签:benchmark, enterprise, temporal-knowledge-base, qa, evaluation

English Abstract

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.

English Summary

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。