ContractScrub: A benchmark for final review of legal contracts
- ID: 50ab10ec
- 原文链接: https://arxiv.org/abs/2608.20204
- PDF: https://arxiv.org/pdf/2608.20204
- 作者: Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean
- 日期: 2026-08-20
- 更新: 2026-08-20
- 分类: industry
- 来源类型: arxiv
- 标签: legal, benchmark, long-context, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-23T05:25:57Z
中文导读
ContractScrub 是面向法律合同终审(scrubbing)的基准:对交易文书做错误与不一致的最终审查,考验长上下文推理一致性检查与命名实体识别,是法律场景落地的高价值测点
为什么值得关注
ContractScrub 是面向法律合同终审(scrubbing)的基准:对交易文书做错误与不一致的最终审查,考验长上下文推理一致性检查与命名实体识别,是法律场景落地的高价值测点
Grounding: first benchmark for contract scrubbing; contracts hand-crafted by experienced lawyers across error categories (misuse of defined terms, incorrect references, inconsistent language); only one frontier model reaches 0.75 macro average recall.
关键信息
- 论文标题: ContractScrub: A benchmark for final review of legal contracts
- 作者: Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean
- arXiv: https://arxiv.org/abs/2608.20204
- 发布时间: 2026-08-20
- arXiv 分类: cs.AI, cs.CL
- 关联标签: legal, benchmark, long-context, arxiv
English Abstract
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.
English Summary
ContractScrub benchmarks contract scrubbing, the final review of transactional agreements for errors and inconsistencies. The routine, long-document task aligns with frontier LLM strengths in long-context reasoning, consistency checking, and NER, making legal final review a measurable automation target.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。