AI 编程 4.0 · 优秀 2026-08-12 · 论文

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

论文推出专为印度及其他中低收入国家临床场景设计的 RAG 系统 VITA,取从疾病-特定指南本土词汇多语言资源在 HealthBench 上与新一代 frontier LLM 相当或更佳,反击通用 LLM 在临床上已超越专业工具的广泛说法

打开原文回到归档

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

  • ID: a2ce6871
  • 原文链接: https://arxiv.org/abs/2608.12138
  • PDF: https://arxiv.org/pdf/2608.12138v1
  • 作者: Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
  • 日期: 2026-08-12
  • 更新: 2026-08-12
  • 分类: coding
  • 来源类型: paper
  • 标签: rag, clinical, healthbench, lmic, frontier-llms, cs.cl, cs.cl-cs.ai-cs.hc-cs.ir-cs.lg
  • 质量评分: 4/5
  • 抓取时间: 2026-08-14T12:20:00Z

中文导读

论文评估 VITA——一个专为印度及其他中低收入国家(LMIC)临床场景构建的检索增强生成(RAG)系统,其检索语料包含疾病特异性指南、印度特异性抗菌素耐药数据、国家处方集约束和资源受限护理协议。在 HealthBench 的 4,023 道英文题(占基准 80.5%)上用 GPT-4.1 做裁判:VITA 以 51.9% 的 rubric 得分率排名第一,领先 GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)、Claude Sonnet 4.6(37.3%),并在 45.4% 的题目上得分最高。为检验对更新模型的稳健性,作者又在 500 题子集上重跑(对比 GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3,用与所有被测系统无血缘关系的开源裁判 DeepSeek-V4-Pro 打分):差距收窄至持平——VITA 与 GPT-5.5 在题均分上统计不可区分,但 VITA 在加权分上领先且赢得题目数最多。VITA 的准确性与完整性优势在裁判中立后保持,沟通分较低。结论:专用临床 RAG 在开放基准上仍可与 frontier LLM 竞争,语料特异性是提升 grounding 的设计变量,代价是沟通打磨。

为什么值得关注

对"通用 LLM 在临床已超越专用工具"的流行说法给出反例:一个语料特异的 RAG 系统在公开基准 + 公开评分 rubric 下与最新 frontier 模型打平甚至局部领先。对做垂直领域 RAG 的人来说,这篇提供了"语料特异性 vs 沟通打磨"的取舍证据,以及用中立裁判做稳健性检验的方法学范式。

关键信息

  • 论文标题:A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
  • 作者:Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
  • arXiv:https://arxiv.org/abs/2608.12138
  • 发布时间:2026-08-12
  • arXiv 分类:cs.CL, cs.AI, cs.HC, cs.IR, cs.LG
  • 关联标签:rag, clinical, healthbench, lmic, frontier-llms
  • 论文备注:2 tables
  • 可复现性:VITA 架构与语料为专有,但基准、医生撰写的 rubric、以及完整响应与评分输出全部公开,可供独立验证

两轮评测关键数字

| 轮次 | 裁判 | VITA | 对照组 | |---|---|---|---| | 4,023 题全量(80.5% of HealthBench) | GPT-4.1 | 51.9%(第一,45.4% 题目最高分) | GPT-5.4 46.1% / o4-mini 44.3% / Gemini 3.1 Pro 42.6% / Claude Sonnet 4.6 37.3% | | 500 题子集(新模型 + 中立裁判) | DeepSeek-V4-Pro | 题均分与 GPT-5.5 统计不可区分;加权分领先、赢题最多 | GPT-5.5 / Claude Opus 4.8 / Gemini 3.5 Pro / Grok 4.3 |

English Abstract

General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

English Summary

The paper evaluates VITA, an LMIC-focused clinical RAG system retrieving from disease-specific guidelines, India-specific AMR data, national formulary constraints, and resource-limited care protocols. On 4,023 HealthBench questions with a GPT-4.1 judge it ranks first at 51.9% rubric points, ahead of GPT-5.4 (46.1%) and other frontier models; a 500-question re-run against newer models with a lineage-neutral judge (DeepSeek-V4-Pro) narrows the gap to statistical parity with GPT-5.5 on mean score while VITA still leads points-weighted and wins the most questions. Accuracy/completeness advantages persist while communication scores lag, supporting corpus specificity as a grounding-improving design variable at some cost in polish.

Obsidian Notes

  • 内容由 opencli arxiv paper 2608.12138 拉取 arXiv 元数据与摘要生成,评测数字均来自论文摘要。
  • 中文导读与价值判断锚定在条目已有摘要与论文摘要、作者、日期、分类信息上;未补充论文摘要之外的实验细节。