Using AI to help physicians diagnose rare genetic diseases affecting children
Source: OpenAI Original URL: https://openai.com/index/diagnose-rare-childhood-diseases Author: OpenAI Published: June 18, 2026 Captured: 2026-06-20 20:18 (Asia/Shanghai)
English
Using AI to help physicians diagnose rare genetic diseases affecting children
In an NEJM AI study, experts used an OpenAI reasoning model to reanalyze 376 previously unsolved cases and surface leads for 18 diagnoses.
Read the study abstract at NEJM AI
Even with genomic sequencing, many people with rare diseases never receive a clear genetic diagnosis. Roughly half remain undiagnosed after extensive testing and specialist review. Their medical data may contain clues but finding them can require sifting through thousands to millions of possible genetic variants, fragmented clinical records, and rapidly changing scientific literature.
As new gene-disease relationships, case reports, and classification evidence accumulate, unsolved cases can become newly interpretable.
Researchers from Boston Children's Hospital's Manton Center for Orphan Disease Research, Harvard University, and OpenAI used the OpenAI o3 Deep Research reasoning model to analyze de-identified clinical and genomic information from 376 previously analyzed cases that remained unsolved. The model surfaced evidence-linked candidate explanations for researchers and clinicians to review. Following expert review, additional testing, and clinical confirmation, physicians established diagnoses in 18 cases—an additional diagnostic yield of 4.8% after earlier analysis by specialists. This study was published on June 18, 2026, in NEJM AI and shows how an AI-assisted research workflow can help experts generate leads when revisiting some of the most difficult cases.
Many of these cases had evaded years of expert analysis. In this study, OpenAI o3 Deep Research helped researchers identify leads that were later assessed through established clinical processes, suggesting that expert-led periodic reanalysis could become more scalable as knowledge evolves. The model did not diagnose any patient or make any clinical decision. It produced evidence-linked hypotheses for specialists to review and, where appropriate, investigate through additional testing and confirm in a clinical laboratory.
Why an old case can contain a new answer
An inconclusive genetic test is not always a permanent finding. A patient's phenotype descriptions, test results, and family history can be split across databases that use different identifiers, formats, and vocabularies. Linking those records is difficult, so even specialists can miss a diagnosis. Experts may also sequence a child's genome before a relevant gene or its variants have been linked to disease. As scientific knowledge advances, the same data can reveal answers that were previously impossible to uncover.
Rare-disease reanalysis is both a scientific and a maintenance problem. The patient's genome may stay the same, but the evidence around it keeps changing: researchers link new genes and variants to disease, labs reclassify old variants, and case databases and papers accumulate new observations. Each update can make an old inconclusive case worth revisiting, so many institutions inherit a growing backlog of genomes to keep in sync with a moving knowledge base.
In this study, researchers designed the workflow so that the model acted as an explanation-first reasoning layer on top of existing genomic pipelines. Instead of returning only a ranked gene, it was asked to connect the clinical features, inheritance pattern, variant evidence, and scientific literature into a justification that a human reviewer could interrogate.
How the reanalysis worked
For each case, the team assembled a de-identified packet containing standardized Human Phenotype Ontology terms to describe the patient's clinical presentation, occasional clinician notes and any descriptive clinical diagnosis, metadata such as age and gender, and a filtered variant table. The table captured each variant's rarity, its predicted effect on the encoded protein, ClinVar classification, and signal quality across available family members. Most cases included data from the child and both biological parents.
The team asked the model to propose the most plausible molecular explanation and to show its work. Researchers then reviewed the outputs using the same ACMG/AMP framework that clinical labs use to classify genetic variants. At least two team members reviewed each candidate, disagreements were resolved by consensus, and a model output was never treated as a diagnosis. A finding counted as a diagnosis only after qualified experts reviewed the evidence, the variant was classified as pathogenic or likely pathogenic, a CLIA-certified laboratory confirmed it, and the clinical team returned the result to the family.
Before analyzing unsolved cases, the team refined the workflow on cases with established diagnoses. It recovered the correct gene and variant in duplicate runs for 48 of 51 cases that included a variety of rare conditions. In a set of 57 neuromuscular cases, the workflow returned the correct diagnosis in duplicate runs for 45 of the cases. In a 15-case long-read genome set, it named the correct gene in every case and both disease-causing alleles in 12 cases. These evaluations aided in prompt development and showed where expert review remained essential.
The model's self-reported confidence scores tracked with correct diagnoses in these previously solved cases: the mean minimum score was 85.6 for consistently correct calls and 42.1 for incorrect or unknown calls. The scores were not calibrated probabilities, and the team did not use them as a substitute for evidence or clinical adjudication. But they were helpful in guiding the expert reviewers to focus on the most promising candidate diagnoses.
What the researchers found
The team then applied the workflow to four groups of previously unsolved cases: children with neurodevelopmental conditions, people with rare neuromuscular disease, children and adolescents with early psychosis, and cases of sudden unexpected death in pediatrics. These were not fresh cases awaiting a first review. Many had already been examined by multiple commercial or institutional pipelines and discussed by multidisciplinary teams.
Results by cohort
| Cohort | Cases | Diagnoses surfaced | Yield | | --- | --- | --- | --- | | Neurodevelopmental | 100 | 10 | 10.0% | | Neuromuscular disease | 61 | 4 | 6.6% | | Sudden unexpected death in pediatrics | 200 | 2 | 1.0% | | Early psychosis | 15 | 2 | 13.3% | | Total | 376 | 18 | 4.8% |
The early psychosis cohort was small, so its percentage has a wide confidence interval. Yield also reflects how likely each cohort was to have a single-gene explanation.
After the model surfaced candidates and experts completed review and clinical confirmation, physicians established diagnoses in 4.8% of the cases. That rate is modest but meaningful in this population because previous expert reviews had not resolved the cases. Similar reanalysis studies report single-digit gains in heavily reviewed cases; higher yields usually come from studies containing new cases or well-known disorders awaiting genetic confirmation.
Of the 18 diagnoses, 7 were rediscoveries: diagnoses established outside the local research workflow but absent from the record the team reviewed. In several cases, the variants were already listed as pathogenic or likely pathogenic in public databases, highlighting the operational challenge of synthesizing information across data sources.
Demonstrating flexibility when identifying variants
In one early-psychosis case, the model inferred a structural event in the genome that was not listed in the input data. It connected a run of low-quality calls on chromosome 22 with the child's cardiac, immune, neurodevelopmental, and psychiatric features, then hypothesized a 22q11.2 deletion associated with DiGeorge syndrome. This hypothesized variant was confirmed with follow-up genome sequencing.
Although the prompt asked for one monogenic cause, the model sometimes surfaced two genes that better explained a complex presentation. Variants in _LAMA2_ and _FOXP1_ together helped account for muscle and neurodevelopmental features in one case; another had a previously unrecognized digenic explanation involving _TTN_ and _SRPK3_.
Producing a testable, biologically coherent hypothesis
In addition to diagnoses, the model also identified a possible novel mechanistic explanation for a condition called vitiligo. In one neurodevelopmental case, the model highlighted an 11-amino-acid deletion in _S1PR1_ in a person with vitiligo. _S1PR1_ encodes a cell-surface receptor involved in signaling, immune-cell movement, and tissue biology. The model integrated evidence suggesting that the deletion could alter receptor structure and signaling in ways that reduce pigment production while also helping immune cells persist in the skin.
The proposed _S1PR1_-vitiligo relationship requires additional experimental validation but it illustrates a powerful role for AI in translating scattered findings from structural biology, immunology, and clinical genetics into concrete, testable hypotheses.
The team also saw possible phenotype expansion in the neuromuscular cohort. Damaging variants in _HSPB8_ and _CDK13_ did not perfectly match the genes' best-known disorders, suggesting a broader clinical spectrum that more cases and laboratory work will need to test.
Case study: Kyra's diagnosis after nearly two decades
It started in karate class, when Kyra's mother noticed that her 9-year-old daughter was not getting as low in her stances as she used to. Kyra was also slowing down during soccer practice and staying up on her toes while walking and running. Her pediatrician could not identify the cause of her muscle weakness, so he referred her to a specialist. What followed was a nearly 20-year journey through tests, treatments, and consultations without a diagnosis.
Kyra's case was one of the four diagnoses surfaced in the neuromuscular cohort. The team linked her condition to a frameshift variant in _HSPB8_ and diagnosed a form of myofibrillar myopathy, in which abnormal protein structures build up in muscle fibers and contribute to weakness. A genetic counselor from the Manton Center called Kyra about a week before her 28th birthday.
By then, Kyra had spent much of her life adapting to the disease. She was dependent on a ventilator and in a wheelchair by the time she was 13, although her condition has since plateaued. Though Kyra's form of myofibrillar myopathy is so rare that little is known about its long-term course, the diagnosis has brought some closure.
Limitations
This study shows that a general-purpose reasoning model can contribute to retrospective genomic reanalysis by combining phenotype, inheritance, variant annotations, data-quality patterns, and scientific literature into reviewable hypotheses. It also shows why periodic reanalysis matters: some answers surface only after knowledge advances or fragmented records are brought together.
This research is not evidence that patients, clinicians, or customers should use OpenAI models to diagnose disease or make medical decisions. It does not describe or endorse an intended customer use of OpenAI o3 Deep Research, ChatGPT, or any other OpenAI product for diagnosis. The model did not diagnose any participant; physicians and other qualified clinical experts made every diagnosis through established review, testing, and clinical-confirmation processes.
The study was retrospective, the cohorts were heterogeneous, and reviewers were not blinded to model confidence. The researchers did not measure time saved, cost, clinician effort, false-positive workload, or changes in care. Nor did they systematically evaluate other forms of genetic variation such as structural variants, repeat expansions, deep-intronic changes, or mosaicism.
Large language models can misread context or produce plausible explanations that fail upon closer inspection. Therefore, every result passed through human adjudication and clinical confirmation. The model widened the search and focused the subsequent human-led analysis; it did not decide what information or diagnosis should be returned to a family.
This study used de-identified information, with no protected health information utilized or transmitted outside approved environments. Broader clinical deployment will require the same attention to privacy, security, auditability, and local regulation that applies to all medical care. Model access does not replace sequencing infrastructure, genetic counseling, confirmatory testing, or specialist judgment.
"The bottleneck is time. An expert can devote only so much of their day to any one particular person."
— Dr. Catherine Brownstein, Boston Children's Hospital's Manton Center for Orphan Disease Research
"Researchers like Catherine and me can't possibly keep 8,000 different diseases in our heads. That's the power of AI."
— Alan Beggs, director of the Manton Center for Orphan Disease Research
What comes next
Prospective, multi-center studies should compare LLM-assisted reanalysis with standard practice on diagnostic yield, time to a candidate, clinician effort, false-positive burden, cost, and effects on care. Versioned prompts, reference checks, audit logs, and calibrated uncertainty will be important for reproducibility and safety. Such studies would still require qualified clinicians to evaluate evidence, order appropriate tests, and make any diagnosis or treatment decision.
This study used OpenAI o3 Deep Research. Newer general-purpose models can search and synthesize more scientific material, while purpose-built systems such as GPT-Rosalind are designed for deeper life-sciences work, including variant effects on protein structure and function. Those capabilities were not tested here and will require their own evaluations and access controls.
While OpenAI helped support this initial research study, the Manton Center will lead the next stage of the work through a grant from the OpenAI Foundation. The grant will support the Center's broader effort to develop a platform-agnostic, low-cost genetics AI copilot that helps clinical teams analyze rare disease cases more quickly and consistently.
The longer-term research opportunity is to explore whether expert-led AI-assisted reanalysis can help scientific understanding keep pace with discovery. The promise is not that AI replaces a doctor's diagnosis, but that carefully evaluated research tools may help specialists identify evidence worth investigating. For thousands of families, today's unanswered questions do not have to remain unanswered forever.
中文
用 AI 协助医生诊断影响儿童的罕见遗传病
在一项发表于《NEJM AI》的研究中,专家使用 OpenAI 推理模型对 376 个此前未能确诊的病例进行了重新分析,为其中 18 个病例找到了诊断线索。
即便有基因测序技术,许多罕见病患者始终无法获得明确的遗传学诊断。在经过大量检测和专家会诊之后,仍有约一半病例得不到确诊。他们的医疗数据中可能藏有线索,但要找到这些线索需要在成千上万乃至上百万种可能的基因变异、碎片化的临床记录以及快速更新的科学文献中反复搜寻。
随着新的基因—疾病关联、个案报告和分类证据不断累积,曾经无解的病例可能变得可以被重新解读。
来自波士顿儿童医院 Manton 孤儿疾病研究中心、哈佛大学以及 OpenAI 的研究人员,使用 OpenAI o3 Deep Research 推理模型对 376 个此前经过分析但仍未确诊的病例的去标识化临床与基因组信息进行了分析。模型为研究者和临床医生输出了有证据支撑的候选解释。在专家审核、补充检测和临床确认之后,医生在 18 个病例中确立了诊断——这意味着在前期专家分析的基础上,又带来了 4.8% 的额外诊断产出。这项研究于 2026 年 6 月 18 日发表于《NEJM AI》,展示了 AI 辅助研究工作流如何帮助专家在最棘手的病例回顾中产生新的线索。
这些病例中很多已经经历过多年的专家分析。在本研究中,OpenAI o3 Deep Research 帮助研究人员识别出线索,随后经由既定的临床流程进行评估。这表明,随着知识的持续演进,由专家主导的周期性再分析有可能变得更加可扩展。模型本身没有为任何患者做出诊断或临床决策,它只是产出有证据支撑的假设,供专家审阅,并在合适的情况下通过额外检测进行验证,再在临床实验室中加以确认。
为什么一个老病例里可能藏着新答案
一次未得出结论的基因检测并不必然意味着永久无解。患者的表型描述、检测结果和家族史可能分散在使用不同标识符、格式和术语体系的多个数据库中。把这些记录关联起来并不容易,因此即便是专家也可能漏掉诊断。有时,专家在测序时,相关基因或其变异尚未与某种疾病建立关联;随着科学知识不断推进,同样一份数据有可能揭示出此前不可能发现的新答案。
罕见病的再分析既是科学问题,也是维护问题。患者基因组本身可能并未改变,但其周围的知识却在持续更新:研究人员不断将新的基因与变异和疾病关联起来;实验室会重新分类旧的变异;病例数据库和论文也在持续累积新的观察结果。每一轮更新都可能让一个原本无解的旧病例变得值得重新审视。因此,许多机构都背上了越来越庞大的"待重审基因组"清单,需要与不断流动的知识库保持同步。
在本研究中,研究人员将工作流设计为:让模型作为叠加在现有基因组分析流程之上的、强调解释优先的推理层。模型并非只返回一个排序后的基因,而是要把临床特征、遗传模式、变异证据和科学文献拼接成一段人类审稿者可以追问的论证。
再分析是如何开展的
对每一个病例,团队都会组装一个去标识化的数据包,其中包含:用标准化的人类表型本体(HPO)术语描述患者临床表现、偶发的临床医生笔记及描述性临床诊断、年龄和性别等元数据,以及一张经过筛选的变异表。变异表记录了每个变异的稀有度、对所编码蛋白的预测影响、ClinVar 分类,以及在可用家系成员中的信号质量。大多数病例都包含患儿及其生物学父母双方的数据。
团队要求模型提出最可能的分子学解释,并展示其推理过程。然后研究人员使用临床实验室用于变异分类的同一套 ACMG/AMP 框架来复核模型输出。每一位候选结果至少由两位团队成员独立审阅,意见不一致时通过协商达成共识;模型的输出从不直接作为诊断结果。一个发现要被算作"诊断",必须满足:经过资质专家审核证据、变异被判定为致病或可能致病、经 CLIA 认证的实验室加以确认,并由临床团队将结果告知家属。
在分析未解病例之前,团队先在已确诊的病例上对工作流进行了调优。在一组包含多种罕见病的 51 个病例中,模型在重复运行中正确还原了 48 例的基因与变异。在一组 57 例神经肌肉疾病病例中,重复运行中有 45 例给出了正确诊断。在一组 15 例的长读长基因组中,模型对每个病例都正确指出了致病基因,并在 12 例中准确命名了两条致病变异。这些评估帮助改进了提示词,并再次说明专家审核始终不可或缺。
在已解决的病例中,模型自报的置信度分数与诊断是否正确呈现出明显的对应关系:始终回答正确的案例最低分平均为 85.6,而回答错误或未知的案例最低分平均为 42.1。这些分数并非经过校准的概率,研究团队也没有把它们当作证据或临床判断的替代品。但它们确实能帮助专家审稿者把注意力集中到最有希望的候选诊断上。
研究人员发现了什么
随后,团队将工作流应用于四组此前未解的病例:神经发育状况患儿、罕见神经肌肉疾病患者、儿童与青少年的早发性精神病病例,以及儿科猝死病例。这些并非等待首次复核的新病例——其中许多已经经过了多个商业或院内流程的检测,并经过多学科团队讨论。
各队列结果
| 队列 | 病例数 | 浮现的诊断数 | 产出率 | | --- | --- | --- | --- | | 神经发育疾病 | 100 | 10 | 10.0% | | 神经肌肉疾病 | 61 | 4 | 6.6% | | 儿科猝死 | 200 | 2 | 1.0% | | 早发性精神病 | 15 | 2 | 13.3% | | 合计 | 376 | 18 | 4.8% |
早发性精神病队列规模较小,因此其百分比置信区间较宽。产出率也反映了各队列中存在单基因解释的可能性高低。
在模型给出候选、专家完成审核和临床确认之后,医生在 4.8% 的病例中确立了诊断。这一比例看起来并不算高,但考虑到此前已经过多轮专家会诊却仍未确诊,其意义不容小觑。类似的再分析研究在那些被反复审阅的病例中通常只能带来个位数的提升;较高的产出往往来自包含新病例或等待基因确诊的常见疾病的研究。
在 18 个新诊断中,7 个属于"再发现"——它们在团队所审阅的本地病历之外已被确诊。在若干病例中,相关变异早就在公共数据库中被列为致病或可能致病,这凸显了跨数据源综合信息的运营挑战。
鉴定变异时展现的灵活性
在一例早发性精神病病例中,模型在输入数据之外推断出了一个基因组层面的结构事件。它把 22 号染色体上一段低质量判读与患儿的 cardiac、免疫、神经发育和精神病学特征关联起来,并由此提出了与 DiGeorge 综合征相关的 22q11.2 缺失。这一假设的变异随后通过追加的全基因组测序得到确认。
虽然提示词要求只给出一个单基因解释,但模型有时会给出两个能更好解释复杂临床表现的基因。在一个病例中,_LAMA2_ 与 _FOXP1_ 的变异共同解释了肌肉和神经发育特征;另一个病例则揭示出涉及 _TTN_ 和 _SRPK3_ 的、此前未被识别的双基因(digenic)解释。
给出可被验证的、生物学上自洽的假设
除诊断之外,模型还针对一种名为白癜风(vitiligo)的疾病提出了一种可能的新机制解释。在一个神经发育病例中,模型标记出 _S1PR1_ 基因中一段长度为 11 个氨基酸的缺失,患者同时患有白癜风。_S1PR1_ 编码一种细胞表面受体,参与信号传导、免疫细胞迁移和组织生物学过程。模型综合证据指出,这一缺失可能通过改变受体结构与信号传导,既减少色素生成,又让免疫细胞更容易在皮肤中驻留。
_S1PR1_ 与白癜风的关联尚需更多实验验证,但这一案例很好地展示了 AI 在将结构生物学、免疫学和临床遗传学的零散发现整合为具体、可验证假设方面的能力。
在神经肌肉队列中,团队也观察到了表型扩展的可能性。_HSPB8_ 和 _CDK13_ 中的有害变异与这两个基因最知名的疾病并不完全吻合,提示其临床表型谱可能更广,仍需更多病例和实验加以验证。
案例研究:Kyra 在近 20 年后终于获得诊断
故事始于空手道课堂。Kyra 的母亲注意到,9 岁的女儿做马步下蹲时,已经不像以前那么低了。Kyra 在足球训练中也越来越慢,走路和跑步时开始踮着脚尖。儿科医生无法确定她肌肉无力的原因,于是把她转诊给专科医生。接下来的近 20 年里,她辗转于各种检查、治疗和会诊之间,始终没有得到明确诊断。
Kyra 的病例是神经肌肉队列中浮现的 4 个诊断之一。团队把她的病情与 _HSPB8_ 中的一个移码变异联系起来,诊断为一种肌原纤维性肌病——在这种疾病中,肌纤维内会堆积异常蛋白结构,进而导致肌力减弱。Manton 中心的一位遗传咨询师在 Kyra 28 岁生日前一周打电话告知了她的诊断结果。
到那时为止,Kyra 已经大半辈子在与疾病相处。她 13 岁时就需要依赖呼吸机并坐上了轮椅,但此后病情趋于稳定。虽然 Kyra 所患的这种肌原纤维性肌病极其罕见、长期病程所知甚少,但拿到诊断本身,已经为这个家庭带来了一些释然。
局限
本研究证明,一个通用推理模型可以通过把表型、遗传模式、变异注释、数据质量信号和科学文献综合起来,生成可供人类审阅的假设,从而对回顾性的基因组再分析做出贡献。它也再次说明周期性再分析为何重要:一些答案只有在知识不断推进,或碎片化的记录被重新整合之后才会浮出水面。
这项研究并不构成患者、临床医生或客户可以用 OpenAI 模型直接诊断疾病或做出医疗决策的证据。它既不描述也不背书 OpenAI o3 Deep Research、ChatGPT 或任何其他 OpenAI 产品用于诊断的预期客户用途。模型本身没有为任何参与者做出诊断;每一项诊断都由医生和其他具备资质的临床专家经过既定审核、检测和临床确认流程得出。
研究为回顾性设计,各队列异质性较强,审稿人也并未对模型置信度设盲。研究并未衡量节省的时间、成本、临床医生的工作量、假阳性负担或诊疗行为的变化,也没有系统评估其他类型的遗传变异,例如结构变异、重复扩展、深内含子变异或嵌合现象。
大语言模型可能会误读上下文,或给出看似合理但在仔细审视后并不成立的解释。因此,所有结果都经过人工裁定和临床确认两道关卡。模型所做的是把搜索范围扩大、把后续的人工分析聚焦;它并不决定哪些信息或诊断应该被告知家属。
研究使用去标识化数据,未在任何获批环境之外使用或传输受保护的健康信息。更广泛的临床部署仍需遵守适用于所有医疗行为的隐私、安全、可审计性和本地监管要求。模型访问并不能替代测序基础设施、遗传咨询、确认性检测和专家判断。
"瓶颈是时间。一位专家一天之内能花在任何一个具体患者身上的时间是有限的。"
— Dr. Catherine Brownstein,波士顿儿童医院 Manton 孤儿疾病研究中心
"Catherine 和我这样的研究者,根本不可能把 8000 种不同的疾病都记在脑子里。这就是 AI 的力量所在。"
— Alan Beggs,Manton 孤儿疾病研究中心主任
接下来
未来需要进行前瞻性、多中心的研究,把 LLM 辅助再分析与标准实践在诊断产出率、得到候选诊断的时间、临床医生工作量、假阳性负担、成本以及对诊疗的影响等方面进行对比。版本化的提示词、参考核对、审计日志以及经过校准的不确定性,都对可重复性和安全性至关重要。即便如此,仍需要具备资质的临床医生来评估证据、安排恰当的检测,并做出最终的诊断或治疗决策。
本研究使用的是 OpenAI o3 Deep Research。更新一代的通用模型可以检索并综合更多科学材料;与此同时,诸如 GPT-Rosalind 这类专用系统则面向更深度的生命科学工作,包括变异对蛋白结构与功能的影响。这些能力在本研究中没有得到测试,未来需要单独的评估与访问控制。
虽然 OpenAI 为这一初步研究提供了支持,但下一阶段的工作将由 Manton 中心通过来自 OpenAI 基金会的资助主导。该笔资助将支持中心打造一个平台无关、低成本的遗传学 AI 副驾驶,帮助临床团队更快、更一致地分析罕见病病例。
更长远的科研机会在于探索:由专家主导、AI 辅助的再分析能否帮助科学理解跟上新发现的步伐。它的意义并不在于 AI 取代医生的诊断,而在于经过严谨评估的研究工具能够帮助专家识别出值得进一步调查的证据。对成千上万个家庭来说,今天悬而未决的问题,并非注定永远无解。
*抓取方式:opencli browser default extract(绕过 Cloudflare Turnstile)* *Entry ID: 2155d7cb*