A research team from Boston Children's Hospital, Harvard University, and OpenAI has announced that it reanalyzed 376 rare-disease cases that had previously gone unsolved and arrived at 18 new diagnoses with the help of AI[1]. The tool was OpenAI's o3 Deep Research reasoning model. The findings were published on June 18, 2026, in the medical journal NEJM AI. Notably, the AI did not make diagnoses in place of physicians; instead, it surfaced leads for experts to review.

Reanalyzing 376 Cases That Even Specialists Could Not Solve

Even after genomic sequencing, roughly half of people with rare diseases never receive a clear genetic diagnosis[1]. The clues may sit within the medical data, but finding them can mean sifting through thousands to millions of genetic variants, fragmented clinical records, and a rapidly changing body of research. As new gene-disease relationships are reported, cases that once resisted analysis can become interpretable later.

The team applied o3 Deep Research to de-identified clinical and genomic information from 376 cases that had been analyzed before but remained unsolved[1]. After experts reviewed the model's candidates and completed additional testing and clinical confirmation, diagnoses were established in 18 cases. That works out to an additional 4.8 percent diagnostic yield after specialist analysis, a modest but meaningful gain for a set of cases already examined many times.

Designed to Produce Reviewable Hypotheses, Not Diagnoses

What matters in this workflow is that the model acts as an explanation-first reasoning layer sitting on top of existing genomic pipelines[1]. Rather than simply returning a ranked list of candidate genes, it was designed to connect clinical features, inheritance patterns, variant evidence, and the scientific literature into evidence-linked hypotheses that a human reviewer can interrogate.

The outputs were never treated as diagnoses on their own[1]. At least two people evaluated each candidate using the ACMG/AMP framework that clinical labs rely on, resolving disagreements by consensus. A finding counted as a diagnosis only after a variant was classified as pathogenic or likely pathogenic, confirmed in a CLIA-certified laboratory, and returned to the family by the clinical team.

In validation, the workflow recovered the correct gene and variant in 48 of 51 already-solved cases, 45 of 57 neuromuscular cases, and named the correct gene in every one of 15 long-read genome cases[1]. The model's self-reported confidence also separated cleanly, averaging 85.6 for correct calls versus 42.1 for incorrect or unknown ones, which helped reviewers prioritize.

A Woman Diagnosed After Nearly Two Decades

One result came from the neuromuscular cohort: the case of Kyra[1]. At age 9, her muscle weakness first surfaced when she could no longer get low into her karate stances, and a diagnosis eluded her for nearly 20 years. She relied on a ventilator and a wheelchair by age 13, but the reanalysis pinpointed a frameshift variant in the HSPB8 gene and diagnosed a form of myofibrillar myopathy, in which abnormal protein clumps build up in muscle fibers. A genetic counselor called her about a week before her 28th birthday.

The model also inferred structural changes that were not in the input data[1]. In an early-psychosis case, it connected a run of low-quality calls on chromosome 22 with the patient's cardiac, immune, neurodevelopmental, and psychiatric features, proposed a 22q11.2 deletion linked to DiGeorge syndrome, and the hypothesis was confirmed with follow-up sequencing. It also suggested a possible new mechanism for vitiligo, a condition that causes loss of skin pigment, involving a deletion in the S1PR1 gene.

Promise, and Limits That Should Not Be Overlooked

The team is careful to bound its claims[1]. This was a retrospective study of past cases, the cohorts were heterogeneous, and reviewers saw the model's confidence scores during evaluation. The researchers did not measure time to diagnosis, cost, false-positive burden, or effects on care. Because large language models can misread context and produce plausible explanations that fail on closer inspection, every result passed through human judgment and clinical confirmation. The AI widened the search; experts narrowed it down.

For the next stage, the Manton Center will lead the work through a grant from the OpenAI Foundation to build a platform-agnostic, low-cost genetics AI copilot[1]. The effort has been widely covered in international media as a concrete example of applying AI to medical research[2]. "Researchers like Catherine and me can't possibly keep 8,000 different diseases in our heads. That's the power of AI," said Alan Beggs, director of the Manton Center[1].

Summary

Using OpenAI's o3 Deep Research reasoning model, a reanalysis of 376 rare-disease cases that even specialists had left unsolved produced 18 new confirmed diagnoses[1]. The model does not make the diagnosis itself; it offers evidence-linked, reviewable hypotheses, with physicians making the final call. Mindful of both its accuracy and its limits, the study marks a step toward making periodic reanalysis, revisiting old cases as knowledge advances, a practical reality.

Source: https://openai.com/index/diagnose-rare-childhood-diseases
Source: https://www.nbcnews.com/tech/innovation/ai-boston-childrens-hospital-diagnose-rare-diseases-kids-openai-rcna350387