OpenAI ships ChatGPT health; o3 re-solves 4.8% of rare

openaichatgptgpt-5-5-instanthealth-airare-diseaseboston-childrens+10high-risk claims
OpenAI health intelligence announcement card with the title 'Using AI to help physicians diagnose rare genetic diseases affecting children' over a dark background
Image: OpenAI / rare disease diagnosis announcement (June 18, 2026)

On June 18, 2026, OpenAI published two health stories. The first is a consumer ChatGPT update on GPT-5.5 Instant rated higher than physician-written responses on a 3,500-response panel, with 71% fewer flagged factuality issues. The second is a peer-reviewed NEJM AI paper in which Boston Children’s Manton Center, Harvard, and OpenAI used o3 Deep Research to reanalyze 376 previously unsolved rare-disease cases, surface candidate diagnoses for 18 of them, and report 4.8% additional diagnostic yield after expert review under ACMG/AMP and CLIA-certified confirmation. Two caveats: neither story is evidence that patients or clinicians should use ChatGPT to diagnose disease; the NEJM AI study is retrospective on heterogeneous cohorts with unblinded reviewers.

Story A — ChatGPT product update

GPT-5.5 Instant is now the default model for free users in ChatGPT. On OpenAI’s hardest health evaluations, including HealthBench Professional, OpenAI reports it reaches performance comparable to its frontier Thinking models. The scale: 230M weekly ChatGPT health users; 3,500 physician-panel reviewed responses; 260+ physicians across 60 countries; 71% drop in flagged factuality issues. The 3,500-response set, criteria, and panel are OpenAI’s own.

Story B — NEJM AI study

Researchers applied o3 Deep Research to 376 previously unsolved cases already through multiple pipelines and multidisciplinary team review at Boston Children’s Manton Center. The de-identified packet per case was standardized HPO terms, clinician notes, age and gender metadata, and a filtered variant table. The workflow acted as an explanation-first reasoning layer on top of existing genomic pipelines: the model had to connect clinical features, inheritance pattern, variant evidence, and the scientific literature into a justification a human reviewer could interrogate. At least two reviewers used the ACMG/AMP framework; a finding counted only after CLIA-certified laboratory confirmation.

Validation on solved cases preceded the unsolved cohort: correct gene and variant in 48 of 51; correct diagnosis in 45 of 57 neuromuscular. Self-reported confidence tracked with correctness: mean 85.6 for correct calls vs 42.1 for incorrect.

CohortCasesDiagnoses surfacedYield
Neurodevelopmental1001010.0%
Neuromuscular disease6146.6%
Sudden unexpected death in pediatrics20021.0%
Early psychosis15213.3%
Total376184.8%

7 of 18 diagnoses were rediscoveries — established outside the local research workflow but absent from the local record. That is a data-integration finding, not a model-capability finding. Worked examples: an early-psychosis case where the model inferred a 22q11.2 deletion (DiGeorge) not in the input data, confirmed by follow-up sequencing; a neurodevelopmental case highlighting an 11-amino-acid deletion in S1PR1 as a “possible novel mechanistic explanation”; a neuromuscular case (Kyra) diagnosed with myofibrillar myopathy linked to a frameshift in HSPB8 after a near-20-year diagnostic journey.

Why it matters

Practical advice

1. Patients and families. ChatGPT is not a clinical diagnostic tool — the NEJM AI study is explicit the model did not diagnose any patient. Real workflow: centers like Boston Children’s Manton Center can re-run unsolved cases through LLM-assisted workflows. Ask the clinical team; the 4.8% yield is on cases that already failed multiple pipelines.

2. Clinicians and informatics teams. Two surfaces, two artifacts. GPT-5.5 Instant is a consumer product; o3 Deep Research is the research workflow. ChatGPT for Clinicians (April 22, 2026) is the BAA-backed clinical surface; HealthBench (May 12, 2025) is the open eval — use it to validate on your specialty, not the headline 3,500-response panel.

3. Researchers. Reproducible scaffold in the paper (DOI 10.1056/AIcs2501343): HPO terms, clinician notes, age/gender, filtered variant table → o3 Deep Research → ACMG/AMP with ≥2 reviewers → CLIA-certified confirmation. The 13.3% early-psychosis sub-cohort has wide CIs (n=15). The S1PR1 / HSPB8 / CDK13 signals are candidate hypotheses, not discoveries.

What to watch

  1. Prospective, multi-center studies comparing LLM-assisted reanalysis with standard practice.
  2. The OpenAI Foundation grant to the Manton Center for a platform-agnostic, low-cost genetics AI copilot.
  3. Regulatory clarity, especially FDA. The NEJM AI study is research, not a cleared device.
  4. Independent reproduction of the 3,500-response physician-panel results.
  5. Generalization of the 4.8% yield beyond Boston Children’s.

Risks and caveats

  1. The model did not diagnose any patient. Every diagnosis passed through physician review using ACMG/AMP and CLIA-certified confirmation.
  2. The 4.8% yield is on previously unsolved cases already through expert pipelines. Not a general-population rate.
  3. The 7 of 18 rediscoveries are operationally important. The challenge is data integration, not a model capability gap.
  4. The “rated higher than physician-written responses” claim is on OpenAI’s own panel, set, and criteria. Treat as OpenAI’s report until independent reproduction.
  5. The study is retrospective on heterogeneous cohorts with unblinded reviewers. No measurement of time saved, cost, false-positive workload, or changes in care.
  6. No FDA clearance; no general HIPAA statement for ChatGPT. ChatGPT for Clinicians and OpenAI for Healthcare carry their own BAA / HIPAA posture.
  7. The early-psychosis 13.3% cohort is small (15 cases, 2 diagnoses, wide CI).

Verdict

OpenAI shipped two health stories on 2026-06-18: a ChatGPT update on GPT-5.5 Instant (3,500-response panel, 71% fewer flagged factuality issues), and a NEJM AI paper in which o3 Deep Research reanalyzed 376 previously unsolved cases at Boston Children’s Manton Center, surfacing 18 diagnoses (4.8% yield; 7 of 18 rediscoveries). The clinical boundary is the load-bearing caveat — physicians made every diagnosis, the study is retrospective with unblinded reviewers. The takeaway is the workflow — explanation-first reasoning, ACMG/AMP review, CLIA-certified confirmation — not the headline number.

Sources