HomeBlog › An AI found what 17 doctors missed
For patients & families

An AI found what 17 doctors missed. The lesson isn’t the AI.

In 2023 a mother did something nobody had managed in three years of appointments: she put her son’s entire medical history in front of a single reader.

Ha Son Nguyen, MD · Neurosurgeon, founder of UploMD · September 2026

You have probably seen the headline by now, in one form or another: ChatGPT diagnosed a boy that seventeen doctors couldn’t. It is a good headline. It is also, I think, the wrong reading of what happened — and the right reading is considerably more useful to you.

What actually happened

A boy named Alex was in pain for three years. Over that time he saw seventeen doctors. He had toothaches. His growth stalled. His balance and posture were off, and his mother noticed he would not sit cross-legged on the floor the way other children did. Each clinician found something, or found nothing, and none of it added up to a single explanation.

His mother, Courtney Hofmann, went line by line through his MRI notes and typed them into ChatGPT, along with everything she had been tracking herself. The model suggested tethered cord syndrome — a condition where the spinal cord is abnormally anchored at its lower end, putting tension on it as a child grows.

She took that to a neurosurgeon, Dr. Holly Gilmer, who reviewed the imaging, agreed, and operated. NEJM AI later brought both the mother and the surgeon onto its Grand Rounds program to work through the case — which tells you the profession treated this as something worth understanding rather than a novelty to wave off.

Why the headline is the wrong lesson

Tethered cord is not an exotic disease. It is a recognized entity with a recognizable picture, and the people who treat it do not generally need three years to get there. What that case required was not superhuman reasoning. It required somebody — anybody — seeing the whole picture at once.

That is the thing that had never happened. Seventeen doctors saw seventeen slices. Someone had the gait. Someone had the teeth. Someone had the MRI. The pattern only exists when those sit next to each other, and for three years they never did.

The model’s advantage was not intelligence. It was completeness. It was the first reader in three years to be handed everything.

The research points the same way

It would be easy to cherry-pick here, so let me give you both directions.

On hard cases, GPT-4 has tested well. In one evaluation it included the correct diagnosis among its top six answers for 61.1% of challenging cases, against a previously reported 49.1% for physicians. That is the number people quote.

Here is the number they don’t. In a randomized trial published in JAMA Network Open, fifty physicians were given an LLM as a diagnostic aid alongside their usual resources. Their diagnostic reasoning did not significantly improve. And yet in the same study, the model working alone outperformed both groups of doctors.

Read those two findings together and you get something more interesting than “AI beats doctors.” You get this: the model does well when it is handed a complete, well-formed case. Drop it into a real clinical workflow — where the information is partial, scattered across systems, and arrives in fragments under time pressure — and the advantage largely evaporates.

Which is Alex’s case again, from the other end.

The detail nobody repeats

She transcribed it by hand.

Line by line, out of MRI reports, into a text box, at her kitchen table. That is the part of this story I cannot get past. The bottleneck was never the reasoning. The bottleneck was that the only way for one reader to see a child’s complete medical history was for his mother to retype it.

That is not a rare failure. It is the ordinary condition of American medical records, and it is measurable:

20.8% vs 13.8% Patients seen across multiple health systems get repeat testing far more often than patients within one system — the same test, run again, because the first result wasn’t visible. Source: research on repeat laboratory testing and EHR interoperability
59% less likely When unaffiliated emergency departments could actually exchange records, patients were 59% less likely to get a redundant CT scan. Across imaging types, participating facilities performed 44–67% fewer duplicate studies. Source: Medical Care (2014), 37 participating EDs vs 410 non-participating, California and Florida
21% of transfers In one trauma series, 21% of transferred patients received a repeat CT. Among the documented reasons: images never arrived, studies were incomplete, and discs that would not open. Source: study of early repeat CT imaging in transferred trauma and neurosurgical patients

A disc that will not open is not a technology problem in any deep sense. It is a child getting a second dose of radiation because a plastic circle failed.

What to actually do about it

Not “use AI.” The useful lesson is upstream of that.

  1. Get the actual images, not just the report. The radiology report is one person’s reading of the study. The study itself is the evidence, and a second reader can come to a different conclusion from it — in one series of outside abdominal imaging reread at a cancer center, 34% of reports were discrepant, and nearly half of the confirmed discrepancies changed the patient’s treatment.
  2. Keep it in one place. Not a folder in email, a drawer of discs, and three patient portals you can’t remember the passwords to. One place, complete, that you control.
  3. Bring the whole file to a second opinion. Not the summary. The whole thing. The value of a second opinion is heavily determined by what you put in front of it.
  4. If you do ask an AI, give it everything — and then take it to a physician. That is exactly what Courtney Hofmann did, and it is the step that actually mattered.

What I am not telling you

I want to be careful here, because this is a subject where enthusiasm does real harm.

No AI examined Alex. None of them took a history, watched him walk, or put hands on him. A language model cannot examine you, cannot order the study that settles the question, and cannot take responsibility for being wrong.

And you are reading a selected story. Nobody writes the article about the family who got a confident, fluent, completely incorrect answer and acted on it. Those cases exist; they are simply not news. Every expert quoted in the coverage of these stories says the same thing, and they are right: these tools can flag and prompt. They cannot diagnose, examine, or treat.

The reason I keep coming back to this case is not that a chatbot was clever. It is that a mother had to do, by hand, the one thing the medical system should have done for her automatically — put her son’s record in one place, complete, and legible to whoever was going to read it next.

That part is fixable. That is the part we work on.

Your records, in one place — including the actual scans

UploMD gathers your records from every facility you’ve been seen at, keeps your real MRI, CT and X-ray images viewable on your phone, and lets you share the whole file securely with any doctor you choose. It doesn’t diagnose anything. It makes sure whoever does is looking at all of it.

Get started free See a live demo

Sources

  1. “A boy saw 17 doctors over 3 years for chronic pain. ChatGPT found the diagnosis.” TODAY.
  2. “Partners in Diagnosis: ChatGPT, a Mother’s Intuition, and a Doctor’s Expertise.” NEJM AI Grand Rounds.
  3. Rutledge GW. “Diagnostic accuracy of GPT-4 on common clinical scenarios and challenging cases.” Learning Health Systems, 2024.
  4. Goh E, et al. “Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial.” JAMA Network Open, 2024.
  5. “Does health information exchange reduce redundant imaging? Evidence from emergency departments.” Medical Care, 2014.
  6. “Early repeat computed tomographic imaging in transferred trauma and neurosurgical patients.” PubMed.
  7. “Clinical importance of second-opinion interpretations of abdominal imaging studies in a cancer hospital.” Clinical Imaging.
  8. ECRI. “Siloed health records are harming patients.” ECRI, September 2026.

This article is general information, not medical advice, and it is not a description of what UploMD does clinically — UploMD does not diagnose, interpret, or recommend treatment. Nothing here should be used to make a medical decision on your own. If you have symptoms that concern you, see a physician.