In 2023 a mother did something nobody had managed in three years of appointments: she put her son’s entire medical history in front of a single reader.
You have probably seen the headline by now, in one form or another: ChatGPT diagnosed a boy that seventeen doctors couldn’t. It is a good headline. It is also, I think, the wrong reading of what happened — and the right reading is considerably more useful to you.
A boy named Alex was in pain for three years. Over that time he saw seventeen doctors. He had toothaches. His growth stalled. His balance and posture were off, and his mother noticed he would not sit cross-legged on the floor the way other children did. Each clinician found something, or found nothing, and none of it added up to a single explanation.
His mother, Courtney Hofmann, went line by line through his MRI notes and typed them into ChatGPT, along with everything she had been tracking herself. The model suggested tethered cord syndrome — a condition where the spinal cord is abnormally anchored at its lower end, putting tension on it as a child grows.
She took that to a neurosurgeon, Dr. Holly Gilmer, who reviewed the imaging, agreed, and operated. NEJM AI later brought both the mother and the surgeon onto its Grand Rounds program to work through the case — which tells you the profession treated this as something worth understanding rather than a novelty to wave off.
Tethered cord is not an exotic disease. It is a recognized entity with a recognizable picture, and the people who treat it do not generally need three years to get there. What that case required was not superhuman reasoning. It required somebody — anybody — seeing the whole picture at once.
That is the thing that had never happened. Seventeen doctors saw seventeen slices. Someone had the gait. Someone had the teeth. Someone had the MRI. The pattern only exists when those sit next to each other, and for three years they never did.
The model’s advantage was not intelligence. It was completeness. It was the first reader in three years to be handed everything.
It would be easy to cherry-pick here, so let me give you both directions.
On hard cases, GPT-4 has tested well. In one evaluation it included the correct diagnosis among its top six answers for 61.1% of challenging cases, against a previously reported 49.1% for physicians. That is the number people quote.
Here is the number they don’t. In a randomized trial published in JAMA Network Open, fifty physicians were given an LLM as a diagnostic aid alongside their usual resources. Their diagnostic reasoning did not significantly improve. And yet in the same study, the model working alone outperformed both groups of doctors.
Read those two findings together and you get something more interesting than “AI beats doctors.” You get this: the model does well when it is handed a complete, well-formed case. Drop it into a real clinical workflow — where the information is partial, scattered across systems, and arrives in fragments under time pressure — and the advantage largely evaporates.
Which is Alex’s case again, from the other end.
She transcribed it by hand.
Line by line, out of MRI reports, into a text box, at her kitchen table. That is the part of this story I cannot get past. The bottleneck was never the reasoning. The bottleneck was that the only way for one reader to see a child’s complete medical history was for his mother to retype it.
That is not a rare failure. It is the ordinary condition of American medical records, and it is measurable:
A disc that will not open is not a technology problem in any deep sense. It is a child getting a second dose of radiation because a plastic circle failed.
Not “use AI.” The useful lesson is upstream of that.
I want to be careful here, because this is a subject where enthusiasm does real harm.
No AI examined Alex. None of them took a history, watched him walk, or put hands on him. A language model cannot examine you, cannot order the study that settles the question, and cannot take responsibility for being wrong.
And you are reading a selected story. Nobody writes the article about the family who got a confident, fluent, completely incorrect answer and acted on it. Those cases exist; they are simply not news. Every expert quoted in the coverage of these stories says the same thing, and they are right: these tools can flag and prompt. They cannot diagnose, examine, or treat.
The reason I keep coming back to this case is not that a chatbot was clever. It is that a mother had to do, by hand, the one thing the medical system should have done for her automatically — put her son’s record in one place, complete, and legible to whoever was going to read it next.
That part is fixable. That is the part we work on.
UploMD gathers your records from every facility you’ve been seen at, keeps your real MRI, CT and X-ray images viewable on your phone, and lets you share the whole file securely with any doctor you choose. It doesn’t diagnose anything. It makes sure whoever does is looking at all of it.
Get started free See a live demoThis article is general information, not medical advice, and it is not a description of what UploMD does clinically — UploMD does not diagnose, interpret, or recommend treatment. Nothing here should be used to make a medical decision on your own. If you have symptoms that concern you, see a physician.