Microsoft claims their new AI framework diagnoses 4x better than...

They took 304 real cases from NEJM and turned them into an interactive game.
The setup:
Step 1: You (human doctor or AI) get a tiny intro like: "52-year-old man with fever and breathing problems." That's it. No test results, no detailed history - just like a patient walking into the ER.
Step 2: There's a "Gatekeeper" (another AI) that has the full case file but won't tell you anything unless you specifically ask.
Step 3: You can do three things:
1. Ask questions ("Any recent travel?" "Is there chest pain?")
2. Order tests ("CBC" "Chest X-ray" "CT scan")
3. Make your final diagnosis ("This is pneumonia")
Step 4: The Gatekeeper then answers the question. BUT it only reveals what you ask for. If you don't think to ask about travel history, you won't find out the patient just returned from a cave expedition (real case - histoplasmosis).
Step 5: Every test costs money (real US hospital prices). Every round of questions = $300 office visit.
How does this framework work?
It asks the LLM to simulate a virtual panel of 5 specialised AI doctors:
Dr. Hypothesis (tracks diagnoses)
Dr. Test-Chooser (selects optimal tests)
Dr. Challenger (plays devil's advocate)
Dr. Stewardship (manages costs)
Dr. Checklist (quality control)
Then argue it out between themselves as to the best path forward.
📊 Accuracy:
Doctors: 20% (ouch)
Standard AI: 30-79%
MAI-DxO: 80-85.5%
💰 Cost per case:
Doctors: $2,963
Standard AI (o3): $7,850
MAI-DxO: $2,397
On paper the AI was 4x more accurate AND cheaper.....
1. They used ZERO healthy patients
95% of sore throats are viral and this AI was only tested on incredibly rare diagnostic cases.
We don't know if it will order biopsies on every patient with a sore throat "just to rule out rhabdomyosarcoma."
Their costs only count lab fees, not:
- 2 weeks of anxiety waiting for biopsy results
- Radiation from "precautionary" CT scans (cancer risk!)
- Complications from unnecessary procedures
- Time off work
- Psychological trauma of false cancer scares
Docs were banned from:
❌ Googling symptoms
❌ Consulting colleagues
❌ Using UpToDate/medical databases
❌ Calling specialists
That's not how we practice!!
It's like testing a chef who can't use recipes or taste their food.
These cases were already SOLVED and published.
Real medicine involves genuine uncertainty - sometimes the diagnosis is never found. Does the AI know when to stop investigating?
Great doctors know when NOT to test. This AI was never evaluated on:
"This headache is just stress"
"Let's wait and see"
"More tests will cause more harm than good"
The benchmark rewards finding zebras, not recognising horses.
But we need:
✓ Testing on actual patient populations (mostly healthy!)
✓ Measuring overdiagnosis harm
✓ Real-world physician comparisons
But what do you think?
It takes me some time to read and write these posts so I'd love to get more people's thoughts on it!
I've also just started a new newsletter on neuroscience:
brainhealthdecoded.substack.com/subscribe


