OpenAI's o1 reasoning model achieved a 67% accuracy rate in diagnosing emergency room cases, compared to 50-55% accuracy among human triage doctors, according to a Harvard-affiliated study reported by the Guardian. The trial evaluated how the AI model processes complex clinical reasoning using standard electronic health records from real ER cases.
Benchmarking Clinical Reasoning in Emergency Triage
Researchers tested the AI against human clinicians using patient data from 76 emergency department arrivals at Beth Israel Deaconess Medical Center in Boston. Both the AI model and pairs of human physicians analyzed identical patient records consisting of vital signs, demographic data, and brief intake notes from nurses. OpenAI's o1 correctly diagnosed 67% of cases, compared to 50-55% for human emergency physicians.
When provided with broader clinical details, the AI's accuracy rose to 82%, while human expert accuracy ranged from 70% to 79%. The model also outperformed 46 human doctors in designing long-term treatment plans across five clinical case studies, scoring 89% on treatment plan quality compared to 34% for doctors using conventional web resources.
How Reasoning Models Process Complex Cases
According to the study, the model was able to cross-reference patient histories that human doctors sometimes missed during brief consultations. In one case cited by researchers, a patient presented with worsening lung symptoms while taking anti-coagulants. Human clinicians assumed the medication was failing, but the AI identified that the patient's underlying lupus was causing lung inflammation — a diagnosis that proved correct.
Practical Caveats and Deferral Risks
Researchers emphasized that the trial does not demonstrate AI can replace emergency physicians. The study evaluated only text-based electronic records and did not test physical visual cues, patient distress levels, or hands-on examinations. Independent experts cited in the reporting also raised concerns about physicians unconsciously deferring to AI recommendations, as well as a lack of performance data for elderly patients, non-English speakers, and unresolved questions around legal liability.
Note: This trial was first reported in late April 2026; the findings described here are not new as of this article's publication.