OpenAI's o1 Model Outperformed Human Doctors on Emergency Room Diagnoses in Harvard Study
-

A peer-reviewed study published in Science this week found that OpenAI's o1 model outperformed two internal medicine attending physicians on emergency room diagnosis accuracy at Beth Israel Deaconess Medical Center, raising significant questions about AI's potential role in clinical medicine. Researchers from Harvard Medical School presented the AI models with the same unprocessed electronic medical record information available to physicians at each diagnostic touchpoint, with assessments conducted blind by two other attending physicians who did not know which diagnoses came from humans and which from AI. The o1 model offered the exact or very close diagnosis in 67% of initial triage cases, compared to 55% and 50% for the two human physicians respectively, with the performance gap most pronounced at the first triage point where information is most limited and urgency is highest.The study's authors were careful to contextualize the findings.
The research does not claim AI is ready to make life-or-death clinical decisions independently, and the authors explicitly called for prospective trials to evaluate these technologies in real-world patient care settings before drawing operational conclusions. Beth Israel physician Adam Rodman noted there is no formal framework for accountability around AI diagnoses, and that patients want human guidance through challenging and life-threatening decisions. The study also only tested text-based information processing, acknowledging that current models are more limited when reasoning over non-text clinical inputs like imaging, physical examination findings, and patient interaction context that experienced physicians incorporate automatically.