Artificial Intelligence Already Outperforms Doctors in Medical Tests

August 9, 2026

Robo-docs are not poised to replace healthcare in the near future, yet they could play a more substantial role in supporting human clinicians—if we permit it.

Artificial intelligence (AI) systems may outperform doctors at diagnosing patients. In a recent sequence of experiments led by researchers at the Beth Israel Deaconess Medical Center and Harvard Medical School, an early iteration of the OpenAI large language model (LLM) known as o1 outperformed human physicians across several assessments of clinical judgment and diagnostic reasoning.

“We evaluated the AI system against nearly every standard, and it eclipsed both earlier models and our physician baselines,” stated Arjun K. Manrai, a professor at Harvard Medical School and one of the study’s senior authors. Employing clinical case studies, the researchers had o1—an older variant since superseded by o3—generate a spectrum of potential diagnoses. The correct diagnosis appeared in about 78 percent of the cases, while physicians succeeded roughly 30 percent of the time.

In a separate examination, o1 was presented with five real-world clinical vignettes and asked what steps should follow. Two physicians reviewed its replies. The average score for o1 reached 89 percent, compared with 34 percent for the human doctors facing the same prompt.

The team also assessed o1’s capacity to diagnose emergency department cases where clinical data may be incomplete. “Overall, o1 outperformed both [an earlier LLM] and two expert attending physicians, as evaluated by two other attending physicians who were unaware of the differential-diagnosis source,” notes the study, which appeared on April 30 in Science. The LLM showed particular advantage at the initial triage stage, accurately pinpointing the “exact or near-exact diagnosis…” in 67.1 percent of cases, whereas the two human doctors achieved 55.3 percent and 50 percent, respectively.

These findings are in line with other recent research. A Swedish study published in The Lancet in January suggests that AI-assisted mammography could improve breast cancer detection.

Analyzing abdominal CT scans from patients who would later be diagnosed with pancreatic cancer, an AI model developed at the Mayo Clinic identified this deadly disease on average 475 days earlier, and up to three years earlier, than clinicians did. “Attaining such early detection would substantially increase the likelihood of cure and better survival,” researchers led by Mayo Clinic’s Sovanlal Mukherjee noted in the journal Gut.

Robo-docs are not likely to take over healthcare in the near term. And “humans should be the ultimate baseline,” as Peter Brodeur, a co-author of the Harvard study, stated in a press release. Yet investigations like these imply that AI models could effectively assist across a variety of diagnostic and medical-management scenarios, and perhaps even yield better patient outcomes—if we allow them to.

Nevada has now prohibited AI systems from making statements that “implicitly indicate” they are “capable of providing professional mental or behavioral health care” and from delivering any service “that would constitute the practice of professional mental or behavioral health care.” An Illinois law enacted last year bars AI from providing therapy, and prohibits therapists from using AI to “make independent therapeutic decisions” or “detect emotions or mental states.” Several states—including Ohio, California, Minnesota, and Kentucky—are weighing similar legislation. (Meanwhile, certain state proposals would require AI chatbots to be able to detect and respond to mental-health issues.)

Policies like these could restrict AI’s capacity to diagnose diseases with higher accuracy than human clinicians on certain occasions.

Natalie Foster

I’m a political writer focused on making complex issues clear, accessible, and worth engaging with. From local dynamics to national debates, I aim to connect facts with context so readers can form their own informed views. I believe strong journalism should challenge, question, and open space for thoughtful discussion rather than amplify noise.