
When will human physicians start dragging down AI in health care?
Key Takeaways
- Large language models are reported to outperform physicians in information gathering, diagnosis, cost-conscious test selection, treatment recommendations, and management of chronic conditions including diabetes, osteoarthritis, hyperlipidemia, and breast cancer.
- Comparative studies cite AMIE outperforming physicians in simulated exams, ChatGPT o3 leading complex-case diagnosis rates, and an AI “diagnostic orchestrator” improving accuracy at lower cost.
AI in medicine is supposed to be a tool used by doctors. What happens when it becomes the doctor?
In the next few years, artificial intelligence (AI) could deliver better medical care than physicians, whether human doctors work alone or with AI.
What’s more, human physicians overseeing AI-assisted care might make things worse for patients, not better.
“Will Autonomous AI Exceed AI-Aided Physicians as the Best in Medical Care?” is
“Data from medicine and other fields suggest that when AI alone performance is consistently superior to human-alone performance, AI alone surpasses human-AI hybrids,” the authors said. “Paradoxically, hybrid care in which humans are in (or on) the loop to correct AI errors is likely to worsen rather than improve AI performance.”
Physicians, patients and policymakers need to think and act now about the best ways to use AI in medicine. Authors Ezekiel J. Emanuel, M.D., Ph.D.; Abe Baker-Butler; Neal Khosla, M.S.; and Vinod Khosla, M.S., MBA, predicted autonomous AI could be ready for real-world deployment in "some, maybe many" clinical workflows by 2030
AI’s quick medical education
Large language models (LLMs) were introduced publicly in November 2022 and already are surpassing human physicians in at least five areas, he authors said:
- Gathering patient information
- Diagnosing conditions
- Selecting tests to pinpoint diagnoses, while staying on budget
- Recommending treatments
- Managing chronic diseases, at least for hyperlipidemia, osteoarthritis, diabetes and breast cancer
The authors cited the preponderance of studies about AI in health care published since 2024. LLMs are getting better and most of those studies did not account for reasoning models, a class of LLMs that began optimizing problem-solving methods since late 2024.
The authors cite several studies to back the claim. In one, Google's Articulate Medical Intelligence Explorer (AMIE) was judged by physicians as more effective than physicians themselves across 159 simulated patient exams covering multiple specialties, eliciting patient complaints (97% vs. 50% favorable), reviewing symptoms across body systems (88% vs. 35%) and taking a medical history (85% vs. 50%), all with statistically significant margins.
In another, OpenAI's ChatGPT o3 identified the correct final diagnosis first in 60% of 377 complex real-world cases, compared with 15.9% for a group of 20 internal medicine physicians reviewing a 302-case subset.
In another, Microsoft AI Diagnostic Orchestrator, working within an $8,000 budget on 56 difficult cases, reached the correct diagnosis roughly four times as often as physicians working without colleagues, textbooks or internet access, at 19.1% lower cost per case.
In a 2023 randomized clinical trial, patients with Type 2 diabetes who used a voice-based AI to adjust insulin dosing reached their optimal dose in a median of 15 days, versus more than 56 days for patients on standard physician-titrated care.
LLMs vs. M.D.s and D.O.s
The authors’ claims run counter to the official position of the American Medical Association (AMA) and the American College of Physicians (ACP), both of which describe AI's proper role in medicine as supportive rather than autonomous.
The AMA has pushed the term "augmented intelligence" specifically to emphasize that AI should assist, not replace, physician judgment. “ACP firmly believes that AI-enabled technologies should complement and not supplant the logic and decision making of physicians and other clinicians,” said the college’s position statement on artificial intelligence.
Physicians may assume "AI-assisted" is inherently the safer choice when treating patients. But once AI alone becomes reliably better than physicians alone at a task, adding a physician back into the loop tends to make outcomes worse, not better.
The authors cited a 2024 meta-analysis of 106 human-AI experiments, with at least 20 of them in medicine, showing that when humans alone already outperformed AI, pairing the two improved results. But when AI alone outperformed humans, adding a human to double-check its work significantly dragged down performance compared with AI working unsupervised. A 2025 review of 52 clinical studies reached a similar conclusion, finding that physician-AI teams "neither outperformed medical AI alone nor surpassed the best of clinicians or medical AI alone."
Distrust and deskilling
The authors point to two contributing forces. One is algorithm aversion, a well-documented tendency, found in roughly 75% of studies on the topic, for people to distrust and override algorithmic advice even after being told it is more accurate. The effect is strongest among highly trained experts, including physicians.
The other is deskilling. As physicians rely more on AI for tasks such as diagnosis or documentation, their own proficiency at those tasks may erode, a pattern already documented among endoscopists exposed to AI-assisted colonoscopy.
It’s possible medicine could follow the arc of humans and computers interacting at chess. The computer program Deep Blue beat world champion Garry Kasparov in 1997. The last time an unaided human beat an unaided computer was 2005, and by 2017, AI alone had surpassed even the best human-AI chess teams.
Proceed with curiosity and caution
The authors themselves list substantial caveats that could affect how physicians and patients respond to the development of AI in health care. Most of the supporting evidence comes from simulated exam scenarios rather than real clinical encounters. Current liability, regulatory and reimbursement barriers limit how much AI-alone performance can be tested in live clinical settings.
A small number of studies published since 2024 still found physicians alone matching or exceeding AI, including in diagnosing paroxysmal atrial fibrillation and cases with atypical features. Adversarial testing also shows LLMs remain unpredictable in some circumstances.
Health care needs a human touch because many cognitive medical tasks are still bound to physical procedures, such as surgery, childbirth, colonoscopies and interventional radiology, that current robotics cannot perform autonomously.
Emanuel is






