News|Articles|July 22, 2026

Cost is the wrong way to judge medical AI, physicians argue

Fact checked by: Keith A. Reynolds
Listen
0:00 / 0:00

Key Takeaways

  • Total Mission Value integrates Balanced Scorecard, Quintuple Aim, and institutional mission statements to evaluate AI beyond cost, prioritizing patient care and embedding ethical obligations across domains.
  • A “scorecard gate” advises postponing AI deployment when systems cannot specify intended impacts on patients, staff, quality, and governance readiness.
SHOW MORE

A framework published in npj Digital Medicine ranks patient care and staff experience alongside price, and says a health system that cannot score a tool across all five domains may not be ready to buy it.

Health systems are adopting artificial intelligence (AI) faster than they have settled on a way to evaluate it, and the measure most of them reach for first is price. Two emergency physicians have published a framework arguing that price is the least useful part of the decision.

R. Andrew Taylor, M.D., M.H.S., vice chair of research and innovation for the University of Virginia School of Medicine's Department of Emergency Medicine, and Arwen B.L. Declan, M.D., Ph.D., clinical assistant professor in Clemson University's School of Health Research, call their model Total Mission Value.

Published in npj Digital Medicine, it sorts an AI tool's worth into five domains: patient care, staff experience, operations, economic impact, and education and research. Patient care sits at the top, economic sustainability forms the base and ethical obligations run through every level.

"AI is being adopted in medicine at a scope and velocity we have never seen before, but hospitals haven't had a good way to weigh these decisions as a whole," Taylor said in a news release. "Typical approaches tend to measure cost, because cost is the easiest thing to measure."

The authors assembled the five domains by merging three existing models: the Balanced Scorecard from business strategy, the Institute for Healthcare Improvement's Quintuple Aim and a review of hospital mission, vision and values statements. What escapes cost-focused analysis, they write, is most of what surfaces later, including bias, opacity, workforce displacement and erosion of the patient-clinician relationship.

The framework's most concrete instruction is a gate. "Organizations unable to define a scorecard for AI may need to postpone adoption," Declan and Taylor write. A system that cannot articulate what a tool is supposed to do for patients, staff and clinical quality, in their reading, is not yet equipped to govern one.

Where physicians fit in the math

Physicians rarely choose the AI that arrives in their exam rooms. That decision usually sits with the health system, the group's leadership or the electronic health record (EHR) vendor, and Total Mission Value is an argument about how it should be made. The framework gives staff experience its own domain, ranked directly below patient care, and holds that AI should support clinicians in their work rather than replace them or add to it.

Its worked example is ambient documentation, the AI tool most primary care physicians already touch every day. The scorecard maps that use case to specific indicators: total EHR time per appointment, same-day note completion rates, after-hours charting, burnout prevalence measured with the Mini Z instrument, workload measured with the NASA Task Load Index, work relative value units, claim denial rates and avoided physician turnover costs.

That list tracks with where the return on scribes has actually materialized. Robert Wachter, M.D., chair of the Department of Medicine at the University of California, San Francisco, told Medical Economics that all of his system's roughly 3,000 to 4,000 physicians have access to an AI scribe and about 70% use one. The economics landed differently than the early projections suggested they would.

"The AI scribe probably kind of pays for itself just in pure throughput and time savings," Wachter said. "But the real benefit is the joy in practice, recruitment, retention." Replacing a departing primary care physician, he said, is often estimated in the range of $1 million, a cost that enters a return-on-investment calculation only if someone put staff experience into it to begin with.

Related content: Who's really liable when your AI scribe makes a mistake?

Sorting out which tools deliver

The volume of tools on the market is part of what the framework is built to handle.

"Hospitals are seeing a huge number of new AI tools marketed to improve healthcare. The challenge is to figure out which ones actually will," Declan said in the release. "That requires weighing an AI tool's impact across clinical, operational and financial dimensions, while keeping patient care at the center of every decision."

For a practicing physician, the narrower version of that question is what is safe to use now. Marc Succi, M.D., executive director of the MESH Incubator at Mass General Brigham, led a study in JAMA Network Open that tested 21 general-purpose large language models across 29 clinical vignettes. The models produced a correct final diagnosis more than 90% of the time when handed complete information, and failed to produce an appropriate differential diagnosis more than 80% of the time.

Succi told Medical Economics that physicians should stay with low-risk, high-feasibility tasks such as documentation, summarization, plain-language explanations for patients and billing support, and slow down at clinical decision support, patient message responses, lab ordering and medication renewals.

"That's where you've got to stop and look at the level of performance very critically in multiple different ways, not just trust what the vendors say," he said.

What the framework does not do yet

Total Mission Value is a conceptual model, and its authors are direct about the gaps. It has no validated metrics, no method for weighting competing stakeholder priorities and no testing across institutions. They call for comparative analyses of AI decisions made with and without mission-value integration, and for longitudinal study of the results.

They also flag an absence in their own source material. Risk management, which they describe as essential to care quality and to responsible stewardship of resources, is poorly represented in the hospital mission and values statements they analyzed, even though AI governance falls under the same operations domain.

And they name trade-offs the framework is designed to surface rather than settle, among them tools that improve the patient experience while adding to clinician workload, and equity-oriented deployments that return little in a fee-for-service environment.

"Our hope is that keeping the mission front and center actually speeds good AI adoption rather than slowing it down, because it builds the trust that patients and clinicians need," Taylor said. "Technology should help us take better care of people. If we keep that as the goal, the efficiency and the savings tend to follow."