
Cost is the wrong way to judge medical AI, physicians argue
Key Takeaways
- Total Mission Value integrates Balanced Scorecard, Quintuple Aim, and institutional mission statements to evaluate AI beyond cost, prioritizing patient care and embedding ethical obligations across domains.
- A “scorecard gate” advises postponing AI deployment when systems cannot specify intended impacts on patients, staff, quality, and governance readiness.
A framework published in npj Digital Medicine ranks patient care and staff experience alongside price, and says a health system that cannot score a tool across all five domains may not be ready to buy it.
Health systems are adopting
R. Andrew Taylor, M.D., M.H.S., vice chair of research and innovation for the University of Virginia School of Medicine's Department of Emergency Medicine, and Arwen B.L. Declan, M.D., Ph.D., clinical assistant professor in Clemson University's School of Health Research, call their model Total Mission Value.
Published in
"AI is being adopted in medicine at a scope and velocity we have never seen before, but hospitals haven't had a good way to weigh these decisions as a whole," Taylor said in a
The authors assembled the five domains by merging three existing models: the Balanced Scorecard from business strategy, the Institute for Healthcare Improvement's Quintuple Aim and a review of hospital mission, vision and values statements. What escapes cost-focused analysis, they write, is most of what surfaces later, including bias, opacity, workforce displacement and erosion of the patient-clinician relationship.
The framework's most concrete instruction is a gate. "Organizations unable to define a scorecard for AI may need to postpone adoption," Declan and Taylor write. A system that cannot articulate what a tool is supposed to do for patients, staff and clinical quality, in their reading, is not yet equipped to govern one.
Where physicians fit in the math
Physicians rarely choose the AI that arrives in their exam rooms. That decision usually sits with the health system, the group's leadership or the
Its worked example is ambient documentation, the AI tool most primary care physicians already touch every day. The scorecard maps that use case to specific indicators: total EHR time per appointment, same-day note completion rates, after-hours charting, burnout prevalence measured with the Mini Z instrument, workload measured with the NASA Task Load Index, work relative value units, claim denial rates and avoided physician turnover costs.
That list tracks with where the return on scribes has actually materialized. Robert Wachter, M.D., chair of the Department of Medicine at the University of California, San Francisco,
"The AI scribe probably kind of pays for itself just in pure throughput and time savings," Wachter said. "But the real benefit is the joy in practice, recruitment, retention." Replacing a departing primary care physician, he said, is often estimated in the range of $1 million, a cost that enters a return-on-investment calculation only if someone put staff experience into it to begin with.
Related content:
Sorting out which tools deliver
The volume of tools on the market is part of what the framework is built to handle.
"Hospitals are seeing a huge number of new AI tools marketed to improve healthcare. The challenge is to figure out which ones actually will," Declan said in the release. "That requires weighing an AI tool's impact across clinical, operational and financial dimensions, while keeping patient care at the center of every decision."
For a practicing physician, the narrower version of that question is what is safe to use now. Marc Succi, M.D., executive director of the MESH Incubator at Mass General Brigham, led a study in
Succi
"That's where you've got to stop and look at the level of performance very critically in multiple different ways, not just trust what the vendors say," he said.
What the framework does not do yet
Total Mission Value is a conceptual model, and its authors are direct about the gaps. It has no validated metrics, no method for weighting competing stakeholder priorities and no testing across institutions. They call for comparative analyses of AI decisions made with and without mission-value integration, and for longitudinal study of the results.
They also flag an absence in their own source material. Risk management, which they describe as essential to care quality and to responsible stewardship of resources, is poorly represented in the hospital mission and values statements they analyzed, even though AI governance falls under the same operations domain.
And they name trade-offs the framework is designed to surface rather than settle, among them tools that improve the patient experience while adding to clinician workload, and equity-oriented deployments that return little in a fee-for-service environment.
"Our hope is that keeping the mission front and center actually speeds good AI adoption rather than slowing it down, because it builds the trust that patients and clinicians need," Taylor said. "Technology should help us take better care of people. If we keep that as the goal, the efficiency and the savings tend to follow."





