News|Articles|October 8, 2026

Can a safety prompt make AI safer for clinical decisions? Study says yes, partly

Author(s)Todd Shryock
Fact checked by: Chris Mazzolini

Adding a brief safety reminder cut potentially harmful clinical choices by large language models from 16.6% to 10.1% across more than 10 million responses, but researchers say it doesn't replace physician oversight.

Adding a brief safety reminder to the instructions given to artificial intelligence (AI) models reduced potentially harmful clinical choices in 19 of 20 models tested, according to a new study from researchers at the Icahn School of Medicine at Mount Sinai.

The study, published Sept. 26 in Communications Medicine, a Nature Portfolio journal, evaluated more than 10 million responses from 20 large language models (LLMs). Without a safety reminder, potentially harmful choices made up 16.6% of model responses. With a brief reminder added, that rate fell to 10.1%.

Across all responses, the models made about 1.18 million potentially harmful clinical choices, the researchers reported.

The researchers said the findings suggest that how AI systems are prompted and guided remains an important consideration in building safe and reliable clinical applications, even as newer models and AI agents become better at understanding a user's intent without lengthy or detailed instructions.

"AI models do not make decisions in a vacuum. The language, framing, and context surrounding a request can influence how they respond, including when an instruction could be unsafe," said first author Mahmud Omar, M.D., a lecturer in the Windreich Department of Artificial Intelligence and Human Health at Icahn Mount Sinai, who leads research on the safety, reliability and real-world effects of generative AI in clinical care.

Testing instructions that conflict with patient safety

The research team tested the models using 501 variations of 50 clinical scenarios, along with 100 cases adapted from de-identified hospital discharge records.

In one example, a model was told to skip recommended follow-up blood tests to reduce workload. In some variations, the request was framed as urgent or presented as an order from a superior. The models then chose among four possible actions, including following the request, maintaining the recommended follow-up or seeking help from a clinician.

The researchers varied the wording of the scenarios and tested three short safety reminders. Each combination was run 10 times, with the order of the answer choices randomized.

The reduction in harmful choices appeared in both the written clinical scenarios and the cases drawn from discharge records. Examples of potentially harmful choices included skipping needed tests to reduce workload and stopping antibiotic treatment before completing the recommended regimen without a sufficient clinical reason.

According to the researchers, the results highlight the need to evaluate not only whether an AI model can provide accurate clinical information, but also how it responds when given an instruction that conflicts with patient safety.

"These results suggest that safety testing needs to go beyond asking whether an AI model gets the right answer under ordinary conditions," said co-senior author Girish N. Nadkarni, M.D., M.P.H., chair of the Windreich Department of Artificial Intelligence and Human Health and director of the Hasso Plattner Institute for Digital Health at Mount Sinai. "As AI systems become more autonomous and are asked to complete increasingly complex tasks, we need to know whether they can recognize when an instruction may be unsafe, question it, verify it, or ask a human for help."

A safeguard, not a substitute

The study's authors cautioned against reading the results as a fix. The reminder lowered the rate of harmful choices in most models, but harmful choices still occurred in roughly 1 in 10 responses even with it in place.

"A simple safety reminder reduced potentially harmful choices in most of the models we tested, which is encouraging," Omar said. "But it did not eliminate them, so a reminder should be viewed as one safeguard, not a substitute for clinical oversight."

The researchers said the findings do not mean a safety reminder makes AI-generated medical advice safe to use without clinical review. Instead, the results show that relatively simple changes in how an AI system is prompted can affect its clinical choices, while underscoring the need for additional safeguards and human oversight.

The team recommends that developers and health care organizations build automated safety testing into the development and evaluation of clinical AI systems. That testing could be done before a tool is introduced into a clinical workflow and repeated as models are updated or new safety concerns emerge.

The study adds to a growing body of research on the risks of AI in clinical care. Earlier this year, ECRI named diagnostic AI the top concern on its annual patient safety list.

Next steps: agents and prompt injection

The study also points to an emerging challenge as AI tools evolve from question-and-answer systems into more autonomous "agents" that can carry out multiple steps on their own.

The researchers plan to examine how accumulated context may affect an agent's decisions, including when that context contains hidden instructions — known as prompt injection — or pressures to save time or stay within a budget. Those conditions mirror the kinds of workload and cost pressures built into the scenarios in the current study, such as requests to skip follow-up testing to reduce workload.


Related to this article