
Why ‘off-the-shelf’ general-purpose LLMs aren’t fit for medication use cases
What clinicians should ask before trusting artificial intelligence with medication decisions
Large language models (LLMs) are quickly moving beyond experimentation to everyday use in health care. Faced with staffing shortages, burnout and relentless cost pressure, organizations are deploying LLMs to take administrative work off clinicians' plates, and for many use cases, that is working.
But as general-purpose LLMs and unauthorized artificial intelligence (AI), or
Why are medication work flows the proving ground for responsible AI?
AI adoption in health care has moved steadily into administrative work, streamlining tasks such as prior authorization, clinical documentation, scheduling and more. These applications offer real value and relief from health care’s most persistent challenges. Medication work flows are emerging as the next frontier for AI adoption, but they also carry a new weight of requirements and considerations.
The difference comes down to proximity to the patient. A missed drug interaction or an incorrect dose can cause serious harm, and medication-related harm already affects
This also affects efficiency. When AI systems generate enough errors that clinicians must continuously validate outputs, the time savings disappear. That erosion of trust tends to move quickly, and the outcome is familiar: Clinical teams revert to the manual work flows the technology was meant to support. Getting medication work flows right addresses that dynamic directly and may be the foundation on which broader AI adoption across clinical care is built.
Why are medication data uniquely complex for AI to manage?
Drug data are not static. Safety signals emerge, formulations change, and regulatory guidance varies across markets and regions, sometimes rapidly. Guidance that was correct six months ago may no longer apply. For a general model trained on a fixed data set or the internet, this rapid change can be a structural issue.
Medication decisions also depend heavily on the individual patient. Kidney function, age, existing conditions and current medications all shape whether a drug is appropriate and how it should be used. Two patients presenting with the same diagnosis may need entirely different approaches. No training data set captures that granularity reliably.
Clinicians and pharmacists have worked through this complexity for years, supported by drug databases built specifically to keep pace with it. Those databases are maintained by domain experts, updated as guidance changes, and designed to account for the kinds of patient-specific variables that determine whether a medication decision is safe. A general-purpose model isn't built with that infrastructure or the knowledge to know whether new research is practice changing. For AI to work reliably in medication work flows, it needs to connect with that clinical foundation.
What are off-the-shelf general-purpose LLMs designed to do, and where do they create problems for medication use cases?
General-purpose LLMs do not inherently maintain continuously updated medication knowledge and therefore require additional retrieval, governance and validation mechanisms to reliably incorporate new safety signals, dosing recommendations and regulatory changes. Drug safety information changes continuously, and a model's underlying training may not reflect those changes unless supported by ongoing updates or external clinical knowledge sources.
A deeper problem is how these models handle uncertainty. General-purpose tools often lack medication-specific constraints, which means they may respond to requests that fall outside their validated scope rather than declining them. In medication work flows, a confident wrong answer is a patient safety failure. In fact,
What does clinical-grade AI actually require?
Clinical-grade AI is not simply a more powerful general-purpose model. It pairs LLM capabilities with medication intelligence built and refined over decades by clinical domain experts, rather than just relying on what a model absorbed during training.
This distinction has real implications. Domain-expert quality assurance must be a fundamental requirement during development. Built-in documentation ensures system behavior is clear and limited, so clinicians understand what the system can and cannot do before they depend on it. The tool's scope needs to be defined and enforced from the start, not through trial and error in a clinical setting.
Consider what patients already expect. A
Why do auditability and testability matter so much in this context?
Clinicians acting on AI-generated guidance need to know where that guidance came from. Auditability makes that possible. It gives clinicians and organizations a way to trace how a system reached a particular output and establishes accountability when questions arise. Testability is about what happens next. A model that performs well today may behave differently after an update or in a different clinical environment.
Consistent verification across those conditions is what distinguishes a system that clinicians can trust from one they have to question. In fact,
What should clinicians ask before trusting an AI system with medication decision support?
The starting point is the source. Where does the system's drug knowledge come from, how often is it reviewed, and who is doing that review? A system built on continuously updated, expert-curated clinical data is a different product from one drawing on general training data, and vendors should be able to answer that question directly.
Scope matters just as much. A well-designed system should clearly decline requests that fall outside its validated range. Clinicians should ask vendors to demonstrate that behavior and prioritize systems with boundaries.
Prescribing decisions carry clinical and legal weight, and the systems supporting them should be held to the same standard. Many of the common general-purpose AI models today are black boxes, producing outputs without explaining how they got there. That’s why AI that’s traceable and transparent is key, so that when a clinician acts on a recommendation, they know where it came from and whether it can be trusted in that specific context.
The standard has to come first
The efficiency gains AI offers in health care are well-documented. But in medication work flows, the cost of moving faster than the evidence supports can show up in patient outcomes.
The organizations getting this right are not the ones deploying the most capable general-purpose models. They are the ones asking harder questions before deployment: about data sources, about validated scope, about what happens when the system encounters something it was not built to handle. Those questions are not obstacles to AI adoption but the foundation it needs to stand on.
Christian Hartman, Pharm.D., MBA, FSMSO, is vice president of product innovation and pharmacy and health technology solutions at Wolters Kluwer Health. Hartman has been an integral part of the Wolters Kluwer Health team since 2012, when he joined as the director of quality and patient safety solutions for the Clinical Surveillance and Compliance business.






