Blog|Articles|September 9, 2026

The AI purchase your practice cannot evaluate (and what to check)

Author(s)Neel Chauhan
Fact checked by: Todd Shryock
Listen
0:00 / 0:00

Key Takeaways

  • FDA clearance must be verified via pathway and submission number, then matched word-for-word to indications for use, because marketing claims frequently exceed the legally cleared scope.
  • Validation claims should be decomposed into discrimination and calibration, with unredacted reports reviewed for baseline event rates and local calibration, since AUROC portability does not ensure usable probabilities.
SHOW MORE

Independent practices carry the same legal, regulatory and clinical exposure as large health systems when they buy AI tools, but none of the informatics oversight — here is the five-question framework to close that gap.

I have occupied both sides of this table. When I was a family physician, software companies presented me with clinical algorithms I had no means of practically evaluating. When I was an operating executive, I approved the enterprise teams who made those pitches.

It is painfully clear from the executive side how much risk is quietly allocated before anyone has signed the order form.

Independent practices suffer from a serious information imbalance. The vendor shows off an ambient scribe, an inbox summarizer or a predictive risk model. The demonstration is perfect, the monthly cost per provider is acceptable and the account executive knows more about the system's architecture than anyone in your clinic will ever know.

We still go ahead and buy it. According to the American Medical Association's 2026 survey of physicians on the use of AI, professional usage stood at 81%, more than double the figure from three years before. Yet when the same survey asked what would make physicians comfortable adopting these tools, 88% named validation of safety and efficacy by a trusted entity.

That answer highlights the fundamental problem facing independent medicine. In a private practice, there is no trusted entity coming to your aid. There is no informatics committee, no chief medical information officer and no dedicated data engineering team reviewing the validation logs.

Your legal, regulatory and clinical exposure is nonetheless just as great as that of a multi-hospital health system. You are subject to the same HIPAA requirements, liable to the same malpractice underwriters, and face the same vulnerable patients.

You will not out-engineer the vendor, and it is a mistake to try. The aim is systematic verification, which means obtaining commitments in writing and noting carefully the things the vendor is unwilling to sign.

Here are five questions that should guide every procurement conversation:

1. Has this device been cleared by the FDA, and what are its precise indications for use?

If the vendor answers yes, ask for the device classification, the regulatory pathway and the 510(k) or De Novo submission number.

The FDA maintains a publicly available list of AI-enabled medical devices that links to each of the clearance records. Check the clearance letter, then compare the FDA's stated indications for use word for word with those in the marketing deck you were just given. They differ very frequently. In such cases, the vendor's own federal regulator is telling you the product has been legally cleared to do considerably less than the sales pitch suggests.

If the vendor answers no, find out what that means. Tools promoted as ambient scribes, administrative summarizers or workflow boosters function outside the scope of device regulation. Neither the software's performance nor its safety record nor its basic logic has been examined by any external body. The only audit it will ever have is the one you carry out in your clinic.

2. On whom was it validated, and were calibration measurements taken?

The statement that "our algorithm is clinically validated" is merely a sales tactic intended to close the discussion. Your first reaction should be to ask which parties carried out the validation, with respect to which particular patient groups, and on the basis of what statistical measures.

Two completely separate metrics are used to assess predictive tools, and companies often conceal one behind the other.

  • Discrimination, frequently given as the AUROC or the C-statistic, indicates whether the model ranks patients in the right order relative to each other. Discrimination tends to generalize across different health systems.
  • Calibration indicates whether the predicted probability matches reality in your own practice setting. Calibration seldom extends to other situations without local adjustment.

Imagine an algorithm trained in a hospital-based tertiary care environment in which 20% of the group experiences deterioration. Used instead in a primary care outpatient panel with a baseline event rate of 4%, the algorithm may retain its relative ranking, but its numerical predictions will be distorted. Your team will be overwhelmed by five times as many false-positive alerts as they can properly deal with.

The outcome that follows is entirely foreseeable. After two months, alert fatigue sets in, clinicians begin to ignore the alerts and the practice keeps paying for shelfware. Ask for the unredacted technical validation report rather than the marketing summary, and look at the baseline event rates to check whether local calibration was evaluated. The fact that a tool has been deployed in 50 health systems is an indication of adoption, not evidence of clinical effectiveness.

3. What is the outcome when the underlying model is altered?

With traditional software, a change takes place when a developer pushes new code. Modern generative and adaptive models are capable of drifting, retraining or altering their underlying weights without notifying your staff.

You need binding answers to three practical questions.

  1. How much advance notice is required by contract when the vendor alters the model, prompt scaffolds, guardrails, or source references?
  2. What validation data is included with that notification to show performance has not deteriorated?
  3. Can your practice stay on a pinned, stable version while you consider the changes?

For products regulated by the FDA, ask for the Predetermined Change Control Plan, which specifies the exact limits within which the manufacturer can update the tool without refiling. When a company promotes continuous real-time updates, recognize that the clinical tool you validated in October may not be the same instrument providing recommendations in your examination rooms next June.

4. What is the complete subprocessor chain, and is our clinical data used to train your foundation models?

Asking whether a vendor is HIPAA compliant is of no use, since every enterprise vendor answers in the affirmative.

Your actual liability usually does not begin with the original vendor. It sits in the chain of downstream subprocessors. An ambient documentation tool picks up confidential patient audio, forwards it to a third-party transcription engine, passes the resulting text to a large language model hosted in the cloud and then sends operational telemetry to an analytics firm.

The patient agreed to have you look after them. You must obtain an explicit, named list of every subprocessor handling protected health information, verify the exact data payload each entity receives and make sure your Business Associate Agreement legally binds every link in that chain.

Just as important is finding out whether the clinical interactions, notes, or patient queries you enter are used to improve the vendor's proprietary models. Where the vendor's terms permit data sharing or model training by default, your patients' encounters are contributing to a commercial product that will be sold to your competitors. If you allow your data to be used in this way, your software licensing fees should take that arrangement into account.

5. What takes place in the case of a failure, and how is historical output checked?

What happens when an upstream EHR field is missing, corrupted or delayed? Does the tool fail in an obvious way, degrade quietly or produce speculative output regardless? Can clinical staff tell the difference immediately from the screen?

You should be able to cut off the software's access across the practice within a few hours, without submitting a support ticket to a remote queue.

From a medico-legal point of view, you need version-locked auditability. Can the platform reconstruct the exact prompt, the raw output, the model version and the clinical context provided for a particular patient encounter three years later? If a negative outcome results in a malpractice deposition, a general system log will not protect you. You will need the precise state of the algorithm on that particular date.

Contractual non-negotiables

Regard the following vendor responses as non-starters.

  1. "The architecture is proprietary." Intellectual property protection is legitimate, but it does not take precedence over clinical safety. If a vendor asks you to assume unquantifiable clinical or regulatory liability without allowing you to examine the validation data, record the refusal and withdraw.
  2. Refusal to commit in writing. A vendor unwilling to put performance baselines, downtime credits or mandatory change-notification timelines into the master services agreement is placing its operational risk on your shoulders. Once the contract has been signed, your commercial leverage is gone.
  3. Disclaimers that contradict the sales pitch. If a product is promoted as a clinical assistant but the underlying agreement excludes all clinical utility and treats the software as administrative or informational only, that inconsistency is deliberate. The vendor's legal team has determined where the malpractice liability should fall.

The reality of deployment

The majority of health care technology companies are not acting dishonestly. Many truly believe in their products, and some of these systems deliver a significant improvement in clinical efficiency.

Enthusiasm is not a replacement for governance. The vendor has compared an idealized model against data from past cases. Nobody has checked your particular deployment, which involves your patients' demographics, your custom templates, your network latency, and your front-desk procedures.

In an academic medical center, narrowing that gap is an entire department's job. In independent practice, that duty is entirely yours. Make sure you have the documentation to support it.

Neel Chauhan, MD, MBA, is a family physician and physician executive, and a former chief operating officer of a U.S. telehealth platform. He has written a ten-volume series on healthcare AI operations and leads The Healthcare AI Institute.