Start with the clinical problem—not the algorithm
An AI product should begin with a clearly defined clinical or operational problem. Ask who will use it, which patient population it concerns, what decision it is intended to inform and what happens if the output is wrong.
A tool that predicts deterioration, drafts documentation or prioritizes a work queue creates different risks. Do not accept broad claims such as “improves care” without a specific intended use and measurable outcome.
- What decision or task is the system intended to support?
- Who is the intended user and who is affected by the output?
- Is it advisory, autonomous or embedded in a device or workflow?
- What is the safe fallback when the system is unavailable or uncertain?
Examine the evidence and its transferability
Performance reported by a vendor is not automatically performance in your patients. Review the study population, comparator, data source, missing-data handling and whether external or prospective validation was performed.
Look beyond a single accuracy number. Sensitivity, specificity, calibration, subgroup performance, prevalence and the consequences of false positives and false negatives may matter more to clinical practice.
- Was the tool evaluated in a population comparable to yours?
- Were clinically meaningful outcomes assessed?
- Is subgroup performance reported across relevant demographic and clinical groups?
- Has performance been independently replicated or prospectively monitored?
Map the human workflow and oversight
Safe performance depends on the people, environment and workflow around the model. Identify where an output appears, who reviews it, how disagreement is documented and when escalation is required.
Human oversight must be operational, not ceremonial. A clinician who lacks time, context or authority to challenge an output is not providing meaningful oversight.
Review governance, privacy and accountability
Before deployment, establish ownership for clinical safety, cybersecurity, privacy, procurement, incident response and change control. Determine whether the system is regulated and what local policies apply.
Document what data enter the system, where they travel, how long they are retained, whether they are reused for model improvement and which contractual safeguards apply.
Monitor the full lifecycle
Approval is not the end of evaluation. Patient mix, clinical practice, software versions and upstream data can change. Define a monitoring plan before launch, including performance thresholds, incident reporting and criteria for pausing or retiring the tool.
The goal is not to prove that AI is perfect. It is to determine whether a specific tool, in a specific workflow, produces a favorable and continuously monitored balance of benefit and risk.
