What Evidence Should Healthcare Organizations Ask for Before Adopting Clinical AI?

EVIDENCE BRIEF

9/13/202614 min read

What Evidence Should Healthcare Organizations Ask for Before Adopting Clinical AI?

A clinical AI vendor may arrive with an impressive presentation.

The system may report high accuracy, cite published studies, demonstrate sophisticated technology, or already be used by hospitals in another country.

All of that may be encouraging.

But none of it, on its own, answers the question a healthcare organization ultimately needs to answer:

Is there enough evidence to use this AI safely and meaningfully for our patients, clinicians, and care environment?

That question is becoming increasingly important in Viet Nam.

The Law on Artificial Intelligence No. 134/2025/QH15, effective from March 1, 2026, gives healthcare specific attention. Article 6 highlights three considerations in particular: patient safety, reliability under real-world conditions of use, and protection of health data.

Internationally, the direction is similar. The WHO, International Telecommunication Union, and World Intellectual Property Organization Global Initiative on AI for Health promotes governance, standards, technical guidance, and evidence-based adoption of AI for health. The NICE Evidence Standards Framework for Digital Health Technologies also illustrates an important principle: the evidence expected from a technology should reflect what it is intended to do and the consequences if it performs poorly.

These international frameworks are not regulatory requirements for healthcare organizations in Viet Nam. But they provide useful reference points for asking better questions.

The practical message is simple:

Do not ask only whether an AI system has evidence. Ask whether it has the right evidence for the way you intend to use it.

First, what do we mean by “clinical AI”?

For the purpose of this article, clinical AI means AI systems whose outputs may influence patient care, including applications supporting:

  • screening;

  • diagnosis;

  • medical imaging interpretation;

  • risk prediction or prognosis;

  • triage;

  • clinical decision support;

  • treatment selection or planning;

  • patient monitoring; and

  • other decisions that may materially affect patient management.

Evidence requirements should be proportionate to the intended use, clinical consequence, and remaining uncertainty.

An AI tool that organizes an administrative inbox should not require the same level of evidence as a system that helps determine whether a patient needs urgent intervention.

The starting question is therefore not:

“How advanced is this AI?”

It is:

“What decision will this AI influence, and what could happen if it is wrong?”

1. What exactly is the system claiming to do?

Before looking at accuracy numbers or research papers, healthcare organizations should define the intended use.

Consider the difference between these descriptions:

“AI for chest X-rays.”

“AI that detects abnormalities on chest X-rays.”

“AI that identifies pulmonary nodules.”

“AI that prioritizes chest X-rays with suspected pneumothorax for earlier radiologist review.”

“AI that independently diagnoses pneumothorax.”

These are very different claims.

They imply different users, different positions in the clinical pathway, different consequences of error, different comparators, and potentially very different evidence requirements.

Healthcare organizations should ask:

Who is the intended patient population?

In which clinical setting should the system be used?

Who is expected to use it?

At what point in the care pathway?

What input does it require?

What output does it produce?

Does it inform, prioritize, support, or directly drive a decision?

What action is expected after the output?

What uses are outside its intended purpose?

Without a precise intended use, it is difficult to know whether the evidence presented by a vendor is actually relevant.

2. Who was the AI tested on?

A headline performance number tells us little without knowing where the data came from and who was represented in them.

Healthcare organizations should understand the population behind the evidence.

Useful questions include:

How many patients were included?

Were the data collected from one hospital or multiple sites?

How were patients selected?

What were their ages and relevant clinical characteristics?

What disease prevalence and severity were represented?

What inclusion and exclusion criteria were used?

Were important patient groups underrepresented?

How similar were the study patients to those who would receive care in our organization?

For imaging AI, factors such as scanner manufacturer, acquisition protocols, image quality, and clinical workflow may also affect performance.

The CLAIM 2024 Update, which applies specifically to AI research in medical imaging, emphasizes transparent reporting of areas such as data sources, patient selection, intended use, image acquisition, and testing datasets.

This is particularly relevant when most evidence comes from another country.

An AI system developed and tested in the United States, Europe, China, Korea, or another market may work very well in Viet Nam.

But that should be assessed rather than assumed.

Differences in disease prevalence, demographics, referral patterns, clinical practice, equipment, language, documentation, laboratory methods, and data quality can all affect real-world performance.

3. Was the system tested independently from the data used to develop it?

This is one of the most important questions in clinical AI evaluation.

AI systems may perform impressively when evaluated on data that are closely related to the data used during development.

A more difficult test is how the system performs when it encounters new patients, new sites, and new conditions outside its development environment.

Healthcare organizations should ask:

Was the model tested only on internal data?

Was a completely separate test dataset used?

Did testing include another hospital or healthcare system?

Did it include a different geographic location?

Was external testing performed only by the developer, or has independent evaluation also occurred?

Did performance remain reasonably stable?

Even the terminology matters.

The CLAIM 2024 Update specifically cautions against vague use of the term “validation” because different audiences may interpret it differently. For medical imaging AI, it recommends clearer terminology such as internal testing and external testing.

So when a vendor says:

“Our AI has been validated.”

A useful response is:

“Where was it tested, on whom, against what reference standard, and how independent was the test data from the data used to develop the model?”

That question usually tells you much more.

4. Are the performance measures clinically meaningful?

Clinical AI presentations often focus on one impressive number.

Perhaps:

95% accuracy.

An area under the receiver operating characteristic curve of 0.94.

90% sensitivity.

High agreement with specialists.

These numbers may be useful.

But no single metric tells the whole story.

Depending on the clinical use case, organizations may need to understand:

  • sensitivity;

  • specificity;

  • positive and negative predictive values;

  • false-positive and false-negative rates;

  • discrimination;

  • calibration;

  • clinically relevant thresholds;

  • uncertainty around the estimates; and

  • performance compared with the current standard of care.

Which measures matter most depends on what the system is meant to do.

For an AI triage tool designed to identify a life-threatening condition, missed positive cases may be especially important.

For another application, excessive false-positive alerts may create unnecessary testing, patient anxiety, additional workload, or alert fatigue.

Prediction models introduce another issue.

A model can distinguish higher-risk patients from lower-risk patients reasonably well while still systematically overestimating or underestimating the actual probability of an outcome.

That is why calibration, not only discrimination, can matter.

For prediction models, TRIPOD+AI, published in 2024, provides updated reporting guidance, while PROBAST+AI, published in 2025, provides a structured approach to assessing quality, risk of bias, and applicability. PROBAST+AI is explicitly intended to be useful not only to researchers but also to people considering whether prediction models should be implemented in healthcare practice.

The practical lesson is straightforward:

Do not accept a performance number without understanding what was measured, how it was measured, and whether it matters clinically.

5. How does the AI perform across relevant patient groups?

Average performance can hide important differences.

An AI system may perform well overall while performing less well in certain groups.

Depending on the technology, relevant differences might involve:

  • age;

  • sex;

  • disease severity;

  • comorbidities;

  • population groups;

  • equipment or facility type;

  • inpatient versus outpatient settings; or

  • other clinically meaningful characteristics.

This does not mean every study can provide reliable estimates for every possible subgroup.

But healthcare organizations should understand where important uncertainty remains.

The FUTURE-AI international consensus guideline, published in 2025 and developed through a consortium of 117 experts from 50 countries, organizes trustworthy healthcare AI around six principles: fairness, universality, traceability, usability, robustness, and explainability. Its recommendations span the AI lifecycle from development through deployment and monitoring.

For a hospital, subgroup performance is not simply an abstract issue of fairness.

It can also be a question of patient safety and generalizability.

If the patients for whom performance is uncertain represent a substantial part of the population you serve, that uncertainty matters.

6. Has the AI been evaluated in real clinical care?

Retrospective testing can tell us whether an algorithm performs on historical data.

It cannot fully tell us what will happen once people start using it.

When an AI system enters clinical practice, the intervention is no longer just the algorithm.

It becomes the AI system plus the clinicians, patients, workflow, interfaces, infrastructure, training, and organizational environment around it.

Clinicians may ignore useful recommendations.

They may over-trust incorrect ones.

An alert may arrive at the wrong point in the workflow.

A technically accurate system may add enough friction that clinicians stop using it.

Unexpected cases may arise.

Users may gradually extend the technology beyond its intended purpose.

This is why real clinical evaluation matters, especially as the consequences of error increase.

The DECIDE-AI guideline was specifically developed for early-stage live clinical evaluation of AI decision-support systems. It emphasizes areas including small-scale clinical utility, safety, human factors, and preparation for larger evaluations, rather than focusing only on offline algorithm performance.

For clinically consequential AI, organizations may therefore ask:

Has the system been evaluated prospectively?

Has it been used in a live clinical workflow?

How did clinicians interact with it?

Did it change decisions?

Were unexpected safety issues identified?

Were recommendations used as intended?

How did the human-AI team perform?

This last point is increasingly important.

The International Medical Device Regulators Forum's 2025 Good Machine Learning Practice principles specifically state that AI-enabled medical devices should be assessed with attention to human-AI interactions in their intended use environment, rather than evaluating the device only in isolation.

7. Does better model performance actually improve care?

This distinction is critical:

Technical performance is not the same as clinical utility.

Imagine an AI diagnostic system with excellent sensitivity and specificity.

That is useful evidence.

But what happens after the result?

Does it lead to earlier diagnosis?

Does it change clinical management appropriately?

Does it prevent unnecessary testing?

Does it improve patient outcomes?

Does it reduce waiting time?

Does it reduce clinician workload?

Does it create new unnecessary interventions?

Does it improve care compared with what clinicians already do?

NICE makes this distinction explicitly in its Evidence Standards Framework. For technologies used to diagnose or guide clinical management, test accuracy alone may not establish clinical utility. Evidence may also need to address the downstream consequences of using the result.

That leads to another useful question:

What would happen if we did not adopt this AI?

The appropriate comparator may be:

current clinical practice,

another technology,

specialist review,

an existing clinical score,

or no additional intervention.

A clinically useful AI system should create value relative to a meaningful alternative, not simply perform well in isolation.

8. Does the system fit the clinical workflow?

Even an accurate AI system can fail if it does not fit how care is actually delivered.

Healthcare organizations should therefore examine implementation evidence.

How long does the system take to use?

Where does its output appear?

Does it integrate with existing clinical information systems?

Does it add extra steps or duplicate work?

Who receives alerts?

What happens outside normal working hours?

How are conflicting results handled?

What happens if the system is unavailable?

How much training is required?

Will clinicians actually use it?

Does it reduce workload, or simply move work to another department?

What new responsibilities does it create?

Evidence about usability and workflow may include both quantitative and qualitative information.

NICE recognizes the value of evidence from patients and healthcare professionals about their experience with digital technologies, as well as real-world evidence showing whether claimed benefits can actually be achieved in practice.

For hospitals, this is important because the value of an AI system is often created, or lost, in the workflow around the algorithm.

9. Is the evidence applicable to our own organization?

For healthcare organizations in Viet Nam, this may be the most important question.

Strong international evidence is valuable.

But international evidence does not automatically remove local uncertainty.

A hospital should ask:

Are our patients sufficiently similar to those studied?

Is disease prevalence comparable?

Do we use similar diagnostic criteria?

Are our scanners, laboratory platforms, or data formats similar?

Is the technology compatible with local clinical workflows?

If language is involved, has performance in Vietnamese been evaluated?

Does the system understand relevant Vietnamese terminology, abbreviations, documentation styles, or conversational patterns?

Will our clinicians use the system in the same way as those in the studies?

Will they receive comparable training?

Do we have the infrastructure and data quality the system requires?

Local evaluation does not necessarily mean repeating the entire research program.

The appropriate approach depends on the risk of the use case, the strength of the existing evidence, the differences between the original setting and the new setting, and the uncertainty that remains.

Options may include:

local technical verification, using representative local data to confirm expected performance;

offline or silent-mode evaluation, where the system runs on local inputs without influencing clinical decisions;

a monitored pilot, allowing the organization to assess workflow, usability, safety, and operational performance before scaling; or

prospective local clinical evaluation, when the use case is sufficiently consequential and uncertainty remains high.

NICE specifically recognizes silent-mode evaluation as one way to examine the performance of AI technologies on local inputs before integrating them into care pathways.

This local focus is also consistent with the direction of Viet Nam's AI Law, which specifically emphasizes reliability under real-world conditions of use for healthcare AI.

The question is therefore not necessarily:

“Has this AI been studied in Viet Nam?”

A better question is:

“What remaining uncertainty is relevant to our setting, and what level of local evaluation is proportionate to that uncertainty and risk?”

10. What happens to the evidence after deployment?

Evidence assessment should not end when the contract is signed.

AI systems and their environments can change.

The model may be updated.

Software may change.

Clinical workflows may evolve.

Input data may shift.

Disease patterns may change.

New patient groups may begin using the service.

Users may change how they interact with the system.

Performance can therefore change over time.

Healthcare organizations should ask:

How will performance be monitored?

Which indicators will be tracked?

How will problems or incidents be identified?

Will we know when the model changes?

How is versioning managed?

What constitutes a significant update?

What testing occurs before updates are released?

Can previous versions be identified?

What happens if performance deteriorates?

What evidence will be generated from real-world use?

Who is responsible for reviewing performance?

The 2025 IMDRF Good Machine Learning Practice principles specifically include ongoing monitoring of deployed models and management of risks associated with retraining, including performance degradation and dataset drift.

The question should therefore not only be:

“Does this AI work?”

It should also be:

“How will we know that it continues to work?”

Regulatory status matters, but it does not answer the evidence question

Healthcare organizations should determine which regulatory requirements apply to an AI product and whether those requirements have been met.

In Viet Nam, an AI-enabled product may also fall within the medical device regulatory framework depending on its intended purpose and characteristics. Circular No. 24/2026/TT-BYT, effective July 1, 2026, addresses risk determination and management measures for medical device products.

But regulatory status and clinical evidence answer different questions.

Meeting applicable regulatory requirements does not, by itself, tell a hospital:

whether the technology addresses an important local problem;

whether the evidence is methodologically strong;

whether the study population is relevant;

whether the workflow fits;

whether clinicians will use the technology appropriately; or

whether the system will create enough clinical or operational value to justify adoption.

The regulatory question is:

“What requirements apply to this product, and have they been met?”

The healthcare organization's evidence question is broader:

“Should we use it here, for this purpose, with these patients, under these conditions?”

Both matter.

They should not be confused.

Peer-reviewed does not automatically mean high-quality

Another common shortcut is to ask whether the technology has been published in a peer-reviewed journal.

Peer-reviewed evidence is usually more informative than an unsupported marketing claim.

But publication alone does not establish quality.

A published study can still involve:

small or highly selected populations;

weak comparators;

inappropriate reference standards;

risk of bias;

inadequate external testing;

poorly defined outcomes;

incomplete reporting; or

limited applicability to another setting.

This is why several AI-specific reporting and appraisal frameworks have emerged.

They do different jobs.

CLAIM focuses on reporting AI studies in medical imaging.

TRIPOD+AI addresses reporting of clinical prediction models.

PROBAST+AI helps assess quality, risk of bias, and applicability of prediction models.

DECIDE-AI focuses on reporting early-stage live clinical evaluation of AI decision-support systems.

STARD-AI, published in Nature Medicine in 2025, focuses on diagnostic accuracy studies involving AI.

These are not certificates of quality, and they are not interchangeable.

They help researchers and decision-makers ask more precise questions about how evidence was generated, reported, and whether it is applicable to the decision at hand.

How much evidence is enough?

There is no single evidence package that is appropriate for every clinical AI system.

Evidence requirements should increase with risk, uncertainty, and the clinical importance of the decision being influenced.

For a lower-risk tool, technical performance, usability, workflow fit, and operational value may form an important part of the evidence.

For a system informing clinical management, stronger evidence of clinical performance and real-world use may be needed.

For systems directly influencing high-stakes decisions, organizations should expect substantially greater assurance around safety, performance, human factors, clinical utility, and real-world reliability.

NICE reflects a similar principle. When accurate and timely diagnosis or treatment is critical to avoiding death, serious deterioration, or long-term disability, less uncertainty in technology performance is acceptable.

A useful rule is:

The greater the potential consequence of being wrong, the stronger the evidence should be.

What should a healthcare organization ask a vendor to provide?

Before adopting clinically consequential AI, an organization might request an evidence package covering:

  1. A precise intended-use statement specifying the target population, setting, intended users, clinical role, output, and important limitations.

  2. The exact product and model version to which the evidence applies.

  3. Applicable regulatory and risk-classification information, including relevant medical device status where appropriate.

  4. Technical performance evidence using clinically meaningful measures and appropriate reference standards.

  5. Independent or external testing evidence, where appropriate for the intended use and risk.

  6. Information about development and evaluation populations, including their relevance to the organization's target population.

  7. Subgroup performance and known limitations, including areas where important uncertainty remains.

  8. Prospective or live clinical evidence when the level of clinical consequence warrants it.

  9. Evidence of clinical utility, where relevant, showing whether use of the AI improves decisions, processes, or patient outcomes compared with a meaningful alternative.

  10. Human factors and workflow evidence, including usability, clinician interaction, training requirements, and safe fallback processes.

  11. Real-world performance evidence, where available.

  12. A post-deployment monitoring and change-management plan, including model updates, versioning, performance drift, incidents, and significant changes.

  13. Data governance and cybersecurity information, because good predictive performance does not establish that data practices are acceptable.

  14. Resource and value information, including integration requirements, staff time, downstream testing, infrastructure, and other operational consequences.

No single document will necessarily contain all of this.

Nor should every AI system be subjected to exactly the same evidence requirements.

But a vendor proposing a technology that may materially influence patient care should be able to engage seriously with these questions.

Some evidence red flags

Healthcare organizations should ask further questions when:

“95% accuracy” is presented without explaining the population, reference standard, comparator, uncertainty, or other clinically relevant performance measures.

The evidence comes only from data closely related to those used to develop the model.

“Validation” is repeatedly claimed without explaining what type of testing was performed.

Only favorable studies are easy to obtain.

The vendor cannot identify which product or model version was used in published research.

Studies demonstrate algorithm performance but provide little information about clinician interaction or workflow.

Performance across clinically important groups is unknown.

The evidence population differs substantially from the proposed local population, but the implications are not discussed.

Important limitations are absent from the presentation.

The model can change after deployment, but change control, versioning, and monitoring are unclear.

Regulatory status is presented as though it proves local clinical benefit.

None of these automatically means that a technology should be rejected.

They mean:

More questions are needed before adoption.

Evidence should continue after procurement

The strongest healthcare organizations will not treat AI evidence assessment as a procurement checklist that is completed once.

They will connect evidence with governance across the technology lifecycle.

Before adoption:
What evidence supports this technology?

During implementation:
Does it work safely and appropriately in our workflow?

After deployment:
Is it performing as expected in our patients and our environment?

When the system or environment changes:
Does the evidence still apply?

This is the difference between purchasing AI and responsibly implementing it.

Healthcare organizations in Viet Nam are likely to encounter a growing number of clinical AI products.

Some will create meaningful value.

Some will be promising but immature.

Some may perform well in one environment and less well in another.

Some may have strong technical performance but limited evidence of clinical utility.

And some claims will simply be stronger than the evidence behind them.

Being able to distinguish between these situations will become an increasingly important organizational capability.

Healthcare leaders do not need to become machine-learning researchers.

But they do need to know which questions to ask, which uncertainties matter, and when the available evidence is not yet strong enough for the decision being considered.

The right question is therefore not simply:

“Does this AI have evidence?”

It is:

“Is the evidence sufficiently robust, relevant, and applicable for the way we intend to use this AI with our patients, clinicians, and organization?”

That is where responsible clinical AI adoption should begin.

Key references

Viet Nam. Law on Artificial Intelligence No. 134/2025/QH15. Effective March 1, 2026. Article 6 specifically addresses healthcare, including patient safety, reliability under real-world conditions of use, and protection of health data.

World Health Organization, International Telecommunication Union, and World Intellectual Property Organization. Global Initiative on AI for Health. Promotes governance, standards, normative technical guidance, and evidence-based adoption of AI for health.

National Institute for Health and Care Excellence. Evidence Standards Framework for Digital Health Technologies. Provides an evidence framework based on technology function, risk, performance, effectiveness, real-world evidence, and other factors relevant to decision-making.

Vasey B, et al. DECIDE-AI. Reporting guideline for early-stage live clinical evaluation of artificial intelligence decision-support systems. BMJ. 2022.

Collins GS, et al. TRIPOD+AI. Updated reporting guidance for clinical prediction models using regression or machine-learning methods. BMJ. 2024.

Tejani AS, et al. CLAIM: 2024 Update. Updated reporting guidance for artificial intelligence studies in medical imaging. Radiology: Artificial Intelligence. 2024.

Lekadir K, et al. FUTURE-AI. International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025.

Moons KGM, et al. PROBAST+AI. Updated tool for assessing quality, risk of bias, and applicability of prediction models using regression or artificial intelligence methods. BMJ. 2025.

Sounderajah V, et al. STARD-AI. Reporting guideline for diagnostic accuracy studies using artificial intelligence. Nature Medicine. 2025.

International Medical Device Regulators Forum. Good Machine Learning Practice for Medical Device Development: Guiding Principles. IMDRF/AIML WG/N88 FINAL:2025, published January 29, 2025.

This article is intended for educational and informational purposes. It does not constitute legal, regulatory, clinical, research methodology, or technology procurement advice. The level and type of evidence required should be determined according to the specific AI system, intended use, clinical consequences, regulatory requirements, patient population, existing evidence, and operating environment.

Updated: September 13, 2026