When AI Assesses AI in Healthcare
Author: Dr Devi Varma1
1. JBI Kalam Institute of Health Technology
I work in health technology assessment as a lead epidemiologist and biostatistician, and a near miss the other day caused me to rethink the use of AI in research.
We often discuss the topic of AI’s suitability as healthcare technology. When we consider an AI-enabled diagnostic tool, our questions are: Is it accurate? Is it safe? Can it support clinicians? Is it worth paying for? And AI is starting to help with the very process of making that assessment.
We are starting to use AI in this way extract evidence, summarise published studies, draft sections of reports, check formulas, assist with meta-analysis, and sometimes, even help build economic models in Excel, Python, and R. These are valuable contributions, as the job is tough, with tight timelines and scattered evidence. Delivering a single report that assesses the viability of an AI-enabled diagnostic tool may require pulling together clinical evidence, evidence on diagnostic accuracy, cost data, epidemiology, assumptions around implementation, and stakeholder inputs.
In those moments, AI can feel like a friendly coworker. But it isn’t a colleague. And that is an important distinction to make, because AI does not have professional responsibility, know the local effect of an incorrect assumption, or feel discomfort when navigating doubt.
A small moment recently crystallised the importance to me of those human contributions.
In a recent healthcare technology assessment discussion, we reviewed an emerging piece of technology for which there is limited evidence from India. There are some international studies, but the local evidence does not cover all the model parameters we needed to complete our assessment.
The model could still be constructed with international inputs, but we judged those to provide only weak evidence. That human judgement call was important to make, as a spreadsheet wouldn’t alert us, ‘This assumption isn’t local enough.’ Instead, it would execute the assessment – a probabilistic sensitivity analysis would generate compelling graphs, a professional cost-effectiveness plan, and a neat report.
The real question to ask, however, was how much trust should an Indian decision-maker put in results partly based on evidence from another health system?
In human-led assessment, we must often make decisions in the absence of perfect evidence. That uncertainty is accounted for and presents a controlled risk. Yet, in the age of AI, we are seeing the introduction of a new phenomenon – hidden uncertainty – which introduces an element of danger.
If we use international evidence for sensitivity, specificity, transition probabilities, uptake, adherence, referral rates, utility values, or resource use, we should state this clearly when reporting our assessment. We should explain why we used this evidence base, call into question whether it can be replicated in India, and apply that evidence in sensitivity and scenario analyses. Most importantly, we should not allow the outputs of the model to appear more certain than our own certainty in the evidence.
Where and where not to take help from AI
I am by no means anti-AI. It has been very helpful in my experience, particularly to organise lengthy documents, find relevant values in a study, compare extracted information across papers, recommend the structure of a report, and identify potential parameters for an economic model – it can even help check code or formulas, if used judiciously.
These contributions are important in areas where teams are small and the workload is heavy.
But I’ve also seen where it falls short. For instance, in my line of work, AI can recognise a number from a paper but not whether the number is appropriate for our decision problem. And it may provide an overview of a study, but that overview may not fully reflect the population, setting, reference standard, threshold, comparator, or implementation context. These types of inputs may help me write a nice paragraph, but they do not tell me whether the evidence is locally meaningful.
That judgement should stay with the human team.
Why we must still put our trust in each other
The biggest concern I have is that AI can make weak material look organised.
A rough assumption turns into a table. A borrowed parameter forms an input. The paragraphs flow. A model is presented as usable.
Fluency, however, is not the same as trustworthiness.
This is the area, I think, where health technology assessment teams need to be careful. If data are obtained through AI, the original article must be checked by a person. When AI performs a meta-analysis, the data and methods should be verified. If it is used to build an economic model, its formulas, distributions, assumptions, and outputs should be reviewed. If it is used to write a report, the final interpretation should be performed by the assessment team and not by the tool.
The World Evidence-Based Healthcare Day 2026 campaign examines how we can retain the human in evidence-based healthcare in the era of AI. And for me, that means the people completing the assessment. We need to be accountable for the process of evidence retrieval and use.
Continuing need for local validation studies
In my line of work, local validation studies are useful not only for regulatory or clinical performance but also to enhance health technology assessment.
An Indian validation study may not always demonstrate the long-term clinical effectiveness of AI-enabled medical devices, but it can still provide us with important model inputs: sensitivity, specificity, false positives, false negatives, test failure rates, referral rates, time saved, repeat testing, clinician override, usability issues, subgroup performance, and implementation costs.
And these can be impactful. For instance, a slightly higher false-positive rate may lead to more patient referrals than are necessary, in turn triggering needless anxiety. Alternatively, a higher false-negative rate can lead to delays in care. The same tool that works well on a controlled dataset may behave differently in a district hospital, private clinic, screening camp, or rural public health setting, which an Indian validation study may serve to identify.
In these ways, health technology assessment can inform innovation, not just evaluate its adoption after the fact. In particular, evidence providing local validation can help reduce uncertainty, as well as make an economic model more realistic and point to evidence needed before broader scale-up. These values continue to be of importance in the age of AI, and may in fact have gained importance given the hidden uncertainty that AI tools unfortunately introduce.
My plan to harness AI
If we are to responsibly proceed with AI-assisted health technology assessment, then I recommend introducing five simple rules:
1. All reports must have an AI usage log, stating whether it was used in screening, extraction, writing, coding, modelling, or visualisation
2. All extracted values should be traceable back to the original source – no more disappearing page numbers, tables, assumptions, or reviewer checks
3. No confidential, unpublished, patient-level, or proprietary data should be uploaded to unsafe public tools
4. Economic models must be accompanied by a register of assumptions – when a parameter is borrowed from another country, it must be clearly marked
5. The individual with ultimate responsibility for the assessment should be named – AI can assist, but the human expert must remain responsible
These rules align with the recommendations on responsible AI use in evidence synthesis, which highlight the need for transparency, human oversight, and accountability, along with the World Health Organization’s guidance on the ethics and governance of AI for health, which call for AI usage that safeguards trust, equity, and human rights. Furthermore, in the Indian context, my simple rules correspond with the Digital Personal Data Protection Act 2023, which reminds us that data protection cannot be an afterthought.
The question steering my responsible work
There is a key question that I now ask when performing health technology assessment work:
If AI helped us reach the conclusion, can we really explain each step?
If the answer is yes, then we have responsibly applied AI to assist with our work. If the answer is no, then the report may appear complete, but its assessment should not yet be trusted. We must go back and identify the steps taken to reach the outcome, as well as verify their correct execution.
That due diligence is important to perform, even if it requires some time and effort. Because AI can allow us to work faster, but speed is not a substitute for judgement.
Ultimately, staying true to the evidence will permit us to innovate while also moving forward safely. Local evidence, as ever, provides solid ground for our decision-making, and human responsibility in applying AI is key to directing our path.
Key messages
1. AI can assist health technology assessment, but judgement must ultimately stay with human experts
2. In the absence of local evidence, international evidence can be used cautiously, but the assumptions guiding its use must be visible, justified, and tested
3. Institutions conducting health technology assessment should develop practical policies on responsible use of AI, including disclosure, verification, data protection, model transparency, and accountability
References
Flemyng, E., Noel-Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., Jordan, Z., Minx, J., Eisele-Metzger, A., Hamel, C., Jemioło, P., Porritt, K., & Grainger, M. (2025). Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence. JBI Evidence Synthesis, 23(11), 2162–2166. https://doi.org/10.11124/JBIES-25-00480
Government of India, Ministry of Electronics and Information Technology. (2023). The Digital Personal Data Protection Act, 2023. https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf
World Evidence-Based Healthcare Day. (2026). World EBHC Day 2026 campaign: Evidence and AI: People at the centre. https://worldebhcday.org/
World Health Organization. (2021). Ethics and governance of artificial intelligence for health: WHO guidance. https://www.who.int/publications/i/item/9789240029200
To link to this article - DOI: https://doi.org/10.70253/JRBU2946
Disclaimer
The views expressed in this World EBHC Day Blog, as well as any errors or omissions, are the sole responsibility of the author and do not represent the views of the World EBHC Day Steering Committee, Official Partners or Sponsors; nor does it imply endorsement by the aforementioned parties.