Verifying AI: An Essential Skill in Evidence-Based Practice
Author: Dr Tiziano Innocenti1 and Dr Nino Cartabellotta1
1. GIMBE Foundation
When an evidence search no longer begins with a study
Picture a clinician seeing a patient with an acute ankle sprain. Twenty years ago, finding out how to help with the pain might have meant searching for a randomised trial. Today, it usually does not. Clinicians are more likely to consult a point-of-care app, check a guideline or ask a generative AI (GenAI) tool. This is not laziness, but an efficient response to an evidence system that now runs through guidelines, systematic reviews and GenAI outputs, rather than primary studies alone. The problem is not that clinicians have stopped reading original research, but that our teaching of evidence-based practice (EBP) has not kept pace with how evidence is actually used.
We write from two positions: as authors of a recent proposal to update EBP competencies, and as educators who spend much of each year in a classroom with practising clinicians. The second role is what convinced us of the need for the first.
A shift the evidence community cannot ignore
This year's World Evidence-Based Health Care (EBHC) Day campaign, Evidence and AI: People at the Centre, names the shift directly: AI is already changing how research evidence is produced, synthesised, translated and implemented, from semi-automated syntheses and living reviews to clinical decision-support tools. For education, this raises an uncomfortable question: if health professionals routinely consult AI-mediated summaries, are we training them to know when to trust what they read?
In a recent paper in BMJ Evidence-Based Medicine, our answer is: usually not. Traditional EBP teaching assumed that most clinicians would search for, appraise and summarise primary studies themselves. Very few ever do – not through lack of diligence, but because the time and specialised training required are rarely available.
Five things worth teaching everyone
Rather than lengthening the list of EBP competencies, we propose narrowing it to five observable, high-yield behaviours that every learner, not only future methodologists, should be able to demonstrate:
- Finding the relevant pre-appraised sources (guidelines, systematic reviews, point-of-care summaries) for a given clinical question
- Assessing their reliability: are benefits and harms clearly reported, and is the certainty of the evidence stated and justified?
- Interpreting effect estimates and certainty, focusing on outcomes that matter to patients rather than surrogate measure.
- Identifying preference-sensitive decisions, where the right choice depends as much on individual values as on the evidence
- Verifying AI-generated outputs by confirming claims using a trustworthy, up-to-date source before acting on them
Trustworthiness is not one item on the list; it runs through all five. Determining whether a review, a guideline or an AI output is transparent, current and independent enough to rely on is now the central judgement in practice.
What this looks like in a classroom
I personally oversee more than ten editions of a residential course on artificial intelligence in healthcare at the GIMBE Foundation, designed for medical doctors and other health professionals. The audience is consistently polarised: at one end, participants who refrain from engaging with these tools; at the opposite end, those who already utilise them extensively without sufficient scrutiny.
The fundamental exercise is straightforward: participants take a real clinical question, employ a large language model to brainstorm and gather evidence, and then verify the outputs. The most common struggle in that regard is not the fabricated citation, which is relatively simple to teach on and increasingly identifiable. It is the reference that exists, appears correctly formatted and yet fails to substantiate the claim to which it is linked. Equally prevalent is the issue of accepting a confident statement purely because it exudes confidence. Participants also encounter a challenge that is difficult to rebut: properly verifying a single answer to the clinical question often requires more time than consulting a guideline would necessitate.
The most challenging aspect to teach is not that these models hallucinate; most participants acknowledge this within minutes. Rather, it is that the model's amiability presents a risk: fluent, agreeable prose lowers the reader's guard precisely when caution should be heightened, and participants consistently underestimate this hazard. The participants who have performed most effectively have not necessarily been those most familiar with the technology. Instead, they have been individuals who have already possessed the skill of interrogating guidelines critically. This observation convinces us that the process of verifying AI outputs should not be taught as a separate, advanced module bolted onto the existing curriculum. Rather, it will function more effectively as an extension of an existing habit that clinicians can be trained to develop.
Historically, we have emphasised the importance of approaching poorly developed guidelines with scepticism; the current task is to extend that scepticism to the confident responses provided by chatbots. The extent to which verification is necessary depends on the stakes involved: for a single chatbot response, it may suffice to trace claims back to a credible source; for a clinical AI system integrated into routine care, verification should involve validation, ongoing monitoring, and well-defined lines of accountability.
Why this matters for governance, not just teaching
The implications reach beyond the classroom. In 2025, Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence issued a joint position statement on AI use in evidence synthesis, the RAISE guidance set out how such tools should be reported, and the Guidelines International Network published eight principles for responsible AI use in guideline development.
All of these updates convey the message at the heart of this year's campaign: AI tools are powerful, but people, not algorithms, are accountable for decisions made based on their outputs. As evidence becomes more mediated, that accountability no longer sits with the clinician alone. It is distributed across the reviewers using semi-automated screening, the guideline panels drafting with AI, the developers who build these tools and the organisations that deploy them.
Accountability without access
This is where we now wish to press harder than we did in our recent paper. The key point we will make is that verification is only a competency that can be developed if the sources you are meant to verify against can actually be opened.
In every edition of the course I oversee, participants are asked whether their institution provides access to a point-of-care resource. The proportion that says yes is roughly one in ten; many Italian institutions do not fund access at all. The consequence of this limitation is visible in the room: participants reach a paywall midway through a verification exercise and stop. Language is a second barrier, since the model answers fluently in Italian while the guideline it should be checked against exists only in English. Access is uneven between professions too: in physiotherapy and much of allied health, trustworthy pre-appraised sources are thinner on the ground than in internal medicine.
Teaching clinicians to verify AI outputs while leaving them without access to the evidence they are verifying against shifts responsibility onto the people least equipped to carry it. Keeping people at the centre does not mean pretending any one person is in control; it means making each responsibility explicit and ensuring it can actually be exercised. Equitable access to trustworthy evidence at the point of decision-making is one of the conditions for building this competency.
Lessons learned
Three things we would rather flag than smooth over.
- The debate we had, and where it landed. Narrowing decades of EBP competency literature to five behaviours involved intentionally excluding certain elements. The most extensive discussion within our author group centred on whether the assessment of risk of bias in primary studies should remain an integral and universally expected component, given that it represents a core commitment of EBP education over the past 20 years. We reached a consensus that clinicians should now begin with pre-appraised evidence, as it is generally impractical for most to directly access and appraise primary research reliably. This consensus was easier to reach than it has been to defend subsequently to educators who are accustomed to traditional, more comprehensive checklists
- What surprised us. Our expectation was that we would encounter resistance towards AI. Instead, we observed polarisation: one group that refrains from utilising these tools entirely and another that employs them uncritically. The commonality between the two groups is their lack of verification. The challenge in education is not merely adoption but rather the requirement for critical scrutiny
- What we still cannot do. Currently, we do not have a validated method of assessing whether a learner can verify an AI output. While it is possible to teach the behaviour and observe it during exercises, we are not yet able to measure if this proficiency transfers to practical application, or if teaching this skill influences anything beyond the learners' confidence in describing the principle
An international effort
We regard our recent paper as a first step, not a conclusion, and certainly not as an update to the Sicily statement on evidence-based practice, agreed by consensus in Sicily in 2003. Instead, our paper is a deliberately partial contribution to a discussion the evidence community needs to have about what that update to the statement should say in an AI-mediated ecosystem.
That discussion will continue at the EBHC International Conference 2026 in Taormina, Sicily, which opens on 21 October, the day after World EBHC Day, and closes on 24 October. On the final day, Dr David Nunan, University of Oxford, will ask what EBP means 20 years on from the Sicily statement.
Key messages
- Verifying AI is not a new skill. It is an old one applied to a new kind of source. The scepticism EBP has always called for towards guidelines and reviews now has to be made explicit and deliberately extended to AI-generated content
- A shorter, sharper competency list beats a longer one. Five observable behaviours give educators something they can actually teach, although we do not yet have a validated way to assess the fifth
- Responsibility is shared, and so is the obligation to make it exercisable. Clinicians cannot verify against evidence they cannot reach, and no competency framework fixes that on its own
References
Dawes, M., Summerskill, W., Glasziou, P., Cartabellotta, A., Martin, J., Hopayian, K., Porzsolt, F., Burls, A., & Osborne, J. (2005). Sicily statement on evidence-based practice. BMC Medical Education, 5, 1. https://doi.org/10.1186/1472-6920-5-1
Flemyng, E., Noel-Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., Jordan, Z., Minx, J., Eisele-Metzger, A., Hamel, C., Jemioło, P., Porritt, K., & Grainger, M. (2025). Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence 2025. Campbell Systematic Reviews, 21, e70074. https://doi.org/10.1002/cl2.70074
GIMBE Foundation. (2026). EBHC International Conference 2026: schedule. https://www.ebhcconference.org/schedule.en-GB.html
Innocenti, T., Guyatt, G., Cartabellotta, N., Glasziou, P., Ilic, D., & Nunan, D. (2026). Educating health professionals in evidence-based practice in 2026. BMJ Evidence-Based Medicine. https://doi.org/10.1136/bmjebm-2025-114184
JBI. Evidence and AI: people at the centre – World EBHC Day 2026 campaign. https://worldebhcday.org/
Sousa Pinto, B., Marquez-Cruz, M., Neumann, I., Chi, Y., Nowak, A. J., Reinap, M., Awad, M., Nothacker, M., Trucl, M., Brozek, J., Alonso-Coello, P., Wiercioch, W., Qaseem, A., Akl, E. A., & Schünemann, H. J. (2025). Guidelines International Network: principles for use of artificial intelligence in the health guideline enterprise. Annals of Internal Medicine, 178(3). https://doi.org/10.7326/ANNALS-24-02338
Thomas, J., Flemyng, E., & Noel-Storr, A. (2025). Responsible AI in Evidence Synthesis (RAISE): guidance and recommendations (version 2). OSF. https://doi.org/10.17605/OSF.IO/FWAUD
To link to this article - DOI: https://doi.org/10.70253/LEYI9029
Disclaimer
The views expressed in this World EBHC Day Blog, as well as any errors or omissions, are the sole responsibility of the author and do not represent the views of the World EBHC Day Steering Committee, Official Partners or Sponsors; nor does it imply endorsement by the aforementioned parties.