Keeping Humans in the Search
Authors: Hugo Neves1,2, Joana Sousa1,2, Adriana Coelho1,2, Carmen Queirós2,3,4, Márcia Pestana-Santos1,2, Rosa Silva1,2,3,4, Vítor Parola1,2
1. Health Sciences Research Unit: Nursing (UICISA: E), Escola Superior de Enfermagem da Universidade de Coimbra [Nursing School of Coimbra], University of Coimbra, Coimbra, Portugal;
2. JBI Portugal Centre for Evidence-Based Practice, UICISA: E, Nursing School of Coimbra, University of Coimbra, Coimbra, Portugal
3. Nursing School of the University of Porto, Porto, Portugal
4. RISE-Health, Porto, Portugal
What we learned from the MCL Search Calibrator
We began with what should have been a simple step. In a qualitative evidence synthesis on experiences of comfort, we wanted to read the evidence through Kolcaba’s comfort theory, so we asked an artificial intelligence (AI)-assisted tool to expand an initial, rough search. The first result looked helpful. It picked up part of the language we expected, including relief, ease, and transcendence, and it handed back a wider set of terms to consider.
The difficulty appeared once we stopped reading the search as a finished answer and started reading it as a set of decisions. Some problems were easy to catch: terms like ‘Kolcaba’ and ‘theory’, when paired with the search operator OR, would only widen the search toward noise, pulling in papers that name the theorist without addressing the experience of comfort. Beyond that, a harder problem to detect, which you would only notice if you know the theory, was what the tool left out. For instance, it barely touched on the four contexts that give Kolcaba’s theory its shape: physical, psychospiritual, sociocultural, and environmental comfort, and it returned barely anything for discomfort, distress, and the other negative experiences that would be covered by a qualitative synthesis.
The problems we identified did not leave the search worthless. It was just uneven: too wide in some places and too thin in others. That was a lesson we kept returning to in our work. A model can produce plenty of language that surrounds a topic and still have no real sense of where the topic ends. We realised we needed a way to examine every suggestion systematically rather than accept or reject it on instinct. That experience became our starting point.
Why we needed a Calibrator
The MCL Search Calibrator was born from that drive to examine AI suggestions. We were not looking for something that would write a strategy and ask us to trust it. We wanted the opposite: a tool that would keep the work visible, including the choices that usually disappear into a clean Boolean string (the final line of terms and operators).
In a search strategy, small choices carry weight. A wildcard meant to catch spelling variants can just as easily drag in noise. A date limit can impose a defensible boundary or just offer a way to reduce an unwieldy result set. The operators are the clearest case: AND makes a strategy more precise and also means it can miss papers, while OR keeps the strategy sensitive and yet can allow it to drift off the question. None of this is automatic, and the reviewer must be able to explain why each choice was made.
With the comfort search, AI gave us a starting frame and nothing more solid than that. Turning it into a strategy still meant stripping out terms that added noise, restoring the contexts required by the theory, and deciding how to capture experiences of discomfort without retrieving studies outside the scope of our review.
And that wasn’t an isolated incident. A second review brought up the same issues with AI from a different angle. This time, we were working on a qualitative evidence synthesis examining athletes with knee injuries and their experiences of empathy, autonomy support, and satisfaction with care during rehabilitation. Compared with the comfort search, this draft looked much more complete, with several concept blocks and plenty of plausible synonyms. And yet it had quietly dropped two ideas that were sitting right there in our own title: autonomy support and satisfaction with care. The fact that they were missing did not jump out on a first read. We only caught it by going through the strategy block by block and asking what each one was actually contributing.

Transparency as part of the method
One of the questions we kept asking ourselves was simple: if someone looked only at the final search strategy, would they know how we arrived there? Usually, the answer was no. The discussions, discarded terms, and AI suggestions disappeared, leaving only the polished Boolean string. That journey toward arriving at the final string is exactly what the MCL Search Calibrator is designed to preserve. Instead of showing only the final strategy, it records what AI suggested, what the reviewer changed, and what the team finally agreed to use. Those decisions become part of the search rather than hidden work.
This is the point where the tool meets the PRISMA-S extension. PRISMA-S asks review teams to describe their searches in sufficient detail that someone else could appraise and rerun them. Once AI is in the loop, there is one more thing worth recording: where each part of the search actually originated. Which lines did the model propose, where did a human step in, and what did the team validate before any of it counted?
We are now working out how to make that kind of reporting possible without turning it into paperwork. The idea we keep returning to is a small set of reason labels you select instead of writing out an explanation for every term. Include a term, and you pick the reason: a thesaurus heading, a construct from the theory, a field synonym, a spelling variant, a deliberate reach for sensitivity. Exclude one, and you state the issue: too broad, ambiguous, out of scope, redundant, bad for specificity. The labels only work if they are light enough that reviewers can actually use them, while still saying enough to land in a methods section.
What still needs work
The Calibrator is still unfinished, and that is worth saying plainly. The early versions handled truncation and wildcards too bluntly, so we often had to correct them by hand. Too much of the suggestion logic still leans on broad, web-style expansion, when what the work actually needs is discovery grounded in the sources themselves: indexed databases, thesauri, established review methods. And the Calibrator also needs proper multilingual support before it is of much use to anyone working outside English.
One surprise has been how often the AI produces searches that look convincing. Most of the important problems only become visible when we inspect concept blocks one by one. We are also building toward a clearer record of how the tool was used: the AI suggestions, human edits, exclusions, validations, perhaps even an adapted PRISMA flow diagram. For us, this is more than a feature request. It is where the methodological and ethical case for the tool actually lives. The 2025 position statement on AI use in evidence synthesis makes clear that responsibility remains with the synthesists, that AI should be operated under human oversight, and that its use should be reported whenever a tool makes or suggests a judgment.
What the Calibrator has taught our team
One thing we did not plan for is how useful the Calibrator has become in teaching. Hand a student a generic chatbot, and they will produce a search string that looks finished, which is exactly the problem. The Calibrator slows them down and makes them treat that same string as a series of choices they have to defend: why this concept and not a broader one, why an AND here, why this term was cut.
Watching students use the tool has been one of the biggest surprises. Something shifts when they work that way. Sensitivity and specificity stop being definitions to memorise for an exam and start being things they can watch happening inside their own search. And ‘AI suggested it’ no longer suffices. If anything, it raises the question of why that suggestion was made, and whether it survives contact with a real review. That is what putting people at the centre has come to mean for us in practice: building searches and tools where human judgment is hard to skip and harder to quietly outsource.
Developing the Calibrator also taught us something about ourselves. The more familiar we were with a review topic, the more comfortable we became with questioning AI’s suggestions. We rarely accepted terms simply because they sounded plausible. Instead, we kept asking whether they reflected the concept we wanted to capture and whether they really belonged in that search strategy.
That is a reminder that AI does not replace expertise. In many ways, it makes expertise even more important. The better we understood the topic, the easier it was to recognise missing concepts, unnecessary terms, or suggestions that looked convincing but did not answer our review question.
For us, keeping humans in the search is no longer a slogan. It means designing tools that make every important decision visible, discussable, and accountable.
Take-home messages
- AI is good at generating the language around a topic, but holding the conceptual boundaries of the question is still the reviewer’s job.
- The MCL Search Calibrator is worth using as it leaves a visible trail: suggestions, edits, exclusions, validations.
- An AI-assisted search is a process worth reporting, rather than just pasting the final string in the appendix.
Acknowledgement
The authors thank Professor João Apóstolo, Director of the JBI Portugal Centre for Evidence Based-Practice, for his support and endorsement of this submission.
References
Flemyng, E., Noel-Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., Jordan, Z., Minx, J., Eisele-Metzger, A., Hamel, C., Jemioło, P., Porritt, K., & Grainger, M. (2025). Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. JBI Evidence Synthesis, 23(11), 2162–2166. https://doi.org/10.11124/JBIES-25-00480
Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., Ayala, A. P., Moher, D., Page, M. J., Koffel, J. B., & PRISMA-S Group. (2021). PRISMA-S: An extension to the PRISMA statement for reporting literature searches in systematic reviews. Systematic Reviews, 10, Article 39. https://doi.org/10.1186/s13643-020-01542-z
To link to this article - DOI: https://doi.org/10.70253/NMZN4132
Links to additional resources
MCL Search Calibrator: https://boolean-calibrator.netlify.app/
Further reading
Bolaños, F., Salatino, A., Osborne, F., & Motta, E. (2024). Artificial intelligence for literature reviews: Opportunities and challenges. Artificial Intelligence Review, 57, Article 259. https://doi.org/10.1007/s10462-024-10902-3
Lieberum, J.-L., Töws, M., Metzendorf, M.-I., Heilmeyer, F., Siemens, W., Haverkamp, C., Böhringer, D., Meerpohl, J. J., & Eisele-Metzger, A. (2025). Large language models for conducting systematic reviews: On the rise, but not yet ready for use: A scoping review. Journal of Clinical Epidemiology, 181, Article 111746. https://doi.org/10.1016/j.jclinepi.2025.111746
O’Connor, A. M., Clark, J., Thomas, J., Spijker, R., Kusa, W., Walker, V. R., & Bond, M. (2024). Large language models, updates, and evaluation of automation tools for systematic reviews: A summary of significant discussions at the eighth meeting of the International Collaboration for the Automation of Systematic Reviews (ICASR). Systematic Reviews, 13, Article 290. https://doi.org/10.1186/s13643-024-02666-2
Wang, S., Scells, H., Koopman, B., & Zuccon, G. (2025). Reassessing large language model Boolean query generation for systematic reviews. arXiv. https://doi.org/10.48550/arXiv.2505.07155
Conflict of interest
The authors are involved in the development and/or use of the MCL Search Calibrator, which is mentioned in the blog as an illustrative example of auditable AI-assisted searching. The blog is not a product endorsement, and the authors declare no pharmaceutical, medical device industry, or commercial conflicts of interest.
Disclaimer
The views expressed in this World EBHC Day Blog, as well as any errors or omissions, are the sole responsibility of the author and do not represent the views of the World EBHC Day Steering Committee, Official Partners or Sponsors; nor does it imply endorsement by the aforementioned parties.