Comprehensive Systematic and Scoping Review Searches and AI Tools
Author: Hailey Wills1,2
1. Maritime SPOR SUPPORT Unit
2. JBI Aligning Health and Evidence for Transformative Change
Comprehensive searches for review projects require a significant amount of time to create and refine. The steps involved include preliminary searching, developing a comprehensive search, having it peer-reviewed and making subsequent edits, translating the syntax and subject headings from one database to another, running the searches, downloading the results and documenting the search methods. As an Evidence Synthesis Coordinator, part of my role is to support review projects such as systematic or scoping reviews by developing search strategies to recover all available evidence on a particular topic.
I have a typical process that I go through when developing a comprehensive search, and there’s a question of whether artificial intelligence (AI) – and, in particular, generative AI (GenAI) tools – could be used to streamline this process to save time or improve the search quality, while maintaining transparency and reproducibility.
The first step I take when starting a systematic search for a new review is to quickly familiarise myself with the topic through Google searching, which also helps to gain an idea of the language people are using when they are writing on this topic, supporting me in brainstorming search terms. The topics are often health-related, and where I don’t have a background in health sciences, it necessitates working with review teams to set search terms that are relevant to the review. Google now has its AI Overview, which frequently appears at the top of your search results, and it may be useful to examine the sources the summary pulled from. The AI Overviews I have seen have sometimes pulled from innovative places, such as YouTube videos created, for instance, for conference presentations, retrieving keywords I otherwise wouldn’t have found in my initial search of webpages. I also find it helpful to visit sources such as Wikipedia pages and thesaurus websites and read background papers on the topic.
One of the goals of my comprehensive searches is for the terms to be evidence-based, so it’s important that there is human oversight in the collection of keywords. In developing the search, I complete extensive testing on different keywords I’m including, to ensure that the results they are retrieving are indeed relevant to the topic and not bringing in too much irrelevant literature. While terms that AI tools can collect may be relevant, it’s important for them to be tested along with other terms collected, or run past a subject matter expert if their meaning or relevance is unclear.
Another part of my process is to look for existing reviews on the same population or intervention as the research question at hand. If there are existing reviews that have a publicly available, comprehensive search, I may adapt the search strings to use in my search and acknowledge this in the final manuscript. There are also existing search filters that are tested, and sometimes there are validated strings of keywords and subject headings for things such as a specific population, geographical location or study design. If I were to use GenAI tools to gather keywords as anything other than a supplementary strategy, how would I know that the keywords and subject headings the model suggested were not directly lifted or adapted from a previously published review?
Another use for AI might be a review of a completed, comprehensive search, to make sure no key terms have been missed or to identify any errors in the syntax or logic of the search. After looking at a search for a long time, you can end up missing obvious spelling mistakes or miss out by not truncating a term correctly and leaving out a key variation on a term. A second set of eyes on the search is an important step in the process. This is already built into the process for information specialists with the Peer Review of Electronic Search Strategies (PRESS), a validated peer-review tool where another expert searcher reviews the search in order to minimise errors and increase its quality. Using an AI tool as an extra step in double-checking your work could be useful, but it does not replace the experience of an expert searcher.
Some aspects of the search process could benefit from streamlining. For instance, when translating the search strategy from one database to another, it takes considerable time to change the syntax of the search to meet the requirements of each database and find equivalent controlled vocabulary. GenAI might be helpful in translating a string of keywords and subject headings, but taking the time to double-check the output would likely equal the time you would have spent translating the strategy yourself. Hallucinations, where the AI model generates false information, are still a common problem with tools, so each subject heading would need to be confirmed in each database, and the databases would also need to be searched to ensure that no subject headings were missed.
While some tools could be helpful in search development or refinement, they are certainly not a replacement for human expertise. While there may be some aspects of the search development process where AI could be used as a supporting tool, it could not be relied upon to replace any of the steps that researchers or information specialists currently carry out without jeopardising the quality of a review. When AI tools are used, it is important that they are reported for transparency and that the search is still reproducible and rigorous.
Key Messages
- We need to keep the human element in developing comprehensive searches for reviews
- We may be able to use some tools in supporting the searching process, but we need human oversight to ensure the integrity of the search
Acknowledgements
Many thanks to Dr. Marilyn Macdonald for suggesting the topic for this blog post and providing edits and feedback.
References
Lieberum, J. L., Toews, M., Metzendorf, M. I., Heilmeyer, F., Siemens, W., Haverkamp, C., Böhringer, D., Meerpohl, J. J., & Eisele-Metzger, A. (2025). Large language models for conducting systematic reviews: On the rise, but not yet ready for use – A scoping review. Journal of Clinical Epidemiology, 181, 111746. https://doi.org/10.1016/j.jclinepi.2025.111746
McGowan, J., Sampson, M., Salzwedel, D. M., Cogo, E., Foerster, V., & Lefebvre, C. (2016). PRESS Peer Review of Electronic Search Strategies: 2015 guideline statement. Journal of Clinical Epidemiology, 75, 40–46. https://doi.org/10.1016/j.jclinepi.2016.01.021
University of Alberta Library. (n.d.). Search filters. https://guides.library.ualberta.ca/systematic-reviews/search-filters
To link to this article - DOI: https://doi.org/10.70253/JNAD3298
Disclaimer
The views expressed in this World EBHC Day Blog, as well as any errors or omissions, are the sole responsibility of the author and do not represent the views of the World EBHC Day Steering Committee, Official Partners or Sponsors; nor does it imply endorsement by the aforementioned parties.