The Umbrella Collaboration®: Quantitative, Living, People-Centred Tertiary Evidence Synthesis
Author: Beltran Carrillo1,2
1. The Umbrella Collaboration®
2. Clinica Beltran Carrillo
The Umbrella Collaboration® did not start as an artificial intelligence (AI) project or as an academic exercise to invent a new method. It started with a practical problem.
In our work exploring the scientific evidence for acupuncture, we repeatedly came across the claim that there was little or no evidence supporting it. So, we started looking for systematic reviews and meta-analyses (SR/MAs). And we found many. But more papers did not necessarily mean a greater understanding of what the evidence was saying. Different SR/MAs could use different comparators, outcomes or statistical measures, and sometimes reach different conclusions. Then came an uncomfortable question: which review should we show as our scientific evidence for acupuncture? The most recent? The largest? The highest-quality? Choosing one could look like cherry-picking, when what we really wanted was to know what the whole published body of evidence was saying.
We are clinicians, not evidence synthesis methodologists. The debate around the evidence for acupuncture pushed us to look much more closely at what ‘evidence’ actually means. We began with SR/MAs, then moved towards umbrella reviews and tertiary syntheses.
The Umbrella Collaboration® grew from this journey, as a small, independent multidisciplinary project. The methodology was conceived by Beltran Carrillo, a physician specialising in geriatrics and medical acupuncture, and Marta Rubinos, a nurse, both primarily involved in clinical practice. We then worked with Jazmin Parellada, a computer engineer; Beltran Carrillo Jr and Alejandra Palacios, biomedical engineers; Andres Carrillo, a telecommunications engineer; and Ana Carrillo, an economist, to turn the idea into working software.
Over time, this became The Umbrella Collaboration®: a quantitative and living tertiary evidence synthesis methodology, operational software (The Umbrella) and a web platform where projects can be created and explored interactively. The system is operational and has undergone initial validation, although further independent evaluation is needed.
Umbrella reviews and overviews of reviews provide established methods for organising and analysing bodies of systematic review evidence. They can include quantitative re-analysis, critical appraisal, assessment of overlap and investigation of heterogeneity or disagreement between reviews. Furthermore, tools such as metaumbrella can support quantitative analyses by recalculating meta-analytic results from primary study data within a common analytical framework.
Yet, in the quest for evidence for acupuncture, our need was more specific. We wanted to retain the published SR/MA estimate itself as the quantitative unit of tertiary synthesis and, for each outcome, obtain a single summary of what those review-level estimates collectively reported.
This is complementary to, rather than a replacement for, conventional umbrella review methods. Understanding why reviews disagree remains important; aggregating their published estimates answers a different question while keeping the individual review-level results visible for exploration.
How does The Umbrella work?
This is where software, automation and AI become practical (see Figure 1).
- The human reviewer creates the project and defines, in natural language, what they want to know
- A large language model, using our Hierarchical Semantic Expansion approach, turns that intent into a high-sensitivity PubMed query. AI helps build the search; it does not decide which evidence to include or interpret clinical findings
- The Umbrella searches PubMed and identifies potentially eligible SR/MAs. The synthesis uses information reported in their abstracts, which makes the process scalable but means that results not reported there cannot contribute
- The human reviewer decides which SR/MAs meet the inclusion criteria. Domain knowledge matters because the same clinical construct may appear under different names, abbreviations or scales
- Effect estimates are mapped to the common RTU scale, and review-level estimates reporting the same outcome are quantitatively aggregated. Results are displayed through an interactive visual layer, with access to the supporting SR/MAs and sources
- Every 24 hours, the system searches again. New candidate SR/MAs remain pending until human review before they can modify the public synthesis
The Umbrella Collaboration®, therefore, combines methodology, software, automation, a defined use of an LLM and human expertise.

Figure 1 The Umbrella Workflow
RTU: a purpose-built tertiary metric
Different SR/MAs report effects using different statistical measures. Because these measures are not comparable, we do not average them directly. We first convert each published estimate to a common scale that keeps its direction and rough size, then combine those converted values. This purpose-built tertiary metric is RTU, or Result The Umbrella.
RTU asks: What is the published body of SR/MAs collectively reporting about this outcome? Because the same primary studies may appear in more than one SR/MA, review-level estimates are not fully independent, and this overlap must be considered when interpreting the result.
In our initial validation against eight traditional umbrella reviews, The Umbrella identified 410 results of interest compared with 86 in the original reviews and reproduced 73 of those 86. We see this as encouraging initial validation, not proof of methodological superiority. The mean execution time across the eight projects was 4 hours 46 minutes, from entering the initial search terms to completing the first version of the project.
A synthesis that stays up to date
There was another practical problem from the beginning: time. In fast-moving fields, a synthesis can become outdated quickly.
A project in The Umbrella does not finish when its first version is completed. The software repeats the search every 24 hours, while new records remain pending until the reviewer confirms eligibility. The reviewer can also update on demand, and the platform displays the date of the latest human review.
There are, therefore, two clocks: automated surveillance, and scientific updating after human validation.
People at the centre: evidence must be understandable
For us, ‘people at the centre’ means both the people who build the synthesis and those who need to use it. Researchers, healthcare professionals, managers, policymakers, patients and caregivers all face a fundamental difficulty: scientific evidence can be hard to navigate and interpret.
The platform organises complexity into layers. The first is visual and intuitive, giving a quick view of the direction and approximate magnitude of an effect, the certainty and the amount of evidence behind it. Those who need more information can go deeper to the supporting SR/MAs. Researchers can also use the overview to identify discordant findings, low-certainty areas or evidence gaps.
The scientific complexity is still there. What we try to make easier is the way into it. Making publications available is not enough; evidence also has to be understandable and usable by the people involved in making healthcare decisions.
Human expertise and accountability
The human reviewer remains central to every project: defining the question, understanding the field, deciding which evidence is eligible and grouping outcomes that may appear under different terminology. AI can broaden the search, and software can automate calculations, presentation and surveillance, but domain expertise remains essential.
This project is not a matter of human versus technology. Instead, it involves both, and we are clear about what each does and who is responsible for the final synthesis. Anyone consulting a project should be able to know who reviewed it, when, what evidence supports the result and, whenever possible, which organisation is responsible. That is what accountability means to us.
From publications to continuously synthesised knowledge
Our longer-term goal is simple: when a clinical question or uncertainty arises, a healthcare professional, researcher, manager, patient or caregiver should be able to rapidly access the highest available level of synthesised evidence relevant to their need. Furthermore, they should easily be able to determine what the evidence suggests, what supports the result, how certain it is, where it comes from and when it was last reviewed.
The Umbrella Collaboration® is already an operational methodology, software system and web platform, but still needs further validation and independent evaluation. Even so, it is supporting a change that we believe matters: moving from searching for individual publications towards consulting continuously synthesised knowledge when healthcare decisions need to be made.
Take-home messages
1. Tertiary synthesis can add a quantitative layer to review-level evidence: The Umbrella Collaboration® aggregates published SR/MA estimates for the same outcome while retaining those estimates as the unit of quantitative analysis
2. Evidence synthesis can stay up to date: The Umbrella Collaboration® projects are automatically re-searched every 24 hours, while changes to the public synthesis require human validation
3. Access alone is not enough: evidence needs to be understandable, traceable and usable by researchers, healthcare professionals, decision-makers, patients and caregivers
References
Aromataris, E., Fernandez, R., Godfrey, C. M., Holly, C., Khalil, H., & Tungpunkom, P. (2015). Summarizing systematic reviews: Methodological development, conduct and reporting of an umbrella review approach. JBI Evidence Implementation, 13(3), 132. https://doi.org/10.1097/XEB.0000000000000055
Biondi-Zoccai, G., editor. (2016). Umbrella reviews: Evidence synthesis with overviews of reviews and meta-epidemiologic studies. Springer International Publishing. https://doi.org/10.1007/978-3-319-25655-9
Gosling, C. J., Solanes, A., Fusar-Poli, P., & Radua, J. (2023). metaumbrella: The first comprehensive suite to perform data analysis in umbrella reviews with stratification of the evidence. BMJ Mental Health, 26(1), e300534. https://doi.org/10.1136/bmjment-2022-300534
Lunny, C., Pieper, D., Thabet, P., & Kanji, S. (2021). Managing overlap of primary study results across systematic reviews: Practical considerations for authors of overviews of reviews. BMC Medical Research Methodology, 21(1), 140. https://doi.org/10.1186/s12874-021-01269-y
Pollock, M., Fernandes, R. M., Becker, L. A., Pieper, D., & Hartling, L. (2023). Chapter V: Overviews of reviews. In: Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., Welch, V., & Flemyng, E. (Eds.), Cochrane handbook for systematic reviews of interventions, version 6.4. Cochrane. https://training.cochrane.org/handbook/current/chapter-v
To link to this article - DOI: https://doi.org/10.70253/PFKN2970
Disclaimer
The views expressed in this World EBHC Day Blog, as well as any errors or omissions, are the sole responsibility of the author and do not represent the views of the World EBHC Day Steering Committee, Official Partners or Sponsors; nor does it imply endorsement by the aforementioned parties.