Network meta-analysis of prediction models using aggregate or individual participant data - A scoping review and recommendations for reporting and conduct
Publication date
2025-12
Editors
Advisors
Supervisors
Document Type
Article
Metadata
Show full item recordCollections
License
cc_by
Abstract
Background Prediction models are essential in clinical decision-making for estimating the probability of current (diagnosis, screening) or future (prognosis) outcomes. Network meta-analysis (NMA) serves as a powerful tool to compare the performance of multiple prediction models simultaneously. However, there is hardly any guidance on methods and reporting for studies employing NMA to evaluate prediction models. Objective To provide an overview of NMAs assessing prediction model (external validation) performance, regardless of whether they use aggregate data (AD) or individual participant data (IPD). In addition, we offer recommendations for improving the reporting and conduct of NMAs in prediction model research. Methods We searched PubMed and Embase up to September 1, 2025, to identify studies that addressed the evaluation of diagnostic or prognostic prediction model performance using NMA. We included articles that employed NMA to compare and assess at least three prediction models. We summarized the identified studies based on, eg, their application (diagnostic vs prognostic), data use (AD vs IPD), medical contexts in which the models were assessed, and evaluation metrics applied (eg, discrimination, calibration, and (re)classification). In addition, we examined the statistical approaches employed, the NMA assumptions (such as consistency, transitivity, and exchangeability), and the ranking methods used for model comparison. Results After screening 2436 articles, 28 were included. Twenty-six studies (92.9%) used AD, while two (7.1%) used IPD. Hospital care was the most common setting (n = 22; 78.6%), with respirology (n = 7; 25.0%) and cardiology (n = 5; 17.9%) as the most frequently studied clinical domains. Key NMA assumptions were addressed differently across the 28 NMAs: 14.3% (n = 4) discussed transitivity, similarity, or exchangeability, and 53.6% (n = 15) tested for consistency. The statistical approach also varied, with 60.7% of studies (n = 17) reporting Bayesian methods and 17.9% (n = 5) reporting frequentist approaches. Surface under the cumulative ranking was the predominant ranking method (n = 18). Most NMAs included 5 to 10 models in the network, with 5 NMAs analyzing more than 20 models. Performance metrics varied, with 39.3% of studies (n = 11) reporting discrimination measures, such as C statistics, while none reported calibration metrics. Sensitivity or specificity was provided in 64.3% of studies (n = 18), and no articles reported advanced decision-analytic metrics like the decision curve analysis. Conclusion This scoping review highlights the limited and diverse use of NMA methods in evaluating prediction models, with a predominant reliance on aggregate rather than IPD, and inconsistent consideration of key NMA assumptions and model performance metrics. We provide recommendations for the reporting and conduct of an NMA of prediction model validation performance. Plain Language Summary Prediction models are tools used in medicine to estimate a patient's risk of developing a disease or experiencing a health outcome. Many different prediction models exist for the same condition, and it can be difficult for doctors and researchers to know which model performs best. One way to compare multiple models is through a statistical method called network meta-analysis (NMA), which is commonly used to compare treatments but has rarely been applied to prediction models. In our study, we reviewed all published NMAs that evaluated prediction models to see how they were conducted and reported. We looked at whether the studies reported important performance measures, such as how well the models could distinguish between patients with and without the outcome (discrimination), and how well the predictions matched actual outcomes (calibration). We also checked if key NMA assumptions were considered and how analyses were conducted. We found that most studies used summary data instead of patient-level data. Many did not report crucial performance measures and rarely checked NMA assumptions. There was also limited transparency in how the models were analyzed, making it difficult for others to reproduce the results. Our findings show that while NMA has great potential to help compare prediction models and identify the most reliable ones, current practice often lacks the detailed reporting needed to make these comparisons fully trustworthy. We recommend better reporting, sharing of data and analysis code, and careful checking of assumptions to help researchers and doctors choose the most reliable models, ultimately improving patient care.
Keywords
Individual participant data, Methodology, Network meta-analysi, Prediction models, Scoping review, Epidemiology, Journal Article
Citation
Yusufujiang, M, Damen, J A A, Idema, D L, Schuit, E, Moons, K G M & de Jong, V M T 2025, 'Network meta-analysis of prediction models using aggregate or individual participant data - A scoping review and recommendations for reporting and conduct', Journal of Clinical Epidemiology, vol. 188, 112006. https://doi.org/10.1016/j.jclinepi.2025.112006