Minimum sample size for developing a multivariable prediction model using multinomial logistic regression

Open Access

19 January 2023

journal article
research article
Published by SAGE Publications in Statistical Methods in Medical Research

Vol. 32 (3), 555-571
https://doi.org/10.1177/09622802231151220

Abstract

Aims Multinomial logistic regression models allow one to predict the risk of a categorical outcome with > 2 categories. When developing such a model, researchers should ensure the number of participants ( n ) is appropriate relative to the number of events ( E k ) and the number of predictor parameters ( p k ) for each category k. We propose three criteria to determine the minimum n required in light of existing criteria developed for binary outcomes.Proposed criteria The first criterion aims to minimise the model overfitting. The second aims to minimise the difference between the observed and adjusted R-2 Nagelkerke. The third criterion aims to ensure the overall risk is estimated precisely. For criterion (i), we show the sample size must be based on the anticipated Cox-snell R-2 of distinct 'one-to-one' logistic regression models corresponding to the sub-models of the multinomial logistic regression, rather than on the overall Cox-snell R(2 )of the multinomial logistic regression.Evaluation of criteria We tested the performance of the proposed criteria (i) through a simulation study and found that it resulted in the desired level of overfitting. Criterion (ii) and (iii) were natural extensions from previously proposed criteria for binary outcomes and did not require evaluation through simulation.We illustrated how to implement the sample size criteria through a worked example considering the development of a multinomial risk prediction model for tumour type when presented with an ovarian mass. Code is provided for the simulation and worked example. We will embed our proposed criteria within the pmsampsize R library and Stata modules.

Funding Information

Medical Research Council (MR/T025085/1)

This publication has 38 references indexed in Scilit:

Assessing the discriminative ability of risk models for more than two outcome categories
European Journal of Epidemiology, 2012
Interpreting the concordance statistic of a logistic regression model: relation to the variance and odds ratio of a continuous explanatory variable
BMC Medical Research Methodology, 2012
A clinical prediction model to assess the risk of operative delivery
BJOG: An International Journal of Obstetrics and Gynaecology, 2012
Polytomous diagnosis of ovarian tumors as benign, borderline, primary invasive or metastatic: development and validation of standard and kernel-based risk prediction models
BMC Medical Research Methodology, 2010
Assessing the Performance of Prediction Models
Epidemiology, 2010
Logistic Regression Model to Distinguish Between the Benign and Malignant Adnexal Mass Before Surgery: A Multicenter Study by the International Ovarian Tumor Analysis Group
Journal of Clinical Oncology, 2005
Sample size calculations for ordered categorical data
Statistics in Medicine, 1993
Linear Combinations of Multiple Diagnostic Markers
Journal of the American Statistical Association, 1993
On Simultaneous Confidence Intervals for Multinomial Proportions
Technometrics, 1965
Large Sample Simultaneous Confidence Intervals for Multinomial Proportions
Technometrics, 1964

Cited by 14 articles