BCSC Blog

Regularized Regression Performs Best for Predicting Advanced Breast Cancer Risk

A new Breast Cancer Surveillance Consortium (BCSC) study finds that complex machine learning methods do not improve prediction of advanced breast cancer compared with more interpretable statistical approaches.

Posted by Chen, Shuai, PhD at 9:37 AM on Aug 12, 2026

Share:


Machine learning methods are increasingly used to develop clinical risk prediction models because they can identify complex patterns and interactions in large datasets. However, more complex methods do not necessarily produce more accurate risk estimates, particularly when the outcome is rare and the number of available predictors is relatively modest.

A new study from the Breast Cancer Surveillance Consortium (BCSC), published in Cancer Epidemiology, Biomarkers & Prevention, compared statistical and machine learning methods for predicting advanced breast cancer. Advanced breast cancer was defined as prognostic pathologic stage II or higher, an outcome that is more strongly associated with breast cancer death than diagnosis of breast cancer overall.

The study included 968,178 women aged 40 to 74 years who underwent more than 3.6 million screening mammograms at BCSC facilities between 2005 and 2019. The models predicted advanced breast cancer within 12 months after an annual screening mammogram or within 24 months after a biennial screening mammogram. Researchers compared conventional logistic regression, two regularized regression approaches called LASSO and elastic net, and two machine learning approaches, random forests and gradient boosting.

All approaches performed similarly in distinguishing between women who did and did not develop advanced breast cancer, with area under the curve values ranging from 0.677 to 0.690. However, the models differed more substantially in calibration, which measures how closely predicted risks agree with the risks actually observed.

LASSO and elastic net provided the most favorable calibration while maintaining discrimination similar to the more complex machine learning approaches. Gradient boosting had comparable discrimination but less favorable calibration. Conventional logistic regression had slightly lower discrimination and somewhat less favorable calibration than the regularized regression approaches. Regression-based approaches were also generally well calibrated across racial and ethnic groups.

Calibration is especially important when risk estimates are used to guide decisions about screening. Even if a model successfully ranks women from lower to higher risk, systematic overestimation or underestimation of absolute risk could lead to inappropriate recommendations for more or less intensive screening.

These findings show that greater model complexity does not necessarily improve clinical risk prediction. When outcomes are rare and prediction models use a modest number of clinical and demographic factors, regularized regression may offer a practical balance of calibration, discrimination, interpretability, and ease of implementation.

Chen S, Kerlikowske K, Su YR, Hubbard RA, Gard CC, Tice JA, Sprague BL, Ahern TP, Miglioretti DL. Performance of Statistical and Machine Learning Risk Prediction Models for Advanced Breast Cancer. Cancer Epidemiol Biomarkers Prev. 2026;35(8):1375-1385. doi: 10.1158/1055-9965.EPI-25-2052. PMID: 42132486; PMCID: PMC13286512. [link]

The full article can be found here

American Association for Cancer Research Journals

By: Chen, Shuai PhD