Publication Year:
Author(s): Mahmad Isaq Karankot, Bradley M. Whitaker
Associated Research Thrusts(s): Artificial Intelligence & Machine Learning
Project Year: Year 3
Abstract
In critical care, accurately predicting length of stay (LOS) for intensive care unit (ICU) patients is paramount for optimizing patient flow, resource allocation, and improving clinical outcomes. This study explores the impact of missing data imputation methods on machine learning models for predicting LOS for ICU patients. We evaluate several methods, using RMSE to measure the effectiveness of the models and ∆RMSE to measure overfitting. Our analysis reveals that some algorithms can achieve robust performance with minimal overfitting even with extensive missing data. However, introducing missing data did not reduce overfitting for any of the intentionally overfit models, and Validation RMSE was marginally worse across all incomplete datasets. The one exception was an underfitting model: a shallow neural network showed reduced underfitting when trained on imputed data, with ∆RMSE moving toward zero — suggesting imputation can act as a corrective regularizer specifically for underfitting architectures, rather than overfitting ones.
Citation
Karankot, M. I., & Whitaker, B. M. (2026). Impact of data imputation on overfitting in ICU length of stay predictions. In 2026 Intermountain Engineering, Technology and Computing (IETC) (pp. 1–5). IEEE. https://doi.org/10.1109/IETC69527.2026.11568670
