Abstract
Missing data is a prevalent issue that can significantly impair model performance and explainability. This paper briefly summarizes the development of the field of missing data with respect to Explainable Artificial Intelligence and experimentally investigates the effects of various imputation methods on SHapley Additive exPlanations (SHAP), a popular technique for explaining the output of complex machine learning models. Using benchmark datasets from the UCI Machine Learning Repository and the MNIST dataset with missing rates ranging from 20% to 80%, we compare different imputation strategies and assess their impact on feature importance and interaction as determined by Shapley values. Moreover, we also theoretically analyze the effects of missing values on Shapley values. Importantly, our findings reveal that the choice of imputation method can introduce biases that could lead to changes in the Shapley values, thereby affecting the explainability of the model. Additionally, we also show that a lower test prediction Mean Squared Error (MSE) does not necessarily imply a lower MSE in Shapley values and vice versa. Furthermore, while eXtreme Gradient Boosting (XGBoost) can directly handle missing data, it can substantially degrade explainability compared to imputing data beforehand. Overall, this study provides a comprehensive evaluation of imputation methods in the context of model explanations, offering practical guidance for selecting appropriate techniques based on dataset characteristics and analysis objectives.
Journal Title
Journal ISSN
Volume Title
Publisher
Elsevier
URL external
External URL
Date
Description
Keywords
Citation
Vo, T. L., Nguyen, T., Lopez-Ramos, L. M., Hammer, H. L., Riegler, M. A., & Halvorsen, P. (2026). Explainability of machine learning models under missing data. Applied Soft Computing, 115105.



