Abstract

Missing data is a prevalent issue that can significantly impair model performance and explainability. This paper briefly summarizes the development of the field of missing data with respect to Explainable Artificial Intelligence and experimentally investigates the effects of various imputation methods on SHapley Additive exPlanations (SHAP), a popular technique for explaining the output of complex machine learning models. Using benchmark datasets from the UCI Machine Learning Repository and the MNIST dataset with missing rates ranging from 20% to 80%, we compare different imputation strategies and assess their impact on feature importance and interaction as determined by Shapley values. Moreover, we also theoretically analyze the effects of missing values on Shapley values. Importantly, our findings reveal that the choice of imputation method can introduce biases that could lead to changes in the Shapley values, thereby affecting the explainability of the model. Additionally, we also show that a lower test prediction Mean Squared Error (MSE) does not necessarily imply a lower MSE in Shapley values and vice versa. Furthermore, while eXtreme Gradient Boosting (XGBoost) can directly handle missing data, it can substantially degrade explainability compared to imputing data beforehand. Overall, this study provides a comprehensive evaluation of imputation methods in the context of model explanations, offering practical guidance for selecting appropriate techniques based on dataset characteristics and analysis objectives.
Loading...

Quotes

0 citations in WOS
0 citations in

Journal Title

Journal ISSN

Volume Title

Publisher

Elsevier

URL external

External URL

Description

Citation

Vo, T. L., Nguyen, T., Lopez-Ramos, L. M., Hammer, H. L., Riegler, M. A., & Halvorsen, P. (2026). Explainability of machine learning models under missing data. Applied Soft Computing, 115105.

Endorsement

Review

Supplemented By

Referenced By

Statistics

Views
3
Downloads
1

Bibliographic managers