Logo image
Enhancing healthcare insurance fraud detection using explainable ensemble machine learning on imbalanced datasets
Thesis   Open access

Enhancing healthcare insurance fraud detection using explainable ensemble machine learning on imbalanced datasets

Given Manyike
Master of Science (MSc), University of Johannesburg
2025
Handle:
https://hdl.handle.net/10210/520724

Abstract

Fraudulent practices in healthcare insurance represent a substantial financial burden, with losses amounting to billions annually. Traditional rule-based detection systems have proven inadequate, as they struggle to adapt to the evolving nature of fraudulent behaviors. Conversely, advanced machine learning models, though highly accurate, often lack transparency, limiting their adoption in regulated healthcare environments where explainability is essential. This study addresses these challenges by developing an explainable stacked ensemble learning framework that achieves a balance between predictive accuracy and interpretability, even under conditions of severe class imbalance, where fraudulent claims comprise only 5% of the dataset. Using 16,000 healthcare reimbursement records, recursive feature elimination identified a refined subset of 20 highly discriminative features. A diverse pool of twelve base classifiers, including Extra Trees, Random Forest, XGBoost, and LightGBM was integrated through a logistic regression meta-classifier. SHapley Additive exPlanations (SHAP) were employed to provide dual-level interpretability, elucidating both the feature-level contributions and the influence of individual base models within the ensemble. To ensure robustness and fairness, model performance was comprehensively assessed using multiple evaluation metrics, including true positive rate (TPR), false positive rate (FPR), balanced accuracy, ROC-AUC, AUC-PR, Matthews correlation coefficient (MCC), precision, and F1-score. The findings demonstrated that the stacked ensemble significantly outperformed all single classifiers. The full stacking model achieved an ROC-AUC of 0.9041, AUC-PR of 0.6206, MCC of 0.5593, TPR of 0.4079, FPR of 0.0049, and a balanced accuracy of 0.7015. Subsequent SHAP-based model interpretability revealed four dominant contributors, extra trees, random forest, bagging, and LightGBM, which formed the basis of the Reduced-4 Ensemble. This lighter configuration matched or marginally exceeded the performance of the full model (ROC-AUC = 0.9051; AUC-PR = 0.6206; MCC = 0.5593; TPR = 0.4079; FPR = 0.0049; Balanced Accuracy = 0.7015) while offering improved computational efficiency. Feature-level SHAP analysis indicated that the primary fraud indicators were financial and consumption-related attributes, particularly total payment amounts and inpatient drug expenditures. In summary, the proposed framework delivers not only superior predictive accuracy but also transparent, auditable explanations at both the feature and model levels, supporting regulatory compliance and enhancing operational decision-making. The Reduced-4 Ensemble thus presents a promising, cost-efficient, and interpretable solution for future healthcare fraud detection systems.
pdf
Manyike G 2170272252.92 MBDownloadView
Open Access

Metrics

1 Record Views

Details

Logo image