Logo image
Enhancing vehicle insurance claims fraud detection with a transparent, shap-weighted voting ensemble
Thesis   Open access

Enhancing vehicle insurance claims fraud detection with a transparent, shap-weighted voting ensemble

Nadia Charlene Erasmus
Master of Science (MSc), University of Johannesburg
2025
Handle:
https://hdl.handle.net/10210/520751

Abstract

Vehicle insurance fraud remains a persistent and economically damaging problem, with global losses estimated to exceed USD 40 billion annually. These losses underscore the urgent need for advanced detection frameworks that not only achieve high predictive accuracy but also comply with stringent regulatory and operational requirements. Traditional machine learning approaches, while widely adopted, often underperform in this domain due to two fundamental challenges: (i) severe class imbalance, which hampers the reliable detection of rare fraudulent cases, and (ii) the black-box nature of complex models, which limits interpretability and impedes adoption in highly regulated sectors such as insurance. To address these limitations, this study introduces a Shapley Additive exPlanations (SHAP)-weighted voting ensemble framework that integrates Logistic Regression and Random Forest classifiers. Class weighting is applied to mitigate imbalance, while SHAP values are leveraged to assign instance-specific weights to base models based on their explanatory confidence. This dual emphasis on predictive performance and model transparency enables the ensemble to achieve a balance between technical robustness and regulatory compliance. Empirical evaluation using real-world vehicle insurance claims data demonstrates that the proposed framework surpasses individual baseline models, achieving a validation AUC of 0.8113 and an F1-score of 0.6667, with strong generalization on unseen test data (AUC: 0.7162). Notably, the SHAP-weighted ensemble not only improves detection accuracy but also significantly reduces false positives, a crucial advantage in fraud analytics to minimize unwarranted claim disputes and preserve customer trust. By embedding explainability directly into the ensemble’s architecture, this research delivers a transparent, auditable, and regulatorily aligned solution consistent with frameworks such as the General Data Protection Regulation (GDPR). Overall, the proposed SHAP-weighted ensemble advances the frontier of trustworthy AI in vehicle insurance fraud detection and establishes a transferable methodological foundation for other high-stakes domains where fairness, accountability, and transparency are paramount.
pdf
Erasmus NC 225271631 amended1.68 MBDownloadView
Open Access

Metrics

1 Record Views

Details

Logo image