Logo image
Integration of machine learning algorithms to predict HIV testing associations using repeated cross-sectional survey data in an adult South African population : an HIV testing predictive model
Dissertation   Open access

Integration of machine learning algorithms to predict HIV testing associations using repeated cross-sectional survey data in an adult South African population : an HIV testing predictive model

Musa Jaiteh
Doctor of Philosophy (PHD), University of Johannesburg
2025
Handle:
https://hdl.handle.net/10210/519433

Abstract

Background: Despite extensive efforts to expand Human Immunodeficiency Virus Acquired Immunodeficiency Syndrome (HIV/AIDS) interventions in South Africa, the country continues to experience a notable annual increase in new HIV cases. The HIV response has been more focused on HIV treatment rather than prevention, impeding progress toward the Joint United Nations Programme on HIV/AIDS (UNAIDS) 2030 goal of eradicating HIV and AIDS as a global health concern. Machine Learning (ML) employs computational algorithms to uncover hidden data correlations, potentially improving predictions by challenging traditional methods. The model identifies those with higher susceptibility to HIV, improves informed screening decisions, and streamlines testing and counselling. By harnessing the power of ML, this study aims to identify patterns and relationships among various socio-demographic, behavioural, and contextual variables that contribute to individuals' decisions to undergo HIV testing. Aim: To integrate ML algorithms for predicting HIV testing associations using repeated cross-sectional survey data in an adult South African population for strengthening HIV testing. Methods: The study employed mixed methods with five harmonised objectives to develop an evidence-based predictive model for strengthening HIV testing in South Africa. Bibliometric analysis: A bibliometric analysis was done to understand the trends, patterns, scholarly impact, and research opportunities on the applications of ML in HIV testing (objective 1). The extracted bibliographic data was exported into the R package (bibliometrix) for performance analysis and science mapping, and VOSviewer will be used for network visualisation. Systematic review: A systematic review was conducted to comprehensively understand the extent to which ML has been applied in HIV testing interventions, as well as the opportunities, successes, gaps, and challenges in their implementations globally (objective 2). This was accomplished by a literature review of published and grey literature using the Preferred Reporting Items for Systematic Reviews and Meta-analysis (PRISMA) guidelines. The review protocol was registered with the International Prospective Register of Systematic Reviews (ID: CRD42023464960). Predictive Modelling: A predictive modelling approach using four supervised ML algorithms was employed in conducting a retrospective analysis of the five cycles of the South African National HIV Prevalence, Incidence, Behaviour, and Communication Survey (SABSSM) data to predict factors associated with HIV testing (objective 3). Logistic regression, support vector machines, random forest, and DT were used. Each dataset from the five cycles of the SABSSM surveys was split into an 80% training sample and a 20% test sample with a 5-fold cross-validation technique. The models were evaluated using accuracy, precision, recall, F1-score, the area under the curve-receiver operating characteristic (AUC-ROC), and a confusion matrix. Additionally, cross-validation v averages were used to assess model consistency. In-depth interviews: In-depth interviews were conducted to explore the views and perceptions of 15 key stakeholders regarding the status, awareness, opportunities, challenges, contextual considerations, implementation roadmaps, and strategic recommendations for integrating ML in HIV testing in South Africa. The Consolidated Framework for Implementation Research (CFIR) was used to map the implementation roadmap for the results. The qualitative data were analysed using thematic content analysis. Evidence-based framework: An evidence-based framework was developed by consolidating all the key findings from objectives 1 to 4 to inform and strengthen HIV testing policies and programs in South Africa (objective 5). The Capability, Opportunity, Behavioural Model (COM-B), Logic Model, and Knowledge Translation framework guided the framework. Results: Bibliometric analysis: The bibliometric analysis included 266 articles to examine the trends, patterns, gaps, and scholarly impacts of the application of ML in HIV testing from 2000 to 2024. The analysis revealed a scientific annual growth rate of 15.68%, with an international co-authorship of 8.22% and an average citation of 17.47 per document. Most articles were conducted in developed countries within top-class universities and published in high-impact public health journals. The discrepancy highlights missed opportunities in strategic partnerships between developed and developing countries. The results further demonstrate that ML and health technologies enhance the effective and efficient implementation of innovative HIV testing methods, including HIV self-testing (HIVST) among priority populations. Systematic review: The systematic review involved screening 845 articles, of which 51 were eligible. More than 75% of the articles included in this review were conducted in the Americas and various parts of Sub-Saharan Africa, and a few were from Europe, Asia, and Australia. The most common algorithms applied were logistic regression, deep learning, support vector machine, random forest, extreme gradient booster, decision tree, and the least absolute shrinkage and selection operator model. The findings demonstrate that ML techniques exhibit higher accuracy in predicting HIV risk/testing compared to traditional approaches. Machine learning models enhance early prediction of HIV transmission, facilitate viable testing strategies to improve the efficiency of testing services, and optimise resource allocation, ultimately leading to improved HIV testing. This review points to the positive impact of ML in enhancing early prediction of HIV spread, optimising HIV testing approaches, improving efficiency, and eventually enhancing the accuracy of HIV diagnosis. Predictive modelling: The findings demonstrate that random forest outperformed the other models across all five SABSSM datasets, with the highest accuracy (81.0%), precision (81.6%), F1-score (80.3%), AUC (88.3%), and cross-validation average (79.1%) in the 2002 data. Random forest receives the highest classification across all the dates, especially in the vi 2017 survey. Support vector machine had a high recall (89.12% in 2005, 86.28% in 2008) but lower precision, leading to a suboptimal F1-score. Logistic regression performed well, but was slightly lower than in the random forest. Decision trees demonstrated moderate accuracy but were prone to overfitting. The topmost consistent predictors of HIV testing are knowledge of HIV testing sites, being a female, a younger adult, having a high socioeconomic status, and being well-informed about HIV through digital platforms. Random Forest is effective in developing data-driven policy initiatives focused on awareness, media engagement, employment, accessibility, and targeting high-risk individuals. In-depth interviews: The study highlights an opportunity to benefit from the capabilities of ML to improve the uptake of HIV testing through risk predictions, accurate and advanced diagnostics, and individualised testing through HIVST among youths. This shows ML is promising in reducing HIV transmission, addressing stigma, and optimising resource allocation. However, the study shows low utilisation of ML compounded by low awareness, policy gaps, significant ethical and structural barriers, as well as existing. Therefore, an equity-focused rollout should prioritise marginalised groups in designing ML interventions. The implementation design should include the perspectives of all the stakeholders involved in HIV testing to address human factors and ethical concerns. There is a need for robust governance to build trust in innovations like ML by providers and users of HIV testing. Conclusion: This study demonstrates that ML techniques are good options for predicting HIV testing using repeated adult population-based survey datasets. Random forest emerged as highly accurate in developing an HIV testing predictive model to inform an evidence-based framework. The diverse methodology, including a bibliometric analysis, systematic review, and in-depth interviews, gives a broader context to the current status quo, as well as gaps and opportunities in the application of ML in HIV testing. The identification of facilitators and barriers to HIV testing, ranging from sociocultural, socioeconomic, behavioural, knowledge and perception of HIV-related factors, will guide policymakers in improving HIV testing. The study suggests that leveraging ML enhances data-driven decisions and addresses testing gaps through policy initiatives focused on awareness, media engagement, employment, accessibility, and targeting high-risk individuals. The evidence-based framework sets a roadmap to the implementation of targeted HIV testing policies and programmes that can significantly upscale HIV testing to fully achieve the 95-95-95 goals intended to eradicate AIDS as a global epidemic by 2030.
pdf
Jaiteh M_Final Doctoral Thesis_31 July 20257.79 MBDownloadView
Open Access

Metrics

1 Record Views

Details

Logo image