Logo image
Evaluating the efficacy of natural language processing to extract prediction data in diffuse large B-cell lymphoma pathology reports
Thesis   Open access

Evaluating the efficacy of natural language processing to extract prediction data in diffuse large B-cell lymphoma pathology reports

Nqobile Ndlazi
Master of Health Sciences , University of Johannesburg
2026
Handle:
https://hdl.handle.net/10210/520529

Abstract

Background: Diffuse Large B-cell Lymphoma (DLBCL) is a heterogeneous and aggressive malignancy requiring the precise extraction of immunohistochemical (IHC) markers and cell-of-origin (COO) signatures for effective management. In public-sector laboratories, manual data extraction from narrative pathology reports is resource-intensive and difficult to scale, creating a barrier to large-scale research and surveillance. Aim: This study aimed to design and evaluate a domain-specific Natural Language Processing (NLP) pipeline capable of automating the extraction of predictive biomarkers from unstructured DLBCL reports. Methods: A dataset of 299 anonymised reports from a South African academic hospital was manually curated to establish a gold-standard reference. A rule-based Python NLP system was developed within the Orange3 platform to extract demographics, nodality, and key IHC markers (CD20, CD3, CD10, BCL6, MUM1, and Ki-67). To assess the internal coherence of the resulting structured dataset, four supervised machine-learning classifiers: Random Forest (RF), Support Vector Machine (SVM), Logistic Regression, and Naïve Bayes were evaluated using 10-fold cross-validation. Results: The NLP pipeline demonstrated exceptional extraction fidelity, achieving accuracies of 99.33% for CD20, 97.66% for CD10, and 99.67% for the Hans classification. Among the machine-learning models, Random Forest emerged as the most robust, achieving an accuracy of 0.928 and an Area Under the Curve (AUC) of 0.95. Feature-importance analysis confirmed that biologically relevant markers, specifically CD10, BCL6, and MUM1, were the primary drivers of model performance. Conclusion: These findings confirm that rule-based NLP can reliably transform unstructured clinical narratives into high-quality structured data. This approach provides a scalable solution for AI-assisted data abstraction, offering the potential to enhance diagnostic efficiency and epidemiological research in resource-limited healthcare environments
pdf
N.Ndlazi Dissertation 13 April 2026 FINAL JM5.45 MBDownloadView
Open Access

Metrics

1 Record Views

Details

Logo image