Abstract
Background: Diffuse Large B-cell Lymphoma (DLBCL) is a heterogeneous and aggressive malignancy requiring the precise extraction of immunohistochemical (IHC) markers and cell-of-origin (COO) signatures for effective management. In public-sector laboratories, manual data extraction from narrative pathology reports is resource-intensive and difficult to scale, creating a barrier to large-scale research and surveillance.
Aim: This study aimed to design and evaluate a domain-specific Natural Language Processing (NLP) pipeline capable of automating the extraction of predictive biomarkers from unstructured DLBCL reports.
Methods: A dataset of 299 anonymised reports from a South African academic hospital was manually curated to establish a gold-standard reference. A rule-based Python NLP system was developed within the Orange3 platform to extract demographics, nodality, and key IHC markers (CD20, CD3, CD10, BCL6, MUM1, and Ki-67). To assess the internal coherence of the resulting structured dataset, four supervised machine-learning classifiers: Random Forest (RF), Support Vector Machine (SVM), Logistic Regression, and Naïve Bayes were evaluated using 10-fold cross-validation.
Results: The NLP pipeline demonstrated exceptional extraction fidelity, achieving accuracies of 99.33% for CD20, 97.66% for CD10, and 99.67% for the Hans classification. Among the machine-learning models, Random Forest emerged as the most robust, achieving an accuracy of 0.928 and an Area Under the Curve (AUC) of 0.95. Feature-importance analysis confirmed that biologically relevant markers, specifically CD10, BCL6, and MUM1, were the primary drivers of model performance.
Conclusion: These findings confirm that rule-based NLP can reliably transform unstructured clinical narratives into high-quality structured data. This approach provides a scalable solution for AI-assisted data abstraction, offering the potential to enhance diagnostic efficiency and epidemiological research in resource-limited healthcare environments