Abstract
Public trust in the justice system in South Africa is influenced by factors such as delays in the
issuance of court rulings, the length of court trials, inconsistent court outcomes (rulings) and
perceived inadequacies in legal aid service delivery, among others. This dissertation proposes
an integrated machine learning and data driven framework aimed at enhancing legal service
delivery in South Africa. By focusing on the prediction of legal outcomes, analysis of sentiment
trends, and identification of jurisdictional inconsistencies, the research combines statistical
methods and artificial intelligence techniques to address inefficiencies, improve public
engagement, and inform strategic policy development within South African justice system.
The first part of this dissertation investigates jurisdictional operations in sentencing outcomes
across different regions of South Africa. A statistical framework using Analysis of Variance
(ANOVA) was developed to evaluate the impact of geographical location, charge type, and case
duration on imprisonment terms. The results reveal significant regional differences in
sentencing, supported by an F-value of 2.347 and a critical value of 3.885 at a 5% confidence
level, with a p-value of 0.138 indicating moderate evidence of inconsistency. This study found
that the time taken to reach a court ruling did not influence the sentence length. These findings
provide a data-driven foundation for identifying systemic inconsistencies and can contribute to
supporting more justifiable and consistent legal outcomes.
The second part of this dissertation evaluates the effectiveness of machine learning algorithms
in predicting legal outcomes within the judicial system. Supervised learning models such as
Logistic Regression, Random Forest, and K-Nearest Neighbours (KNN) were applied to legal
case datasets sourced from a South African state law firm. After applying standard preprocessing
techniques such as tokenization and lemmatisation, model performance was
evaluated using accuracy metrics. Logistic Regression and Random Forest achieved similar
performance, with accuracies of 75.05% and 75.08% respectively, while the KNN algorithm
underperformed with an accuracy of 62.76%. These results demonstrate the practicality of
employing machine learning to predict legal case outcomes and inform judicial processes. The
research also highlights the ethical considerations of AI adoption in legal contexts and highlights
the limitations of using data from a single Law firm.