Classifying lymphoma and tuberculosis case reports using machine learning algorithms

Open Access

1 October 2021

journal article
Published by Institute of Advanced Engineering and Science in Bulletin of Electrical Engineering and Informatics

Vol. 10 (5), 2857-2865
https://doi.org/10.11591/eei.v10i5.3132

Abstract

Available literature reports several lymphoma cases misdiagnosed as tuberculosis, especially in countries with a heavy TB burden. This frequent misdiagnosis is due to the fact that the two diseases can present with similar symptoms. The present study therefore aims to analyse and explore TB as well as lymphoma case reports using Natural Language Processing tools and evaluate the use of machine learning to differentiate between the two diseases. As a starting point in the study, case reports were collected for each disease using web scraping. Natural language processing tools and text clustering were then used to explore the created dataset. Finally, six machine learning algorithms were trained and tested on the collected data, which contained 765 lymphoma and 546 tuberculosis case reports. Each method was evaluated using various performance metrics. The results indicated that the multi-layer perceptron model achieved the best accuracy (93.1%), recall (91.9%) and precision score (93.7%), thus outperforming other algorithms in terms of correctly classifying the different case reports.

Classifying lymphoma and tuberculosis case reports using machine learning algorithms

Abstract

Keywords