Predicting Diabetes Mellitus With Machine Learning Techniques
Top Cited Papers
Open Access
- 6 November 2018
- journal article
- research article
- Published by Frontiers Media SA in Frontiers in Genetics
- Vol. 9, 515
- https://doi.org/10.3389/fgene.2018.00515
Abstract
Diabetes mellitus is a chronic disease characterized by hyperglycemia. It may cause many complications. According to the growing morbidity in recent years, in 2040, the world’s diabetic patients will reach 642 million, which means that one of the ten adults in the future is suffering from diabetes. There is no doubt that this alarming figure needs great attention. With the rapid development of machine learning, machine learning has been applied to many aspects of medical health. In this study, we used decision tree, random forest and neural network to predict diabetes mellitus. The dataset is the hospital physical examination data in Luzhou, China. It contains 14 attributes. In this study, five-fold cross validation was used to examine the models. In order to verity the universal applicability of the methods, we chose some methods that have the better performance to conduct independent test experiments. We randomly selected 68994 healthy people and diabetic patients’ data, respectively as training set. Due to the data unbalance, we randomly extracted 5 times data. And the result is the average of these five experiments. In this study, we used principal component analysis (PCA) and minimum redundancy maximum relevance (mRMR) to reduce the dimensionality. The results showed that prediction with random forest could reach the highest accuracy (ACC = 0.8084) when all the attributes were used.This publication has 53 references indexed in Scilit:
- Diagnosis and Classification of Diabetes MellitusDiabetes Care, 2011
- An automatic diabetes diagnosis system based on LDA-Wavelet Support Vector Machine ClassifierExpert Systems with Applications, 2011
- Standards of Medical Care in Diabetes—2010Diabetes Care, 2010
- Tests for Screening and Diagnosis of Type 2 DiabetesClinical Diabetes, 2009
- An expert system approach based on principal component analysis and adaptive neuro-fuzzy inference system to diagnosis of diabetes diseaseDigital Signal Processing, 2007
- Image structure clustering for image quality verification of color retina images in diabetic retinopathy screeningMedical Image Analysis, 2006
- Random Forest: A Classification and Regression Tool for Compound Classification and QSAR ModelingJournal of Chemical Information and Computer Sciences, 2003
- Feature extraction and dimensionality reduction algorithms and their applications in vowel recognitionPattern Recognition, 2003
- Decision tree classification of land cover from remotely sensed dataRemote Sensing of Environment, 1997
- C4.5: Programs for Machine Learning by J. Ross Quinlan. Morgan Kaufmann Publishers, Inc., 1993Machine Learning, 1994