Optimization of Diabetes Mellitus Classification Using the Random Forest and SMOTE-ENN Methods
DOI:
https://doi.org/10.30871/jaic.v10i4.11387Keywords:
Classification, Diabetes Mellitus, Machine Learning, Random Forest, SMOTE-ENNAbstract
Diabetes Mellitus is a non-communicable disease (NCD) that has turned into a worldwide health issue with a steadily rising prevalence. Timely identification is essential for minimizing the risk of complications and the financial strain on the healthcare system. This research focuses on creating a precise and dependable diabetes classification model through the Random Forest algorithm by implementing a series of systematic data preprocessing methods. This methodology utilizes a dataset obtained from Kaggle, consisting of 768 samples. The steps taken include addressing missing values using Multiple Imputation by Chained Equations (MICE) , removing outliers with the Z-Score and Interquartile Range (IQR) techniques , selecting features based on ANOVA F-value to identify the eight most significant features , and balancing classes through the Synthetic Minority Over-sampling Technique-Edited Nearest Neighbours (SMOTE-ENN) to correct dataset imbalance. The assessment of the Random Forest model revealed outstanding performance, attaining an accuracy of 94.3% and an Area Under Curve (AUC) value of 0.98. These findings suggest that the model possesses strong discriminative capability to differentiate between diabetic and non-diabetic individuals. This research concludes that the Random Forest algorithm, when backed by suitable data preprocessing, is very efficient and could be utilized in clinical decision support systems as an early tool for diabetes screening.
Downloads
References
[1] N. N. Rosyidah and E. A. Cahyono, “Diabetes Melitus Tipe 2,” Enfermeria Ciencia, vol. 3, no. 1, pp. 44–63, Feb. 2025, doi: 10.56586/ec.v3i1.74.
[2] W. Ode Nurmila, A. Edy Dawu, P. S. Studi, and I. Teknologi adan KesehatanaAvicenna, “Hubungan Dukungan Keluarga, Pengetahuan dan Sikap Dengan Tingkat Kepatuhan Pasien dalam Pengobatan Diabetes Melitus di Wilayah Kerja Puskesmas Poasia Kota Kendari Tahun 2024,” 2025. doi: https://doi.org/10.69677/avicenna.v4i2.152.
[3] N. Hikmah and P. Yuwono, “Hubungan Efikasi Diri Terhadap Kualitas Hidup Pasien Diabetes Melitus Tipe 2 Di RS PKU Muhammadiyah Gombong,” Jurnal Ilmiah Kesehatan Keperawatan, vol. 21, no. 1, p. 6, Jul. 2025, doi: 10.26753/jikk.v21i1.1449.
[4] E. Elsa, “Hubungan Kepatuhan Diet Terhadap Kadar Glukosa Pada Penderita Diabetes Melitus Tipe II di Wilayah Kerja UPTD Puskesmas Selajambe Tahun 2025,” Abdimas Awang Long, vol. 8, no. 2, pp. 230–237, Jun. 2025, doi: 10.56301/awal.v8i2.1684.
[5] M. R. Maulana, A. Sucipto, and H. Mulyo, “Optimisasi Parameter Support Vector Machine Dengan Particle Swarm Optimization Untuk Peningkatan Klasifikasi Diabetes,” Nov. 2024.
[6] E. R. Subhiyakto et al., “Evaluation of Resampling Techniques in CNN-Based Heartbeat Classification,” Ingénierie des systèmes d information, vol. 29, no. 4, pp. 1323–1332, Aug. 2024, doi: 10.18280/isi.290408.
[7] M. Salsabil, N. Lutvi, and A. Eviyanti, “Implementasi Data Mining Dalam Melakukan Prediksi Penyakit Diabetes Menggunakan Metode Random Forest Dan Xgboost,” Jurnal Ilmiah Komputasi, vol. 23, no. 1, Mar. 2024, doi: 10.32409/jikstik.23.1.3507.
[8] T. Hidayat, S. S. Anelia, R. I. Pratiwi, N. Salsabila, and D. S. Prasvita, Perbandingan Akurasi Klasifikasi Penyakit Diabetes Menggunakan Algoritma Adaboost-Random Forest Dan Adaboost-Decision Tree Dengan Imputasi Median Dan KNN. 2021. [Online]. Available: https://www.kaggle.com/uciml/pima-indians-diabetes-database
[9] E. Cahya, P. Witjaksana, R. Rohmat Saedudin, and V. P. Widartha, “Perbandingan Akurasi Algoritma Random Forest Dan Algoritma Artificial Neural Network Untuk Klasifikasi Penyakit Diabetes,” Oct. 2021.
[10] W. Nugraha, “Resampling Technique for Handling Class Imbalance in the Classification of Diabetes using C4.5, Random Forest, and SVM,” Aug. 2021, doi: 10.33633/tc.v20i3.
[11] L. Maretva Cendani and A. Wibowo, “Perbandingan Metode Ensemble Learning pada Klasifikasi Penyakit Diabetes,” 2022. doi: https://doi.org/10.14710/jmasif.13.1.42912.
[12] S. P. Nainggolan and A. Sinaga, “Comparative Analysis Of Accuracy Of Random Forest And Gradient Boosting Classifier Algorithm For Diabetes Classification,” Sebatik, vol. 27, no. 1, pp. 97–102, Jun. 2023, doi: 10.46984/sebatik.v27i1.2157.
[13] A. B. Alpiansah and Y. Ramdhani, “Optimasi Fitur dengan Forward Selection pada Estimasi Tingkat Obesitas menggunakan Random Forest,” SISTEMASI, vol. 12, no. 3, p. 860, Sep. 2023, doi: 10.32520/stmsi.v12i3.3125.
[14] Y. Tasya, “Kombinasi Metode Imputasi Mean Dan Multiple Imputation By Chained Equations (Mice) Untuk Penanganan Data Hilang Dan Peningkatan Evaluasi Kinerja Klasifikasi Prediksi Penyakit Diabetes Melitus,” 2023.
[15] F. Kurniawan, “Analisis Pengaruh Seleksi Fitur Anova Terhadap Performa Model Klasifikasi Gaussian Naïve Bayes Pada Dataset Pima Indians Diabetes,” 2023.
[16] N. Nur Muttaqin, “Klasifikasi Penyakit Diabetes Menggunakan Metode Random Forest Dan Adaboost,” 2024.
[17] C. Paramita, C. S. Simbolon, A. S. Pamungkas, J. M. Triono, E. P. Widi Utomo, and E. R. Subhiyakto, “Analisis Pengaruh SMOTE terhadap Kinerja Model KNN untuk Prediksi Risiko Stroke,” Jurnal Informatika: Jurnal Pengembangan IT, vol. 10, no. 4, pp. 978–988, Sep. 2025, doi: 10.30591/jpit.v10i4.8809.
[18] M. A.-Z. Faradeya and E. R. Subhiyakto, “Klasifikasi Penyakit Gagal Jantung Menggunakan Algoritma Naive Bayes,” Jurnal Algoritma, vol. 22, no. 1, pp. 115–127, May 2025, doi: 10.33364/algoritma/v.22-1.2178.
[19] M. Fadli and R. A. Saputra, “Klasifikasi Dan Evaluasi Performa Model Random Forest Untuk Prediksi Stroke Classification And Evaluation Of Performance Models Random Forest For Stroke Prediction,” vol. 12, Oct. 2023, doi: http://dx.doi.org/10.31000/jt.v12i2.9099.
[20] N. H. Alfajr and S. Defiyanti, “Prediksi Penyakit Jantung Menggunakan Metode Random Forest Dan Penerapan Principal Component Analysis (PCA),” Jurnal Informatika dan Teknik Elektro Terapan, vol. 12, no. 3S1, Oct. 2024, doi: 10.23960/jitet.v12i3S1.5055.
[21] A. Latif and S. Khotimatul Wildah, “Analisis Kinerja Algoritma Ensemble Dalam Prediksi Perilaku Pembelian Pelanggan,” JATI (Jurnal Mahasiswa Teknik Informatika), vol. 9, no. 1, pp. 557–563, Dec. 2024, doi: 10.36040/jati.v9i1.12413.
[22] S. Ernawati and I. Maulana, “Meningkatkan Klasifikasi Penyakit Diabetes Menggunakan Metode Ensemble Softvoting Dengan SMOTE-ENN dan Optimasi Bayesian,” Evolusi : Jurnal Sains dan Manajemen, vol. 13, no. 1, pp. 71–86, Mar. 2025, doi: 10.31294/evolusi.v13i1.8267.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Imam Fadhur Rahman, Egia Rosi Subhiyakto, Cinantya Paramita

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).



