Comparison of Feature Selection Methods in Classifying Poverty Levels in Indonesia Using Comparative Machine Learning Methods
DOI:
https://doi.org/10.30871/jaic.v10i3.12604Keywords:
Chi-Square, Feature Selection, Normalization, Poverty Classification, Random ForestAbstract
Poverty classification requires models capable of handling multidimensional data and imbalanced class distributions. This study aims to develop and compare several machine learning algorithms for classifying poverty levels in Indonesia, as well as to analyze the impact of feature selection and reduction methods on model performance. The study employs a comparative approach using a secondary dataset consisting of 514 districts/cities with socio-economic indicators and a binary target variable. The methodology includes data preprocessing, the application of Chi-Square, Pearson Correlation, and Principal Component Analysis (PCA), and the handling of imbalanced data using the Synthetic Minority Oversampling Technique (SMOTE). Modelling is conducted using Random Forest, Support Vector Machine (SVM), Logistic Regression, and Artificial Neural Network (ANN), with evaluation performed using Stratified K-Fold Cross Validation and metrics including accuracy, precision, recall, and F1-score. The results indicate that Chi-Square and Pearson Correlation outperform PCA, with Random Forest achieving the best performance, attaining an accuracy of 0.9854 and an F1-score of 0.9507, while effectively detecting the minority class. Therefore, the combination of Chi-Square and Random Forest is identified as the most effective approach in this study, as it produces a model that is accurate, stable, and capable of handling imbalanced data.
Downloads
References
[1] World Bank, “the World Bank Has Set a Clear Mission : Ending Extreme Poverty and Boosting Shared,” 2024.
[2] Badan Pusat Statistik (BPS), “Profil Kemiskinan di Indonesia Maret 2024,” Present. Pendud. miskin Maret 2010-2024, no. 50, 2024.
[3] Badan Pusat Statistik, Indonesian Sustainable Development Goals Indicators 2024 - BPS-Statistics Indonesia. 2024. Accessed: Dec. 19, 2025. [Online]. Available: https://www.bps.go.id/en/publication/2024/12/31/936a26d5d2b168b9971d3b02/indonesian-sustainable-development-goals-indicators-2024.html
[4] UNDP, “Unstacking Global Poverty: Data for high impact action,” Glob. Multi-dimensional Poverty Index 2023, pp. 1–2, 2023, [Online]. Available: https://hdr.undp.org/content/2023-global-multidimensional-poverty-index-mpi#/indicies/MPI
[5] UNDP & OPHI, “2023 Global Multidimensional Poverty Index (MPI) | Human Development Reports,” Jul. 2023. Accessed: Dec. 18, 2025. [Online]. Available: https://hdr.undp.org/content/2023-global-multidimensional-poverty-index-mpi
[6] S. Alkire et al., “The global Multidimensional Poverty Index ( MPI ) 2023 disaggregation results and methodological note Sabina Alkire *, Usha Kanagaratnam **, Nicolai Suppa ***,” no. July, 2023.
[7] S. Alkire, U. Kanagaratnam, and N. Suppa, “The global Multidimensional Poverty Index,” no. October, pp. 2–17, 2023, doi: 10.18356/9789210028356c002.
[8] L. Nuzula, A. Prahutama, and A. R. Hakim, “Klasifikasi Status Kemiskinan Rumah Tangga Dengan Metode Support Vector Machines (Svm) Dan Classification And Regression Trees (Cart) Menggunakan Gui R (Studi Kasus di Kabupaten Wonosobo Tahun 2018),” J. Gaussian, vol. 9, no. 4, pp. 525–534, 2020, doi: 10.14710/j.gauss.v9i4.29449.
[9] D. V. Ramadhanti, R. Santoso, and T. Widiharih, “Perbandingan Smote Dan Adasyn Pada Data Imbalance Untuk Klasifikasi Rumah Tangga Miskin Di Kabupaten Temanggung Dengan Algoritma K-Nearest Neighbor,” J. Gaussian, vol. 11, no. 4, pp. 499–505, 2023, doi: 10.14710/j.gauss.11.4.499-505.
[10] N. N. Sholihah and A. Hermawan, “Implementation of Random Forest and Smote Methods for Economic Status Classification in Cirebon City,” J. Tek. Inform., vol. 4, no. 6, pp. 1387–1397, 2023, doi: 10.52436/1.jutif.2023.4.6.1135.
[11] R. Siringoringo, D. Arisandi, E. Kurniawan, and E. B. Nababan, “Model Klasifikasi Dengan Logistic Regression Dan Recursive Classification Model Using Logistic Regression And Recursive,” vol. 11, no. 4, 2024, doi: 10.25126/jtiik.1148198.
[12] S. Rahayu and Y. Yamasari, “Klasifikasi Penyakit Stroke dengan Metode Support Vector Machine ( SVM ),” vol. 05, pp. 440–446, 2024.
[13] V. J. S. Vimal, M. K. Mi, and Y. Lee, “AI ‑ based smart prediction of clinical disease using random forest classifier and Naive Bayes,” J. Supercomput., vol. 77, no. 5, pp. 5198–5219, 2021, doi: 10.1007/s11227-020-03481-x.
[14] Y. Muamar and A. Muhajirin, “Penerapan Jaringan Saraf Tiruan Dengan Metode Backpropagation Untuk Memprediksi Tingkat Kelulusan Mahasiswa Perguruan Tinggi,” Digit. Transform. Technol., vol. 4, no. 1, pp. 214–224, 2024, doi: 10.47709/digitech.v4i1.3810.
[15] L. Mardiana, D. Kusnandar, and N. Satyahadewi, “Analisis Diskriminan Dengan K Fold Cross Validation Untuk Klasifikasi Kualitas Air Di Kota Pontianak,” vol. 11, no. 1, pp. 97–102, 2022.
[16] A. P. Sabiq Sofyan, “Penerapan Synthetic Minority Oversampling Technique ( SMOTE ) Terhadap Data Tidak Seimbang Pada Tingkat Pendapatan Pekerja,” vol. 2019, pp. 868–877, 2021.
[17] P. R. Sihombing and A. M. Arsani, “Comparison of Machine Learning Methods in Classifying Poverty in Indonesia in 2018,” J. Tek. Inform., vol. 2, no. 1, pp. 51–56, 2021, doi: 10.20884/1.jutif.2021.2.1.52.
[18] R. Ariefudin, M. Alfarizzi, N. Machmudah, E. Aminah, Y. Febiyanti, and Wasono, “Analisis Diskriminan untuk Klasifikasi Tingkat Kemiskinan di Perkotaan Menurut Provinsi Berdasarkan Bagian Wilayah di Indonesia Tahun 2022,” Pros. Semin. Nas. Mat. Stat. dan Apl., pp. 214–223, 2023.
[19] H. A. Ahmed, P. J. M. Ali, A. K. Faeq, and S. M. Abdullah, “An Investigation on Disparity Responds of Machine Learning Algorithms to Data Normalization Method,” pp. 29–37, 2022, doi: 10.14500/aro.10970.
[20] R. Firliana, R. Wulanningrum, and W. Sasongko, “Implementasi Principal Component Analysis (PCA) Untuk Pengenalan Wajah Manusia,” Nusant. Eng. ISSN 2355-6684, vol. 2, no. 1, pp. 65–69, 2005.
[21] D. Leni, A. Dwiharzandis, R. Sumiati, and S. Afriyani, “Seleksi Fitur Berdasarkan Korelasi Pearson dalam Pemodelan Efisiensi Energi Bangunan Feature Selection Based on Pearson Correlation in Building Energy Efficiency Modeling,” vol. 08, 2023.
[22] A. Ratna, R. Rakhmat, and I. Novita, “Perbandingan Metode Seleksi Fitur Chi-Square dan Information Gain untuk Peningkatan Interpretabilitas dan Optimasi Kinerja Model TabNet,” vol. 03, pp. 253–262, 2025.
[23] P. Putu, N. Ardhaneswari, I. W. C. Suwitra, and J. J. I. S. Siwirabuda, “Analisis Korelasi Pearson Dalam Menentukan Hubungan Harga Dengan Volume Penjualan Wardah Matte Lip Cream Pada Platform E-Commerce Shopee,” vol. 02, no. 02, pp. 151–156, 2024.
[24] R. A. Nugraha, E. W. Hidayat, N. I. Kurniati, and R. N. Shofa, “Klasifikasi Jenis Buah Jambu Biji Menggunakan Algoritma Principal Component Analysis dan K-Nearest Neighbor,” vol. 7, no. 1, pp. 1–7, 2023.
[25] D. Herinanto, B. H. S. Utami, D. Arif, and M. Gumanti, “Analisis Chi Square Zona Wilayah Marketing Terhadap Penjualan Produk Ekonomi Kreatif,” vol. 6, no. 09, pp. 1626–1637, 2024.
[26] J. R. J.M. Gorriz, R. Martin Clemente, F. Segovia, “I S K- Fold Cross Validation The Best Model Selection Method For M Achine L Earning ?,” 2024.
[27] T. Wahyuningsih and E. Rahwanto, “Comparison of Min-Max normalization and Z-Score Normalization in the K-nearest neighbor ( kNN ) Algorithm to Test the Accuracy of Types of Breast Cancer,” vol. 4, no. 1, pp. 13–20, 2021.
[28] W. Zhang, A. Bifet, X. Zhang, and L. G. Aug, “FARF: A Fair and Adaptive Random Forests Classifier,” pp. 1–12.
[29] M. Arya and C. S. S. Bedi, “Survey on SVM and their application in image classification,” 2018.
[30] S. Rahmah, H. Azis, D. Widyawati, and A. U. Tenripada, “Prediksi potensi donatur menggunakan model Logistic Regression,” vol. 4, no. 1, pp. 31–37, 2023.
[31] R. Mawarni, “Analisis Determinan Financialdistress Dengan Menggunakan Metode Artificial Neural Network Dalam Perspektif Ekonomi Islam,” 2022.
[32] E. Hasibuan and E. Allistair, “Analisis Sentimen Pada Ulasan Aplikasi Amazon Shopping Di Google Play,” vol. 1, no. 3, pp. 13–24, 2022.
[33] T. Hidayati and B. F. Putra, “Scientia Sacra : Jurnal Sains , Teknologi dan Masyarakat Implementasi Deep Learning Untuk Image Classification menggunakan Convolutional Neural Network Pada Citra Wayang ( Studi Kasus : SDN Leuwibatu 03 ),” vol. 4, no. 1, pp. 1–7, 2024.
[34] T. Z. Jasman, M. A. Fadhlullah, A. L. Pratama, and Rismayanti, “Analisis Algoritma Gradient Boosting , AdaBoost dan CatBoost dalam Klasifikasi Kualitas Air Analysis of Gradient Boosting , Adaboost , Catboost Algorithms in Water Quality Classification,” vol. 8, pp. 392–402, 2022.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nuniska Dwi Kamayanti, Ifnu Wisma Dwi Prastya, Sahri Sahri

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








