Implementation of a Hybrid Model Using Principal Component Analysis, K-Means, and Naïve Bayes for Tuition Fee Category Prediction

Authors

  • Nurdin Nurdin Universitas Malikussaleh
  • Jessika Jessika Universitas Malikussaleh
  • Munirul Ula Universitas Malikussaleh

DOI:

https://doi.org/10.30871/jaic.v10i4.13184

Keywords:

Tuition Fee Categories, Principal Component Analysis, K-Means Clustering, Naïve Bayes Classifier, Machine Learning

Abstract

The determination of Tuition Fee Categories in higher education institutions is commonly conducted through manual verification of students’ socioeconomic documents, which may lead to subjectivity and inconsistencies in decision-making. This study proposes a hybrid machine learning approach that integrates Principal Component Analysis (PCA), K-Means Clustering, and Naïve Bayes Classifier within a semi-supervised learning framework for student socioeconomic classification based on pseudo-labels generated from clustering results. The dataset used in this study consists of 452 student records with 12 socioeconomic attributes obtained from the New Student Admission system of STAIN Teungku Dirundeng Meulaboh in 2025. Data preprocessing includes attribute selection, categorical transformation using One Hot Encoding, and feature standardization. PCA is applied to reduce dimensionality from 18 features to 12 principal components while retaining 95% of the total variance. The processed data are clustered using K-Means with the optimal number of clusters determined as 8 based on Elbow and Silhouette Score analysis. These clusters are used as pseudo-labels for training the Naïve Bayes classifier. Experimental results show that the proposed model achieves 98.89% training accuracy and 97.80% testing accuracy, with a weighted average F1-score of 0.98. The results indicate that the proposed hybrid approach is effective in capturing underlying socioeconomic patterns and provides a stable classification performance. However, the model is based on pseudo-labels rather than official tuition fee categories. Therefore, further validation using real labeled data is recommended to enhance generalizability and practical applicability.

Downloads

Download data is not yet available.

References

[1] G. Al-Tameemi, J. Xue, I. H. Ali, and S. Ajit, “A Hybrid Machine Learning Approach for Predicting Student Performance Using Multi-class Educational Datasets,” Procedia Comput. Sci., vol. 238, pp. 888–895, 2024, doi: 10.1016/j.procs.2024.06.108.

[2] A. Khaidar, N. Nurdin, F. Fajriana, T. Taufiq, and D. Hamdhana, “Classification Analysis of Single Tuition Fees Using the Random Forest Method with K-Fold Cross Validation,” J. Appl. Informatics Comput., vol. 10, no. 1, pp. 125–133, 2026.

[3] Kementerian Pendidikan, Kebudayaan, Riset, dan Teknologi, “Peraturan Menteri Pendidikan, Kebudayaan, Riset, dan Teknologi Nomor 2 Tahun 2024 tentang Standar Satuan Biaya Operasional Pendidikan Tinggi (SSBOPT).” 2024.

[4] I. Papadogiannis, M. Wallace, and G. Karountzou, “Educational Data Mining: A Foundational Overview,” Mach. Learn. Knowl. Extr., vol. 6, no. 4, pp. 1644–1664, 2024, doi: https://doi.org/10.3390/encyclopedia4040108.

[5] A. F. M. Nafuri, N. S. Sani, N. F. A. Zainudin, A. H. A. Rahman, and M. Aliff, “Clustering Analysis for Classifying Student Academic Performance in Higher Education,” Appl. Sci., vol. 12, no. 13, pp. 1–22, 2022, doi: https://doi.org/10.3390/app12199467.

[6] D. Wulandari, T. Prahasto, and V. Gunawan, “Penerapan Principal Component Analysis (PCA) Untuk Mereduksi Dimensi Data Penerapan Teknologi Informasi dan Komunikasi untuk Pendidikan di Sekolah,” J. Sist. Inf. Bisnis, vol. 02, pp. 91–96, 2016, doi: 10.21456/vol6iss2pp91-96.

[7] A. Zaki, Irwan, and I. A. Sembe, “Penerapan K-Means Clustering dalam Pengelompokan Data ( Studi Kasus Profil Mahasiswa Matematika FMIPA UNM ),” J. Math. Comput. Stat., vol. 5, no. 2, pp. 163–176, 2022.

[8] S. Hartati, N. A. Ramdhan, and H. A. SAN, “Prediksi Kelulusan Mahasiswa dengan Naïve Bayes dan Feature Selection Information Gain,” J. Ilm. Intech Inf. Technol. J. UMUS, vol. 4, no. 02, pp. 223–235, 2022.

[9] D. E. Prayogo and Kusrini, “Application of K-Means and Naïve Bayes Algorithms for Prediction Model of Student Interest Concentration (Case Study: Amikom University Yogyakarta),” G-Tech J. Teknol. Terap., vol. 9, no. 1, pp. 511–519, 2025, doi: https://doi.org/10.70609/gtech.v9i1.6523.

[10] M. R. Gusmansyah, Rahmaddeni, Rohid, S. Daulay, and A. Rivaldi, “Perbandingan Metode Machine Learning untuk Klastering Penerima Bantuan Pendidikan Siswa,” J. Pustaka AI, vol. 5, no. 2, pp. 193–203, 2025, doi: https://doi.org/10.55382/jurnalpustakaai.v5i2.1139.

[11] A. Sapitri, N. Nurdin, and Y. Afrilia, “Implementation of Clustering Method Using K-Means Algorithm for Grouping BPJS Health Patient Medical Record Data,” J. Appl. Informatics Comput., vol. 9, no. 5, pp. 2391–2398, 2025, [Online]. Available: http://jurnal.polibatam.ac.id/index.php/JAIC

[12] A. Alfitra, Nurdin, and R. Meiyanti, “Comparison of K-Means and K-Medoids Methods in Clustering High Population Density Areas in Bireuen Regency,” JITE (Journal Informatics Telecommun. Eng., vol. 9, no. 1, pp. 292–302, 2025, doi: 10.31289/jite.v9i1.15602.

[13] B. A. Fauzan, M. Jamaris, Junadhi, and H. Asnal, “Implementation of K-Means Clustering Algorithm for Grouping Traffic Violation Levels in Siak,” J. Teknol. dan Open Source, vol. 5, no. 1, pp. 81–88, 2022, doi: 10.36378/jtos.v5i1.2427.

[14] S. N. Hariono, N. Nurdin, and L. Rosnita, “Comparison of K-Nearest Neighbors Method and Naïve Bayes Method in Classifying the Quality of Oil Palm Seed Varieties,” J. Inov. Teknol. dan Rekayasa, vol. 10, no. 2, pp. 413–425, 2025, doi: 10.31572/inotera.Vol10.Iss2.2025.ID539.

[15] M. S. Samosir and L. Wati, “Penerapan Naive Bayes Untuk Memprediksi Kelulusan Mahasiswa Rekayasa Perangkat Lunak Politeknik Negeri Bengkalis,” Remik Ris. dan E-Jurnal Manaj. Inform. Komput., vol. 8, no. 3, pp. 838–848, 2024, doi: http://doi.org/10.33395/remik.v8i3.13964.

[16] N. Nurdin, M. Suhendri, and Y. Afrilia, “Klasifikasi Karya Ilmiah ( Tugas Akhir ) Mahasiswa Menggunakan Metode Naive Bayes Classifier ( Nbc ),” Sist. J. Sist. Inf., vol. 10, pp. 268–279, 2021, doi: 10.32520/stmsi.v10i2.1193.

[17] Y. Safrina and Nurdin, “Analisis Perbandingan Metode Naïve Bayes dengan Forward Chaining Pada Sistem Pakar Diagnosa Kanker Servik,” J. Sist. Inf. Kaputama, vol. 9, no. 2, pp. 139–145, 2025.

[18] N. Amalia, N. Nurdin, and F. Fajriana, “Z-Score Based Initialization for K-Medoids Clustering : Application on QSAR Toxicity Data,” J. Appl. Informatics Comput., vol. 9, no. 5, pp. 2410–2417, 2025, doi: https://doi.org/10.30871/jaic.v9i5.10448.

[19] A. A. Alharbi and J. Allohibi, “A New Hybrid Classification Algorithm for Predicting Student Performance,” AIMS Math., vol. 9, no. 7, pp. 18308–18323, 2024, doi: 10.3934/math.2024893.

[20] A. Khaidar, N. Nurdin, and F. Fajriana, “Single Tuition Fee Classification Using Light Gradient Boosting Machine with Confusion Matrix Analysis,” Journal of Artificial Intelligence and Software Engineering, vol. 5, no. 4, pp. 1444–1454, 2025, doi: 10.30811/jaise.v5i4.847.

[21] V. Popovych and M. Drlik, “Identification of Students with Similar Performances in Micro-Learning Programming Courses with Automatically Evaluated Student Assignments,” Appl. Sci., vol. 14, no. 3615, pp. 1–26, 2024, doi: https://doi.org/10.3390/app14093615.

Downloads

Published

2026-08-11

How to Cite

[1]
N. Nurdin, J. Jessika, and M. Ula, “Implementation of a Hybrid Model Using Principal Component Analysis, K-Means, and Naïve Bayes for Tuition Fee Category Prediction”, JAIC, vol. 10, no. 4, pp. 3753–3760, Aug. 2026.

Most read articles by the same author(s)

1 2 > >> 

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.