Implementation of a Hybrid Model Using Principal Component Analysis, K-Means, and Naïve Bayes for Tuition Fee Category Prediction
DOI:
https://doi.org/10.30871/jaic.v10i4.13184Keywords:
Tuition Fee Categories, Principal Component Analysis, K-Means Clustering, Naïve Bayes Classifier, Machine LearningAbstract
The determination of Tuition Fee Categories in higher education institutions is commonly conducted through manual verification of students’ socioeconomic documents, which may lead to subjectivity and inconsistencies in decision-making. This study proposes a hybrid machine learning approach that integrates Principal Component Analysis (PCA), K-Means Clustering, and Naïve Bayes Classifier within a semi-supervised learning framework for student socioeconomic classification based on pseudo-labels generated from clustering results. The dataset used in this study consists of 452 student records with 12 socioeconomic attributes obtained from the New Student Admission system of STAIN Teungku Dirundeng Meulaboh in 2025. Data preprocessing includes attribute selection, categorical transformation using One Hot Encoding, and feature standardization. PCA is applied to reduce dimensionality from 18 features to 12 principal components while retaining 95% of the total variance. The processed data are clustered using K-Means with the optimal number of clusters determined as 8 based on Elbow and Silhouette Score analysis. These clusters are used as pseudo-labels for training the Naïve Bayes classifier. Experimental results show that the proposed model achieves 98.89% training accuracy and 97.80% testing accuracy, with a weighted average F1-score of 0.98. The results indicate that the proposed hybrid approach is effective in capturing underlying socioeconomic patterns and provides a stable classification performance. However, the model is based on pseudo-labels rather than official tuition fee categories. Therefore, further validation using real labeled data is recommended to enhance generalizability and practical applicability.
Downloads
References
[1] G. Al-Tameemi, J. Xue, I. H. Ali, and S. Ajit, “A Hybrid Machine Learning Approach for Predicting Student Performance Using Multi-class Educational Datasets,” Procedia Comput. Sci., vol. 238, pp. 888–895, 2024, doi: 10.1016/j.procs.2024.06.108.
[2] A. Khaidar, N. Nurdin, F. Fajriana, T. Taufiq, and D. Hamdhana, “Classification Analysis of Single Tuition Fees Using the Random Forest Method with K-Fold Cross Validation,” J. Appl. Informatics Comput., vol. 10, no. 1, pp. 125–133, 2026.
[3] Kementerian Pendidikan, Kebudayaan, Riset, dan Teknologi, “Peraturan Menteri Pendidikan, Kebudayaan, Riset, dan Teknologi Nomor 2 Tahun 2024 tentang Standar Satuan Biaya Operasional Pendidikan Tinggi (SSBOPT).” 2024.
[4] I. Papadogiannis, M. Wallace, and G. Karountzou, “Educational Data Mining: A Foundational Overview,” Mach. Learn. Knowl. Extr., vol. 6, no. 4, pp. 1644–1664, 2024, doi: https://doi.org/10.3390/encyclopedia4040108.
[5] A. F. M. Nafuri, N. S. Sani, N. F. A. Zainudin, A. H. A. Rahman, and M. Aliff, “Clustering Analysis for Classifying Student Academic Performance in Higher Education,” Appl. Sci., vol. 12, no. 13, pp. 1–22, 2022, doi: https://doi.org/10.3390/app12199467.
[6] D. Wulandari, T. Prahasto, and V. Gunawan, “Penerapan Principal Component Analysis (PCA) Untuk Mereduksi Dimensi Data Penerapan Teknologi Informasi dan Komunikasi untuk Pendidikan di Sekolah,” J. Sist. Inf. Bisnis, vol. 02, pp. 91–96, 2016, doi: 10.21456/vol6iss2pp91-96.
[7] A. Zaki, Irwan, and I. A. Sembe, “Penerapan K-Means Clustering dalam Pengelompokan Data ( Studi Kasus Profil Mahasiswa Matematika FMIPA UNM ),” J. Math. Comput. Stat., vol. 5, no. 2, pp. 163–176, 2022.
[8] S. Hartati, N. A. Ramdhan, and H. A. SAN, “Prediksi Kelulusan Mahasiswa dengan Naïve Bayes dan Feature Selection Information Gain,” J. Ilm. Intech Inf. Technol. J. UMUS, vol. 4, no. 02, pp. 223–235, 2022.
[9] D. E. Prayogo and Kusrini, “Application of K-Means and Naïve Bayes Algorithms for Prediction Model of Student Interest Concentration (Case Study: Amikom University Yogyakarta),” G-Tech J. Teknol. Terap., vol. 9, no. 1, pp. 511–519, 2025, doi: https://doi.org/10.70609/gtech.v9i1.6523.
[10] M. R. Gusmansyah, Rahmaddeni, Rohid, S. Daulay, and A. Rivaldi, “Perbandingan Metode Machine Learning untuk Klastering Penerima Bantuan Pendidikan Siswa,” J. Pustaka AI, vol. 5, no. 2, pp. 193–203, 2025, doi: https://doi.org/10.55382/jurnalpustakaai.v5i2.1139.
[11] A. Sapitri, N. Nurdin, and Y. Afrilia, “Implementation of Clustering Method Using K-Means Algorithm for Grouping BPJS Health Patient Medical Record Data,” J. Appl. Informatics Comput., vol. 9, no. 5, pp. 2391–2398, 2025, [Online]. Available: http://jurnal.polibatam.ac.id/index.php/JAIC
[12] A. Alfitra, Nurdin, and R. Meiyanti, “Comparison of K-Means and K-Medoids Methods in Clustering High Population Density Areas in Bireuen Regency,” JITE (Journal Informatics Telecommun. Eng., vol. 9, no. 1, pp. 292–302, 2025, doi: 10.31289/jite.v9i1.15602.
[13] B. A. Fauzan, M. Jamaris, Junadhi, and H. Asnal, “Implementation of K-Means Clustering Algorithm for Grouping Traffic Violation Levels in Siak,” J. Teknol. dan Open Source, vol. 5, no. 1, pp. 81–88, 2022, doi: 10.36378/jtos.v5i1.2427.
[14] S. N. Hariono, N. Nurdin, and L. Rosnita, “Comparison of K-Nearest Neighbors Method and Naïve Bayes Method in Classifying the Quality of Oil Palm Seed Varieties,” J. Inov. Teknol. dan Rekayasa, vol. 10, no. 2, pp. 413–425, 2025, doi: 10.31572/inotera.Vol10.Iss2.2025.ID539.
[15] M. S. Samosir and L. Wati, “Penerapan Naive Bayes Untuk Memprediksi Kelulusan Mahasiswa Rekayasa Perangkat Lunak Politeknik Negeri Bengkalis,” Remik Ris. dan E-Jurnal Manaj. Inform. Komput., vol. 8, no. 3, pp. 838–848, 2024, doi: http://doi.org/10.33395/remik.v8i3.13964.
[16] N. Nurdin, M. Suhendri, and Y. Afrilia, “Klasifikasi Karya Ilmiah ( Tugas Akhir ) Mahasiswa Menggunakan Metode Naive Bayes Classifier ( Nbc ),” Sist. J. Sist. Inf., vol. 10, pp. 268–279, 2021, doi: 10.32520/stmsi.v10i2.1193.
[17] Y. Safrina and Nurdin, “Analisis Perbandingan Metode Naïve Bayes dengan Forward Chaining Pada Sistem Pakar Diagnosa Kanker Servik,” J. Sist. Inf. Kaputama, vol. 9, no. 2, pp. 139–145, 2025.
[18] N. Amalia, N. Nurdin, and F. Fajriana, “Z-Score Based Initialization for K-Medoids Clustering : Application on QSAR Toxicity Data,” J. Appl. Informatics Comput., vol. 9, no. 5, pp. 2410–2417, 2025, doi: https://doi.org/10.30871/jaic.v9i5.10448.
[19] A. A. Alharbi and J. Allohibi, “A New Hybrid Classification Algorithm for Predicting Student Performance,” AIMS Math., vol. 9, no. 7, pp. 18308–18323, 2024, doi: 10.3934/math.2024893.
[20] A. Khaidar, N. Nurdin, and F. Fajriana, “Single Tuition Fee Classification Using Light Gradient Boosting Machine with Confusion Matrix Analysis,” Journal of Artificial Intelligence and Software Engineering, vol. 5, no. 4, pp. 1444–1454, 2025, doi: 10.30811/jaise.v5i4.847.
[21] V. Popovych and M. Drlik, “Identification of Students with Similar Performances in Micro-Learning Programming Courses with Automatically Evaluated Student Assignments,” Appl. Sci., vol. 14, no. 3615, pp. 1–26, 2024, doi: https://doi.org/10.3390/app14093615.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Nurdin Nurdin, Jessika Jessika, Munirul Ula

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).



