Work Readiness Prediction of Vocational High School Students Using a Genetic Algorithm-Optimized Random Forest and XGBoost as a Benchmark Model with Explainable AI (SHAP)

Authors

  • Syifa Amalia Sistem Informasi, Fakultas Teknik, Universitas Muria Kudus
  • Noor Latifah Sistem Informasi, Fakultas Teknik, Universitas Muria Kudus
  • Andy Prasetyo Utomo Sistem Informasi, Fakultas Teknik, Universitas Muria Kudus

DOI:

https://doi.org/10.30871/jaic.v10i4.13435

Keywords:

Explainable AI, Genetic Algorithm, Random Forest, Vocational School Students, XGBoost

Abstract

Work readiness among vocational high school (SMK) students remains a critical challenge in aligning educational outcomes with industry expectations. This study proposes a binary classification system using Genetic Algorithm (GA)-optimized Random Forest as the primary model and XGBoost as a benchmark. Two predictor variables are used: average practical score (X1) and attitude score (X2). The classification target is based on the minimum competency threshold (KKM = 80) regulated by the Indonesian Ministry of Education: a student is labeled Ready to Work when X1 >= 80 AND X2 >= 80. A boundary noise injection mechanism is applied to mitigate label determinism at the borderline zone. The dataset comprises 1,211 student records split into 968 training and 243 testing records. Evaluation employs hold-out testing, Stratified 5-Fold Cross Validation, overfitting detection, and McNemar's statistical significance test. SHAP interpretability is performed at three levels: global feature importance, beeswarm summary plot, and dependence plots. RF+GA achieved 90.12% accuracy, 0.9391 precision, 0.8640 recall, 0.9000 F1-score, and 0.9582 AUC-ROC. Cross-validation confirmed stable performance at 91.53% +/- 2.51%. Overfitting analysis showed a training-testing gap of only 1.72% (Not Detected). McNemar's test (p = 1.000) indicated statistically equivalent performance between both models. SHAP analysis revealed that attitude score (X2) is the dominant predictor (mean SHAP = 0.2611), followed by average practical score (X1, mean SHAP = 0.1628). Practical implications include an early warning system, targeted remedial programs, and individual guidance recommendations generated by the deployed Streamlit application.

Downloads

Download data is not yet available.

References

[1] “Educational Data Mining Using Cluster Analysis Methods and Decision Trees based on Log Mining,” Ikat. Ahli Inform. Indones., vol. Vol 6 No 3 (2022): Juni 2022, 2022.

[2] I. Nurjanah, A. Ana, and A. Masek, “Systematic Literature Review : Work readiness of vocational high school,” J. Pendidik. Teknol. dan Kejuru., vol. 28, no. 2, pp. 139–153, 2022.

[3] A. Agustiningsih, Y. Findawati, and I. Alnarus Kautsar, “Classification Of Vocational High School Graduates’ Ability In Industry Using Extreme Gradient Boosting (Xgboost), Random Forest, And Logistic Regression,” J. Tek. Inform., vol. 4, no. 4 SE-Articles, pp. 977–985, Sep. 2023, doi: 10.52436/1.jutif.2023.4.4.945.

[4] S. Jayachandran and B. Joshi, “Customized support vector machine for predicting the employability of students pursuing engineering,” Int. J. Inf. Technol., vol. 16, no. 5, pp. 3193–3204, 2024, doi: 10.1007/s41870-024-01818-w.

[5] R. Haque, A. Quek, C. Y. Ting, H. N. Goh, and M. R. Hasan, “Classification Techniques Using Machine Learning for Graduate Student Employability Predictions,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 14, no. 1, pp. 45–56, 2024, doi: 10.18517/ijaseit.14.1.19549.

[6] L. G. R. Putra, D. D. Prasetya, and M. Mayadi, “Student Dropout Prediction Using Random Forest and XGBoost Method,” INTENSIF J. Ilm. Penelit. dan Penerapan Teknol. Sist. Inf., vol. 9, no. 1, pp. 147–157, 2025, doi: 10.29407/intensif.v9i1.21191.

[7] P. R. Togatorop, M. Sianturi, D. Simamora, and D. Silaen, “Optimizing Random Forest using Genetic Algorithm for Heart Disease Classification,” Lontar Komput. J. Ilm. Teknol. Inf., vol. 13, no. 1, p. 60, 2022, doi: 10.24843/lkjiti.2022.v13.i01.p06.

[8] Y. Guan, F. Wang, and S. Song, “Interpretable machine learning for academic performance prediction: A SHAP-based analysis of key influencing factors,” Innov. Educ. Teach. Int., pp. 1–20, Jul. 2025, doi: 10.1080/14703297.2025.2532050.

[9] M. M. Islam, F. H. Sojib, M. F. H. Mihad, M. Hasan, and M. Rahman, “The integration of explainable AI in Educational Data Mining for student academic performance prediction and support system,” Telemat. Informatics Reports, vol. 18, no. February, p. 100203, 2025, doi: 10.1016/j.teler.2025.100203.

[10] G. Caroline, P. Ningsih, F. Liantoni, and Y. Sujana, “Telematika A Systematic Analysis of the Impact of Non-Academic Factors on Student Academic Performance Prediction using Data Mining,” vol. 19, no. 1, pp. 116–125, 2026.

[11] Swono Sibagariang, “Interpretable Machine Learning for Job Placement Prediction: A SHAP-Based Feature Analysis,” J. Nas. Tek. Elektro dan Teknol. Inf., vol. 14, no. 3, pp. 190–198, 2025, doi: 10.22146/jnteti.v14i3.20516.

[12] C. Surya, E. Yubarda, K. Ameliza, and W. Simatupang, “Analyzing the Influence of Academic Competence and Soft Skills on Vocational Students ’ Work Readiness Using Regression and Machine Learning Approaches,” vol. 13, no. 1, pp. 32–43, 2026.

[13] O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Sci. Rep., vol. 14, no. 1, pp. 1–14, 2024, doi: 10.1038/s41598-024-56706-x.

[14] M. F. Al Hakim, S. Wahyuni, K. Budiman, A. Marianti, and B. E. Susilo, “Optimization of Machine Learning Model using Grid and Random Search Algorithms for Predicting Student Dropout,” J. Tek. Inform., vol. 7, no. 3, pp. 3012–3024, 2026.

[15] M. Martinović, K. Dokic, and D. Pudić, “Comparative Analysis of Machine Learning Models for Predicting Innovation Outcomes: An Applied AI Approach,” Appl. Sci., vol. 15, no. 7, pp. 1–44, 2025, doi: 10.3390/app15073636.

[16] G. R. Saputra and M. Mujiyono, “Evaluating the implementation of the mechanical engineering student skills competency,” J. Pendidik. Vokasi, vol. 14, no. 2, pp. 194–207, 2024, doi: 10.21831/jpv.v14i2.71359.

[17] S. E. Davis, M. E. Matheny, S. Balu, and M. P. Sendak, “A framework for understanding label leakage in machine learning for health care,” J. Am. Med. Informatics Assoc., vol. 31, no. 1, pp. 274–280, 2024, doi: 10.1093/jamia/ocad178.

[18] S. Studer et al., “Towards CRISP-ML(Q): A Machine Learning Process Model with Quality Assurance Methodology,” Mach. Learn. Knowl. Extr., vol. 3, no. 2, pp. 392–413, 2021, doi: 10.3390/make3020020.

[19] I. V. Tetko, R. van Deursen, and G. Godin, “Be aware of overfitting by hyperparameter optimization!,” J. Cheminform., vol. 16, no. 1, 2024, doi: 10.1186/s13321-024-00934-w.

[20] S. Widodo, H. Brawijaya, and S. Samudi, “Stratified K-fold cross validation optimization on machine learning for prediction,” Sinkron, vol. 7, no. 4, pp. 2407–2414, 2022, doi: 10.33395/sinkron.v7i4.11792.

Downloads

Published

2026-08-08

How to Cite

[1]
S. Amalia, N. Latifah, and A. P. Utomo, “Work Readiness Prediction of Vocational High School Students Using a Genetic Algorithm-Optimized Random Forest and XGBoost as a Benchmark Model with Explainable AI (SHAP)”, JAIC, vol. 10, no. 4, pp. 3369–3374, Aug. 2026.

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.