Benchmarking Random Forest, Support Vector Machine, and XGBoost for Flood Risk Classification Using a Synthetic Dataset: A Case Study of Indramayu Regency

Authors

  • Nur Budi Nugraha Department of Informatics Engineering, Politeknik Negeri Indramayu
  • Rendi Rendi Department of Informatics Engineering, Politeknik Negeri Indramayu
  • Yaqutina Marjani Santosa Department of Informatics Engineering, Politeknik Negeri Indramayu

DOI:

https://doi.org/10.30871/jaic.v10i4.13708

Keywords:

Random Forest, Support Vector Machine, XGBoost, Flood Risk Classification, Benchmarking, Machine Learning

Abstract

Flood risk assessment plays a critical role in disaster mitigation planning, yet flood classification studies in Indonesia have largely relied on a single machine learning algorithm without systematically evaluating alternative classifiers under identical experimental conditions. This study benchmarks Random Forest (RF), Support Vector Machine (SVM), and XGBoost for three-class flood risk classification (Low, Medium, and High) in Indramayu Regency, Indonesia, using a literature-informed synthetic dataset of 4,500 samples generated from seven flood-related features. To ensure a fair comparison, all models were trained and evaluated under an identical preprocessing, hyperparameter optimization, and validation framework, with Logistic Regression and Decision Tree included as baseline classifiers. Experimental results show that SVM achieved the highest predictive performance with an accuracy of 88.56% and a macro F1-score of 86.83%, followed closely by XGBoost (88.44% accuracy, 86.76% macro F1), while RF obtained 85.00% accuracy and an 83.57% macro F1-score. Statistical significance testing confirmed that SVM and XGBoost significantly outperformed RF, whereas no significant difference was observed between SVM and XGBoost. Feature importance analysis consistently identified rainfall and river distance as the two most influential predictors across all models. Although SVM provided the strongest overall classification performance, RF demonstrated competitive predictive capability with better generalization characteristics than XGBoost, supporting its suitability for operational flood mitigation decision-support systems where model interpretability and robustness are important. Because the benchmark is based on a synthetic dataset, further validation using real observational flood data is recommended before operational deployment.

Downloads

Download data is not yet available.

References

[1] R. Rahayu, S. A. Mathias, S. Reaney, G. Vesuviano, R. Suwarman, and A. M. Ramdhan, “Impact of land cover, rainfall and topography on flood risk in West Java,” Natural Hazards, vol. 116, no. 2, pp. 1735–1758, 2023, doi: 10.1007/s11069-022-05737-6.

[2] M. Ardiansyah, R. A. Nugraha, L. O. S. Iman, and S. D. Djatmiko, “Impact of Land Use and Climate Changes on Flood Inundation Areas in the Lower Cimanuk Watershed, West Java Province,” Jurnal Ilmu Tanah dan Lingkungan, vol. 23, no. 2, pp. 53–60, Dec. 2021, doi: 10.29244/jitl.23.2.53-60.

[3] S. Putiamini, M. P. Patria, T. E. B. Soesilo, and A. Karsidi, “Coastal Vulnerability Assessment To Tidal (Rob) Flooding In Indramayu Coast, West Java, Indonesia,” Indonesian Journal of Geography, vol. 55, no. 3, pp. 517–526, 2023, doi: 10.22146/ijg.65549.

[4] A. Mosavi, P. Ozturk, and K. W. Chau, “Flood prediction using machine learning models: Literature review,” Oct. 27, 2018, MDPI AG. doi: 10.3390/w10111536.

[5] J. H. Danumah, W. A. Ataba, V. C. Jofack Sokeng, Y. L. Akpa, M. B. Saley, and A. Ogilvie, “Assessing Urban Flood Susceptibility Using Random Forest Machine Learning and Geospatial Technologies: Application to the Bonoumin-Palmeraie Watershed, Abidjan (Côte d’Ivoire),” Water (Switzerland), vol. 18, no. 3, Feb. 2026, doi: 10.3390/w18030402.

[6] R. Abedi, R.-D. Costache, H. Shafizadeh-Moghadam, and Q. Pham, “Flash-flood susceptibility mapping based on XGBoost, Random Forest and Boosted Regression Trees,” Geocarto Int., vol. 37, Apr. 2021, doi: 10.1080/10106049.2021.1920636.

[7] L. Breiman, “Random Forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.

[8] D. T. Bui, P. Tsangaratos, V.-T. Nguyen, N. Van Liem, and P. T. Trinh, “Comparing the prediction performance of a Deep Learning Neural Network model with conventional machine learning models in landslide susceptibility assessment,” Catena (Amst)., vol. 188, p. 104426, 2020, doi: https://doi.org/10.1016/j.catena.2019.104426.

[9] A. Salvati et al., “Flood susceptibility mapping using support vector regression and hyper-parameter optimization,” J. Flood Risk Manag., vol. 16, no. 4, p. e12920, Dec. 2023, doi: https://doi.org/10.1111/jfr3.12920.

[10] N. Tepetidis, I. Benekos, T. Iliopoulou, P. Dimitriadis, and D. Koutsoyiannis, “Combining Machine Learning Models and Satellite Data of an Extreme Flood Event for Flood Susceptibility Mapping,” Water (Switzerland), vol. 17, no. 18, Sep. 2025, doi: 10.3390/w17182678.

[11] R. Bentivoglio, E. Isufi, S. N. Jonkman, and R. Taormina, “Deep Learning Methods for Flood Mapping: A Review of Existing Applications and Future Research Directions,” Mar. 02, 2022. doi: 10.5194/hess-2022-83.

[12] P. K. Das, R. L. Sahu, and P. C. Swain, “Comparative machine learning for flood susceptibility in Subarnarekha Basin, Odisha,” J. Atmos. Sol. Terr. Phys., vol. 274, p. 106578, 2025, doi: https://doi.org/10.1016/j.jastp.2025.106578.

[13] S. Ramayanti et al., “Performance comparison of two deep learning models for flood susceptibility map in Beira area, Mozambique,” The Egyptian Journal of Remote Sensing and Space Science, vol. 25, no. 4, pp. 1025–1036, 2022, doi: https://doi.org/10.1016/j.ejrs.2022.11.003.

[14] T. M. Asrade, S. A. Abebe, K. B. Tadesse, M. S. Kerebih, and T. M. Meshesha, “Flood susceptibility assessment using three machine learning techniques and comparison of their performance,” Sci. Rep., vol. 16, no. 1, Dec. 2026, doi: 10.1038/s41598-026-38391-0.

[15] S. T. Seydi, Y. Kanani-Sadat, M. Hasanlou, R. Sahraei, J. Chanussot, and M. Amani, “Comparison of Machine Learning Algorithms for Flood Susceptibility Mapping,” Remote Sens. (Basel)., vol. 15, no. 1, Jan. 2023, doi: 10.3390/rs15010192.

[16] K. M. Kurugama, S. Kazama, Y. Hiraga, and C. Samarasuriya, “A comparative spatial analysis of flood susceptibility mapping using boosting machine learning algorithms in Rathnapura, Sri Lanka,” J. Flood Risk Manag., vol. 17, no. 2, p. e12980, Jun. 2024, doi: https://doi.org/10.1111/jfr3.12980.

[17] M. Feizbahr, N. Brake, H. Arbabkhah, H. Hariri Asli, and K. Woods, “Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems,” Remote Sens. (Basel)., vol. 17, no. 20, Oct. 2025, doi: 10.3390/rs17203471.

[18] C. Meng and H. Jin, “A Comparison of Machine Learning Models for Predicting Flood Susceptibility Based on the Enhanced NHAND Method,” Sustainability (Switzerland), vol. 15, no. 20, Oct. 2023, doi: 10.3390/su152014928.

[19] Z. U. Rahman et al., “Flood susceptibility mapping using supervised machine learning models: insights into predictors’ significance and models’ performance,” Geomatics, Natural Hazards and Risk, vol. 16, no. 1, p. 2516728, Jun. 2025, doi: 10.1080/19475705.2025.2516728.

[20] D. Chicco and G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, no. 1, p. 6, 2020, doi: 10.1186/s12864-019-6413-7.

[21] C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995, doi: 10.1007/BF00994018.

[22] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” Jun. 2016, doi: 10.1145/2939672.2939785.

[23] S. M. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 4765–4774.

Downloads

Published

2026-08-13

How to Cite

[1]
N. B. Nugraha, R. Rendi, and Y. M. Santosa, “Benchmarking Random Forest, Support Vector Machine, and XGBoost for Flood Risk Classification Using a Synthetic Dataset: A Case Study of Indramayu Regency”, JAIC, vol. 10, no. 4, pp. 4091–4099, Aug. 2026.

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.