The Application of Naïve Bayes Algorithm in Detecting Hoaxes on National News Portals in Indonesia
DOI:
https://doi.org/10.30871/jaic.v10i4.13306Keywords:
Disinformation, Hoax Detection, Multinomial Naïve Bayes, Natural Language Processing, TF-IDFAbstract
The rapid advancement of information technology in Indonesia has led to a massive spread of digital disinformation, commonly known as an infodemic. The inability to filter inaccurate information manually necessitates a reliable, automated hoax detection system. This study aims to implement and evaluate the Multinomial Naïve Bayes algorithm combined with Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction to classify news articles as either factual or hoax. The research utilizes a dataset of 2,910 Indonesian news articles published in 2025, collected from verified national news portals and fact-checking websites. The text data underwent comprehensive preprocessing—including case folding, cleansing, stopword removal, and stemming—before being evaluated using 5-Fold Cross-Validation and an 80:20 data split. Experimental results demonstrate that the Naïve Bayes model achieves highly stable and competitive performance, recording an accuracy of 93.81%, a precision of 93.84%, a recall of 93.81%, an F1-Score of 93.82%, and a 5-Fold Cross-Validation F1-Score of 93.39%. Notably, the algorithm exhibited a significantly low False Negative rate, missing only 15 hoax documents out of 582 test samples. Furthermore, the trained model was successfully integrated into a real-time, web-based user interface using Streamlit. This practical implementation provides an accessible and efficient initial screening tool for the general public and journalists to assist in verifying news authenticity, thereby supporting efforts to mitigate the impact of digital hoaxes.
Downloads
References
[1] Muhammad Salim Albana, Alif Dava Mahesa, Indriani Putri, and Noerma Kurnia Fajarwati, “Interaksi Komunikasi Hoax Di Media Sosial Serta Antisipasinya,” SABER J. Tek. Inform. Sains dan Ilmu Komun., vol. 2, no. 2, pp. 34–39, 2024, doi: 10.59841/saber.v2i2.958.
[2] M. Solikhah, “the Effect of Machine Learning Algorithms on Hoax Detection on Social Media: Implications for National Information Security,” J. Artif. Intell. Res., vol. 1, no. 1, pp. 11–20, 2025, doi: 10.64910/jouair.v1i1.7.
[3] A. Sandu, L. A. Cotfas, C. Delcea, C. Ioanăș, M. S. Florescu, and M. Orzan, “Machine Learning and Deep Learning Applications in Disinformation Detection: A Bibliometric Assessment,” Electron., vol. 13, no. 22, 2024, doi: 10.3390/electronics13224352.
[4] P. S. More, A. Jadhav, and S. Yadav, “Detection of Fake News using Machine Learning,” Ijarcce, vol. 10, no. 4, pp. 485–490, 2021, doi: 10.17148/ijarcce.2021.10484.
[5] S. Pandey, S. Prabhakaran, N. V. S. Reddy, and D. Acharya, “Fake News Detection from Online media using Machine learning Classifiers,” J. Phys. Conf. Ser., vol. 2161, no. 1, 2022, doi: 10.1088/1742-6596/2161/1/012027.
[6] H. F. Villela, F. Corrêa, J. S. de A. N. Ribeiro, A. Rabelo, and D. B. F. Carvalho, “Fake news detection: a systematic literature review of machine learning algorithms and datasets,” J. Interact. Syst., vol. 14, no. 1, pp. 47–58, 2023, doi: 10.5753/jis.2023.3020.
[7] A. Badawi, “the Effectiveness of Natural Language Processing (Nlp) As a Processing Solution and Semantic Improvement,” Int. J. Econ. Technol. Soc. Sci., vol. 2, no. 1, pp. 36–44, 2021, doi: 10.53695/injects.v2i1.194.
[8] A. Khaidar, S. Muliana, and P. Studi Magister Teknologi Informasi, “Sentiment Analysis of Instagram Comments on the BPS Province X Account Using the Naive Bayes Algorithm Based on Machine Learning,” J. Artif. Intell. Softw. Eng., vol. 5, no. 3, pp. 1231–1237, 2025, doi: 10.30811/jaise.v5i3.7815.
[9] E. Y. Hidayat and M. A. Rizqi, “Klasifikasi Dokumen Berita Menggunakan Algoritma Enhanced Confix Stripping Stemmer dan Naïve Bayes Classifier,” J. Nas. Teknol. dan Sist. Inf., vol. 6, no. 2, pp. 90–99, 2020, doi: 10.25077/teknosi.v6i2.2020.90-99.
[10] M. Suhendri and Y. Afrilia, “SISTEMASI: Jurnal Sistem Informasi Klasifikasi Karya Ilmiah (Tugas Akhir) Mahasiswa Menggunakan Metode Naive Bayes Classifier (Nbc),” vol. 10, pp. 268–279, 2021, [Online]. Available: http://sistemasi.ftik.unisi.ac.id
[11] O. Peretz, M. Koren, and O. Koren, “Naive Bayes classifier – An ensemble procedure for recall and precision enrichment,” Eng. Appl. Artif. Intell., vol. 136, no. PB, p. 108972, 2024, doi: 10.1016/j.engappai.2024.108972.
[12] A. F. Watratan, A. P. B, and D. Moeis, “Implementation of the Naive Bayes Algorithm to Predict the Spread of Covid-19 in Indonesia,” J. Appl. Comput. Sci. Technol., vol. 1, no. 1, pp. 7–14, 2020, doi: https://doi.org/10.52158/jacost.v1i1.9.
[13] S. Nasyira and L. Rosnita, “Comparison of K-Nearest Neighbors Method and Naïve Bayes Method in Classifying the Quality of Oil Palm Seed Varieties,” vol. 10, no. 2, pp. 413–425, 2025, doi: 10.31572/inotera.Vol10.Iss2.2025.ID539.
[14] H. D. Abubakar and M. Umar, “Sentiment Classification: Review of Text Vectorization Methods: Bag of Words, Tf-Idf, Word2vec and Doc2vec,” SLU J. Sci. Technol., vol. 4, no. 1&2, pp. 27–33, 2022, doi: 10.56471/slujst.v4i.266.
[15] R. Rahmadani, A. Rahim, and R. Rudiman, “Analisis Sentimen Ulasan ‘Ojol the Game’ Di Google Play Store Menggunakan Algoritma Naive Bayes Dan Model Ekstraksi Fitur Tf-Idf Untuk Meningkatkan Kualitas Game,” J. Inform. dan Tek. Elektro Terap., vol. 12, no. 3, 2024, doi: 10.23960/jitet.v12i3.4988.
[16] F. Rahutomo, I. Y. R. Pratiwi, and D. M. Ramadhani, “Eksperimen Naïve Bayes Pada Deteksi Berita Hoax Berbahasa Indonesia,” J. Penelit. Komun. Dan Opini Publik, vol. 23, no. 1, 2019, doi: 10.33299/jpkop.23.1.1805.
[17] Rianto, A. B. Mutiara, E. P. Wibowo, and P. I. Santosa, “Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation,” J. Big Data, vol. 8, no. 1, pp. 1–16, 2021, doi: 10.1186/s40537-021-00413-1.
[18] N. V. Pusean, N. Charibaldi, and B. Santosa, “Comparison of Scenario Pre-processing Performance on Support Vector Machine and Naïve Bayes Algorithms for Sentiment Analysis,” Inf. J. Ilm. Bid. Teknol. Inf. dan Komun., vol. 8, no. 1, pp. 57–63, 2023, doi: 10.25139/inform.v8i1.5667.
[19] M. N. Raza, “Sistem Deteksi Berita Hoax Menggunakan Algoritma Naïve Bayes Dan Random Forest Pada Machine Learning,” Pondasi J. Appl. Sci. Eng., vol. 1, no. 2, pp. 43–57, 2024, [Online]. Available: https://journal.alshobar.or.id/index.php/pondasi/article/view/221
[20] Febriyanty Nur Elyta, “Deteksi Berita Hoax Dari Media Online Indonesia Menggunakan Algoritma Naive Bayes dan Support Vector Machine,” pp. 1–128, 2023.
[21] J. S. Wibowo, E. N. Wahyudi, and H. Listiyono, “Performance Comparison of SVM, Naive Bayes, and Random Forest Models in Fake News Classification,” Eng. Technol. J., vol. 09, no. 08, pp. 4799–4804, 2024, doi: 10.47191/etj/v9i08.27.
[22] A. P. Wibawa, F. A. Dwiyanto, I. A. E. Zaeni, R. K. Nurrohman, and A. Afandi, “Stemming javanese affix words using nazief and adriani modifications,” J. Inform., vol. 14, no. 1, p. 36, 2020, doi: 10.26555/jifo.v14i1.a17106.
[23] A. Alaiya, N. Nurdin, and C. Agusniar, “Sentiment Analysis of E-Commerce Product Reviews on Tokopedia Using Support Vector Machine,” J. Appl. Informatics Comput., vol. 9, no. 5, pp. 2869–2878, 2025, doi: 10.30871/jaic.v9i5.10977.
[24] F. Nurwanda and J. R. Rizkiani, “Perbandingan Metode Naive Bayes Classifier dan Support Vector Machine pada Analisis Sentimen Twitter Topik Lifestyle,” J. Ilm. Wahana Pendidik., vol. 9, no. 21, pp. 314–323, 2023, [Online]. Available: https://doi.org/10.5281/zenodo.10077023
[25] I. A. Ropikoh, R. Abdulhakim, U. Enri, and N. Sulistiyowati, “Penerapan Algoritma Support Vector Machine (SVM) untuk Klasifikasi Berita Hoax Covid-19,” J. Appl. Informatics Comput., vol. 5, no. 1, pp. 64–73, 2021, doi: 10.30871/jaic.v5i1.3167.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Cut Rifa Salsabil, Nurdin Nurdin, Rizki Suwanda

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








