Indonesian Cyberbullying Detection Using IndoBERTweet-BiGRU Model on Class-Imbalanced X (Twitter) Data
DOI:
https://doi.org/10.30871/jaic.v10i4.13686Keywords:
Cyberbullying, IndoBERTweet, BiGRU, Focal Loss, X (Twitter)Abstract
Cyberbullying on social media platforms, particularly X (formerly Twitter), has become a serious issue that negatively affects users' mental health and well-being. Automatic cyberbullying detection in Indonesian remains challenging due to the widespread use of informal language, slang, abbreviations, and highly imbalanced class distributions. This study proposes a hybrid deep learning model that integrates IndoBERTweet with a Bidirectional Gated Recurrent Unit (BiGRU) to improve cyberbullying detection performance on Indonesian tweets. A dataset of Indonesian tweets was collected from X and annotated using a multi-stage dual large language model (LLM) labeling strategy to reduce the time and effort required for manual annotation while maintaining label consistency. To address class imbalance, this study investigates the effectiveness of Focal Loss and label distribution modification through multiple experimental scenarios. The proposed approach was evaluated using accuracy, precision, recall, and F1-score. The best performance was achieved by combining Focal Loss with a modified four-class label configuration consisting of Rude and Vulgar Words, Sexual Harassment, Body Shaming and Hate Speech, and Non-Cyberbullying. This configuration obtained an accuracy of 0.93, precision of 0.90, recall of 0.90, and F1-score of 0.90. These findings demonstrate that integrating contextual language representations with sequential modeling, supported by an efficient LLM-assisted labeling strategy and class imbalance handling, provides an effective approach for Indonesian cyberbullying detection and offers a practical solution for large-scale social media content moderation.
Downloads
References
[1] Haseeba, “The Impact of Social Media on Social Interactions: A Sociological Perspective,” 2023.
[2] R. D. Setyaningsih and F. A. Nur, “Social revolution: The impact of social media in our lives,” Symposium of Literature, Culture, and Communication (SYLECTION) 2022, vol. 3, no. 1, p. 59, Nov. 2023, doi: 10.12928/sylection.v3i1.13933.
[3] I. Putri, “Media Sosial Sebagai Media Pergeseran Interaksi Sosial Remaja,” Jurnal Ilmu Komunikasi Balayudha, vol. 2, no. 2, p. 1, Dec. 2022.
[4] S. Kemp, “Digital 2025: Indonesia,” Data Reportal. Accessed: Jun. 09, 2025. [Online]. Available: https://datareportal.com/reports/digital-2025-indonesia?rq=indonesia%202025
[5] S. Kemp, “Digital 2024: Indonesia,” Data Reportal. Accessed: Jun. 09, 2025. [Online]. Available: https://datareportal.com/reports/digital-2024-indonesia?rq=indonesia%202024
[6] L. Fazry and N. C. Apsari, “Pengaruh Media Sosial Terhadap Perilaku Cyberbullying Di Kalangan Remaja,” Jurnal Penelitian dan Pengabdian Kepada Masyarakat (JPPM), vol. 2, no. 2, p. 272, Aug. 2021, doi: 10.24198/jppm.v2i2.34679.
[7] Y. Zhang, B. Zhou, Y. Hu, and K. Zhai, “From Individual Expression to Group Polarization: A Study on Twitter’s Emotional Diffusion Patterns in the German Election,” Behavioral Sciences, vol. 15, no. 3, p. 360, Mar. 2025, doi: 10.3390/bs15030360.
[8] M. E. Rao and D. M. Rao, “The Mental Health of High School Students During the COVID-19 Pandemic,” Front. Educ. (Lausanne)., vol. 6, Jul. 2021, doi: 10.3389/feduc.2021.719539.
[9] S. Bansal, N. Garg, J. Singh, and F. Van Der Walt, “Cyberbullying and mental health: past, present and future,” Front. Psychol., vol. 14, Jan. 2024, doi: 10.3389/fpsyg.2023.1279234.
[10] M. T. Chamizo-Nieto and L. Rey, “Cybervictimization and suicidal ideation in adolescents: A prospective view through gratitude and life satisfaction,” J. Health Psychol., vol. 28, no. 7, pp. 620–632, Jun. 2023, doi: 10.1177/13591053221140259.
[11] I. Planellas Kirchner and C. Calderon Garrido, “Do Cybervictimizations Predict Suicide-Related Behaviors in Adolescents? Mediating Role of the ‘Escaping’ Coping Strategy,” J. Interpers. Violence, vol. 40, no. 5–6, pp. 1015–1036, Mar. 2025, doi: 10.1177/08862605241256384.
[12] P. S. Reddy, “Natural Language Processing (NLP) and Understanding,” International Journal of Advanced Research in Electrical, Electronics and Instrumentation Engineering, 2025, doi: 10.15662/IJAREEIE.2025.1402022.
[13] S. N. Afandy, K. M. Hindrayani, and A. T. Damaliana, “Comparison of the Effectiveness IndoBERT and mBERT for Sentiment Analysis of SME Customer Reviews,” bit-Tech, vol. 8, no. 3, pp. 3383–3394, Apr. 2026, doi: 10.32877/bt.v8i3.3501.
[14] A. J. Andika, Y. Kristian, and E. I. Setiawan, “Deteksi Komentar Cyberbullying Pada YouTube Dengan Metode Convolutional Neural Network – Long Short-Term Memory Network (CNN-LSTM),” Teknika, vol. 12, no. 3, pp. 183–188, Oct. 2023, doi: 10.34148/teknika.v12i3.677.
[15] Y. D. Novandian et al., “IndoBERT-based Indonesian Cyberbullying Detection with Multi-stage Labeling,” in 2024 International Seminar on Application for Technology of Information and Communication (iSemantic), IEEE, Sep. 2024, pp. 515–521. doi: 10.1109/iSemantic63362.2024.10762553.
[16] F. Koto, J. H. Lau, and T. Baldwin, “IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 10660–10668. doi: 10.18653/v1/2021.emnlp-main.833.
[17] F. Indriani, R. A. Nugroho, M. R. Faisal, and D. Kartini, “Comparative Evaluation of IndoBERT, IndoBERTweet, and mBERT for Multilabel Student Feedback Classification,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 8, no. 6, pp. 748–757, Dec. 2024, doi: 10.29207/resti.v8i6.6100.
[18] J. F. Kusuma and A. Chowanda, “Indonesian Hate Speech Detection Using IndoBERTweet and BiLSTM on Twitter,” JOIV : International Journal on Informatics Visualization, vol. 7, no. 3, pp. 773–780, Sep. 2023, doi: 10.30630/joiv.7.3.1035.
[19] M. R. R. Rana, A. Nawaz, S. U. Rehman, M. A. Abid, M. Garayevi, and J. Kajanová, “BERT-BiGRU-Senti-GCN: An Advanced NLP Framework for Analyzing Customer Sentiments in E-Commerce,” International Journal of Computational Intelligence Systems, vol. 18, no. 1, p. 21, Feb. 2025, doi: 10.1007/s44196-025-00747-1.
[20] K. L. Tan, C. P. Lee, and K. M. Lim, “RoBERTa-GRU: A Hybrid Deep Learning Model for Enhanced Sentiment Analysis,” Applied Sciences, vol. 13, no. 6, Mar. 2023, doi: 10.3390/app13063915.
[21] A. S. Talaat, “Hybrid Deep Learning Models for Text Classification: Performance Evaluation of TriDistilBERT and BiGRU Architectures,” Journal of Computer Science, vol. 21, no. 9, pp. 1983–1992, Sep. 2025, doi: 10.3844/jcssp.2025.1983.1992.
[22] A. Muzakir, K. Adi, and R. Kusumaningrum, “Short Text Classification Based on Hybrid Semantic Expansion and Bidirectional GRU (BiGRU) based Method to Improve Hate Speech Detection,” Revue d’Intelligence Artificielle, vol. 37, no. 6, pp. 1471–1481, Dec. 2023, doi: 10.18280/ria.370611.
[23] F. Koto, T. Beck, Z. Talat, I. Gurevych, and T. Baldwin, “Zero-shot Sentiment Analysis in Low-Resource Languages Using a Multilingual Sentiment Lexicon,” Feb. 2024.
[24] C. Padurariu and M. E. Breaban, “Dealing with Data Imbalance in Text Classification,” Procedia Comput. Sci., vol. 159, pp. 736–745, 2019, doi: 10.1016/j.procs.2019.09.229.
[25] A. C. Zahro et al., “Fairer Public Complaint Classification on LaporGub: Integrating XLM-RoBERTa with Focal Loss for Imbalance Data,” sinkron, vol. 9, no. 4, pp. 1850–1862, Oct. 2025, doi: 10.33395/sinkron.v9i4.15260.
[26] D. Jiang and J. He, “Text Semantic Classification of Long Discourses Based on Neural Networks with Improved Focal Loss,” Comput. Intell. Neurosci., vol. 2021, no. 1, Jan. 2021, doi: 10.1155/2021/8845362.
[27] EriSetyawan166, “ScraperTwitter,” GitHub. Accessed: Jul. 07, 2026. [Online]. Available: https://github.com/EriSetyawan166/ScraperTwitter
[28] M. Y. Pratiwi, “Analisis Pengaruh Data Normalisasi dan Non Normalisasi pada Deteksi Cyberbullying Berbahasa Indonesia dengan Metode IndoBERT,” Unversitas Lambung Mangkurat, Banjarmasin, 2023.
[29] Y. Xu and P. Trzaskawka, “Towards Descriptive Adequacy of Cyberbullying: Interdisciplinary Studies on Features, Cases and Legislative Concerns of Cyberbullying,” Int. J. Semiot. Law, vol. 34, no. 4, pp. 929–943, Sep. 2021, doi: 10.1007/s11196-021-09856-4.
[30] A. C. Ehman and A. M. Gross, “Sexual cyberbullying: Review, critique, & future directions,” Aggress. Violent Behav., vol. 44, pp. 80–87, Jan. 2019, doi: 10.1016/j.avb.2018.11.001.
[31] Y. W. Riyayanatasya and R. Rahayu, “Involvement of Teenage-Students in Cyberbullying on WhatsApp,” Jurnal Komunikasi Indonesia, vol. 9, no. 1, Jun. 2020, doi: 10.7454/jki.v9i1.11824.
[32] L. H. Collantes, F. A. Saputra, P. R. Solikhah, T. C. Laksana, N. K. Yakti, and J. Tipagau, “Cyberbullying Body-Shaming Levels in Adolescence,” Bulletin of Social Informatics Theory and Application, vol. 6, no. 2, pp. 111–119, Dec. 2022, doi: 10.31763/businta.v6i2.603.
[33] H. S. Jo, S. H. Lee, and M. G. Na, “Prediction of small-scale leak flow rate in LOCA situations using bidirectional GRU,” Nuclear Engineering and Technology, vol. 56, no. 9, pp. 3594–3601, Sep. 2024, doi: 10.1016/j.net.2024.04.009.
[34] Z. Labd, S. Bahassine, and K. Housni, “AFL-BERT : Enhancing Minority Class Detection in Multi-Label Text Classification with Adaptive Focal Loss and BERT,” International Journal of Advanced Computer Science and Applications, vol. 16, no. 7, 2025, doi: 10.14569/IJACSA.2025.0160749.
[35] C. Miller, T. Portlock, D. M. Nyaga, and J. M. O’Sullivan, “A review of model evaluation metrics for machine learning in genetics and genomics,” Frontiers in Bioinformatics, vol. 4, Sep. 2024, doi: 10.3389/fbinf.2024.1457619.
[36] J. R. Landis and G. G. Koch, “The Measurement of Observer Agreement for Categorical Data,” Biometrics, vol. 33, no. 1, p. 159, Mar. 1977, doi: 10.2307/2529310.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Fajria Ulumin Nafiah, Aviolla Terza Damaliana, Kartika Maulida Hindrayani

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








