Comparative Analysis Of Automatic Labeling With Cohen's Kappa Validation For Cyberbullying On Tiktok Using Pre-Trained Transformers

Authors

  • David Rian Prabowo Universitas PGRI Semarang
  • Khoiriya Latifah Universitas PGRI Semarang
  • Nugroho Dwi Saputro Universitas PGRI Semarang

DOI:

https://doi.org/10.30871/jaic.v10i4.13108

Keywords:

Cyberbullying, InSet Lexicon, TikTok, Transformer, Zero Shot Classification

Abstract

The rapid growth of TikTok users in Indonesia increases the risk of cyberbullying, making manual content moderation inefficient. This study compares two automatic labeling methods, InSet Lexicon and Zero-Shot Classification, to address the problem of limited labeled data in cyberbullying detection. A total of 5,864 comment data were collected through scraping techniques and processed through comprehensive text preprocessing stages, including slang normalization. To evaluate labeling reliability, a validation test using Cohen's Kappa Score metric was conducted against a human annotated Gold Standard of 1,141 comments. The results show that Zero-Shot Classification achieves high reliability with a Kappa score of 0.9132 (almost perfect agreement), outperforming InSet Lexicon which drops to 0.2585 (fair agreement) due to lexical rigidity and high false positives on casual slang. The automatically labeled datasets were balanced using Random Oversampling (for the Zero-Shot Classification labeling scenario) and split via a stratified 80:10:10 ratio to fine-tune IndoBERT and RoBERTa. On independent test data, IndoBERT trained on Zero-Shot labels delivers the best performance, reaching an Accuracy and F1-Score of 91.10%, outperforming RoBERTa under the same scenario (85.83%). Conversely, training on InSet Lexicon labels reduces performance, limiting IndoBERT to an 86.94% F1-Score and RoBERTa to 80.21%. This study concludes that using Zero-Shot Classification and fine-tuning IndoBERT is more optimal for application to informal social media comment moderation systems.

Downloads

Download data is not yet available.

References

[1] “Digital 2026: Indonesia,” DataReportal – Global Digital Insights. Accessed: Apr. 23, 2026. [Online]. Available: https://datareportal.com/reports/digital-2026-indonesia

[2] H. D. Jayanti and A. Rohman, “Cyberbullying Detection in Indonesian TikTok Comments Using IndoBERT with Fairness Evaluation,” J. Inf. Syst. Inform., vol. 8, no. 1, pp. 907–927, Mar. 2026, doi: 10.63158/journalisi.v8i1.1448.

[3] E. Asalnaije, Y. Bete, M. A. Manikin, R. A. Labu, S. A. D. Tira, and Y. P. Lian, “Bentuk-Bentuk Cyberbullying Di Indonesia,” Innov. J. Soc. Sci. Res., vol. 4, no. 4, pp. 6465–6473, Jul. 2024, doi: 10.31004/innovative.v4i4.12471.

[4] F. Husain, H. Alostad, and H. Omar, “Q8SentiLabeler Automated Labeling System for Arabic Sentiment Analysis,” ResearchGate. Accessed: Apr. 23, 2026. [Online]. Available: https://www.researchgate.net/figure/Q8SentiLabeler-System-Architecture_fig4_378081726

[5] H. Firda, P. Putra, N. R. Oktadini, P. E. Sevtiyuni, and A. Meiriza, “Comparison of Rating-based and Inset Lexicon-based Labeling in Sentiment Analysis using SVM (Case Study: GoBiz Application Reviews on Google Play Store),” SISTEMASI, vol. 14, no. 2, p. 516, Mar. 2025, doi: 10.32520/stmsi.v14i2.4795.

[6] F. Koto and G. Y. Rahmaningtyas, “Inset Lexicon: Evaluation of A Word List for Indonesian Sentiment Analysis in Microblogs,” in Proc. Int. Conf. Asian Lang. Process. (IALP), 2017, pp. 391–394. doi: 10.1109/IALP.2017.8300625.

[7] P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing,” Mar. 24, 2023, arXiv: arXiv:2111.09543. doi: 10.48550/arXiv.2111.09543.

[8] J. Cohen, “A Coefficient of Agreement for Nominal Scales,” Educ. Psychol. Meas., vol. 20, no. 1, pp. 37–46, Apr. 1960, doi: 10.1177/001316446002000104.

[9] J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,” Biometrics, vol. 33, no. 1, pp. 159–174, Mar. 1977.

[10] B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” Oct. 08, 2020, arXiv: arXiv:2009.05387. doi: 10.48550/arXiv.2009.05387.

[11] Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” Jul. 26, 2019, arXiv: arXiv:1907.11692. doi: 10.48550/arXiv.1907.11692.

[12] A. S. Rizkia, W. Wufron, and F. F. Roji, “Analisis Sentimen Coretax: Perbandingan Pelabelan Data Manual, Transformers-Based, dan Lexicon-Based pada Performa IndoBERT: Sentiment Analysis of Coretax: A Comparison of Manual, Transformers-Based, and Lexicon-Based Data Labeling on IndoBERT Performance,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 5, no. 3, pp. 1037–1048, Jul. 2025, doi: 10.57152/malcom.v5i3.2151.

[13] “MoritzLaurer/mDeBERTa-v3-base-mnli-xnli • Hugging Face.” Accessed: May 04, 2026. [Online]. Available: https://huggingface.co/MoritzLaurer/mDeBERTa-v3-base-mnli-xnli

[14] T. Wolf et al., “Transformers: State-of-the-art Natural Language Processing,” Jul. 14, 2020, arXiv: arXiv:1910.03771. doi: 10.48550/arXiv.1910.03771.

[15] Muhammad Fernanda Naufal Fathoni, Eva Yulia Puspaningrum, and Andreas Nugroho Sihananto, “Perbandingan Performa Labeling Lexicon InSet dan VADER pada Analisa Sentimen Rohingya di Aplikasi X dengan SVM,” Modem J. Inform. Dan Sains Teknol., vol. 2, no. 3, pp. 62–76, Jul. 2024, doi: 10.62951/modem.v2i3.112.

[16] null Riccosan, K. E. Saputra, G. D. Pratama, and A. Chowanda, “Emotion dataset from Indonesian public opinion,” Data Brief, vol. 43, p. 108465, Aug. 2022, doi: 10.1016/j.dib.2022.108465.

[17] K. H. Prastiawan and D. Yuniarto, “Analisis Sentimen Publik terhadap Program Makan Bergizi Gratis dengan Algoritma Naive Bayes,” RIGGS J. Artif. Intell. Digit. Bus., vol. 4, no. 4, pp. 5412–5419, Dec. 2025, doi: 10.31004/riggs.v4i4.3652.

[18] P. A. A. Wijaya, I. M. Dika Anggara, I. K. Hendra Trinium Jaya, G. Indrawan, and I. M. Agus Oka Gunawan, “Comprehensive Analysis of Teacher Teaching Performance Through Sentiment and POS Tagging,” J. Ilm. Merpati Menara Penelit. Akad. Teknol. Inf., vol. 12, no. 3, p. 147, Dec. 2024, doi: 10.24843/JIM.2024.v12.i03.p02.

[19] R. Erama, “Pemanfaatan Platform Cloud Google Colab Untuk Scraping Komentar Tiktok Pada Konten Gorontalo sebagai Dasar Analisis Respons Warganet,” J. Appl. Eng. Sci., vol. 1, no. 2, pp. 124–134, Dec. 2025, doi: 10.65177/jaes.v1i2.38.

[20] A. F. Setyawan and R. I. Nugraha, “Optimizing GPT And IndoBERT for sentiment analysis and consumer trend prediction on Lazada product reviews,” JIKO J. Inform. Dan Komput., vol. 8, no. 2, pp. 113–120, Jul. 2025, doi: 10.33387/jiko.v8i2.10066.

[21] U. Khairani, V. Mutiawani, and H. Ahmadian, “Pengaruh Tahapan Preprocessing Terhadap Model Indobert Dan Indobertweet Untuk Mendeteksi Emosi Pada Komentar Akun Berita Instagram,” J. Teknol. Inf. Dan Ilmu Komput., vol. 11, no. 4, pp. 887–894, Aug. 2024, doi: 10.25126/jtiik.1148315.

[22] N. Jain, H. Suh, S. Adeyinka, L. Roseman, and A. Allsop, “Multi-LLM Thematic Analysis with Dual Reliability Metrics: Combining Cohen’s Kappa and Semantic Similarity for Qualitative Research Validation,” Feb. 14, 2026, arXiv: arXiv:2512.20352. doi: 10.48550/arXiv.2512.20352.

[23] M. L. McHugh, “Interrater reliability: the kappa statistic,” Biochem. Medica, vol. 22, no. 3, pp. 276–282, Oct. 2012.

[24] D. R. Alfinsyah and B. P. Hartato, “Evaluating the Impact of Random Over Sampling on IndoBERT Performance for Indonesian Sentiment Analysis,” vol. 9, no. 6, pp. 3270–3282, 2025, doi: https://doi.org/10.30871/jaic.v9i6.11488.

[25] Z. Bami, A. Behnampour, A. Bora, and H. Doosti, “A New Flexible Train-Test Split Algorithm, an approach for choosing among the Hold-out, K-fold cross-validation, and Hold-out iteration,” Jan. 01, 2026, arXiv: arXiv:2501.06492. doi: 10.48550/arXiv.2501.06492.

[26] C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to Fine-Tune BERT for Text Classification?,” Feb. 05, 2020, arXiv: arXiv:1905.05583. doi: 10.48550/arXiv.1905.05583.

[27] A. Vaswani et al., “Attention Is All You Need,” Aug. 02, 2023, arXiv: arXiv:1706.03762. doi: 10.48550/arXiv.1706.03762.

[28] A. M. Putri, W. K. Nofa, and D. A. P. Hapsari, “Penerapan Metode BERT Untuk Analisis Sentimen Ulasan Pengguna Aplikasi Segari Di Google Play Store,” vol. 4, no. 1, 2025, doi: https://doi.org/10.56127/juit.v4i1.1902.

[29] M. Jazzar and T. Duridi, “A Comprehensive Review of Machine Learning and Deep Learning Techniques for Cyberbullying Detection | Request PDF,” in ResearchGate. doi: 10.1007/978-981-96-2182-8_1.

[30] S. S. Sabrina, D. F. Shiddieq, and F. F. Roji, “Comparative Analysis of SVM and BERT for Sentiment and Sarcasm Detection in the Boycott of Israeli Products on Platform X,” Sinkron, vol. 9, no. 2, pp. 872–883, May 2025, doi: 10.33395/sinkron.v9i2.14723.

[31] D. Ananda, I. Budi, A. B. Santoso, and A. A. Qureshi, “Sentiment analysis of public health app reviews using IndoBERT and XLM-RoBERTa: A study on SATUSEHAT mobile app,” JIKO J. Inform. Dan Komput., vol. 8, no. 3, pp. 161–172, Nov. 2025, doi: 10.33387/jiko.v8i3.10083.

[32] E. Boiy and M.-F. Moens, “A machine learning approach to sentiment analysis in multilingual Web texts,” Inf. Retr., vol. 12, no. 5, pp. 526–558, Oct. 2009, doi: 10.1007/s10791-008-9070-z.

[33] Arif Fitra Setyawan, Amelia Devi Putri Ariyanto, Fari Katul Fikriah, and Rozaq Isnaini Nugraha, “Analisis Sentimen Ulasan iPhone di Amazon Menggunakan Model Deep Learning BERT Berbasis Transformer,” Elkom J. Elektron. Dan Komput., vol. 17, no. 2, pp. 447–452, Dec. 2024, doi: 10.51903/elkom.v17i2.2150.

Downloads

Published

2026-08-08

How to Cite

[1]
D. R. Prabowo, K. Latifah, and N. Dwi Saputro, “Comparative Analysis Of Automatic Labeling With Cohen’s Kappa Validation For Cyberbullying On Tiktok Using Pre-Trained Transformers”, JAIC, vol. 10, no. 4, pp. 3262–3269, Aug. 2026.

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.