Real-Time BISINDO Gesture Detection Using YOLOv8 for a Web-Based Text-to-Speech Prototype
DOI:
https://doi.org/10.30871/jaic.v10i4.13362Keywords:
BISINDO, YOLOv8, Object Detection, Gesture Sequence Accumulation, Text-to-SpeechAbstract
Communication challenges persist between the deaf and hard-of-hearing community and the general public, largely due to limited awareness and understanding of Indonesian Sign Language (BISINDO) in the broader population. This study develops a real-time web-based system that identifies BISINDO gestures using the YOLOv8 object detection model and converts the recognized gesture sequence into spoken words via a text-to-speech function. The study is based on the CRISP-ML(Q) framework, which includes stages such as understanding the data, preparing the data, building models, assessing their performance, implementing them in real-world applications, and continuously tracking their effectiveness. A total of 1,550 images were independently collected using a laptop camera and categorized into 31 classes, including 30 BISINDO gesture classes and 1 class for negative samples. The dataset was split using a stratified method, allocating 80% for training, 10% for validation, and 10% for testing. The YOLOv8 model was trained on Google Colaboratory using a T4 GPU runtime. The evaluation results indicate that the model achieved a precision of 0.979, a recall of 0.976, an mAP50 of 0.993, and an mAP50-95 of 0.839, which demonstrates robust performance in detecting BISINDO gestures. The developed model was incorporated into a web prototype, with FastAPI serving as the backend and HTML, CSS, and JavaScript utilized for the frontend. The system can identify hand gestures using a webcam, show bounding boxes with corresponding labels, compile the detected gestures into simple text sequences, and produce speech from that information. Latency testing revealed an average response time of around 210 milliseconds when the model was deployed locally and approximately 670 milliseconds when deployed online via Hugging Face Spaces. The system still faces challenges in identifying similar-looking gestures and is influenced by factors such as lighting, hand placement, and the availability of hosting resources. These results indicate that YOLOv8s is effective for detecting the 30 static BISINDO gesture classes evaluated in this study within a controlled data collection setting, though further validation is required before the system can be considered suitable for broader real-world assistive communication use.
Downloads
References
[1] M. M. Fresmanda, Istiadi, and S. W. Iriananda, “Deteksi Objek Video Bahasa Isyarat Untuk Anak Tuna Rungu dan Tuna Wicara Menggunakan YOLOv8,” Jurnal Komputer, Informasi dan Teknologi, vol. 4, no. 2, Oct. 2024, doi: 10.53697/jkomitek.v4i2.1895.
[2] F. I. Hadinata and S. A. Sanjaya, “BISINDO Sign Language Recognition: A Systematic Literature Review of Deep Learning Techniques for Image Processing,” Indonesian Journal of Computer Science Attribution, vol. 12, no. 6, p. 3281, Dec. 2023.
[3] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-Time Object Detection,” 2021. [Online]. Available: http://pjreddie.com/yolo/
[4] H. Y. A. Swasono, A. R. Himamunanto, and F. Maedjaja, “Implementasi YOLO11 dan OpenCV untuk Pengenalan Frasa dalam Video Real-Time Bahasa Isyarat Tangan,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, vol. 5, no. 3, Jul. 2025, doi: 10.57152/malcom.v5i3.2130.
[5] A. Ahmad Ilham and I. Nurtanio, “Dynamic Sign Language Recognition Using Mediapipe Library and Modified LSTM Method,” vol. 13, no. 6, 2023.
[6] B. Subramanian, B. Olimov, S. M. Naik, S. Kim, K. H. Park, and J. Kim, “An integrated mediapipe-optimized GRU model for Indian sign language recognition,” Sci. Rep., vol. 12, no. 1, Dec. 2022, doi: 10.1038/s41598-022-15998-7.
[7] N. A. Megantara and E. Utami, “Sistemasi: Jurnal Sistem Informasi Object Detection Using YOLOv8 : A Systematic Review,” 2025. [Online]. Available: http://sistemasi.ftik.unisi.ac.id
[8] Y. T. Buana and D. Ariatmanto, “Comparative Evaluation of Optimizer in YOLOv8 for BISINDO Alphabet Detection,” 2025. [Online]. Available: http://publishing-widyagama.ac.id/ejournal-v2/index.php/jointecs
[9] E. L. Kelana, M. R. A. Prasetya, Mambang, and M. Zulfadhilah, “Integrating the CNN Model with the Web for Indonesian Sign Language (BISINDO) Recognition,” 2025. [Online]. Available: http://jurnal.polibatam.ac.id/index.php/JAIC
[10] N. Thalbiatul et al., “Pengembangan Aplikasi Web Pengenalan Huruf Bahasa Isyarat Indonesia (BISINDO) Real-Time Menggunakan Mobilenetv2,” vol. 10, no. 2, 2025.
[11] Mangai V and Kalaimagal R, “Comparative Analysis of Various Yolo Models for Sign Language Recognition with a specific dataset,” 2024. [Online]. Available: https://www.jisem-journal.com/
[12] A. B. Pangestu, R. Muttaqin, and A. Sunandar, “Sistem Deteksi Bahasa Isyarat Indonesia (BISINDO) Menggunakan Algoritma You Only Look Once (YOLO)V8,” 2024.
[13] N. Renaningtias, F. Putra Utama, A. Nur, and A. Sobri, “Deteksi Bahasa Isyarat Indonesia (BISINDO) Pada Video dengan YOLOv7,” JSAI: Journal Scientific and Applied Informatics, vol. 8, no. 1, 2025, doi: 10.36085.
[14] S. Putra and E. Rakun, “End-to-end system for translating bahasa isyarat Indonesia sign language gestures into Indonesian text,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 40, no. 2, p. 719, Nov. 2025, doi: 10.11591/ijeecs.v40.i2.pp719-734.
[15] S. Studer et al., “Towards CRISP-ML(Q): A Machine Learning Process Model with Quality Assurance Methodology,” Mach. Learn. Knowl. Extr., vol. 3, no. 2, pp. 392–413, Jun. 2021, doi: 10.3390/make3020020.
[16] J. Murrugarra-Llerena and C. R. Jung, “Can we trust bounding box annotations for object detection?,” 2022. [Online]. Available: http://www.cvlibs.net/datasets/kitti/
[17] J. Terven, D. M. Córdova-Esparza, and J. A. Romero-González, “A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS,” Dec. 01, 2023, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/make5040083.
[18] M. Hussain, “YOLO-v1 to YOLO-v8, the Rise of YOLO and Its Complementary Nature toward Digital Manufacturing and Industrial Defect Detection,” Jul. 01, 2023, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/machines11070677.
[19] M. Hasan, B. K. Paul, N. Islam, and R. Mostafiz, “Advancing real-time sign language detection for deaf and hearing-impaired communities: a customized YOLOv8 approach with tailored annotations in computer vision,” BMC Artificial Intelligence, vol. 1, no. 1, Oct. 2025, doi: 10.1186/s44398-025-00010-9.
[20] W. Jia and C. Li, “SLR-YOLO: An improved YOLOv8 network for real-time sign language recognition,” Journal of Intelligent & Fuzzy Systems, vol. 46, no. 1, pp. 1663–1680, Jan. 2024, doi: 10.3233/JIFS-235132.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Riska Dewi Yuliyanti, M. Rafi Muttaqin, Teguh Iman Hermanto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








