Construction of a Dialect-Sensitive Javanese Semantic Lexicon to Support Machine Translation Systems

Authors

  • Musthofa Galih Pradana Universitas Pembangunan Nasional Veteran Jakarta
  • Ridwan Raafi'udin Universitas Pembangunan Nasional Veteran Jakarta
  • Nurul Afifah Arifuddin Universitas Pembangunan Nasional Veteran Jakarta
  • Mohammad Asaduzzaman Rasel Monash University

DOI:

https://doi.org/10.30871/jaic.v10i4.13379

Keywords:

Javanese Semantic Lexicon, Speech Level Variation, Lexical Resources, Low-Resources Languages, Polysemy Exploration

Abstract

The development of linguistic resources for natural language processing (NLP) in Javanese remains limited, especially regarding the representation of semantic relationships between different speech levels. This study aims to construct a Javanese semantic lexicon that integrates Indonesian lexical equivalents with three Javanese speech levels: ngoko, krama alus, and krama inggil. A research design based on lexical resource construction was employed, using a Javanese digital dictionary as the primary data source. The methodology included data extraction, preprocessing, semantic lexicon construction, analysis of speech level variation, and a preliminary exploration of polysemous lexical entries using automatic identification, followed by validation by native speakers. The resulting semantic lexicon successfully represents lexical relationships between levels in a structured manner. Analysis of speech-level variation revealed that partially distinct lexical patterns were the most dominant, with 733 entries, followed by fully distinct patterns (193 entries) and identical patterns (21 entries). These findings indicate that speech-level differences in Javanese are selectively realized and should be explicitly considered in the development of linguistic resources. Furthermore, preliminary exploration of polysemous candidates demonstrated that dictionary-based automatic identification can overestimate polysemy without linguistic validation. Only a limited number of lexical entries exhibited features consistent with genuine polysemous relationships. This study provides an initial basis for the development of Javanese semantic resources that are sensitive to speech-level variation and semantic complexity. The constructed semantic lexicon has the potential to support future research in NLP applications in Javanese, including politeness identification, lexical normalization, word sense disambiguation, and machine translation.

Downloads

Download data is not yet available.

References

[1] A. Z. Munibi, “Assessing Neural Machine Translation in Speech: Problems and Solutions in AI-Powered Translations,” Journal of Computing Innovations and Emerging Technologies, vol. 2, no. 1, pp. 25–32, Jun. 2026, doi: https://doi.org/10.64472/jciet.v2i1.27

[2] M. R. Farhansyah, I. Darmawan, A. Kusumawardhana, G. I. Winata, A. F. Aji, and D. T. Wijaya, “Do Language Models Understand Honorific Systems in Javanese?,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Stroudsburg, PA, USA: Association for Computational Linguistics, 2025, pp. 26732–26754. doi: https://doi.org/10.18653/v1/2025.acl-long.1296

[3] W. Tan and K. Zhu, “NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models,” Oct. 2024. doi : https://doi.org/10.48550

[4] F. I. Putri, A. P. Wibawa, and L. H. Collante, “Refining the Performance of Indonesian-Javanese Bilingual Neural Machine Translation Using Adam Optimizer,” ILKOM Jurnal Ilmiah, vol. 16, no. 3, pp. 271–282, Dec. 2024, doi: https://doi.org/10.33096/ilkom.v16i3.2467.271-282

[5] E. Rippeth, M. Carpuat, K. Duh, and M. Post, “Improving Word Sense Disambiguation in Neural Machine Translation with Salient Document Context,” Nov. 2023. doi : https://doi.org/10.18653/v1/2024.emnlp-demo.34

[6] H. D. Masethe, M. A. Masethe, S. O. Ojo, P. A. Owolawi, and F. Giunchiglia, “Hybrid Transformer-Based Large Language Models for Word Sense Disambiguation in the Low-Resource Sesotho sa Leboa Language,” Applied Sciences, vol. 15, no. 7, p. 3608, Mar. 2025, doi: https://doi.org/10.3390/app15073608

[7] R. Goworek et al., “SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods,” in Proceedings of the 7th Workshop on Research in Computational Linguistic Typology and Multilingual NLP, Stroudsburg, PA, USA: Association for Computational Linguistics, 2025, pp. 61–74. doi: https://doi.org/10.18653/v1/2025.sigtyp-1.7

[8] M. B. Prayoga, B. R. Ermawan, A. R. Fadhillah, M. N. Farizi, and E. Utami, “Evaluating Machine Translation Models and LLMs for Indonesian–Javanese Translation Across Speech Levels,” Journal of Applied Informatics and Computing, vol. 10, no. 2, pp. 1549–1560, Apr. 2026, doi: https://doi.org/10.30871/jaic.v10i2.12326

[9] D. A. Sulistyo, A. P. Wibawa, D. D. Prasetya, and F. A. Ahda, “An enhanced pivot-based neural machine translation for low-resource languages,” International Journal of Advances in Intelligent Informatics, vol. 11, no. 2, p. 258, May 2025, doi: https://doi.org/10.26555/ijain.v11i2.2115

[10] C. I. Ratnasari and D. D. Ramadhan, “Neural Machine Translation from Indonesian to Javanese Ngoko using RNN-LSTM Approach,” in 2026 International Conference on Artificial Intelligence, Computer, Data Sciences and Applications (ACDSA), IEEE, Feb. 2026, pp. 1–6. doi: https://doi.org/10.1109/ACDSA67686.2026.11468163

[11] M. G. Pradana, H. B. Seta, N. Irzavika, P. H. Saputro, and R. Rusiyono, “Levenshtein Distance Algorithm in Javanese Character Translation Machine Based on Optical Character Recognition,” JOIV : International Journal on Informatics Visualization, vol. 9, no. 4, p. 1411, Jul. 2025, doi: https://doi.org/10.62527/joiv.9.4.3151

[12] A. Nurhadiyatna, D. Rahman, and T. Hidayat, “Improving Javanese OCR via Transliteration and Levenshtein-based Post-Processing,” Journal of Intelligent Informatics and Visualization (JOIV), vol. 4, no. 2, pp. 89–98, 2023, doi: https://doi.org/10.35735/joiv.v4i2.1579

[13] H. D. Masethe, M. A. Masethe, S. O. Ojo, P. A. Owolawi, and F. Giunchiglia, “Hybrid Transformer-Based Large Language Models for Word Sense Disambiguation in the Low-Resource Sesotho sa Leboa Language,” Applied Sciences, vol. 15, no. 7, p. 3608, Mar. 2025, doi: https://doi.org/10.3390/app15073608

[14] D. A. Sulistyo, D. D. Prasetya, F. A. Ahda, and A. P. Wibawa, “Pivoted Low Resource Multilingual Translation with NER Optimization,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 5, pp. 1–16, May 2025, doi: https://doi.org/10.1145/3727876

[15] F. Almu’iini Ahda, A. Prasetya Wibawa, D. Dwi Prasetya, D. Arbian Sulistyo, and A. Nafalski, “Advanced Machine Translation-Based Stammer for Preserving the Minangkabau Traditional Medicine Domain,” JOIV : International Journal on Informatics Visualization, vol. 10, no. 3, p. 959, May 2026, doi: https://doi.org/10.62527/joiv.10.3.4042

[16] “Word Sense Disambiguation and Location Named Entity Recognition for Madurese-Indonesian Rule-based Machine Translation,” International Journal of Intelligent Engineering and Systems, vol. 18, no. 2, pp. 572–588, Mar. 2025, doi: https://doi.org/10.22266/ijies2025.0331.42

Downloads

Published

2026-08-08

How to Cite

[1]
M. G. Pradana, R. Raafi'udin, N. A. Arifuddin, and M. Asaduzzaman Rasel, “Construction of a Dialect-Sensitive Javanese Semantic Lexicon to Support Machine Translation Systems”, JAIC, vol. 10, no. 4, pp. 3270–3276, Aug. 2026.

Similar Articles

<< < 2 3 4 5 6 > >> 

You may also start an advanced similarity search for this article.