Comparative Analysis of Siamese BiLSTM and IndoBERT for Semantic Textual Similarity Detection in Functional Scientific Papers of Meteorology, Climatology, and Geophysics Officers at BMKG

Authors

  • Dhedy Listyawan Universitas Pamulang
  • Ahmad Musyafa Master of Informatics Engineering Program, Universitas Pamulang
  • Choirul Basir Department of Mathematics, Faculty of Mathematics and Natural Science, Universitas Pamulang

DOI:

https://doi.org/10.30871/jaic.v10i4.13773

Keywords:

Semantic Textual Similarity, Siamese BiLSTM, IndoBERT, Document Similarity, PMG BMKG

Abstract

The competency assessment process for Meteorology, Climatology, and Geophysics (PMG) functional officers at BMKG requires manual review of Scientific Paper (KTI) similarity, a task that is time-consuming, prone to subjectivity, and increasingly burdensome as submission volume grows. While Semantic Textual Similarity (STS) research has advanced considerably for high-resource languages, empirical evidence for Indonesian-language technical documents in specialized scientific domains remains limited, and prior work has not established which neural architecture is preferable under such data-constrained conditions. This study addresses that gap by empirically comparing two Siamese Network architectures, Siamese BiLSTM and IndoBERT, for automatic STS detection on 87 PMG KTI documents, yielding 3,741 document pairs automatically labeled using the 90th percentile of TF-IDF cosine similarity scores and validated against manual annotation (Cohen's Kappa κ=0.82). Siamese BiLSTM employs Word2Vec embeddings with Focal Loss, while IndoBERT fine-tunes the pretrained indobert-base-p1 model with Contrastive Loss; both apply class weighting to address the 90:10 class imbalance, with classification thresholds independently calibrated via grid search. Evaluated on 1,123 held-out test pairs, Siamese BiLSTM achieves an F1-Score of 64.84% (Accuracy 91.99%, Precision 58.04%, Recall 73.45%) at threshold 0.60, outperforming IndoBERT's F1-Score of 57.73% (Accuracy 89.05%, Precision 47.19%, Recall 74.34%) at threshold 0.935, a difference confirmed statistically significant by McNemar's test (χ²=7.6992; p=0.0055). This result runs counter to the common assumption that Transformer-based models universally outperform recurrent architectures, suggesting that smaller, domain-tuned embeddings can be more effective under limited-data, domain-specific conditions. The better-performing Siamese BiLSTM model, requiring only 16.77 MB with no GPU dependency, was deployed as a Streamlit web application, enabling the BMKG PMG Assessment Team to perform similarity detection quickly and consistently within their existing workflow.

Downloads

Download data is not yet available.

References

[1] J. Mueller and A. Thyagarajan, "Siamese recurrent architectures for learning sentence similarity," in Proc. 30th AAAI Conf. Artif. Intell., 2016, pp. 2786–2792.

[2] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.

[3] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," arXiv:1810.04805, 2019.

[4] Y. Li, C. L. P. Chen, and T. Zhang, "A survey on Siamese network: Methodologies, applications, and opportunities," IEEE Transactions on Artificial Intelligence, vol. 3, no. 6, pp. 994–1014, 2022.

[5] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddings using Siamese BERT-networks," in Proc. 2019 Conf. Empirical Methods in Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 3982–3992.

[6] L. D. Krisnawati, A. W. Mahastama, S.-C. Haw, K.-W. Ng, and P. Naveen, "Indonesian-English textual similarity detection using Universal Sentence Encoder (USE) and Facebook AI Similarity Search (FAISS)," CommIT Journal, vol. 18, no. 2, pp. 183–195, 2024.

[7] G. Salton and C. Buckley, "Term-weighting approaches in automatic text retrieval," Information Processing & Management, vol. 24, no. 5, pp. 513–523, 1988.

[8] D. Viji and S. Revathy, "Analyzing semantic similarity amongst textual documents to suggest near duplicates," Indonesian Journal of Electrical Engineering and Computer Science, vol. 25, no. 3, pp. 1703–1711, 2022.

[9] J. Cohen, "A coefficient of agreement for nominal scales," Educational and Psychological Measurement, vol. 20, no. 1, pp. 37–46, 1960.

[10] B. Wilie et al., "IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding," arXiv:2009.05387, 2020.

[11] T. Mikolov, K. Chen, G. Corrado, and J. Dean, "Efficient estimation of word representations in vector space," arXiv:1301.3781, 2013.

[12] R. Hadsell, S. Chopra, and Y. LeCun, "Dimensionality reduction by learning an invariant mapping," in Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., vol. 2, 2006, pp. 1735–1742.

[13] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, "Focal loss for dense object detection," in Proc. IEEE Int. Conf. Comput. Vis., 2017, pp. 2980–2988.

[14] I. Loshchilov and F. Hutter, "Decoupled weight decay regularization," in Proc. Int. Conf. Learn. Represent., 2019.

[15] M. Sokolova and G. Lapalme, "A systematic analysis of performance measures for classification tasks," Information Processing & Management, vol. 45, no. 4, pp. 427–437, 2009.

[16] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," arXiv:1412.6980, 2015.

[17] Q. McNemar, "Note on the sampling error of the difference between correlated proportions or percentages," Psychometrika, vol. 12, no. 2, pp. 153–157, 1947.

[18] B. Efron and R. J. Tibshirani, An Introduction to the Bootstrap. New York: Chapman & Hall/CRC, 1993, pp. 168–177.

[19] C. R. Harris et al., "Array programming with NumPy," Nature, vol. 585, no. 7825, pp. 357–362, 2020.

[20] T. Gao, X. Yao, and D. Chen, "SimCSE: Simple contrastive learning of sentence embeddings," in Proc. 2021 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2021, pp. 6894–6910.

Downloads

Published

2026-08-12

How to Cite

[1]
D. Listyawan, A. Musyafa, and C. Basir, “Comparative Analysis of Siamese BiLSTM and IndoBERT for Semantic Textual Similarity Detection in Functional Scientific Papers of Meteorology, Climatology, and Geophysics Officers at BMKG”, JAIC, vol. 10, no. 4, pp. 3884–3890, Aug. 2026.

Similar Articles

<< < 4 5 6 7 8 > >> 

You may also start an advanced similarity search for this article.