Indonesia–English Bilingual Visual Question Answering Using Partial Fine-Tuning on a ViT-GPT2 Architecture

Authors

  • Anas Anas Operations and Service Control (POP), Regional VI, PT Pos Indonesia
  • Hazriani Hazriani Computer Systems, Handayani University Makassar
  • Yuyun Yuyun National Research and Innovation Agency (BRIN), Republic of Indonesia
  • Syamsul Rijal Computer Systems, Handayani University Makassar
  • Tirta Chiantalia Sharief Computer Systems, Handayani University Makassar

DOI:

https://doi.org/10.30871/jaic.v10i4.13336

Keywords:

Visual Question Answering, Vision-Language Model, Vision Transformer, GPT-2, Fine-Tuning

Abstract

Visual Question Answering (VQA) is a multimodal task that integrates visual understanding and natural language processing to generate answers based on information contained in an image. Most existing VQA research focuses on the English language and general-domain datasets, limiting its applicability to bilingual environments and domain-specific scenarios. This study proposes an Indonesia–English bilingual VQA model based on the VisionEncoderDecoderModel architecture, which combines a Vision Transformer (ViT) as the visual encoder and an Indonesian GPT-2 model as the language decoder. Bilingual capability is achieved through the introduction of special language tokens, <id> and <en>. The model is trained using a combination of the bilingual VQAv2 and bilingual LosariVQAv1 datasets, representing general-domain and local tourism-domain knowledge, respectively. Four fine-tuning strategies are evaluated: Encoder Freeze, Full Fine-Tuning, Partial-4, and Partial-6. Experimental results show that the Partial-6 strategy achieves the best performance on LosariVQAv1, obtaining an Exact Match score of 26.67%, a BLEU score of 26.80%, and a CIDEr score of 292.67, while maintaining competitive performance on VQAv2 with an Exact Match score of 41.26%. Cross-language evaluation reveals only a small performance gap between Indonesian and English. The findings indicate that partial fine-tuning provides a better balance between generalization capability and domain adaptation than the other fine-tuning strategies evaluated in this study.

Downloads

Download data is not yet available.

References

[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017.

[2] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” International Conference on Learning Representations (ICLR), 2021.

[3] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Technical Report, 2019.

[4] A. Goyal, Y. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 6904–6913, 2017.

[5] K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” Proc. 40th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 311–318, 2002.

[6] R. Vedantam, C. L. Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 4566–4575, 2015.

[7] H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,” Proc. EMNLP-IJCNLP, pp. 5100–5111, 2019.

[8] W. Kim, B. Son, and I. Kim, “ViLT: Vision-and-language transformer without convolution or region supervision,” International Conference on Machine Learning (ICML), 2021.

[9] J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” International Conference on Machine Learning (ICML), 2022.

[10] P. Wang et al., “OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” International Conference on Machine Learning (ICML), 2022.

[11] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” International Conference on Machine Learning (ICML), 2023.

[12] D. Dai et al., “InstructBLIP: Towards general-purpose vision-language models with instruction tuning,” Advances in Neural Information Processing Systems (NeurIPS), 2023.

[13] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in Neural Information Processing Systems (NeurIPS), 2023.

[14] B. Bai et al., “Qwen-VL: A frontier large vision-language model with versatile abilities,” 2023.

[15] Z. Wang et al., “Vision-language models for vision tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5475–5498, 2024.

[16] P. Xu et al., “Multimodal foundation models: From specialists to general-purpose assistants,” ACM Computing Surveys, vol. 57, no. 1, pp. 1–44, 2024.

[17] Y. Tang et al., “Multilingual translation with extensible multilingual pretraining and finetuning,” Proc. 59th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 863–884, 2021

[18] T. Wolf et al., “Transformers: State-of-the-art natural language processing,” Proc. EMNLP System Demonstrations, pp. 38–45, 2020.

[19] M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov, “Label Studio: Data labeling software,” HumanSignal, 2020.

[20] Y. Du et al., “Vision-language pre-training: Basics, recent advances, and future trends,” Foundations and Trends in Computer Graphics and Vision, vol. 18, no. 1–2, pp. 1–214, 2024.

Downloads

Published

2026-08-07

How to Cite

[1]
A. Anas, H. Hazriani, Y. Yuyun, S. Rijal, and T. C. Sharief, “Indonesia–English Bilingual Visual Question Answering Using Partial Fine-Tuning on a ViT-GPT2 Architecture”, JAIC, vol. 10, no. 4, pp. 3241–3252, Aug. 2026.

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.