Indonesia–English Bilingual Visual Question Answering Using Partial Fine-Tuning on a ViT-GPT2 Architecture
DOI:
https://doi.org/10.30871/jaic.v10i4.13336Keywords:
Visual Question Answering, Vision-Language Model, Vision Transformer, GPT-2, Fine-TuningAbstract
Visual Question Answering (VQA) is a multimodal task that integrates visual understanding and natural language processing to generate answers based on information contained in an image. Most existing VQA research focuses on the English language and general-domain datasets, limiting its applicability to bilingual environments and domain-specific scenarios. This study proposes an Indonesia–English bilingual VQA model based on the VisionEncoderDecoderModel architecture, which combines a Vision Transformer (ViT) as the visual encoder and an Indonesian GPT-2 model as the language decoder. Bilingual capability is achieved through the introduction of special language tokens, <id> and <en>. The model is trained using a combination of the bilingual VQAv2 and bilingual LosariVQAv1 datasets, representing general-domain and local tourism-domain knowledge, respectively. Four fine-tuning strategies are evaluated: Encoder Freeze, Full Fine-Tuning, Partial-4, and Partial-6. Experimental results show that the Partial-6 strategy achieves the best performance on LosariVQAv1, obtaining an Exact Match score of 26.67%, a BLEU score of 26.80%, and a CIDEr score of 292.67, while maintaining competitive performance on VQAv2 with an Exact Match score of 41.26%. Cross-language evaluation reveals only a small performance gap between Indonesian and English. The findings indicate that partial fine-tuning provides a better balance between generalization capability and domain adaptation than the other fine-tuning strategies evaluated in this study.
Downloads
References
[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017.
[2] A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” International Conference on Learning Representations (ICLR), 2021.
[3] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Technical Report, 2019.
[4] A. Goyal, Y. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 6904–6913, 2017.
[5] K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” Proc. 40th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 311–318, 2002.
[6] R. Vedantam, C. L. Zitnick, and D. Parikh, “CIDEr: Consensus-based image description evaluation,” Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), pp. 4566–4575, 2015.
[7] H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,” Proc. EMNLP-IJCNLP, pp. 5100–5111, 2019.
[8] W. Kim, B. Son, and I. Kim, “ViLT: Vision-and-language transformer without convolution or region supervision,” International Conference on Machine Learning (ICML), 2021.
[9] J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” International Conference on Machine Learning (ICML), 2022.
[10] P. Wang et al., “OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” International Conference on Machine Learning (ICML), 2022.
[11] J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” International Conference on Machine Learning (ICML), 2023.
[12] D. Dai et al., “InstructBLIP: Towards general-purpose vision-language models with instruction tuning,” Advances in Neural Information Processing Systems (NeurIPS), 2023.
[13] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in Neural Information Processing Systems (NeurIPS), 2023.
[14] B. Bai et al., “Qwen-VL: A frontier large vision-language model with versatile abilities,” 2023.
[15] Z. Wang et al., “Vision-language models for vision tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 8, pp. 5475–5498, 2024.
[16] P. Xu et al., “Multimodal foundation models: From specialists to general-purpose assistants,” ACM Computing Surveys, vol. 57, no. 1, pp. 1–44, 2024.
[17] Y. Tang et al., “Multilingual translation with extensible multilingual pretraining and finetuning,” Proc. 59th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 863–884, 2021
[18] T. Wolf et al., “Transformers: State-of-the-art natural language processing,” Proc. EMNLP System Demonstrations, pp. 38–45, 2020.
[19] M. Tkachenko, M. Malyuk, A. Holmanyuk, and N. Liubimov, “Label Studio: Data labeling software,” HumanSignal, 2020.
[20] Y. Du et al., “Vision-language pre-training: Basics, recent advances, and future trends,” Foundations and Trends in Computer Graphics and Vision, vol. 18, no. 1–2, pp. 1–214, 2024.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Anas Anas, Hazriani Hazriani, Yuyun Yuyun, Syamsul Rijal, Tirta Chiantalia Sharief

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








