Retrieval-Augmented Local Llama 3.2 for Dynamic Quest Generation and Consistent NPC Personalities in Role-Playing Games
DOI:
https://doi.org/10.30871/jaic.v10i4.13560Keywords:
Llama3.2, NPC Personality, Procedural Narrative, RAG, SLMAbstract
The integration of Large Language Models into game engines is hindered by cloud dependency and high latency. This study proposes a fully localized Retrieval-Augmented Generation (RAG) framework using the Llama 3.2 Small Language Model to generate role-playing game quests and maintain character personalities without parameter fine-tuning. Operating within a C++ environment under strict hardware constraints (primarily CPU-bound), the methodology evaluates three retrieval methods (MiniLM, FastText, and BPE Tokenizer) combined with a memory-efficient JSON Vector Database. System effectiveness was measured using BLEU, ROUGE, and user evaluations. Results show the initial BPE Tokenizer achieved the lowest quest generation time of 152.01 seconds, the fastest average response time of 23.53 seconds, the highest BLEU score of 0.0118, and a peak persona consistency score of 3.70. However, the relatively long response times remain a primary weakness hindering real-time immersion. Furthermore, a critical "Stopping Condition Dilemma" emerged; lacking engine awareness, the model failed to detect narrative conclusions. This caused generation loops that spiked processing times up to 220.59 seconds. Future research must integrate strict, state-based logic triggers from the game engine to prevent context collapse and optimize inference to reduce latency.
Downloads
References
[1] R. Gallotta et al., “Large Language Models and Games: A Survey and Roadmap,” IEEE Trans. Games, pp. 1–18, 2024, doi: 10.1109/TG.2024.3461510.
[2] S. Värtinen, P. Hämäläinen, and C. Guckelsberger, “Generating Role-Playing Game Quests With GPT Language Models,” IEEE Trans. Games, vol. 16, no. 1, pp. 127–139, Mar. 2024, doi: 10.1109/TG.2022.3228480.
[3] S. H. Alavi et al., “Game Plot Design With an LLM-Powered Assistant: An Empirical Study With Game Designers,” IEEE Trans. Games, pp. 1–10, 2026, doi: 10.1109/TG.2026.3663566.
[4] Y. Sun et al., “Bring Game Characters to the Social Space: Developing Storytelling Community Ai Agents Driven by Llms,” 2024. doi: 10.2139/ssrn.4806067.
[5] J. P. W. Hardiman, D. C. Thio, A. Y. Zakiyyah, and Meiliana, “AI-powered dialogues and quests generation in role-playing games using Google’s Gemini and Sentence BERT framework,” Procedia Comput. Sci., vol. 245, pp. 1111–1119, 2024, doi: 10.1016/j.procs.2024.10.340.
[6] M. A. Ferrag, N. Tihanyi, and M. Debbah, “Reasoning beyond limits: Advances and open problems for LLMs,” ICT Express, vol. 11, no. 6, pp. 1054–1096, Dec. 2025, doi: 10.1016/j.icte.2025.09.003.
[7] P. P. Ray and M. P. Pradhan, “An evaluation framework for measuring prompt wise metrics for large language models in resource-constrained edge,” BenchCouncil Trans. Benchmarks Stand. Eval., vol. 5, no. 4, p. 100249, Dec. 2025, doi: 10.1016/j.tbench.2025.100249.
[8] G. Michelet and F. Breitinger, “ChatGPT, Llama, can you write my report? An experiment on assisted digital forensics reports written using (local) large language models,” Forensic Sci. Int. Digit. Investig., vol. 48, p. 301683, Mar. 2024, doi: 10.1016/j.fsidi.2023.301683.
[9] H. T. Kesgin and M. F. Amasyali, “A cyclical loss-based optimization algorithm for pretraining LLMs on noisy data,” Knowl.-Based Syst., vol. 328, p. 114189, Oct. 2025, doi: 10.1016/j.knosys.2025.114189.
[10] D. Rohrschneider, M. Pehlke, U. Handmann, and M. Jansen, “LLM-based JSON Mapping and Blockchain Integration for Digital Product Passports,” Digit. Bus., vol. 6, no. 1, p. 100167, Jun. 2026, doi: 10.1016/j.digbus.2026.100167.
[11] B. Gao et al., “RAQAG : A framework for automatically generating Q&A datasets with retrieval-augmented generation,” Knowl.-Based Syst., vol. 342, p. 115842, Jun. 2026, doi: 10.1016/j.knosys.2026.115842.
[12] L.-C. Chen, M. S. Pardeshi, Y.-X. Liao, and K.-C. Pai, “Application of retrieval-augmented generation for interactive industrial knowledge management via a large language model,” Comput. Stand. Interfaces, vol. 94, p. 103995, Aug. 2025, doi: 10.1016/j.csi.2025.103995.
[13] A. Alansari and H. Luqman, “Large language models hallucination: A comprehensive survey,” Comput. Sci. Rev., vol. 61, p. 100970, Aug. 2026, doi: 10.1016/j.cosrev.2026.100970.
[14] Y. Zhao, Y. He, X. Zhang, and W. Wen, “Combining retrieved with generated contexts via a listwise reranker for open-domain question answering,” Neurocomputing, vol. 670, p. 132577, Mar. 2026, doi: 10.1016/j.neucom.2025.132577.
[15] M. Kořínek and K. Štekerová, “Words Matter: How Prompt Framing Shapes Strategic Behaviour in Large Language Models,” J. Cases Inf. Technol., vol. 28, no. 1, pp. 1–28, Jan. 2026, doi: 10.4018/JCIT.398628.
[16] S. Saleem, M. N. Asim, S. Zulfiqar, and A. Dengel, “The evolution of natural language processing: How prompt optimization and language models are shaping the future,” Comput. Sci. Rev., vol. 61, p. 100938, Aug. 2026, doi: 10.1016/j.cosrev.2026.100938.
[17] X. Li, C. Guo, Q. He, M. Dong, and K. Ota, “PRLoRA: Pyramid-Structured Low-Rank Adaptation Balancing Global Context and Local Precision in Large Language Models,” Knowl.-Based Syst., vol. 343, p. 116045, Jun. 2026, doi: 10.1016/j.knosys.2026.116045.
[18] D. Martens, J. Hinns, C. Dams, M. Vergouwen, and T. Evgeniou, “Tell me a story! Narrative-driven XAI with Large Language Models,” Decis. Support Syst., vol. 191, p. 114402, Apr. 2025, doi: 10.1016/j.dss.2025.114402.
[19] R. Kataishi, “Enhancing Retrieval-Augmented Generation with topic-enriched embeddings: A hybrid approach integrating traditional NLP techniques,” Nat. Lang. Process. J., vol. 14, p. 100200, Mar. 2026, doi: 10.1016/j.nlp.2026.100200.
[20] F. Li, X. Li, S. Wen, H. Huang, and J. Bao, “SAMAC-R3-MED: Semantic alignment and multi-agent collaboration of retriever-reranker-responder models for multimodal engineering documents,” Comput. Ind., vol. 171, p. 104336, Oct. 2025, doi: 10.1016/j.compind.2025.104336.
[21] N. H. Phung, C. T. Nguyen, M.-T. Nguyen, T. H. Nguyen, H. L. Le, and T.-P. Nguyen, “A fine-tuning framework based on question, context, and answer relationships for enhancing legal information retrieval,” Eng. Appl. Artif. Intell., vol. 159, p. 111570, Nov. 2025, doi: 10.1016/j.engappai.2025.111570.
[22] T. Capel and M. Brereton, “What is Human-Centered about Human-Centered AI? A Map of the Research Landscape,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Hamburg Germany: ACM, Apr. 2023, pp. 1–23. doi: 10.1145/3544548.3580959.
[23] N. Pourhaji Aghayengejeh, M. A. Balafar, J. Tanha, and N. Nikzad Khasmakhi, “From embeddings to interpretations: A comprehensive review of language models in clustering,” Comput. Sci. Rev., vol. 61, p. 100974, Aug. 2026, doi: 10.1016/j.cosrev.2026.100974.
[24] E. N. Shaday, V. J. L. Engel, and H. Heryanto, “Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics,” Procedia Comput. Sci., vol. 245, pp. 137–146, 2024, doi: 10.1016/j.procs.2024.10.237.
[25] P. Singh, M. Gaspari, N. Meratnia, and M. Presser, “LLM-based harmonized data ingestion for dataspace: a novel system for automating data ingestion across heterogeneous data sources,” Knowl.-Based Syst., vol. 342, p. 115800, Jun. 2026, doi: 10.1016/j.knosys.2026.115800.
[26] R. Schmitt, “From raw text to fairseq RoBERTa: A modular snakemake-based framework enabling language-specific BPE tokenization,” Softw. Impacts, vol. 27, p. 100824, Apr. 2026, doi: 10.1016/j.simpa.2026.100824.
[27] D. Karapiperis, L. Akritidis, and P. Bozanis, “The impact of fine-tuning on entity resolution: An experimental evaluation,” Knowl.-Based Syst., vol. 338, p. 115427, Apr. 2026, doi: 10.1016/j.knosys.2026.115427.
[28] M. Rehman, A. Petrillo, and K. Awasare, “An LLM-Based System for Accessible and Personalized Scientific Communication,” Procedia Comput. Sci., vol. 274, pp. 939–952, 2025, doi: 10.1016/j.procs.2025.12.092.
[29] L. Attouche et al., “Elimination of annotation dependencies in validation for Modern JSON Schema,” Theor. Comput. Sci., vol. 1063, p. 115645, Feb. 2026, doi: 10.1016/j.tcs.2025.115645.
[30] M. Kamat, J. Jagasia, A. Vaidya, and O. Surve, “Embedding-Based decision support framework for large-scale content analysis,” Knowl.-Based Syst., vol. 332, p. 114926, Jan. 2026, doi: 10.1016/j.knosys.2025.114926.
[31] Y. Xiong, X. Tu, and W. Zhao, “AFR-Rank: An effective and highly efficient LLM-based listwise reranking framework via filtering noise documents,” Inf. Process. Manag., vol. 62, no. 6, p. 104232, Nov. 2025, doi: 10.1016/j.ipm.2025.104232.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Benaya Friyandi Siahaan, Hanny Haryanto

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








