Retrieval-Augmented Local Llama 3.2 for Dynamic Quest Generation and Consistent NPC Personalities in Role-Playing Games

Authors

  • Benaya Friyandi Siahaan Universitas Dian Nuswantoro
  • Hanny Haryanto Universitas Dian Nuswantoro

DOI:

https://doi.org/10.30871/jaic.v10i4.13560

Keywords:

Llama3.2, NPC Personality, Procedural Narrative, RAG, SLM

Abstract

The integration of Large Language Models into game engines is hindered by cloud dependency and high latency. This study proposes a fully localized Retrieval-Augmented Generation (RAG) framework using the Llama 3.2 Small Language Model to generate role-playing game quests and maintain character personalities without parameter fine-tuning. Operating within a C++ environment under strict hardware constraints (primarily CPU-bound), the methodology evaluates three retrieval methods (MiniLM, FastText, and BPE Tokenizer) combined with a memory-efficient JSON Vector Database. System effectiveness was measured using BLEU, ROUGE, and user evaluations. Results show the initial BPE Tokenizer achieved the lowest quest generation time of 152.01 seconds, the fastest average response time of 23.53 seconds, the highest BLEU score of 0.0118, and a peak persona consistency score of 3.70. However, the relatively long response times remain a primary weakness hindering real-time immersion. Furthermore, a critical "Stopping Condition Dilemma" emerged; lacking engine awareness, the model failed to detect narrative conclusions. This caused generation loops that spiked processing times up to 220.59 seconds. Future research must integrate strict, state-based logic triggers from the game engine to prevent context collapse and optimize inference to reduce latency.

Downloads

Download data is not yet available.

References

[1] R. Gallotta et al., “Large Language Models and Games: A Survey and Roadmap,” IEEE Trans. Games, pp. 1–18, 2024, doi: 10.1109/TG.2024.3461510.

[2] S. Värtinen, P. Hämäläinen, and C. Guckelsberger, “Generating Role-Playing Game Quests With GPT Language Models,” IEEE Trans. Games, vol. 16, no. 1, pp. 127–139, Mar. 2024, doi: 10.1109/TG.2022.3228480.

[3] S. H. Alavi et al., “Game Plot Design With an LLM-Powered Assistant: An Empirical Study With Game Designers,” IEEE Trans. Games, pp. 1–10, 2026, doi: 10.1109/TG.2026.3663566.

[4] Y. Sun et al., “Bring Game Characters to the Social Space: Developing Storytelling Community Ai Agents Driven by Llms,” 2024. doi: 10.2139/ssrn.4806067.

[5] J. P. W. Hardiman, D. C. Thio, A. Y. Zakiyyah, and Meiliana, “AI-powered dialogues and quests generation in role-playing games using Google’s Gemini and Sentence BERT framework,” Procedia Comput. Sci., vol. 245, pp. 1111–1119, 2024, doi: 10.1016/j.procs.2024.10.340.

[6] M. A. Ferrag, N. Tihanyi, and M. Debbah, “Reasoning beyond limits: Advances and open problems for LLMs,” ICT Express, vol. 11, no. 6, pp. 1054–1096, Dec. 2025, doi: 10.1016/j.icte.2025.09.003.

[7] P. P. Ray and M. P. Pradhan, “An evaluation framework for measuring prompt wise metrics for large language models in resource-constrained edge,” BenchCouncil Trans. Benchmarks Stand. Eval., vol. 5, no. 4, p. 100249, Dec. 2025, doi: 10.1016/j.tbench.2025.100249.

[8] G. Michelet and F. Breitinger, “ChatGPT, Llama, can you write my report? An experiment on assisted digital forensics reports written using (local) large language models,” Forensic Sci. Int. Digit. Investig., vol. 48, p. 301683, Mar. 2024, doi: 10.1016/j.fsidi.2023.301683.

[9] H. T. Kesgin and M. F. Amasyali, “A cyclical loss-based optimization algorithm for pretraining LLMs on noisy data,” Knowl.-Based Syst., vol. 328, p. 114189, Oct. 2025, doi: 10.1016/j.knosys.2025.114189.

[10] D. Rohrschneider, M. Pehlke, U. Handmann, and M. Jansen, “LLM-based JSON Mapping and Blockchain Integration for Digital Product Passports,” Digit. Bus., vol. 6, no. 1, p. 100167, Jun. 2026, doi: 10.1016/j.digbus.2026.100167.

[11] B. Gao et al., “RAQAG : A framework for automatically generating Q&A datasets with retrieval-augmented generation,” Knowl.-Based Syst., vol. 342, p. 115842, Jun. 2026, doi: 10.1016/j.knosys.2026.115842.

[12] L.-C. Chen, M. S. Pardeshi, Y.-X. Liao, and K.-C. Pai, “Application of retrieval-augmented generation for interactive industrial knowledge management via a large language model,” Comput. Stand. Interfaces, vol. 94, p. 103995, Aug. 2025, doi: 10.1016/j.csi.2025.103995.

[13] A. Alansari and H. Luqman, “Large language models hallucination: A comprehensive survey,” Comput. Sci. Rev., vol. 61, p. 100970, Aug. 2026, doi: 10.1016/j.cosrev.2026.100970.

[14] Y. Zhao, Y. He, X. Zhang, and W. Wen, “Combining retrieved with generated contexts via a listwise reranker for open-domain question answering,” Neurocomputing, vol. 670, p. 132577, Mar. 2026, doi: 10.1016/j.neucom.2025.132577.

[15] M. Kořínek and K. Štekerová, “Words Matter: How Prompt Framing Shapes Strategic Behaviour in Large Language Models,” J. Cases Inf. Technol., vol. 28, no. 1, pp. 1–28, Jan. 2026, doi: 10.4018/JCIT.398628.

[16] S. Saleem, M. N. Asim, S. Zulfiqar, and A. Dengel, “The evolution of natural language processing: How prompt optimization and language models are shaping the future,” Comput. Sci. Rev., vol. 61, p. 100938, Aug. 2026, doi: 10.1016/j.cosrev.2026.100938.

[17] X. Li, C. Guo, Q. He, M. Dong, and K. Ota, “PRLoRA: Pyramid-Structured Low-Rank Adaptation Balancing Global Context and Local Precision in Large Language Models,” Knowl.-Based Syst., vol. 343, p. 116045, Jun. 2026, doi: 10.1016/j.knosys.2026.116045.

[18] D. Martens, J. Hinns, C. Dams, M. Vergouwen, and T. Evgeniou, “Tell me a story! Narrative-driven XAI with Large Language Models,” Decis. Support Syst., vol. 191, p. 114402, Apr. 2025, doi: 10.1016/j.dss.2025.114402.

[19] R. Kataishi, “Enhancing Retrieval-Augmented Generation with topic-enriched embeddings: A hybrid approach integrating traditional NLP techniques,” Nat. Lang. Process. J., vol. 14, p. 100200, Mar. 2026, doi: 10.1016/j.nlp.2026.100200.

[20] F. Li, X. Li, S. Wen, H. Huang, and J. Bao, “SAMAC-R3-MED: Semantic alignment and multi-agent collaboration of retriever-reranker-responder models for multimodal engineering documents,” Comput. Ind., vol. 171, p. 104336, Oct. 2025, doi: 10.1016/j.compind.2025.104336.

[21] N. H. Phung, C. T. Nguyen, M.-T. Nguyen, T. H. Nguyen, H. L. Le, and T.-P. Nguyen, “A fine-tuning framework based on question, context, and answer relationships for enhancing legal information retrieval,” Eng. Appl. Artif. Intell., vol. 159, p. 111570, Nov. 2025, doi: 10.1016/j.engappai.2025.111570.

[22] T. Capel and M. Brereton, “What is Human-Centered about Human-Centered AI? A Map of the Research Landscape,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Hamburg Germany: ACM, Apr. 2023, pp. 1–23. doi: 10.1145/3544548.3580959.

[23] N. Pourhaji Aghayengejeh, M. A. Balafar, J. Tanha, and N. Nikzad Khasmakhi, “From embeddings to interpretations: A comprehensive review of language models in clustering,” Comput. Sci. Rev., vol. 61, p. 100974, Aug. 2026, doi: 10.1016/j.cosrev.2026.100974.

[24] E. N. Shaday, V. J. L. Engel, and H. Heryanto, “Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics,” Procedia Comput. Sci., vol. 245, pp. 137–146, 2024, doi: 10.1016/j.procs.2024.10.237.

[25] P. Singh, M. Gaspari, N. Meratnia, and M. Presser, “LLM-based harmonized data ingestion for dataspace: a novel system for automating data ingestion across heterogeneous data sources,” Knowl.-Based Syst., vol. 342, p. 115800, Jun. 2026, doi: 10.1016/j.knosys.2026.115800.

[26] R. Schmitt, “From raw text to fairseq RoBERTa: A modular snakemake-based framework enabling language-specific BPE tokenization,” Softw. Impacts, vol. 27, p. 100824, Apr. 2026, doi: 10.1016/j.simpa.2026.100824.

[27] D. Karapiperis, L. Akritidis, and P. Bozanis, “The impact of fine-tuning on entity resolution: An experimental evaluation,” Knowl.-Based Syst., vol. 338, p. 115427, Apr. 2026, doi: 10.1016/j.knosys.2026.115427.

[28] M. Rehman, A. Petrillo, and K. Awasare, “An LLM-Based System for Accessible and Personalized Scientific Communication,” Procedia Comput. Sci., vol. 274, pp. 939–952, 2025, doi: 10.1016/j.procs.2025.12.092.

[29] L. Attouche et al., “Elimination of annotation dependencies in validation for Modern JSON Schema,” Theor. Comput. Sci., vol. 1063, p. 115645, Feb. 2026, doi: 10.1016/j.tcs.2025.115645.

[30] M. Kamat, J. Jagasia, A. Vaidya, and O. Surve, “Embedding-Based decision support framework for large-scale content analysis,” Knowl.-Based Syst., vol. 332, p. 114926, Jan. 2026, doi: 10.1016/j.knosys.2025.114926.

[31] Y. Xiong, X. Tu, and W. Zhao, “AFR-Rank: An effective and highly efficient LLM-based listwise reranking framework via filtering noise documents,” Inf. Process. Manag., vol. 62, no. 6, p. 104232, Nov. 2025, doi: 10.1016/j.ipm.2025.104232.

Downloads

Published

2026-08-10

How to Cite

[1]
B. F. Siahaan and H. Haryanto, “Retrieval-Augmented Local Llama 3.2 for Dynamic Quest Generation and Consistent NPC Personalities in Role-Playing Games”, JAIC, vol. 10, no. 4, pp. 3567–3575, Aug. 2026.

Similar Articles

1 2 > >> 

You may also start an advanced similarity search for this article.