Improving Retrieval-Augmented Generation Grounding Using Document Hierarchy Chunk Graph

Authors

  • Putri Cristin Institut Teknologi Sepuluh Nopember Surabaya
  • Hilmil Pradana Institut Teknologi Sepuluh Nopember

DOI:

https://doi.org/10.30871/jaic.v10i4.13368

Keywords:

Retrieval Augmented Generation(RAG),, Document Hierarchy, Information Retrieval, Grounding, Large Language Models

Abstract

Retrieval-Augmented Generation (RAG) is widely used to enhance question-answering systems across various domains. However, while real-world source documents are inherently structured, conventional RAG approaches primarily rely on semantic similarity between isolated text chunks, which can overlook document hierarchy and limit retrieval effectiveness. To address this issue, this study introduces a document hierarchy-based Chunk Graph approach to improve retrieval grounding in RAG systems. The proposed framework preserves document hierarchy during chunking and models inter-chunk relationships using a weighted graph that combines structural proximity and semantic similarity. The approach was evaluated using the StructuredQA and CUAD benchmark datasets, with performance measured via Precision, Recall, and F1-Score. Experimental results demonstrate that the effectiveness of the Chunk Graph depends heavily on the source document format. On the highly structured StructuredQA dataset, the proposed method successfully connects fragmented information, improving the F1-Score from 47.62% to 50.23%. Conversely, on the CUAD dataset which consists of raw text with implicit hierarchy and no nested structure, the model becomes redundant and does not yield performance gains. These findings conclude that integrating structural-semantic relationships significantly improves context selection, specifically for documents with explicit hierarchical structures.

Downloads

Download data is not yet available.

References

[1] P. Lewis and others, “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems, 2020, pp. 9459–9474. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf

[2] W. X. Zhao et al., “A Survey of Large Language Models,” Front. Comput. Sci., vol. 20, no. 12, p. 2012627, Dec. 2026, doi: 10.1007/s11704-026-60308-3.

[3] Y. Chang and others, “A Survey on Evaluation of Large Language Models,” ACM Trans. Intell. Syst. Technol., vol. 15, no. 3, pp. 1–45, 2024, doi: 10.1145/3641289.

[4] T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto, “Benchmarking Large Language Models for News Summarization,” Trans. Assoc. Comput. Linguist., vol. 12, pp. 39–57, Jan. 2024, doi: 10.1162/tacl_a_00632.

[5] A. Mansurova, A. Tleubayeva, A. Nugumanova, A. Shomanov, and S. E. Seker, “A Systematic Evaluation of Large Language Models and Retrieval-Augmented Generation for the Task of Kazakh Question Answering,” Information, vol. 16, no. 11, p. 943, Oct. 2025, doi: 10.3390/info16110943.

[6] H. Xiong et al., “When Search Engine Services Meet Large Language Models: Visions and Challenges,” IEEE Trans. Serv. Comput., vol. 17, no. 6, pp. 4558–4577, Nov. 2024, doi: 10.1109/TSC.2024.3451185.

[7] L. Huang and others, “A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,” ACM Trans. Inf. Syst., vol. 43, no. 2, pp. 1–55, 2025, doi: 10.1145/3703155.

[8] L. M. Amugongo, P. Mascheroni, S. Brooks, S. Doering, and J. Seidel, “Retrieval augmented generation for large language models in healthcare: A systematic review,” PLOS Digital Health, vol. 4, no. 6, p. e0000877, Jun. 2025, doi: 10.1371/journal.pdig.0000877.

[9] Y. Pu et al., “Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 44, no. 12, pp. 4615–4628, Dec. 2025, doi: 10.1109/TCAD.2025.3568776.

[10] P. Zhao et al., “Retrieval-Augmented Generation for AI-Generated Content: A Survey,” Data Sci. Eng., vol. 11, no. 1, pp. 1–29, Mar. 2026, doi: 10.1007/s41019-025-00335-5.

[11] S. Liu, A. B. McCoy, and A. Wright, “Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines,” J. Am. Med. Inform. Assoc., vol. 32, pp. 605–615, 2025, doi: doi:10.1093/jamia/ocaf008.

[12] V. Karpukhin and others, “Dense Passage Retrieval for Open-Domain Question Answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Association for Computational Linguistics, 2020, pp. 6769–6781. doi: 10.18653/v1/2020.emnlp-main.550.

[13] S. Borgeaud et al., “Improving Language Models by Retrieving from Trillions of Tokens,” in Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., in Proceedings of Machine Learning Research, vol. 162. PMLR, Jun. 2022, pp. 2206–2240. [Online]. Available: https://proceedings.mlr.press/v162/borgeaud22a.html

[14] F. Shi et al., “Large language models can be easily distracted by irrelevant context,” in Proceedings of the 40th International Conference on Machine Learning, in ICML’23. JMLR.org, 2023.

[15] C. A. Gomez-Cabello et al., “Comparative Evaluation of Advanced Chunking for Retrieval-Augmented Generation in Large Language Models for Clinical Decision Support,” Bioengineering, vol. 12, no. 11, p. 1194, Nov. 2025, doi: 10.3390/bioengineering12111194.

[16] H. Brådland, M. Goodwin, P.-A. Andersen, A. S. Nossum, and A. Gupta, “A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking,” in Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025, pp. 170–179. doi: 10.1145/3726302.3729882.

[17] Z. Wang et al., “Document Segmentation Matters for Retrieval-Augmented Generation,” in Findings of the Association for Computational Linguistics: ACL 2025, Stroudsburg, PA, USA: Association for Computational Linguistics, 2025, pp. 8063–8075. doi: 10.18653/v1/2025.findings-acl.422.

[18] M. R. Hajar, E. Utami, and A. Hendi Muhammad, “A Systematic Literature Review of Retrieval-Augmented Generation: Methods, Applications, and Future Research Directions,” Journal of Applied Computer Science and Technology, vol. 6, no. 2, pp. 115–128, Dec. 2025, doi: 10.52158/jacost.v6i2.1170.

[19] X. Zhu, X. Guo, S. Cao, S. Li, and J. Gong, “StructuGraphRAG: Structured Document-Informed Knowledge Graphs for Retrieval-Augmented Generation,” Proceedings of the AAAI Symposium Series, vol. 4, no. 1, pp. 242–251, 2024, doi: 10.1609/aaaiss.v4i1.31798.

[20] D. Edge and others, “From Local to Global: A Graph RAG Approach to Query-Focused Summarization,” arXiv preprint arXiv:2404.16130, 2025, doi: 10.48550/arXiv.2404.16130.

[21] Mozilla AI, “Structured QA,” 2025, Mozilla AI Blog. Accessed: Jan. 14, 2026. [Online]. Available: https://blog.mozilla.ai/structured-question-answering-2/

[22] D. Hendrycks, C. Burns, A. Chen, and S. Ball, “CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review,” arXiv preprint arXiv:2103.06268, 2021, doi: 10.48550/ARXIV.2103.06268.

[23] J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu, “M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation,” in Findings of the Association for Computational Linguistics ACL 2024, Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 2318–2335. doi: 10.18653/v1/2024.findings-acl.137.

[24] Q. A. Yang et al., “Qwen2.5 Technical Report,” ArXiv, vol. abs/2412.15115, 2024, [Online]. Available: https://api.semanticscholar.org/CorpusID:274859421

[25] S. Toro et al., “Dynamic Retrieval Augmented Generation of Ontologies using Artificial Intelligence (DRAGON-AI),” J. Biomed. Semantics, vol. 15, no. 1, p. 19, Oct. 2024, doi: 10.1186/s13326-024-00320-3.

[26] M. Hindi, L. Mohammed, O. Maaz, and A. Alwarafy, “Enhancing the Precision and Interpretability of Retrieval-Augmented Generation (RAG) in Legal Technology: A Survey,” IEEE Access, vol. 13, pp. 46171–46189, 2025, doi: 10.1109/ACCESS.2025.3550145.

[27] K. Rangan and Y. Yin, “A fine-tuning enhanced RAG system with quantized influence measure as AI judge,” Sci. Rep., vol. 14, no. 1, p. 27446, Nov. 2024, doi: 10.1038/s41598-024-79110-x.

[28] M. Abo El-Enen, S. Saad, and T. Nazmy, “A survey on retrieval-augmentation generation (RAG) models for healthcare applications,” Neural Comput. Appl., vol. 37, no. 33, pp. 28191–28267, Nov. 2025, doi: 10.1007/s00521-025-11666-9.

Downloads

Published

2026-08-08

How to Cite

[1]
P. Cristin and H. Pradana, “Improving Retrieval-Augmented Generation Grounding Using Document Hierarchy Chunk Graph”, JAIC, vol. 10, no. 4, pp. 3394–3401, Aug. 2026.

Similar Articles

1 2 3 4 5 > >> 

You may also start an advanced similarity search for this article.