[1] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S. and Kiela, D., et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, Advances in Neural Information Processing Systems, Vol. 33, pp. 9459-9474, 2020.
[2] Guu, K., Lee, K., Tung, Z., Pasupat, P. and Chang, M., “REALM: Retrieval-Augmented Language Model Pre-Training”, Proceedings of the 37th International Conference on Machine Learning, pp. 3929-3938, 2020.
[3] Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D. and Yih, W., “Dense Passage Retrieval for Open-Domain Question Answering”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pp. 6769-6781, 2020.
[4] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L. and Polosukhin, I., “Attention Is All You Need”, Advances in Neural Information Processing Systems, Vol. 30, pp. 5998-6008, 2017.
[5] Devlin, J., Chang, M., Lee, K. and Toutanova, K., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, pp. 4171-4186, 2019.
[6] Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Madotto, A. and Fung, P., “Survey of Hallucination in Natural Language Generation”, ACM Computing Surveys, Vol. 55, No. 12, pp. 1-38, 2023.
[7] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G. and Askell, A., et al., “Language Models are Few-Shot Learners”, Advances in Neural Information Processing Systems, Vol. 33, pp. 1877-1901, 2020.
[8] Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D., “Scaling Laws for Neural Language Models”, arXiv preprint, arXiv:2001.08361, 2020.
[9] Reimers, N. and Gurevych, I., “Sentence-BERT: Sentence Embeddings using Siamese BERTNetworks”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pp. 3982-3992, 2019.
[10] Wang, W., Wei, F., Dong, L., Bao, H., Yang, N. and Zhou, M., “MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-trained Transformers”, Advances in Neural Information Processing Systems, Vol. 33, pp. 5776-5788, 2020.
[11] Johnson, J., Douze, M. and Jégou, H., “Billion-Scale Similarity Search with GPUs”, IEEE Transactions on Big Data, Vol. 7, No. 3, pp. 535-547, 2019.
[12] Thakur, N., Reimers, N., Rücklé, A., Srivastava, A. and Gurevych, I., “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models”, Proceedings of the 35th Conference on Neural Information Processing Systems, Datasets and Benchmarks Track, 2021.
[13] Papineni, K., Roukos, S., Ward, T. and Zhu, W., “BLEU: a Method for Automatic Evaluation of Machine Translation”, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311-318, 2002.
[14] Lin, C., “ROUGE: A Package for Automatic Evaluation of Summaries”, Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, pp. 74-81, 2004.
[15] Izacard, G. and Grave, E., “Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering”, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pp. 874-880, 2021.
[16] Asai, A., Wu, Z., Wang, Y., Sil, A. and Hajishirzi, H., “Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection”, arXiv preprint, arXiv:2310.11511, 2023.
[17] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K. and Ray, A., et al., “Training Language Models to Follow Instructions with Human Feedback”, Advances in Neural Information Processing Systems, Vol. 35, pp. 27730-27744, 2022.
[18] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. and Zhou, D., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, Advances in Neural Information Processing Systems, Vol. 35, pp. 24824-24837, 2022.
[19] Rajpurkar, P., Zhang, J., Lopyrev, K. and Liang, P., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 2383-2392, 2016.
[20] Rajpurkar, P., Jia, R. and Liang, P., “Know What You Don’t Know: Unanswerable Questions for SQuAD”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Vol. 2, pp. 784-789, 2018.
[21] Petroni, F., Piktus, A., Fan, A., Lewis, P., Yazdani, M., De Cao, N., Thorne, J., Jernite, Y., Karpukhin, V. and Maillard, J., et al., “KILT: a Benchmark for Knowledge Intensive Language Tasks”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, pp. 2523-2544, 2021.
[22] Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Dodds, Z.H., DasSarma, N. and Tran-Johnson, E., et al., “Language Models (Mostly) Know What They Know”, arXiv preprint, arXiv:2207.05221, 2022.
[23] Zhang, H., Diao, S., Lin, Y., Fung, Y., Lian, Q., Wang, X., Chen, Y., Ji, H. and Zhang, T., “R-Tuning: Instructing Large Language Models to Say ‘I Don’t Know”’, arXiv preprint, arXiv:2311.09677, 2023.
[24] Chase, H., “LangChain: Building Applications with LLMs through Composability”, GitHub repository, available at: https://github.com/langchain-ai/langchain, 2022.
[25] Qwen Team, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D. and Huang, F., et al., “Qwen2.5 Technical Report”, arXiv preprint, arXiv:2412.15115, 2025.
[26] Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J. and Wang, H., “Retrieval-Augmented Generation for Large Language Models: A Survey”, arXiv preprint, arXiv:2312.10997, 2023.
[27] Sanchez, G., “Harry Potter Books Dataset”, GitHub repository, available at: https://github.com/gastonstat/harry-potter-data, 2024.
[28] vapit, “HarryPotterQA”, Hugging Face Datasets, available at: https://huggingface.co/datasets/vapit/HarryPotterQA, 2024.
[29] McNemar, Q., “Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages”, Psychometrika, Vol. 12, No. 2, pp. 153-157, 1947.
[30] Dietterich, T.G., “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms”, Neural Computation, Vol. 10, No. 7, pp. 1895-1923, 1998.