
MSc, Computing (Software Engineering)
The Open University
Overview
With a focus on Software Engineering, this degree was awarded in June 2026, covering Data Management, Software Development, Project Management and Software Engineering.
A comparative analysis of the accuracy of retrieval-augmented generation based knowledge chatbots and baseline LLM-based knowledge chatbots.
In the field of knowledge chatbots, the issue of accuracy remains a significant challenge, particularly when it comes to fabrications in generated responses, commonly known as “hallucinations”. While previous studies have focused on hallucinations within LLMs, little research has been done to compare the responses of Retrieval-Augmented Generation based (RAG) knowledge chatbots with baseline LLM-based knowledge chatbots. This research aims to compare these two architectures of knowledge chatbots to assist individuals in determining which architecture fits their requirements for accuracy.
Using a post-positivist, quantitative approach, data was collected through an experiment to compare two knowledge chatbot architectures using three sets of three questions that were used to prompt the LLMs to generate a response. The analysis was done using Grice’s Conversational Maxims to test the knowledge boundaries and to determine the truthfulness of the responses and to determine if hallucinations were present, primarily by assessing the responses against the maxim of quality.
Analysis on the responses provided by each architecture found that the RAG-based solution used within this research performed better when it came to assessing the quality maxim (accuracy) but performed worse overall when assessing the quantity and relation maxims. As expected within the experiment, both architectures scored the same number of points when assessing for the maxim of manner as the model version was a controlled variable. Findings from the experimental phase of this research project revealed a significant performance divergence in low-context, high-difficult queries, where the baseline LLM’s quality score plummeted to 4/9, where the RAG architecture maintained a score of 9/9.
The data identified a critical confidence-accuracy gap in baseline LLMs; despite high scores in the maxims of relation and quantity, this architecture produced authoritative hallucinations, including non-existent API endpoints. Conversely, the RAG-based solution demonstrated superior factual robustness by grounding responses in external documentation. It is concluded that while baseline LLMs are suitable for low-stakes conversational tasks, RAG is indispensable for technical environments, such as developer tools and documentation, where mitigating the deceptive authority of hallucinations is a business imperative.
Modules
- Software Engineering
- Software Development
- Data Management
- Project Management
- Information Security