The Power of Noise: Redefining Retrieval for RAG Systems
Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, Fabrizio Silvestri
Introduction
Large Language Models (LLMs) (Brown et al. 2020) have demonstrated unprecedented proficiency in various tasks, ranging from text generation and complex question answering (Beeching et al. 2023), to information retrieval (IR) tasks (Kenton and Toutanova 2019; Yates et al. 2021). However, LLMs have limitations in the handling of long contexts (Vaswani et al. 2017), a constraint that leads to an increased reliance on their pre-trained knowledge. This limitation not only confines their ability to effectively manage extended discourse, such as in books or long conversations, but also increases the probability of generating hallucinations, instances for which the model produces factually incorrect or nonsensical information (Roller et al. 2021). To improve the accuracy of responses generated by LLMs, Retrieval-Augmented Generation (RAG) has emerged as a promising solution (Lewis et al. 2020). RAG is primarily designed to improve factual accuracy by providing the model access to auxiliary information, thereby augmenting the original prompt with information not necessarily memorized in the LLM. A key benefit of this approach is that it helps ground the prompt with relevant information that might help the LLM generate more accurate answers at inference time. At their core, RAG systems consist of two fundamental components: a retriever and a generator. The retriever is responsible for invoking an external IR system (dense and/or sparse) and feeding the selected results to a generator component.
This study focuses on the IR aspect of RAG, posing the following research question: “What characteristics are desirable in a retriever to optimize prompt construction for RAG systems? Are current retrievers ideal?". We focus on the three main types of documents (or passages We interchangeably use here the terms “passage” or “document” to represent the indexing/retrieval unit of the IR system.) that a retriever can return: relevant, distracting, and random. Relevant documents contain pertinent information that either directly answers or might inform the query. Distracting documents, while not directly answering the query, are semantically or contextually linked to the topic. For instance, if one asks for the color of Napoléon’s horse, a passage describing the color of Joséphine de Beauharnais’ (Napoléon’s first wife) horse, while not containing the right information, would be highly related. Random documents have no relation whatsoever to the query and can be seen as a kind of informational noise within the retrieval process. One of the key goals of our study is to determine the role of each type of document and the relative value they bring to the LLM effectiveness. In particular, we verify whether there is a need to revisit some of the commonly accepted assumptions in IR systems when used in the context of LLMs. The main contributions of our work are the following:
We conduct the first comprehensive study examining the impact of the type of retrieved documents in RAG on the LLM effectiveness.
We propose retrieval RAG heuristics that leverage the unexpected results of this study.
We release all associated code and data to the community to encourage further research.
Related Works
The inception of the modern LLM era can be traced back to the seminal paper titled “Attention Is All You Need" (Vaswani et al. 2017). This work introduced the transformer architecture, a framework that adopts an attention mechanism instead of recurrent layers, enabling the model to capture global dependencies within the data. The following year, BERT (Bidirectional Encoder Representations from Transformers) (Kenton and Toutanova 2019) offered a significant improvement over the state-of-the-art via a novel bidirectional, unsupervised language representation. The evolution of transformer-based models continued with the development of the Generative Pre-trained Transformer (GPT) (Radford et al. 2018). Its successor, GPT-2 (Radford et al. 2019), expanded upon this foundation with a larger scale model and demonstrated improved performance across a variety of language tasks without task-specific training. The subsequent iteration, GPT-3 (Brown et al. 2020), represented a further enhancement in model scale and capabilities, particularly in the realm of few-shot learning. Finally, recent times have seen a surge in the production of large, publicly available language models. Several actors have released their models, most notably, Llama (Touvron et al. 2023a; Touvron et al. 2023b), Falcon (Almazrouei et al. 2023), Mosaic MPT (Team et al. 2023), and Phi (Li et al. 2023; Javaheripi et al. 2023). There are also versions of these models that have been fine-tuned on specific languages (Garrachonr 2023; jphme 2023; Cui et al. 2023; Santilli and Rodolà 2023; Bacciu et al. 2023b). The proliferation and quality of these models are expanding the range of tasks and the vision they address (Wang et al. 2024; Xie et al. 2024; Tolomei et al. 2023).
2. Information Retrieval
Foundational information retrieval methodologies, such as the Vector Space Model and the TF-IDF scoring (Salton and McGill 1983) introduced in the 1980s are the basis for quantifying textual similarity. These retrieval methods are characterized by their use of high-dimensional and sparse feature vectors and have been essential in developing a full generation of IR systems. BM25 represents the most famous current iteration (Robertson et al. 2009). A significant evolution in IR is the introduction of dense retrievers, which emerged from advancements in deep learning; they utilize low-dimensional dense vectors for textual representation, and allow to capture semantic relationships. This is in contrast to traditional IR methods (referred to as sparse in opposition to dense), which typically rely on lexical match and struggle with semantic match (Manning et al. 2008). In the last few years, dense methods such as DPR (Karpukhin et al. 2020a) and others (Khattab and Zaharia 2020; Izacard et al. 2021) have demonstrated that they can compete with sparse methods.
3. Retrieve and Generate
RAG introduces a new approach in AI, combining the strengths of both retrieval-based and generative models. The concept of RAG was coined and popularized in (Lewis et al. 2020), which introduced a model that combines a dense passage retriever with a sequence-to-sequence model, demonstrating substantial improvements in knowledge-intensive tasks. Similar methods/variations have also been proposed concurrently or soon after, such as (Guu et al. 2020; Borgeaud et al. 2022; Asai et al. 2023; Ke et al. 2024; Bacciu et al. 2023a); see (Mialon et al. 2023) for a survey on augmented language models. Researchers and practitioners have recently started to explore these RAG systems’ inner workings. Notably, (Sauchuk et al. 2022; Trappolini et al. 2023) analyzed the impact of different types of documents on cascading IR/NLP systems. Other works have tried to study how attentive transformers are to their input (Sun et al. 2021; Khandelwal et al. 2018; Liu et al. 2023; Ram et al. 2023; Lu et al. 2022). (BehnamGhader et al. 2023) studied the effect of the retriever’s similarity metric, which was found to be insufficient for reasoning. In (Xie et al. 2023; Koopman and Zuccon 2023), authors analyzed LLM’s receptiveness to external evidence against internal memory. In (Zuccon et al. 2023), they test the model’s (in)ability to ground references.
In this paper, we want to provide the first comprehensive analysis of the implications of using a retriever module in a RAG system, studying the impact of several key factors, like the type, number, and position of documents that should augment the prompt to the LLM.
RAG
In this paper, we explore the application of RAG in the context of Question Answering, arguably its most popular application.
Open-domain Question Answering (OpenQA) refers to the task of developing systems capable of providing accurate and contextually relevant answers to a broad range of questions posed in natural language without limitations to specific domains or predefined datasets. In general, we want to find an answer to a query . To do so, we draw information from a corpus of documents , which is usually assumed to be large in size. A prevalent approach for this task involves a two-step architecture, typically comprising a retriever and a reasoner (typically a generator). This methodology addresses the inherent complexities of OpenQA by dividing the process into distinct phases: first finding the appropriate set of documents that can potentially address the query and then synthesizing an answer, which can be consumed by the user of the QA system.
2. Retriever
The retriever plays a critical role in the OpenQA task. Its goal is to find a sufficiently small subset of documents to allow the reasoner to answer the query correctly. Among the various retrieval methodologies, the use of a dense retriever has gained prominence due to its effectiveness in handling semantic matches. Dense retrieval requires transforming textual data into vector representations, which is typically achieved with a neural network, often a transformer-based encoder, like BERT (Kenton and Toutanova 2019). The dense retriever processes both the query and potential source documents to generate corresponding embeddings for the query and for each document . The embedding process can be represented as:
where and are neural network-based encoders, potentially sharing weights or architecture, designed to map the textual data into a vector space. Once the embeddings are generated, the retrieval process involves computing the similarity between the query embedding and each document embedding. The most common approach is to use dot product (Karpukhin et al. 2020b), defined as: . This score quantifies the relevance of each document to the query by measuring their similarity in the embedded vector space, with higher scores indicating greater relevance. According to these scores, the top-ranked documents are selected for further processing in the generator component.
3. Reasoner
The second step involves a generator component in charge of synthesizing an answer, typically implemented via an LLM. Generative language models operate by predicting the probability distribution of the next token, given the previous tokens. For a given sequence of words , a generative language model aims to maximize the likelihood of this sequence, expressed using the chain rule of probability:
where is the conditional probability of the word given the preceding sequence of words . In RAG, the generative language model takes a query and the retrieved documents as input and generates a response by sequentially predicting the next token in the sequence. More formally,
where is the retrieval component that provides a (truncated) probability distribution for the top-scoring documents, and is a probability distribution parameterized by that generates a current token based on the previously generated tokens, the query, and the retrieved document; this role is filled by the LLM. In the case of dense retrieval, the probability distribution for the top-scoring documents may assume a functional form of the kind . Given our formalization of the RAG task, we notice how the generative component depends on a given text, that is the query, and a dynamic text, that is the set of retrieved documents. We study in the next two sections the impact of changing the set of retrieved documents on the generator and, consequently, the whole end-to-end system. In particular, we aim to find the best set of documents that a retriever should feed the generator to maximize the system’s effectiveness.
Experimental Methodology
In this section, we detail the experimental framework. We start by describing the data used in the experiments and then discuss the type of documents that a retriever can return and pass to the LLM.
The Natural Questions (NQ) dataset (Kwiatkowski et al. 2019) is a large-scale collection of real-world queries derived from Google search data. Each entry in the dataset consists of a user query and the corresponding Wikipedia page containing the answer. The NQ-open dataset (Lee et al. 2019), a subset of the NQ dataset, differs by removing the restriction of linking answers to specific Wikipedia passages, thereby mimicking a more general information retrieval scenario similar to web searches. This open-domain nature significantly impacts our experimental design, particularly in the selection and categorization of documents. Following the methodology of Lee et al. 2019, our primary source for answering queries is the English Wikipedia dump as of 20 December 2018. Consistently with the Dense Passage Retrieval (DPR) approach (Karpukhin et al. 2020b), each Wikipedia article in this dump was segmented into non-overlapping passages of words. A significant challenge in open-domain question answering is the potential temporal mismatch between the Wikipedia dump and the question-answer pairs in the dataset, which can lead to missing answers in the dataset, as highlighted in the AmbigQA study (Min et al. 2020). To mitigate this, we integrated the gold documents from the original NQ dataset into our Wikipedia document set. Given the open-domain nature of our task, there may be additional documents relevant to the query, i.e., containing the answer, but we will not consider them as gold. The final dataset comprises documents, with queries in the train set and in the test set.
2. Types of Documents
In our study, we categorize documents into four distinct types, each represented by a unique symbol, based on their relevance and relationship to the queries:
The gold document, identified by , refers to the original context in the NQ dataset, specifically the passage of a Wikipedia page containing the answer and contextually relevant to a given query.
Denoted by , relevant documents are passages that, akin to the gold document, contain the correct answer and are contextually useful for answering the query. They provide additional sources of information that are correct and pertinent to the query. Notably, the gold document is a relevant document.
Symbolized by , distracting documents are semantically similar to the query but do not contain the correct answer. They serve a crucial role in evaluating the generator’s proficiency in discerning between relevant and non-relevant information. In practice, these are the top-scoring retrieved documents that are not relevant.
Indicated by \dice6, random documents are neither related to the query nor contain the answer. They are instrumental in assessing the model’s ability to handle completely unrelated information. In practice, in our tests, we will randomly sample these documents from the corpus. In our analysis, the entire set of documents fetched by the retriever is represented by the symbol . This possibly encompasses all document types — gold, relevant, distracting, or random — and serves to discuss the retrieval output in a generalized manner without specifying individual document categories.
3. Document Retrieval
Our methodology utilizes a two-step approach in line with a typical RAG setting, as explained in Section 3.2. As the first component, our experiments use Contriever (Izacard et al. 2021), a BERT-based dense retriever, as the default retriever. It is trained without supervision using a contrastive loss. To enhance the efficiency of similarity searches within our corpus, comprising about 21 million documents, we also employ the FAISS IndexFlatIP indexing system (Douze et al. 2024). The embedding of each document and query is obtained by averaging the hidden state of the last layer of the model.
4. LLM Input
Upon receiving a query, the retriever selects the top- documents from the corpus according to a given similarity measure. These documents, in conjunction with the task instruction and the query, constitute the input for the LLM to generate a response. The NQ-open dataset was structured to include only those queries whose answers consist of no more than five tokens (Lee et al. 2019). Consequently, the LLM is tasked with extracting a query response, confined to a maximum of five tokens, from the provided documents. The input is encoded into a prompt, whose template is shown in Figure 1, beginning with the task instruction, presented in italics for clarity. This is followed by the context, which comprises the selected documents followed by the query string. This prompt design aligns with the methodological approach outlined in (Liu et al. 2023).
While the composition of the context will vary according to the single experiment, the instruction will always be placed at the beginning of the prompt and the query always at the end.
5. LLMs Tested
We consider several LLMs in our experiments. Consistently across all models, we adopt a greedy generation approach with a maximum response length of 15 tokens. Acknowledging the constraints imposed by memory and computational resources, we have implemented a model quantization strategy, reducing all models to a 4-bit representation. Besides the above prompt, the models are not provided with additional exemplars for few-shot learning, which, while of interest, is outside the scope of this paper. We conduct tests on both the base and the instruct versions of the LLMs. However, we only report on the latter, as while the behavior is consistent across both, the instruct versions demonstrate superior performance.
Llama2. The 7B parameters version of the Llama2 family (Touvron et al. 2023b) shows state-of-the-art performance on most downstream tasks compared to models of the same size. It was trained with a 4096 tokens context window and uses multi-query attention (Shazeer 2019).
Falcon. Falcon 7B, the smallest model of the Falcon series, (Almazrouei et al. 2023) was trained on the RefinedWeb dataset (Penedo et al. 2023), a large, filtered, and deduplicated corpus. Similarly to Llama2, it uses multi-query attention, with a context length of 2048 tokens.
Phi-2. This is the smallest model used in this work (2.7B parameters). Despite its modest size, it achieves performance comparable to the other models (Li et al. 2023; Javaheripi et al. 2023), thanks to its pre-training on “textbook-quality” data. It has a context window of 2048 tokens.
MPT. This 7B parameters model uses ALiBi attention (Press et al. 2022; Team et al. 2023) for a virtually unlimited context length. In our experiments, to leverage the model’s full potential, we set the limit to 2048 tokens, i.e., the same used for the model’s pre-training.
6. Accuracy
The NQ-open dataset allows a range of potential answers for each query. Frequently, these answers are different variants of the same concept (e.g., “President D. Roosevelt” or “President Roosevelt”), while in some cases, a single query may accept multiple distinct correct answers. To evaluate the accuracy of responses generated by LLMs, we use an assessment technique in line with (Kandpal et al. 2023; Liu et al. 2023). This methodology examines whether at least one of the predefined correct answers is contained within the response produced by the LLM. We measure the correctness of the LLM’s responses as either accurate or inaccurate based on the presence of the answer in a binary fashion. Nevertheless, this evaluation strategy is not without challenges. A principal issue arises in determining response correctness, particularly in instances involving date representations or varying phrasings conveying identical meanings. For example, if the LLM generates “Roosevelt” in response to a query where the established correct answer is “President Roosevelt”, the response would be deemed incorrect under our current evaluation schema. Recognizing this limitation, we acknowledge the necessity for a more advanced analysis of answer variations, which we leave to future research.
Results
Studying the characteristics of optimal prompts for RAG systems corresponds to answering our research question (RQ): "What characteristics are desirable in a retriever to optimize prompt construction for RAG systems in order to increase the LLM effectiveness?". More specifically, we focus on three essential elements of the configuration: type, number, and positioning of the documents, and for each, we test various prompt combinations. To facilitate the understanding of our experimental setup, we employ a streamlined schema for representing the composition of prompts via the following symbols: [I, , , , \dice6, Q]. The task instruction (I) and the query (Q) are consistently positioned at the beginning and end, respectively. The middle section varies and represents different contextual elements - in this instance, these are gold, relevant, distracting, and random, appearing in that specific sequence. Additionally, the number of contextual documents is a variable in its own right and will be reported in the results tables below.
In our first set of experiments, we use a selection of 10K queries from the training set of the NQ-open dataset and assume an oracle setup in which the gold document for the query is known. To this effect, we add to the gold document a set of distracting documents, i.e., documents with high retrieval scores but not containing the answer, in order to measure their impact on the system; schematically [I, , , Q]. Figure 2 shows an example of this setup’s visualization. Results of this experiment are shown in Table 1 (far, mid, and near relate to the distance between the gold document and the query; more details in the following sub-section). A critical observation emerging from this analysis is a clear pattern of progressive accuracy degradation as the number of distracting documents included in the context increases. This was observed across all LLMs, with accuracy deteriorating by more than 0.38 () in some cases. Even more importantly, adding just one distracting document causes a sharp reduction in accuracy, with peaks of 0.24 (), as can be seen by comparing the row with distracting documents (only gold scenario, as seen in Figure 1) with that of distracting document. This experiment highlights a critical issue for RAG systems, particularly in real-world IR settings where related but non-answer-containing documents are commonplace. Our empirical analysis suggests that introducing semantically aligned yet non-relevant documents adds a layer of complexity, potentially misguiding LLMs away from the correct response. A visual explanation can be seen in Figure 3, which illustrates the attention scores within the prompt’s context for a specific example in which the LLM incorrectly answers. This figure highlights the model’s disproportionate focus on a distracting document (leftmost) at the expense of the gold document (rightmost), likely contributing to the erroneous response. Note that for consistency of results across LLMs, we need to account for their various input token capabilities: Llama2 can process up to 4096 tokens, but other models are limited to 2048 tokens. This led to the exclusion of evaluations with a higher number of distracting documents (namely greater than 10) as reflected by the empty values in the tables.
In addition, we wanted to verify that our results were not overly dependent on the type of dense retrieval system we used. We wanted, in particular, to check whether another dense retriever specifically trained on “hard negatives" would better distinguish between directly relevant and distracting documents, potentially leading to different results. To explore this hypothesis, we used ADORE (Zhan et al. 2021), a state-of-the-art retriever trained with “dynamic hard negatives”, to select the distracting documents. In scenarios with 1, 2, and 4 distracting documents in the [I, , , Q] setting with Llama2, we obtain an accuracy of 0.4068, 0.3815, and 0.3626, respectively. This is significantly lower than the baseline accuracy of 0.5642, where no distracting documents were included, and than the results obtained with Contriever in the same settings. We conclude from this that distinguishing between relevant and distracting information is a hard problem that cannot be mitigated simply by changing the dense retrieval method at this stage.
2. Impact of Gold Positioning
We conduct here another experiment where we systematically shift the position of the gold document within the context to study its impact on the model’s effectiveness. We define the positions of the gold document as follows:
Near: placed adjacent to the query in the prompt [I, , , Q] (as in Figure 2)
Mid: inserted in the middle of the context [I, , , , Q]
Far: positioned as far as possible from the query in the context [I, , , Q]
Results in these settings partially corroborate evidence from (Liu et al. 2023). The accuracy is higher when the gold document is near the query, lower when the gold document is furthest from it, and lowest when the gold document is placed in the middle of the context. For instance, Llama2, with 18 distracting documents, reaches an accuracy of 0.37, 0.23, and 0.17, respectively. These results are consistent across all models tested in the setting with distracting documents.
3. Impact of Noise
We devise an additional experimental setting aimed at evaluating the robustness of the RAG system against noise. To this effect, we take the gold document and add to it a certain number of documents picked at random from the corpus; see an example in Figure 4. Against our expectations, the performance does not deteriorate in the presence of noise, as can be seen in Table 2. Instead, we observe an improvement in performance under the best-performing setting (near [I, \dice6, , Q]), with an improvement of 0.08 () in the case of MPT. Furthermore, we observe that different models exhibit distinct behaviors. Both Llama2 and Phi-2 showed improvements in this setting when the noise is introduced furthest from the query. However, when the noise is positioned in the far [I, , \dice6, Q] and mid [I, \dice6, , \dice6, Q] settings, these models exhibit a decline in performance. Notably, this performance degradation is much less accentuated when compared to the earlier setting with distracting documents. This suggests that while Llama2 and Phi-2 can effectively handle noise far from the query, their ability to sift through irrelevant information diminishes as the noise is placed closer to it. The MPT model presented a unique response; it showed an improvement in performance under all settings. Standing out from the rest, the Falcon model did not exhibit an improvement in performance as observed in other models with the introduction of noise. Peculiarly enough, Falcon and Llama2 do not consistently exhibit a “lost in the middle” phenomenon, having in some instances better accuracy in the mid than far setting, for instance, in the case with noisy documents added.
4. RAG in Practice
To address our primary Research Question (RQ) about the characteristics of an effective RAG retriever, and following the results reported above, we now consider a more realistic scenario than an oracle setup. Namely, given a query, we retrieve a set of documents that can be either relevant or distracting. We then add random documents to this set of retrieved ones, schematically: [I, \dice6, , Q]. For this second set of experiments, we use the test set of the NQ-open dataset. Results for this experiment, using Llama2, can be seen on the left side of Table 3. These results show that, regardless of the number of retrieved documents, adding random documents up until the context length is filled is almost always beneficial, with gains in terms of accuracy up to 0.07 () in the case of 4 retrieved documents.
In an effort to validate our initial observations, we replicate our experiment using a sparse retrieval approach, specifically BM25. The corresponding results are outlined in the right section of Table 3. Consistent with earlier findings, we observe that including random documents leads to an improvement in the effectiveness of the LLM. Notably, the use of BM25 yields an average increase in accuracy of 3-4 percentage points. This improvement is attributed to the quality of documents retrieved by BM25. We quantitatively evaluate the effectiveness of the retrieval methods by computing the top- accuracy for varying numbers of retrieved documents. Note that this heuristic, while indicative, does not capture the full spectrum of relevance. Our evaluation, based on the presence of correct answers within documents, might overlook the context-specific relevance due to potential lexical matches of the answer string in documents. Despite this limitation, this method aligns with established computational practices in literature (Karpukhin et al. 2020a; Izacard et al. 2021). In our analysis, BM25 demonstrated higher relative top- accuracy (0.2966, 0.4105, 0.5237, 0.6663 for ) compared to those of Contriever (0.2502, 0.3569, 0.4784, 0.6085 for the same ), underscoring its effectiveness in retrieving more relevant documents in our experimental setup.
4.2. Increasing The Randomness
While our previous experiments show the benefits of adding random documents, one might argue that these documents are not totally random as they originate from the same corpus (Wikipedia) and that they might help the LLM answer in a fashion that is consistent with the corpus. For this reason, we carry out another experiment in which random documents are drawn from a drastically different corpus in terms of tone and style, namely Reddit Webis-TLDR-17 dataset (Völske et al. 2017). The results are outlined on the left of Table 4. The inclusion of documents from the Reddit corpus not only maintains the observed increase in accuracy but even enhances it, with an improvement of 0.023 ( accuracy) when compared to the previous best score. Pushing the randomness even further, we carry out another test where we consider nonsensical sentences made up of random words as random documents. Remarkably, even in this scenario, we observe a performance improvement when compared to the base case of Wikipedia random documents, as shown in the right side of Table 4.
4.3. Falcon
As shown in Table 2, Falcon does not reach the same performance increase when random documents are added to the gold document [I, \dice6, , Q]. Accordingly, we want to verify whether it behaves differently when adding retrieved rather than gold documents. We find that the addition of random documents on top of retrieved documents [I, \dice6, , Q] does improve the effectiveness of Falcon; see detailed results in Table 5. These results are in contrast with the ones obtained in the oracle setting, where Falcon was robust to noise. This new finding further validates our experimental evidence, namely that, outside the oracle setting, all the tested models show an improvement when a certain amount of noise is added.
5. Retriever Trade-Off
The experimental evidence detailed above not only contradicts the common perception that semantically close documents are helpful for LLMs but also highlights the need for a delicate balance between relevant and random documents. When arranged as described, random documents seem to exert a positive influence on LLM accuracy. However, for the LLM to generate accurate answers, some degree of relevant information must exist in the context. On the other hand, an overabundance of retrieved documents increases the likelihood of including distracting and non-relevant information, leading to a sharp decline in performance. While establishing a formal or comprehensive theory behind these findings remains an open research challenge, we can still infer that there seems to be a trade-off between the number of relevant and totally irrelevant documents. More specifically, we observed that the best effectiveness is achieved when a minimal set of documents is initially retrieved and then supplemented with random documents until the context limit is reached. For the queries examined in this study, retrieving between 3 and 5 documents is the most effective choice. Adding more increases the risk of including too many distracting, thus counterproductive, documents. We argue here that there is a pressing need for further research towards investigating how these initial findings can be exploited. More importantly, it is evident that we have yet to refine our understanding of the retriever’s role within a RAG system.
We cannot close this paper without attempting to explain the results shown up to this point. We refer back to our RAG formulation, particularly the conditioned function . In hindsight, we can now state that by adding random documents to the context, we are better conditioning this function, inducing enhanced accuracy. Previous research (Attanasio et al. 2022; Hoffmann et al. 2023), particularly (Zhai et al. 2023), hints that there might be cases in which a pathologically low attention entropy causes the LLM to generate degenerate outputs with a sharp decrease in performance. These episodes are named entropy collapse. Following this line of research, we measure the entropy of the attention scores in the case where only the gold document is supplied [I, , Q] against the case in which random documents are added [I, \dice6, , Q]. We find that when we introduce random documents, the entropy of the systems has a 3X increase. Although these experiments show a pattern, we cannot yet answer this question in a definitive manner. While out of the scope of this work, which focuses on the retriever component of RAG systems, we believe it is highly important to investigate the reasons for which the LLM shows this behavior. Future studies should aim to elucidate why this noisy state is more advantageous and identify the characteristics that contribute to its effectiveness.
Conclusions
In this paper, we conducted the first comprehensive study focusing on the impact of retrieved documents on the RAG framework, aiming to understand the traits required in a retriever to optimize prompt construction for a RAG system. This study led to several important findings, including two unexpected ones. First, the position of relevant information should be placed near the query; otherwise, the model seriously struggles to attend to it. Second, in contrast to common perception, top-scoring retrieved documents that do not contain the answer, when added to a prompt, negatively impact the LLM effectiveness. Finally, and even more surprisingly, random, noisy documents are actually helpful in increasing the accuracy of these systems when correctly positioned within a prompt. While we have proposed heuristics to exploit these findings, further research is needed both to uncover the inner mechanisms behind this behavior and to develop a new generation of information retrieval techniques that are specifically designed to interact with the generative component.
Acknowledgments
This work is supported by the Spoke “FutureHPC & BigData” of the ICSC – Centro Nazionale di Ricerca in High-Performance Computing, Big Data and Quantum Computing, the Spoke “Human-centered AI” of the M4C2 - Investimento 1.3, Partenariato Esteso PE00000013 - "FAIR - Future Artificial Intelligence Research", SERICS (PE00000014), IR0000013 - SoBigData.it, funded by European Union – NextGenerationEU, the FoReLab project (Departments of Excellence), and the NEREO PRIN project funded by the Italian Ministry of Education and Research Grant no. 2022AEFHAZ. This work was carried out while Florin Cuconasu was enrolled in the Italian National Doctorate on Artificial Intelligence run by the Sapienza University of Rome.