Large Language Models for Information Retrieval: A Survey

Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Haonan Chen, Zheng Liu, Zhicheng Dou, Ji-Rong Wen

Introduction

Information access is one of the fundamental daily needs of human beings. To fulfill the need for rapid acquisition of desired information, various information retrieval (IR) systems have been developed . Prominent examples include search engines such as Google, Bing, and Baidu, which serve as IR systems on the Internet, adept at retrieving relevant web pages in response to user queries, and provide convenient and efficient access to information on the Internet. It is worth noting that IR extends beyond web page retrieval. In dialogue systems (chatbots) , such as Microsoft Xiaoice , Apple Siri,Apple Siri, https://www.apple.com/siri/ and Google Assistant,Google Assistant, https://assistant.google.com/ IR systems play a crucial role in retrieving appropriate responses to user input utterances, thereby producing natural and fluent human-machine conversations. Similarly, in question-answering systems , IR systems are employed to select relevant clues essential for addressing user questions effectively. In image search engines , IR systems excel at returning images that align with user input queries. Given the exponential growth of information, research and industry have become increasingly interested in the development of effective IR systems.

The core function of an IR system is retrieval, which aims to determine the relevance between a user-issued query and the content to be retrieved, including various types of information such as texts, images, music, and more. For the scope of this survey, we concentrate solely on reviewing those text retrieval systems, in which query-document relevance is commonly measured by their matching score.The term “document” will henceforth refer to any text-based content subject to retrieve, including both long articles and short passages. Given that IR systems operate on extensive repositories, the efficiency of retrieval algorithms becomes of paramount importance. To improve the user experience, the retrieval performance is enhanced from both the upstream (query reformulation) and downstream (reranking and reading) perspectives. As an upstream technique, query reformulation is designed to refine user queries so that they are more effective at retrieving relevant documents . With the recent surge in the popularity of conversational search, this technique has received increasing attention. On the downstream side, reranking approaches are developed to further adjust the document ranking . In contrast to the retrieval stage, reranking is performed only on a limited set of relevant documents, already retrieved by the retriever. Under this circumstance, the emphasis is placed on achieving higher performance rather than keeping higher efficiency, allowing for the application of more complex approaches in the reranking process. Additionally, reranking can accommodate other specific requirements, such as personalization and diversification . Following the retrieval and reranking stages, a reading component is incorporated to summarize the retrieved documents and deliver a concise document to users . While traditional IR systems typically require users to gather and organize relevant information themselves; however, the reading component is an integral part of new IR systems such as New Bing,New Bing, https://www.bing.com/new streamlining users’ browsing experience and saving valuable time.

The trajectory of IR has traversed a dynamic evolution, transitioning from its origins in term-based methods to the integration of neural models. Initially, IR was anchored in term-based methods and Boolean logic, focusing on keyword matching for document retrieval. The paradigm gradually shifted with the introduction of vector space models , unlocking the potential to capture nuanced semantic relationships between terms. This progression continued with statistical language models , refining relevance estimation through contextual and probabilistic considerations. The influential BM25 algorithm played an important role during this phase, revolutionizing relevance ranking by accounting for term frequency and document length variations. The most recent chapter in IR’s journey is marked by the ascendancy of neural models . These models excel at capturing intricate contextual cues and semantic nuances, reshaping the landscape of IR. However, these neural models still face challenges such as data scarcity, interpretability, and the potential generation of plausible yet inaccurate responses. Thus, the evolution of IR continues to be a journey of balancing traditional strengths (such as the BM25 algorithm’s high efficiency) with the remarkable capability (such as semantic understanding) brought about by modern neural architectures.

Large language models (LLMs) have recently emerged as transformative forces across various research fields, such as natural language processing (NLP) , recommender systems , finance , and even molecule discovery . These cutting-edge LLMs are primarily based on the Transformer architecture and undergo extensive pre-training on diverse textual sources, including web pages, research articles, books, and codes. As their scale continues to expand (including both model size and data volume), LLMs have demonstrated remarkable advances in their capabilities. On the one hand, LLMs have exhibited unprecedented proficiency in language understanding and generation, resulting in responses that are more human-like and better align with human intentions. On the other hand, the larger LLMs have shown impressive emergent abilities when dealing with complex tasks , such as generalization and reasoning skills. Notably, LLMs can effectively apply their learned knowledge and reasoning abilities to tackle new tasks with just a few task-specific demonstrations or appropriate instructions . Furthermore, advanced techniques, such as in-context learning, have significantly enhanced the generalization performance of LLMs without requiring fine-tuning on specific downstream tasks . This breakthrough is particularly valuable, as it reduces the need for extensive fine-tuning while attaining remarkable task performance. Powered by prompting strategies such as chain-of-thought, LLMs can generate outputs with step-by-step reasoning, navigating complex decision-making processes . Leveraging the impressive power of LLMs can undoubtedly improve the performance of IR systems. By incorporating these sophisticated language models, IR systems can provide users with more accurate responses, ultimately reshaping the landscape of information access and retrieval.

Initial efforts have been made to utilize the potential of LLMs in the development of novel IR systems. Notably, in terms of practical applications, New Bing is designed to improve the users’ experience of using search engines by extracting information from disparate web pages and condensing it into concise summaries that serve as responses to user-generated queries. In the research community, LLMs have proven useful within specific modules of IR systems (such as retrievers), thereby enhancing the overall performance of these systems. Due to the rapid evolution of LLM-enhanced IR systems, it is essential to comprehensively review their most recent advancements and challenges.

Our survey provides an insightful exploration of the intersection between LLMs and IR systems, covering key perspectives such as query rewriters, retrievers, rerankers, and readers (as shown in Figure 1).As yet, there has not been a formal definition for LLMs. In this paper, we mainly focus on models with more than 1B parameters. We also notice that some methods do not rely on such strictly defined LLMs, but due to their representativeness, we still include an introduction to them in this survey. We also include some recent studies that leverage LLMs as search agents to perform various IR tasks. This analysis enhances our understanding of LLMs’ potential and limitations in advancing the IR field. For this survey, we create a Github repository by collecting the relevant papers and resources about LLM4IR.https://github.com/RUC-NLPIR/LLM4IR-Survey We will continue to update the repository with newer papers. This survey will also be periodically updated according to the development of this area. We notice that there are several surveys for PLMs, LLMs, and their applications (e.g., AIGC or recommender systems) . Among these, we highly recommend the survey of LLMs , which provides a systematic and comprehensive reference to many important aspects of LLMs. Compared with them, we focus on the techniques and methods for developing and applying LLMs for IR systems. In addition, we notice a perspective paper discussing the opportunity of IR when meeting LLMs . It would be an excellent supplement to this survey regarding future directions.

The remaining part of this survey is organized as follows: Section 2 introduces the background for IR and LLMs. Section 3, 4, 5, VII respectively review recent progress from the four perspectives of query rewriter, retriever, reranker, and reader, which are four key components of an IR system. Then, Section 8 discusses some potential directions in future research. Finally, we conclude the survey in Section 9 by summarizing the major findings.

Background

Information retrieval (IR), as an essential branch of computer science, aims to efficiently retrieve information relevant to user queries from a large repository. Generally, users interact with the system by submitting their queries in textual form. Subsequently, IR systems undertake the task of matching and ranking these user-supplied queries against an indexed database, thereby facilitating the retrieval of the most pertinent results.

The field of IR has witnessed significant advancement with the emergence of various models over time. One such early model is the Boolean model, which employs Boolean logic operators to combine query terms and retrieve documents that satisfy specific conditions . Based on the “bag-of-words” assumption, the vector space model represents documents and queries as vectors in term-based space. Relevance estimation is then performed by assessing the lexical similarity between the query and document vectors. The efficiency of this model is further improved through the effective organization of text content using the inverted index. Moving towards more sophisticated approaches, statistical language models have been introduced to estimate the likelihood of term occurrences and incorporate context information, leading to more accurate and context-aware retrieval . In recent years, the neural IR paradigm has gained considerable attention in the research community. By harnessing the powerful representation capabilities of neural networks, this paradigm can capture semantic relationships between queries and documents, thereby significantly enhancing retrieval performance.

Researchers have identified several challenges with implications for the performance and effectiveness of IR systems, such as query ambiguity and retrieval efficiency. In light of these challenges, researchers have directed their attention toward crucial modules within the retrieval process, aiming to address specific issues and effectuate corresponding enhancements. The pivotal role of these modules in ameliorating the IR pipeline and elevating system performance cannot be overstated. In this survey, we focus on the following four modules, which have been greatly enhanced by LLMs.

Query Rewriter is an essential IR module that seeks to improve the precision and expressiveness of user queries. Positioned at the early stage of the IR pipeline, this module assumes the crucial role of refining or modifying the initial query to align more accurately with the user’s information requirements. As an integral part of query rewriting, query expansion techniques, with pseudo relevance feedback being a prominent example, represent the mainstream approach to achieving query expression refinement. In addition to its utility in improving search effectiveness across general scenarios, the query rewriter finds application in diverse specialized retrieval contexts, such as personalized search and conversational search, thus further demonstrating its significance.

Retriever, as discussed here, is typically employed in the early stages of IR for document recall. The evolution of retrieval technologies reflects a constant pursuit of more effective and efficient methods to address the challenges posed by ever-growing text collections. In numerous experiments on IR systems over the years, the classical “bag-of-words” model BM25 has demonstrated its robust performance and high efficiency. In the wake of the neural IR paradigm’s ascendancy, prevalent approaches have primarily revolved around projecting queries and documents into high-dimensional vector spaces, and subsequently computing their relevance scores through inner product calculations. This paradigmatic shift enables a more efficient understanding of query-document relationships, leveraging the power of vector representations to capture semantic similarities.

Reranker, as another crucial module in the retrieval pipeline, primarily focuses on fine-grained reordering of documents within the retrieved document set. Different from the retriever, which emphasizes the balance of efficiency and effectiveness, the reranker module places a greater emphasis on the quality of document ranking. In pursuit of enhancing the search result quality, researchers delve into more complex matching methods than the traditional vector inner product, thereby furnishing richer matching signals to the reranker. Moreover, the reranker facilitates the adoption of specialized ranking strategies tailored to meet distinct user requirements, such as personalized and diversified search results. By integrating domain-specific objectives, the reranker module can deliver tailored and purposeful search results, enhancing the overall user experience.

Reader has evolved as a crucial module with the rapid development of LLM technologies. Its ability to comprehend real-time user intent and generate dynamic responses based on the retrieved text has revolutionized the presentation of IR results. In comparison to presenting a list of candidate documents, the reader module organizes answer texts more intuitively, simulating the natural way humans access information. To enhance the credibility of generated responses, the integration of references into generated responses has been an effective technique of the reader module.

Furthermore, researchers explore unifying the above modules to develop a novel LLM-driven search model known as the Search Agent. The search agent is distinguished by its simulation of an automated search and result understanding process, which furnishes users with accurate and readily comprehensible answers. WebGPT serves as a pioneering work in this category, which models the search process as a sequence of actions of an LLM-based agent within a search engine environment, autonomously accomplishing the whole search pipeline. By integrating the existing search stack, search agents have the potential to become a new paradigm in future IR.

2 Large Language Models

Language models (LMs) are designed to calculate the generative likelihood of word sequences by taking into account the contextual information from preceding words, thereby predicting the probability of subsequent words. Consequently, by employing certain word selection strategies (such as greedy decoding or random sampling), LMs can proficiently generate natural language texts. Although the primary objective of LMs lies in text generation, recent studies have revealed that a wide array of natural language processing problems can be effectively reformulated into a text-to-text format, thus rendering them amenable to resolution through text generation. This has led to LMs becoming the de facto solution for the majority of text-related problems.

The evolution of LMs can be categorized into four primary stages, as discussed in prior literature . Initially, LMs were rooted in statistical learning techniques and were termed statistical language models. These models tackled the issue of word prediction by employing the Markov assumption to predict the subsequent word based on preceding words. Thereafter, neural networks, particularly recurrent neural networks (RNNs), were introduced to calculate the likelihood of text sequences and establish neural language models. These advancements made it feasible to utilize LMs for representation learning beyond mere word sequence modeling. ELMo first proposed to learn contextualized word representations through pre-training a bidirectional LSTM (biLSTM) network on large-scale corpora, followed by fine-tuning on specific downstream tasks. Similarly, BERT proposed to pre-train a Transformer encoder with a specially designed Masked Language Modeling (MLM) task and Next Sentence Prediction (NSP) task on large corpora. These studies initiated a new era of pre-trained language models (PLMs), with the “pre-training then fine-tuning” paradigm emerging as the prevailing learning approach. Along this line, numerous generative PLMs (e.g., GPT-2 , BART , and T5 ) have been developed for text generation problems including summarization, machine translation, and dialogue generation. Recently, researchers have observed that increasing the scale of PLMs (e.g., model size or data amount) can consistently improve their performance on downstream tasks (a phenomenon commonly referred to as the scaling law ). Moreover, large-sized PLMs exhibit promising abilities (termed emergent abilities ) in addressing complex tasks, which are not evident in their smaller counterparts. Therefore, the research community refers to these large-sized PLMs as large language models (LLMs).

As shown in Figure 2, existing LLMs can be categorized into two groups based on their architectures: encoder-decoder and decoder-only models. The encoder-decoder models incorporate an encoder component to transform the input text into vectors, which are then employed for producing output texts. For example, T5 is an encoder-decoder model that converts each natural language processing problem into a text-to-text form and resolves it as a text generation problem. In contrast, decoder-only models, typified by GPT, rely on the Transformer decoder architecture. It uses a self-attention mechanism with a diagonal attention mask to generate a sequence of words from left to right. Building upon the success of GPT-3 , which is the first model to encompass over 100B parameters, several noteworthy models have been inspired, including GPT-J, BLOOM , OPT , Chinchilla , and LLaMA . These models follow the similar Transformer decoder structure as GPT-3 and are trained on various combinations of datasets.

Owing to their vast number of parameters, fine-tuning LLMs for specific tasks, such as IR, is often deemed impractical. Consequently, two prevailing methods for applying LLMs have been established: in-context learning (ICL) and parameter-efficient fine-tuning. ICL is one of the emergent abilities of LLMs empowering them to comprehend and furnish answers based on the provided input context, rather than relying merely on their pre-training knowledge. This method requires only the formulation of the task description and demonstrations in natural language, which are then fed as input to the LLM. Notably, parameter tuning is not required for ICL. Additionally, the efficacy of ICL can be further augmented through the adoption of chain-of-thought prompting, involving multiple demonstrations (describe the chain of thought examples) to guide the model’s reasoning process. ICL is the most commonly used method for applying LLMs to IR. Parameter-efficient fine-tuning aims to reduce the number of trainable parameters while maintaining satisfactory performance. LoRA , for example, has been widely applied to open-source LLMs (e.g., LLaMA and BLOOM) for this purpose. Recently, QLoRA has been proposed to further reduce memory usage by leveraging a frozen 4-bit quantized LLM for gradient computation. Despite the exploration of parameter-efficient fine-tuning for various NLP tasks, its implementation in IR tasks remains relatively limited, representing a potential avenue for future research.

Query Rewriter

Query rewriting in modern IR systems is essential for improving search query effectiveness and accuracy. It reformulates users’ original queries to better match search results, alleviating issues like vague queries or vocabulary mismatches between the query and target documents. This task goes beyond mere synonym replacement, requiring an understanding of user intent and query context, particularly in complex searches like conversational queries. Effective query rewriting enhances search engine performance.

Traditional methods for query rewriting improve retrieval performance by expanding the initial query with information from highly-ranked relevant documents. Mainly-used methods include relevance feedback , word-embedding based methods etc. However, the limited ability of semantic understanding and comprehension of user search intent limits their performance in capturing the full scope of user intent.

Recent advancements in LLMs present promising opportunities to boost query rewriting capabilities. On one hand, given the context and subtleties of a query, LLMs can provide more accurate and contextually relevant rewrites. On the other hand, LLMs can leverage their extensive knowledge to generate synonyms and related concepts, enhancing queries to cover a broader range of relevant documents, thereby effectively addressing the vocabulary mismatch problem. In the following sections, we will introduce the recent works that employ LLMs in query rewriting.

Query rewriting typically serves two scenarios: ad-hoc retrieval, which mainly addresses vocabulary mismatches between queries and candidate documents, and conversational search, which refines queries based on evolving conversations. The upcoming section will delve into the role of query rewriting in these two domains and explore how LLMs enhance this process.

In ad-hoc retrieval, queries are often short and ambiguous. In such scenarios, the main objectives of query rewriting include adding synonyms or related terms to address vocabulary mismatches and clarifying ambiguous queries to more accurately align with user intent. From this perspective, LLMs have inherent advantages in query rewriting. Primarily, LLMs have a deep understanding of language semantics, allowing them to capture the meaning of queries more effectively. Besides, LLMs can leverage their extensive training on diverse datasets to generate contextually relevant synonyms and expand queries, ensuring broader and more precise search result coverage. Additionally, studies have shown that LLMs’ integration of external factual corpora and thoughtful model design further enhance their accuracy in generating effective query rewrites, especially for specific tasks.

Currently, there are many studies leveraging LLMs to rewrite queries in adhoc retrieval. We introduce the typical method Query2Doc as an example. As shown in Figure 3, Query2Doc prompts the LLMs to generate a relevant passage according to the original query (“when was pokemon green released?”). Subsequently, the original query is expanded by incorporating the generated passage. The retriever module uses this new query to retrieve a list of relevant documents. Notably, the generated passage contains additional detailed information, such as “Pokemon Green was released in Japan on February 27th”, which effectively mitigates the “vocabulary mismatch” issue to some extent.

In addition to addressing the “vocabulary mismatch” problem , other works utilize LLMs for different challenges in ad-hoc retrieval. For instance, PromptCase leverages LLMs in legal case retrieval to simplify complex queries into more searchable forms. This involves using LLMs to identify legal facts and issues, followed by a prompt-based encoding scheme for effective language model encoding.

1.2 Conversational Search

Query rewrites in conversational search play a pivotal role in enhancing the search experience. Unlike traditional queries in ad-hoc retrieval, conversational search involves a dialogue-like interaction, where the context and user intent evolve with each interaction. In conversational search, query rewriting involves understanding the entire conversation’s context, clarifying any ambiguities, and personalizing responses based on user history. The process includes dynamic query expansion and refinement based on dialogue information. This makes conversational query rewriting a sophisticated task that goes beyond traditional search, focusing on natural language understanding and user-centric interaction.

In the era of LLMs, leveraging LLMs in conversational search tasks offers several advantages. First, LLMs possess strong contextual understanding capabilities, enabling them to better comprehend users’ search intent within the context of multi-turn conversations between users and the system. Second, LLMs exhibit powerful generation abilities, allowing them to simulate dialogues between users and the system, thereby facilitating more robust search intent modeling.

The LLMCS framework is a pioneering approach that employs LLMs to effectively extract and understand user search intent within conversational contexts. As illustrated in their work, LLMCS uses LLMs to produce both query rewrites and extensive hypothetical system responses from various perspectives. These outputs are combined into a comprehensive representation that effectively captures the user’s full search intent. The experimental results show that including detailed hypothetical responses with concise query rewrites markedly improves search performance by adding more plausible search intent. Ye et al. claims that human query rewrite may lack sufficient information for optimal retrieval performance. It defines four essential properties for well-formed LLM-generated query rewrites. Results show that their method informative query rewrites can yield substantially improved retrieval performance compared to human rewrites.

Besides, LLMs can be used as a data expansion tool in conversational dense retrieval. Attributed to the high cost of producing hand-written dialogues, data scarcity presents a significant challenge in the domain of conversational search. To address this problem, CONVERSER employs LLMs to generate synthetic passage-dialogue pairs through few-shot demonstrations. Furthermore, it efficiently trains a dense retriever using a minimal dataset of six in-domain dialogues, thus mitigating the issue of data sparsity.

2 Rewriting Knowledge

Query rewriting typically necessitates additional corpora for refining initial queries. Considering that LLMs incorporate world knowledge in their parameters, they are naturally capable of rewriting queries. We refer to these methods, which rely exclusively on the intrinsic knowledge of LLMs, as LLM-only methods. While LLMs encompass a broad spectrum of knowledge, they may be inadequate in specialized areas. Furthermore, LLMs can introduce concept drift, leading to noisy relevance signals. To address this issue, some methods incorporate domain-specific corpora to provide more detailed and relevant information in query rewriting. We refer to methods enhanced by domain-specific corpora to boost LLM performance as corpus-enhanced LLM-based methods. In this section, we will introduce these two methods in detail.

LLMs are capable of storing knowledge within their parameters, making it a natural choice to capitalize on this knowledge for the purpose of query rewriting. As a pioneering work in LLM-based query rewriting, HyDE generates a hypothetical document by LLMs according to the given query and then uses a dense retriever to retrieve relevant documents from the corpus that are relevant to the generated document. Query2doc generates pseudo documents via prompting LLMs with few-shot demonstrations, and then expands the query with the generated pseudo document. Furthermore, the influence of different prompting methods and various model sizes on query rewriting has also been investigated . To better accommodate the frozen retriever and the LLM-based reader, a small language model is employed as the rewriter that is trained using reinforcement learning techniques with the rewards provided by the LLM-based reader . GFF presents a “Generate, Filter, and Fuse” method for query expansion. It employs an LLM to create a set of related keywords via a reasoning chain. Then, a self-consistency filter is used to identify the most important keywords, which are concatenated with the original queries for the downstream reranking task.

It is worth noting that though the designs of these methods are different, all of them rely on the world knowledge stored in LLMs without additional corpora.

2.2 Corpus-enhanced LLM-based methods

Although LLMs exhibit remarkable capabilities, the lack of domain-specific knowledge may lead to the generation of hallucinatory or irrelevant queries. To address this issue, recent studies have proposed a hybrid approach that enhances LLM-based query rewriting methods with an external document corpus.

Why incorporate a document corpus? The integration of a document corpus offers several notable advantages. Firstly, it boosts relevance by using relevant documents to refine query generation, reducing irrelevant content and improving contextually appropriate outputs. Second, enhancing LLMs with up-to-date information and specialized knowledge in specific fields enables them to effectively deal with queries that are both current and specific to certain domains.

How to incorporate a document corpus? Thanks to the flexibility of LLMs, various paradigms have been proposed to incorporate a document corpus into LLM-based query rewriting, which can be summarized as follows.

∙\bullet Late fusion of LLM-based re-writing and pseudo relevance feedback (PRF) retrieval results. Traditional PRF methods leverage relevant documents retrieved from a document corpus to rewrite queries, which restricts the query to the information contained in the target corpus. On the contrary, LLM-based rewriting methods provide external context not present in the corpus, which is more diverse. Both approaches have the potential to independently enhance retrieval performance. Therefore, a straightforward strategy for combining them is using a weighted fusion method for retrieval results .

∙\bullet Combining retrieved relevant documents in the prompts of LLMs. In the era of LLMs, incorporating instructions within the prompts is the most flexible method for achieving specific functionalities. QUILL and CAR illustrate how retrieval augmentation of queries can provide LLMs with context that significantly enhances query understanding. LameR takes this further by using LLM expansion to improve the simple BM25 retriever, introducing a retrieve-rewrite-retrieve framework. Experimental results reveal that even basic term-based retrievers can achieve comparable performance when paired with LLM-based rewriters. Additionally, InteR proposes a multi-turn interaction framework between search engines and LLMs. This enables search engines to expand queries using LLM-generated insights, while LLMs refine prompts using relevant documents sourced from the search engines.

∙\bullet Enhancing factuality of generative relevance feedback (GRF) by pseudo relevance feedback (PRF). Although generative documents are often relevant and diverse, they exhibit hallucinatory characteristics. In contrast, traditional documents are generally regarded as reliable sources of factual information. Motivated by this observation, GRM proposes a novel technique known as relevance-aware sample estimation (RASE). RASE leverages relevant documents retrieved from the collection to assign weights to generated documents. In this way, GRM ensures that relevance feedback is not only diverse but also maintains a high degree of factuality.

3 Rewriting Approaches

There are three main approaches used for leveraging LLMs in query rewriting: prompting methods, fine-tuning, and knowledge distillation. Prompting methods involve using specific prompts to direct LLM output, providing flexibility and interpretability. Fine-tuning adjusts pre-trained LLMs on specific datasets or tasks to improve domain-specific performance, mitigating the general nature of LLM world knowledge. Knowledge distillation, on the other hand, transfers LLM knowledge to lightweight models, simplifying the complexity associated with retrieval augmentation. In the following section, we will introduce these three methods in detail.

Prompting in LLMs refers to the technique of providing a specific instruction or context to guide the model’s generation of text. The prompt serves as a conditioning signal and influences the language generation process of the model. Existing prompting strategies can be roughly categorized into three groups: zero-shot prompting, few-shot prompting, and chain-of-thought (CoT) prompting .

∙\bullet Zero-shot prompting. Zero-shot prompting involves instructing the model to generate texts on a specific topic without any prior exposure to training examples in that domain or topic. The model relies on its pre-existing knowledge and language understanding to generate coherent and contextually relevant expanded terms for original queries. Experiments show that zero-shot prompting is a simple yet effective method for query rewriting .

∙\bullet Few-shot prompting. Few-shot prompting, also known as in-context learning, involves providing the model with a limited set of examples or demonstrations related to the desired task or domain . These examples serve as a form of explicit instruction, allowing the model to adapt its language generation to the specific task or domain at hand. Query2Doc prompts LLMs to write a document that answers the query with some demo query-document pairs provided by the ranking dataset, such as MSMARCO and NQ . This work experiments with a single prompt. To further study the impact of different prompt designing, recent works have explored eight different prompts, such as prompting LLMs to generate query expansion terms instead of entire pseudo documents and CoT prompting. There are some illustrative prompts in Table I. This work conducts more experiments than Query2Doc, but the results show that the proposed prompt is less effective than Query2Doc.

∙\bullet Chain-of-thought prompting. CoT prompting is a strategy that involves iterative prompting, where the model is provided with a sequence of instructions or partial outputs . In conversational search, the process of query re-writing is multi-turn, which means queries should be refined step-by-step with the interaction between search engines and users. This process is naturally coincided with CoT process. As shown in 4, users can conduct the CoT process through adding some instructions during each turn, such as “Based on all previous turns, xxx”. While in ad-hoc search, there is only one-round in query re-writing, so CoT could only be accomplished in a simple and coarse way. For example, as shown in Table I, researchers add “Give the rationale before answering” in the instructions to prompt LLMs think deeply .

3.2 Fine-tuning

Fine-tuning is an effective approach for adapting LLMs to specific domains. This process usually starts with a pre-trained language model, like GPT-3, which is then further trained on a dataset tailored to the target domain. This domain-specific training enables the LLM to learn unique patterns, terminology, and context relevant to the domain, which is able to improve its capacity to produce high-quality query rewrites.

BEQUE leverages LLMs for rewriting queries in e-commerce product searches. It designs three Supervised Fine-Tuning (SFT) tasks: quality classification of e-commerce query rewrites, product title prediction, and CoT query rewriting. To our knowledge, it is the first model to directly fine-tune LLMs, including ChatGLM , ChatGLM2.0 , Baichuan , and Qwen , specifically for the query rewriting task. After the SFT stage, BEQUE uses an offline system to gather feedback on the rewrites and further aligns the rewriters with e-commerce search objectives through an object alignment stage. Online A/B testing demonstrates the effectiveness of the method.

3.3 Knowledge Distillation

Although LLM-based methods have demonstrated significant improvements in query rewriting tasks, their practical implementation for online deployment is hindered by the substantial latency caused by the computational requirements of LLMs. To address this challenge, knowledge distillation has emerged as a prominent technique in the industry. In the QUILL framework, a two-stage distillation method is proposed. This approach entails utilizing a retrieval-augmented LLM as the professor model, a vanilla LLM as the teacher model, and a lightweight BERT model as the student model. The professor model is trained on two extensive datasets, namely Orcas-I and EComm , which are specifically curated for query intent understanding. Subsequently, a two-stage distillation process is employed to transfer knowledge from the professor model to the teacher model, followed by knowledge transfer from the teacher model to the student model. Empirical findings demonstrate that this knowledge distillation methodology surpasses the simple scaling up of model size from base to XXL, resulting in even more substantial improvements. In a recently proposed “rewrite-retrieve-read” framework , an LLM is first used to rewrite the queries by prompting, followed by a retrieval-augmented reading process. To improve framework effectiveness, a trainable rewriter, implemented as a small language model, is incorporated to further adapt search queries to align with both the frozen retriever and the LLM reader’s requirements. The rewriter’s refinement involves a two-step training process. Initially, supervised warm-up training is conducted using pseudo data. Then, the retrieve-then-read pipeline is described as a reinforcement learning scenario, with the rewriter’s training acting as a policy model to maximize pipeline performance rewards.

4 Limitations

While LLMs offer promising capabilities for query rewriting, they also meet several challenges. Here, we outline two main limitations of LLM-based query rewriters.

When using LLMs for query rewriting, they may introduce unrelated information, known as concept drift, due to their extensive knowledge base and tendency to produce detailed and redundant content. While this can enrich the query, it also risks generating irrelevant or off-target results.

This phenomenon has been reported in several studies These studies highlight the need for a balanced approach in LLM-based query rewriting, ensuring that the essence and focus of the original query are maintained while leveraging the LLM’s ability to enhance and clarify the query. This balance is crucial for effective search and IR applications.

4.2 Correlation between Retrieval Performance and Expansion Effects

Recently, a comprehensive study conduct experiments on various expansion techniques and downstream ranking models, which reveals a notable negative correlation between retriever performance and the benefits of expansion. Specifically, while expansion tends to enhance the scores of weaker models, it generally hurts stronger models. This observation suggests a strategic approach: employ expansions with weaker models or in scenarios where the target dataset substantially differs in format from the training corpus. In other cases, it is advisable to avoid expansions to maintain clarity of the relevance signal.

Retriever

In an IR system, the retriever serves as the first-pass document filter to collect broadly relevant documents for user queries. Given the enormous amounts of documents in an IR system, the retriever’s efficiency in locating relevant documents is essential for maintaining search engine performance. Meanwhile, a high recall is also important for the retriever, as the retrieved documents are then fed into the ranker to generate final results for users, which determines the ranking quality of search engines.

In recent years, retrieval models have shifted from relying on statistic algorithms to neural models . The latter approaches exhibit superior semantic capability and excel at understanding complicated user intent. The success of neural retrievers relies on two key factors: data and model. From the data perspective, a large amount of high-quality training data is essential. This enables retrievers to acquire comprehensive knowledge and accurate matching patterns. Furthermore, the intrinsic quality of search data, i.e., issued queries and document corpus, significantly influences retrieval performance. From the model perspective, a strongly representational neural architecture allows retrievers to effectively store and apply knowledge obtained from the training data.

Unfortunately, there are some long-term challenges that hinder the advancement of retrieval models. First, user queries are usually short and ambiguous, making it difficult to precisely understand the user’s search intents for retrievers. Second, documents typically contain lengthy content and substantial noise, posing challenges in encoding long documents and extracting relevant information for retrieval models. Additionally, the collection of human-annotated relevance labels is time-consuming and costly. It restricts the retrievers’ knowledge boundaries and their ability to generalize across different application domains. Moreover, existing model architectures, primarily built on BERT , exhibit inherent limitations, thereby constraining the performance potential of retrievers. Recently, LLMs have exhibited extraordinary abilities in language understanding, text generation, and reasoning. This has motivated researchers to use these abilities to tackle the aforementioned challenges and aid in developing superior retrieval models. Roughly, these studies can be categorized into two groups, i.e., (1) leveraging LLMs to generate search data, and (2) employing LLMs to enhance model architecture.

In light of the quality and quantity of search data, there are two prevalent perspectives on how to improve retrieval performance via LLMs. The first perspective revolves around search data refinement methods, which concentrate on reformulating input queries to precisely present user intents. The second perspective involves training data augmentation methods, which leverage LLMs’ generation ability to enlarge the training data for dense retrieval models, particularly in zero- or few-shot scenarios.

Typically, input queries consist of short sentences or keyword-based phrases that may be ambiguous and contain multiple possible user intents. Accurately determining the specific user intent is essential in such cases. Moreover, documents usually contain redundant or noisy information, which poses a challenge for retrievers to extract relevance signals between queries and documents. Leveraging the strong text understanding and generation capabilities of LLMs offers a promising solution to these challenges. As yet, research efforts in this domain primarily concentrate on employing LLMs as query rewriters, aiming to refine input queries for more precise expressions of the user’s search intent. Section 3 has provided a comprehensive overview of these studies, so this section refrains from further elaboration. In addition to query rewriting, an intriguing avenue for exploration involves using LLMs to enhance the effectiveness of retrieval by refining lengthy documents. This intriguing area remains open for further investigation and advancement.

1.2 Training Data Augmentation

Due to the expensive economic and time costs of human-annotated labels, a common problem in training neural retrieval models is the lack of training data. Fortunately, the excellent capability of LLMs in text generation offers a potential solution. A key research focus lies in devising strategies to leverage LLMs’ capabilities to generate pseudo-relevant signals and augment the training dataset for the retrieval task.

Why do we need data augmentation? Previous studies of neural retrieval models focused on supervised learning, namely training retrieval models using labeled data from specific domains. For example, MS MARCO provides a vast repository, containing a million passages, more than 200,000 documents, and 100,000 queries with human-annotated relevance labels, which has greatly facilitated the development of supervised retrieval models. However, this paradigm inherently constrains the retriever’s generalization ability for out-of-distribution data from other domains. The application spectrum of retrieval models varies from natural question-answering to biomedical IR, and it is expensive to annotate relevance labels for data from different domains. As a result, there is an emerging need for zero-shot and few-shot learning models to address this problem . A common practice to improve the models’ effectiveness in a target domain without adequate label signals is through data augmentation.

How to apply LLMs for data augmentation? In the scenario of IR, it is easy to collect numerous documents. However, the challenging and costly task lies in gathering real user queries and labeling the relevant documents accordingly. Considering the strong text generation capability of LLMs, many researchers suggest using LLM-driven processes to create pseudo queries or relevance labels based on existing collections. These approaches facilitate the construction of relevant query-document pairs, enlarging the training data for retrieval models. According to the type of generated data, there are two mainstream approaches that complement the LLM-based data augmentation for retrieval models, i.e., pseudo query generation and relevance label generation. Their frameworks are visualized in Figure 5. Next, we will give an overview of the related studies.

∙\bullet Pseudo query generation. Given the abundance of documents, a straightforward idea is to use LLMs for generating their corresponding pseudo queries. One such illustration is presented by inPairs , which leverages the in-context learning capability of GPT-3. This method employs a collection of query-document pairs as demonstrations. These pairs are combined with a document and presented as input to GPT-3, which subsequently generates possible relevant queries for the given document. By combining the same demonstration with various documents, it is easy to create a vast pool of synthetic training samples and support the fine-tuning of retrievers on specific target domains. Recent studies have also leveraged open-sourced LLMs, such as Alpaca-LLaMA and tk-Instruct, to produce sufficient pseudo queries and applied curriculum learning to pre-train dense retrievers. To enhance the reliability of these synthetic samples, a fine-tuned model (e.g., a monoT5-3B model fine-tuned on MSMARCO ) is employed to filter the generated queries. Only the top pairs with the highest estimated relevance scores are kept for training. This “generating-then-filtering” paradigm can be conducted iteratively in a round-trip filtering manner, i.e., by first fine-tuning a retriever on the generated samples and then filtering the generated samples using this retriever. Repeating these EM-like steps until convergence can produce high-quality training sets . Furthermore, by adjusting the prompt given to LLMs, they can generate queries of different types. This capability allows for a more accurate simulation of real queries with various patterns .

In practice, it is costly to generate a substantial number of pseudo queries through LLMs. Balancing the generation costs and the quality of generated samples has become an urgent problem. To tackle this, UDAPDR is proposed, which first produces a limited set of synthetic queries using LLMs for the target domain. These high-quality examples are subsequently used as prompts for a smaller model to generate a large number of queries, thereby constructing the training set for that specific domain. It is worth noting that the aforementioned studies primarily rely on fixed LLMs with frozen parameters. Empirically, optimizing LLMs’ parameters can significantly improve their performance on downstream tasks. Unfortunately, this pursuit is impeded by the prohibitively high demand for computational resources. To overcome this obstacle, SPTAR introduces a soft prompt tuning technique that only optimizes the prompts’ embedding layer during the training process. This approach allows LLMs to better adapt to the task of generating pseudo-queries, striking a favorable balance between training cost and generation quality.

In addition to the above studies, pseudo query generation methods are also introduced in other application scenarios, such as conversational dense retrieval and multilingual dense retrieval .

∙\bullet Relevance label generation. In some downstream tasks of retrieval, such as question-answering, the collection of questions is also sufficient. However, the relevance labels connecting these questions with the passages of supporting evidence are very limited. In this context, leveraging the capability of LLMs for relevance label generation is a promising approach that can augment the training corpus for retrievers. A recent method, ART , exemplifies this approach. It first retrieves the top-relevant passages for each question. Then, it employs an LLM to produce the generation probabilities of the question conditioned on these top passages. After a normalization process, these probabilities serve as soft relevance labels for the training of the retriever.

Additionally, to highlight the similarities and differences among the corresponding methods, we present a comparative result in Table III. It compares the aforementioned methods from various perspectives, including the number of examples, the generator employed, the type of synthetic data produced, the method applied to filter synthetic data, and whether LLMs are fine-tuned. This table serves to facilitate a clearer understanding of the landscape of these methods.

2 Employing LLMs to Enhance Model Architecture

Leveraging the excellent text encoding and decoding capabilities of LLMs, it is feasible to understand queries and documents with greater precision compared to earlier smaller-sized models . Researchers have endeavored to utilize LLMs as the foundation for constructing advanced retrieval models. These methods can be grouped into two categories, i.e., dense retrievers and generative retrievers.

In addition to the quantity and quality of the data, the representative capability of models also greatly influences the efficacy of retrievers. Inspired by the LLM’s excellent capability to encode and comprehend natural language, some researchers leverage LLMs as retrieval encoders and investigate the impact of model scale on retriever performance.

General Retriever. Since the effectiveness of retrievers primarily relies on the capability of text embedding, the evolution of text embedding models often has a significant impact on the progress of retriever development. In the era of LLMs, a pioneer work is made by OpenAI . They view the adjacent text segments as positive pairs to facilitate the unsupervised pre-training of a set of text embedding models, denoted as cpt-text, whose parameter values vary from 300M to 175B. Experiments conducted on the MS MARCO and BEIR datasets indicate that larger model scales have the potential to yield improved performance in unsupervised learning and transfer learning for text search tasks. Nevertheless, pre-training LLMs from scratch is prohibitively expensive for most researchers. To overcome this limitation, some studies use pre-trained LLMs to initialize the bi-encoder of dense retriever. Specifically, GTR adopts T5-family models, including T5-base, Large, XL, and XXL, to initialize and fine-tune dense retrievers. RepLLaMA further fine-tunes the LLaMA model on multiple stages of IR, including retrieval and reranking. For the dense retrieval task, RepLLaMA appends an end-of-sequence token “<</s>>” to the input sequences, i.e., queries or documents, and regards its output embeddings as the representation of queries or documents. The experiments confirm again that larger model sizes can lead to better performance, particularly in zero-shot settings. Notably, the researchers of RepLLaMA also study the effectiveness of applying LLaMA in the reranking stage, which will be introduced in Section 5.1.3.

Task-aware Retriever. While the aforementioned studies primarily focus on using LLMs as text embedding models for downstream retrieval tasks, retrieval performance can be greatly enhanced when task-specific instructions are integrated. For example, TART devises a task-aware retrieval model that introduces a task-specific instruction before the question. This instruction includes descriptions of the task’s intent, domain, and desired retrieved unit. For instance, given that the task is question-answering, an effective prompt might be “Retrieve a Wikipedia text that answers this question. {question}”. Here, “Wikipedia” (domain) indicates the expected source of retrieved documents, “text” (unit) suggests the type of content to retrieve, and “answers this question” (intent) demonstrates the intended relationship between the retrieved texts and the question. This approach can take advantage of the powerful language modeling capability and extensive knowledge of LLMs to precisely capture the user’s search intents across various retrieval tasks. Considering the efficiency of retrievers, it first fine-tunes a TART-full model with cross-encoder architecture, which is initialized from LLMs (e.g., T0-3B, Flan-T5). Then, a TART-dull model initialized from Contriever is learned by distillating knowledge from the TART-full.

2.2 Generative Retriever

Traditional IR systems typically follow the “index-retrieval-rank” paradigm to locate relevant documents based on user queries, which has proven effective in practice. However, these systems usually consist of three separate modules: the index module, the retrieval module, and the reranking module. Therefore, optimizing these modules collectively can be challenging, potentially resulting in sub-optimal retrieval outcomes. Additionally, this paradigm demands additional space for storing pre-built indexes, further burdening storage resources. Recently, model-based generative retrieval methods have emerged to address these challenges. These methods move away from the traditional “index-retrieval-rank” paradigm and instead use a unified model to directly generate document identifiers (i.e., DocIDs) relevant to the queries. In these model-based generative retrieval methods, the knowledge of the document corpus is stored in the model parameters, eliminating the need for additional storage space for the index. Existing methods have explored generating document identifiers through fine-tuning and prompting of LLMs

Fine-tuning LLMs. Given the vast amount of world knowledge contained in LLMs, it is intuitive to leverage them for building model-based generative retrievers. DSI is a typical method that fine-tunes the pre-trained T5 models on retrieval datasets. The approach involves encoding queries and decoding document identifiers directly to perform retrieval. They explore multiple techniques for generating document identifiers and find that constructing semantically structured identifiers yields optimal results. In this strategy, DSI applies hierarchical clustering to group documents according to their semantic embeddings and assigns a semantic DocID to each document based on its hierarchical group. To ensure the output DocIDs are valid and do represent actual documents in the corpus, DSI constructs a trie using all DocIDs and utilizes a constraint beam search during the decoding process. Furthermore, this approach observes that the scaling law, which suggests that larger LMs lead to improved performance, is also applied to generative retrievers.

Prompting LLMs. In addition to fine-tuning LLMs for retrieval, it has been found that LLMs (e.g., GPT-series models) can directly generate relevant web URLs for user queries with a few in-context demonstrations . This unique capability of LLMs is believed to arise from their training exposure to various HTML resources. As a result, LLMs can naturally serve as generative retrievers that directly generate document identifiers to retrieve relevant documents for input queries. To achieve this, an LLM-URL model is proposed. It utilizes the GPT-3 text-davinci-003 model to yield candidate URLs. Furthermore, it designs regular expressions to extract valid URLs from these candidates to locate the retrieved documents.

To provide a comprehensive understanding of this topic, Table IV summarizes the common and unique characteristics of the LLM-based retrievers discussed above.

3 Limitations

Though some efforts have been made for LLM-augmented retrieval, there are still many areas that require more detailed investigation. For example, a critical requirement for retrievers is fast response, while the main problem of existing LLMs is the huge model parameters and overlong inference time. Addressing this limitation of LLMs to ensure the response time of retrievers is a critical task. Moreover, even when employing LLMs to augment datasets (a context with lower inference time demands), the potential mismatch between LLM-generated texts and real user queries could impact retrieval effectiveness. Furthermore, as LLMs usually lack domain-specific knowledge, they need to be fine-tuned on task-specific datasets before applying them to downstream tasks. Therefore, developing efficient strategies to fine-tune these LLMs with numerous parameters emerges as a key concern.

Reranker

Reranker, as the second-pass document filter in IR, aims to rerank a document list retrieved by the retriever (e.g., BM25) based on the query-document relevance. Based on the usage of LLMs, the existing LLM-based reranking methods can be divided into three paradigms: utilizing LLMs as supervised rerankers, utilizing LLMs as unsupervised rerankers, and utilizing LLMs for training data augmentation. These paradigms are summarized in Table V and will be elaborated upon in the following sections. Recall that we will use the term document to refer to the text retrieved in general IR scenarios, including instances such as passages (e.g., passages in MS MARCO passage ranking dataset ).

Supervised fine-tuning is an important step in applying pre-trained LLMs to a reranking task. Due to the lack of awareness of ranking during pre-training, LLMs cannot appropriately measure the query-document relevance and fully understand the reranking tasks. By fine-tuning LLMs on task-specific ranking datasets, such as the MS MARCO passage ranking dataset , which includes signals of both relevance and irrelevance, LLMs can adjust their parameters to yield better performance in the reranking tasks. Based on the backbone model structure, we can categorize existing supervised rerankers as: (1) encoder-only, (2) encoder-decoder, and (3) decoder-only.

The encoder-based rerankers represent a significant turning point in applying LLMs to document ranking tasks. They demonstrate how some pre-trained language models (e.g., BERT ) can be finetuned to deliver highly accurate relevance predictions. A representative approach is monoBERT , which transforms a query-document pair into a sequence “[CLS] query [SEP] document [SEP]” as the model input and calculates the relevance score by feeding the “[CLS]” representation into a linear layer. The reranking model is optimized based on the cross-entropy loss.

1.2 Encoder-Decoder

In this field, existing studies mainly formulate document ranking as a generation task and optimize an encoder-decoder-based reranking model . Specifically, given the query and the document, reranking models are usually fine-tuned to generate a single token, such as “true” or “false”. During inference, the query-document relevance score is determined based on the logit of the generated token. For example, a T5 model can be fine-tuned to generate classification tokens for relevant or irrelevant query-document pairs . At inference time, a softmax function is applied to the logits of “true” and “false” tokens, and the relevance score is calculated as the probability of the “true” token. The following method involves a multi-view learning approach based on the T5 model. This approach simultaneously considers two tasks: generating classification tokens for a given query-document pair and generating the corresponding query conditioned on the provided document. DuoT5 considers a triple (q,di,dj)(q,d_{i},d_{j}) as the input of the T5 model and is fine-tuned to generate token “true” if document did_{i} is more relevant to query qiq_{i} than document djd_{j}, and “false” otherwise. During inference, for each document did_{i}, it enumerates all other documents djd_{j} and uses global aggregation functions to generate the relevance score sis_{i} for document did_{i} (e.g., si=∑jpi,js_{i}=\sum_{j}p_{i,j}, where pi,jp_{i,j} represents the probability of generating “true” when taking (q,di,dj)(q,d_{i},d_{j}) as the model input).

Although these generative loss-based methods outperform several strong ranking baselines, they are not optimal for reranking tasks. This stems from two primary reasons. First, it is commonly expected that a reranking model will yield a numerical relevance score for each query-document pair rather than text tokens. Second, compared to generation losses, it is more reasonable to optimize the reranking model using ranking losses (e.g., RankNet ). Recently, RankT5 has directly calculated the relevance score for a query-document pair and optimized the ranking performance with “pairwise” or “listwise” ranking losses. An avenue for potential performance enhancement lies in the substitution of the base-sized T5 model with its larger-scale counterpart.

1.3 Decoder-only

Recently, there have been some attempts to rerank documents by fine-tuning decoder-only models (such as LLaMA). For example, RankLLaMA proposes formatting the query-document pair into a prompt “query: {query} document: {document} [EOS]” and utilizes the last token representation for relevance calculation. Besides, RankingGPT has been proposed to bridge the gap between LLMs’ conventional training objectives and the specific needs of document ranking through two-stage training. The first stage involves continuously pretraining LLMs using a large number of relevant text pairs collected from web resources, helping the LLMs to naturally generate queries relevant to the input document. The second stage focuses on improving the model’s text ranking performance using high-quality supervised data and well-designed loss functions. Different from these pointwise rerankers , Rank-without-GPT proposes to train a listwise reranker that directly outputs a reranked document list. The authors first demonstrate that existing pointwise datasets (such as MS MARCO ), which only contain binary query-document labels, are insufficient for training efficient listwise rerankers. Then, they propose to use the ranking results of existing ranking systems (such as Cohere rerank API) as gold rankings to train a listwise reranker based on Code-LLaMA-Instruct.

2 Utilizing LLMs as Unsupervised Rerankers

As the size of LLMs scales up (e.g., exceeding 10 billion parameters), it becomes increasingly difficult to fine-tune the reranking model. Addressing this challenge, recent efforts have attempted to prompt LLMs to directly enhance document reranking in an unsupervised way. In general, these prompting strategies can be divided into three categories: pointwise, listwise, and pairwise methods. A comprehensive exploration of these strategies follows in the subsequent sections.

The pointwise methods measure the relevance between a query and a single document, and can be categorized into two types: relevance generation and query generation .

The upper part in Figure 6 (a) shows an example of relevance generation based on a given prompt, where LLMs output a binary label (“Yes” or “No”) based on whether the document is relevant to the query. Following , the query-document relevance score f(q,d)f(q,d) can be calculated based on the log-likelihood of token “Yes” and “No” with a softmax function:

where SYS_{Y} and SNS_{N} represent the LLM’s log-likelihood scores of “Yes” and “No” respectively. In addition to binary labels, Zhuang et al. propose to incorporate fine-grained relevance labels (e.g., “highly relevant”, “somewhat relevant” and “not relevant”) into the prompt, which helps LLMs more effectively differentiate among documents with varying levels of relevance to a query.

As for the query generation shown in the lower part of Figure 6 (a), the query-document relevance score is determined by the average log-likelihood of generating the actual query tokens based on the document:

where ∣q∣|q| denotes the token number of query qq, dd denotes the document, and P\mathcal{P} represents the provided prompt. The documents are then reranked based on their relevance scores. It has been proven that some LLMs (such as T0) yield significant performance in zero-shot document reranking based on the query generation method . Recently, research has also shown that the LLMs that are pre-trained without any supervised instruction fine-tuning (such as LLaMA) also yield robust zero-shot ranking ability.

Although effective, these methods primarily rely on a handcrafted prompt (e.g., “Please write a query based on this document”), which may not be optimal. As prompt is a key factor in instructing LLMs to perform various NLP tasks, it is important to optimize prompt for better performance. Along this line, a discrete prompt optimization method Co-Prompt is proposed for better prompt generation in reranking tasks. Besides, PaRaDe proposes a difficulty-based method to select few-show demonstrations to include in the prompt, proving significant improvements compared with zero-shot prompts.

Note that these pointwise methods rely on accessing the output logits of LLMs to calculate the query-document relevance scores. As a result, they are not applicable to closed-sourced LLMs, whose API-returned results do not include logits.

2.2 Listwise Methods

Listwise methods aim to directly rank a list of documents (see Figure 6 (b)). These methods insert the query and a document list into the prompt and instruct the LLMs to output the reranked document identifiers. Due to the limited input length of LLMs, it is not feasible to insert all candidate documents into the prompt. To alleviate this issue, these methods employ a sliding window strategy to rerank a subset of candidate documents each time. This strategy involves ranking from back to front using a sliding window, re-ranking only the documents within the window at a time.

Although listwise methods have yielded promising performance, they still suffer from some weaknesses. First, according to the experimental results , only the GPT-4-based method can achieve competitive performance. When using smaller parameterized language models (e.g., FLAN-UL2 with 20B parameters), listwise methods may produce very few usable results and underperform many supervised methods. Second, the performance of listwise methods is highly sensitive to the document order in the prompt. When the document order is randomly shuffled, listwise methods perform even worse than BM25 , revealing positional bias issues in the listwise ranking of LLMs. To alleviate this issue, Tang et al. introduce a permutation self-consistency method, which involves shuffling the list in the prompt and aggregating the generated results to achieve a more accurate and unbiased ranking.

2.3 Pairwise Methods

In pairwise methods , LLMs are given a prompt that consists of a query and a document pair (see Figure 6 (c)). Then, they are instructed to generate the identifier of the document with higher relevance. To rerank all candidate documents, aggregation methods like AllPairs are used. AllPairs first generates all possible document pairs and aggregates a final relevance score for each document. To speed up the ranking process, efficient sorting algorithms, such as heap sort and bubble sort, are usually employed . These sorting algorithms utilize efficient data structures to compare document pairs selectively and elevate the most relevant documents to the top of the ranking list, which is particularly useful in top-kk ranking. Experimental results show the state-of-the-art performance on the standard benchmarks using moderate-size LLMs (e.g., Flan-UL2 with 20B parameters), which are much smaller than those typically employed in listwise methods (e.g., GPT3.5).

Although effective, pairwise methods still suffer from high time complexity. To alleviate the efficiency problem, a setwise approach has been proposed to compare a set of documents at a time and select the most relevant one from them. This approach allows the sorting algorithms (such as heap sort) to compare more than two documents at each step, thereby reducing the total number of comparisons and speeding up the sorting process.

2.4 Comparison and Discussion

In this part, we will compare different unsupervised methods from various aspects to better illustrate the strengths and weaknesses of each method, which is summarized in Table VI. We choose representative methods in pointwise, listwise and pairwise ranking, and include several supervised methods mentioned in Section 5.1 for performance comparison.

The pointwise methods (Query Generation and Relevance Generation) judge the relevance of each query-document pair independently, thus offering lower time complexity and enabling batch inference. However, compared to other methods, it does not have an advantage in terms of performance. The listwise method yields significant performance especially when calling GPT-4, but suffers from expensive API cost and non-reproducibility . Compared with the listwise method, the pairwise method shows competitive results based on a much smaller model FLAN-UL2 (20B). Stemming from the necessity to compare an extensive number of document pairs, its primary drawback is low efficiency.

3 Utilizing LLMs for Training Data Augmentation

Furthermore, in the realm of reranking, researchers have explored the integration of LLMs for training data augmentation . For example, ExaRanker generates explanations for retrieval datasets using GPT-3.5, and subsequently trains a seq2seq ranking model to generate relevance labels along with corresponding explanations for given query-document pairs. InPars-Light is proposed as a cost-effective method to synthesize queries for documents by prompting LLMs. Contrary to InPars-Light , a new dataset ChatGPT-RetrievalQA is constructed by generating synthetic documents based on LLMs in response to user queries.

Recently, many studies have also attempted to distill the document ranking capability of LLMs into a specialized model. RankVicuna proposes to use the ranking list of RankGPT3.5\text{RankGPT}_{3.5} as the gold list to train a 7B parameter Vicuna model. RankZephyr introduces a two-stage training strategy for distillation: initially applying the RankVicuna recipe to train Zephyrγ\text{Zephyr}{\gamma} in the first stage, and then further finetuning it in the second stage with the ranking results from RankGPT4\text{RankGPT}_{4}. These two studies not only demonstrate competitive results but also alleviate the issue of ranking results non-reproducibility of black-box LLMs. Besides, researchers have also tried to distill the ranking ability of a pairwise ranker, which is computationally demanding, into a simpler but more efficient pointwise ranker.

4 Limitations

Although recent research on utilizing LLMs for document reranking has made significant progress, it still faces some challenges. For example, considering the cost and efficiency, minimizing the number of calls to LLM APIs is a problem worth studying. Besides, while existing studies mainly focus on applying LLMs to open-domain datasets (such as MSMARCO ) or relevance-based text ranking tasks, their adaptability to in-domain datasets and non-standard ranking datasets remains an area that demands more comprehensive exploration.

Reader

With the impressive capabilities of LLMs in understanding, extracting, and processing textual data, researchers explore expanding the scope of IR systems beyond content ranking to answer generation. In this evolution, a reader module has been introduced to generate answers based on the document corpus in IR systems. By integrating a reader module, IR systems can directly present conclusive passages to users. Compared with providing a list of documents, users can simply comprehend the answering passages instead of analyzing the ranking list in this new paradigm. Furthermore, by repeatedly providing documents to LLMs based on their generating texts, the final generated answers can potentially be more accurate and information-rich than the original retrieved lists.

A naive strategy for implementing this function is to heuristically provide LLMs with documents relevant to the user queries or the previously generated texts to support the following generation. However, this passive approach limits LLMs to merely collecting documents from IR systems without active engagement. An alternative solution is to train LLMs to interact proactively with search engines. For example, LLMs can formulate their own queries instead of relying solely on user queries or generated texts for references. According to the way LLMs utilize IR systems in the reader module, we can categorize them into passive readers and active readers. Each approach has its advantages and challenges for implementing LLM-powered answer generation in IR systems. Furthermore, since the documents provided by upstream IR systems are sometimes too long to directly feed as input for LLMs, some compression modules are proposed to extractively or abstractively compress the retrieved contexts for LLMs to understand and generate answers for queries. We will present these reader and compressor modules in the following parts and briefly introduce the existing analysis work on retrieval-augmented generation strategy and their applications.

To generate answers for users, a straightforward strategy is to supply the retrieved documents according to the queries or previously generated texts from IR systems as inputs to LLMs for creating passages . By this means, these approaches use the LLMs and IR systems separately, with LLMs functioning as passive recipients of documents from the IR systems. The strategies for utilizing LLMs within IR systems’ reader modules can be categorized into the following three groups according to the frequency of retrieving documents for LLMs.

To obtain useful references for LLMs to generate responses for user queries, an intuitive way is to retrieve the top documents based on the queries themselves in the beginning. For example, REALM adopts this strategy by directly attending the document contents to the original queries to predict the final answers based on masked language modeling. RAG follows this strategy but applies the generative language modeling paradigm. However, these two approaches only use language models with limited parameters, such as BERT and BART. Recent approaches such as REPLUG and Atlas have improved them by leveraging LLMs such as GPTs, T5s, and LLaMAs for response generation. To yield better answer generation performances, these models usually fine-tune LLMs on QA tasks. However, due to the limited computing resources, many methods choose to prompt LLMs for generation as they could use larger LMs in this way. Furthermore, to improve the quality of the generated answers, several approaches also try to train or prompt the LLMs to generate contexts such as citations or notes in addition to the answers to force LLMs to understand and assess the relevance of retrieved passages to the user queries. Some approaches evaluate the importance of each retrieved reference using policy gradients to indicate which reference is more useful for generating. Besides, researchers explore instruction tuning LLMs such LLaMAs to improve their abilities to generate conclusive passages relying on retrieved knowledge .

1.2 Periodic-Retrieval Reader

However, while generating long conclusive answers, it is shown that only using the references retrieved by the original user intents as in once-retrieval readers may be inadequate. For example, when providing a passage about “Barack Obama”, language models may need additional knowledge about his university, which may not be included in the results of simply searching the initial query. In conclusion, language models may need extra references to support the following generation during the generating process, where multiple retrieval processes may be required. To address this, solutions such as RETRO and RALM have emerged, emphasizing the periodic collection of documents based on both the original queries and the concurrently generated texts (triggering a retrieval every nn generated tokens). In this manner, when generating the text about the university career of Barack Obama, the LLM can receive additional documents as supplementary materials. This need for additional references highlights the necessity for multiple retrieval iterations to ensure robustness in subsequent answer generation. Notably, RETRO introduces a novel approach incorporating cross-attention between the generating texts and the references within the Transformer attention calculation, as opposed to directly embedding references into the input texts of LLMs. Since it involves additional cross-attention modules in the Transformer’s structure, RETRO trains this model from scratch. However, these two approaches mainly rely on the successive nn tokens to separate generation and retrieve documents, which may not be semantically continuous and may cause the collected references noisy and useless. To solve this problem, some approaches such as IRCoT also explore retrieving documents for every generated sentence, which is a more complete semantic structure. Furthermore, researchers find that the whole generated passages can be considered as conclusive contexts for current queries and can be used to find more relevant knowledge to generate more thorough answers. Consequently, many recent approaches have also tried to extend this periodic-retrieval paradigm to iteratively using the whole generated passages to retrieve references to re-generate the answers, until the iterations reach a pre-defined limitation. Particularly, these methods can be regarded as special periodic-retrieval readers that retrieve passages when every answer is (re)-generated. Since the LLMs can receive more comprehensive and relevant references with the iterations increase, these methods that combine retrieval-augmented-generation and generation-augmented-retrieval strategies can generate more accurate answers but consume more computation costs.

1.3 Aperiodic-Retrieval Reader

In the above strategy, the retrieval systems supply documents to LLMs in a periodic manner. However, retrieving documents in a mandatory frequency may mismatch the retrieval timing and can be costly. Recently, FLARE has addressed this problem by automatically determining the timing of retrieval according to the probability of generating texts. Since the probability can serve as an indicator of LLMs’ confidence during text generation , a low probability for a generated term could suggest that LLMs require additional knowledge. Specifically, when the probability of a term falls below a predefined threshold, FLARE employs IR systems to retrieve references in accordance with the ongoing generated sentences, while removing these low-probability terms. FLARE adopts this strategy of prompting LLMs for answer generation solely based on the probabilities of generating terms, avoiding the need for fine-tuning while still maintaining effectiveness. Besides, self-RAG tends to solve this problem by training LLMs such as LlaMA to generate specific tokens when they need additional knowledge to support following generations. Another critical model is introduced to judge whether the retrieved references are beneficial for generating.

We summarize representative passive reader approaches in Table VII, considering various aspects such as the backbone language models, the insertion point for retrieved references, the timing of using retrieval models, and the tuning strategy employed for LLMs.

2 Active Reader

However, the passive reader-based approaches separate IR systems and generative language models. This signifies that LLMs can only submissively utilize references provided by IR systems and are unable to interactively engage with the IR systems in a manner akin to human interaction such as issuing queries to seek information.

To allow LLMs to actively use search engines, Self-Ask and DSP try to employ few-shot prompts for LLMs, triggering them to search queries when they believe it is required. For example, in a scenario where the query is “When was the existing tallest wooden lattice tower built?”, these prompted LLMs can decide to search a query “What is the existing tallest wooden lattice tower” to gather necessary references as they find the query cannot be directly answered. Once acquired information about the tower, they can iteratively query IR systems for more details until they determine to generate the final answers instead of asking questions. Notably, these methods involve IR systems to construct a single reasoning chain for LLMs. MRC further improves these methods by prompting LLMs to explore multiple reasoning chains and subsequently combining all generated answers using LLMs.

3 Compressor

Existing LLMs, especially open-sourced ones, such as LLaMA and Flan-T5, have limited input lengths (usually 4,096 or 8,192 tokens). However, the documents or web pages retrieved by upstream IR systems are usually long. Therefore, it is difficult to concatenate all the retrieved documents and feed them into LLMs to generate answers. Though some approaches manage to solve these problems by aggregating the answers supported by each reference as the final answers, this strategy neglects the potential relations between retrieved passages. A more straightforward way is to directly compress the retrieved documents into short input tokens or even dense vectors .

To compress the retrieved references, an intuitive idea is to extract the most useful KK sentences from the retrieved documents. LeanContext applies this method and trains a small model by reinforcement learning (RL) to select the top KK similar sentences to the queries. The researchers also augment this strategy by using a free open-sourced text reduction method for the rest sentences as a supplement. Instead of using RL-based methods, RECOMP directly uses the probability or the match ratio of the generated answers to the golden answers as signals to build training datasets and tune the compressor model. For example, the sentence corresponding to the highest generating probability is the positive one while others are negative ones. Furthermore, FILCO applies the “hindsight” methods, which directly align the prior distribution (the predicted importance probability distribution of sentences without knowing the gold answer) to the posterior distribution (the same distribution of sentences within knowing the gold answer) to tune language models to select sentences.

However, these extractive methods may lose potential intent among all references. Therefore, abstractive methods are proposed to summarize retrieved documents into short but concise summaries for downstream generation. These methods usually distill the summarizing abilities of LLMs to small models. For example, TCRA leverages GPT-3.5-turbo to build abstractive compression datasets for MT5 model.

4 Analysis

With the rapid development of the above reader approaches, many researchers have begun to analyze the characteristics of retrieval-augmented LLMs:

∙\bullet Liu et al. find that the position of the relevant/golden reference has significant influences on the final generation performance. The performance is always better when the relevant reference is at the beginning or the end, which indicates the necessity of introducing a ranking module to order the retrieved knowledge.

∙\bullet Ren et al. observe that by applying retrieval augmentation generation strategy, LLMs can have a better awareness of their knowledge boundaries.

∙\bullet Liu et al. analyze different strategies of integrating retrieval systems and LLMs such as concatenate (i.e., concatenating all references for answer generation) and post fusion (i.e., aggregating the answers corresponding to each reference). They also explore several ways of combining these two strategies.

∙\bullet Aksitov et al. demonstrate that there exists an attribution and fluency tradeoff for retrieval-augmented LLMs: with more received references, the attribution of generated answers increases while the fluency decreases.

∙\bullet Mallen et al. argue that always retrieving references to support LLMs to generate answers hurts the question-answering performance. The reason is that LLMs themselves may have adequate knowledge while answering questions about popular entities and the retrieved noisy passages may interfere and bias the answering process. To overcome this challenge, they devise a simple strategy that only retrieves references while the popularity of entities in the query is quite low. By this means, the efficacy and efficiency of retrieval-augmented generation both improve.

5 Applications

Recently, researchers have applied the retrieval-augmented generation strategy to areas such as clinical QA, medical QA, and financial QA to enhance LLMs with external knowledge and to develop domain-specific applications. For example, ATLANTIC adapts Atlas to the scientific domain to derive a science QA system. Besides, some approaches also apply techniques in federated learning such as multi-party computation to perform personal retrieval-augmented generation with privacy protection.

Furthermore, to better facilitate the deployment of these retrieval-augmented generation systems, some tools or frameworks are proposed . For example, RETA-LLM breaks down the whole complex generation task into several simple modules in the reader pipeline. These modules include a query rewriting module for refining query intents, a passage extraction module for aligning reference lengths with LLM limitations, and a fact verification module for confirming the absence of fabricated information in the generated answers.

6 Limitations

Several IR systems applying the retrieval-augmented generation strategy, such as New Bing and Langchain, have already entered commercial use. However, there are also some challenges in this novel retrieval-augmented content generation system. These include challenges such as effective query reformulation, optimal retrieval frequency, correct document comprehension, accurate passage extraction, and effective content summarization. It is crucial to address these challenges to effectively realize the potential of LLMs in this paradigm.

Search Agent

With the development of LLMs, IR systems are also facing new changes. Among them, developing LLMs as intelligent agents has attracted more and more attention. This conceptual shift aims to mimic human browsing patterns, thereby enhancing the capability of these models to handle complex retrieval tasks. Empowered by the advanced natural language understanding and generation capabilities of LLMs, these agents can autonomously search, interpret, and synthesize information from a wide range of sources.

One way to achieve this ability is to design a pipeline that combines a series of modules and assigns different roles to them. Such a pre-defined pipeline mimics users’ behaviors on the web by breaking it into several sub-tasks which are performed by different modules. However, this kind of static agent cannot deal with the complex nature of users’ behavior sequences on the web and may face challenges when interacting with real-world environments. An alternative solution is to allow LLMs to freely explore the web and make interactions themselves, namely letting the LLM itself decide what action it will take next based on the feedback from the environment (or humans). These agents have more flexibility and act more like human beings.

To mimic human search patterns, a straightforward approach is to design a static system to browse the web and synthesize information step by step . By breaking the information-seeking process into multiple subtasks, they design a pipeline that contains various LLM-based modules in advance and assigns different subtasks to them.

LaMDA serves as an early work of the static agent. It consists of a family of Transformer-based neural language models specialized for dialog, with up to 137B parameters, pre-trained on 1.56T tokens from public dialogue data and web text. The study emphasizes the model’s development through a static pipeline, encompassing large-scale pre-training, followed by strategic fine-tuning stages aimed at enhancing three critical aspects: dialogue quality, safety, and groundedness. It can integrate external IR systems for factual grounding. This integration allows LaMDA to access and use external and authoritative sources when generating responses. SeeKeR also incorporates the Internet search into its modular architecture for generating more factual responses. It performs three sequential tasks: generating a search query, generating knowledge from search results, and generating a final response. GopherCite uses a search engine like Google Search to find relevant sources. It then synthesizes a response that includes verbatim quotes from these sources as evidence, aligning the Gopher’s output with verified information. WebAgent develops a series of tasks, including instruction decomposition and planning, action programming, and HTML summarization. It can navigate the web, understand and synthesize information from multiple sources, and execute web-based tasks, effectively functioning as an advanced search and interaction agent. WebGLM designs an LLM-augmented retriever, a bootstrapped generator, and a human preference-aware scorer. These components work together to provide accurate web-enhanced question-answering capabilities that are sensitive to human preferences. Shi et al. focus on enhancing the relevance, responsibility, and trustworthiness of LLMs in web search applications via an intent-aware generator, an evidence-sensitive validator, and a multi-strategy supported optimizer.

2 Dynamic Agent

Instead of statically arranging LLMs in a pipeline, WebGPT takes an alternate approach by training LLMs to use search engines automatically. This is achieved through the application of a reinforcement learning framework, within which a simulated environment is constructed for GPT-3 models. Specifically, the WebGPT model employs special tokens to execute actions such as querying, scrolling through rankings, and quoting references on search engines. This innovative approach allows the GPT-3 model to use search engines for text generation, enhancing the reliability and real-time capability of the generated texts. A following study has extended this paradigm to the domain of Chinese question answering. Besides, some works develop important benchmarks for interactive web-based agents . For example, WebShop aims to provide a scalable, interactive web-based environment for language understanding and decision-making, focusing on the task of online shopping. ASH (Actor-Summarizer-Hierarchical) prompting significantly enhances the ability of LLMs on WebShop benchmark. It first takes a raw observation from the environment and produces a new, more meaningful representation that aligns with the specific goal. Then, it dynamically predicts the next action based on the summarized observation and the interaction history.

3 Limitations

Though the aspect of static search agents has been thoroughly studied, the literature on dynamic search agents remains limited. Some agents may lack mechanisms for real-time fact-checking or verification against authoritative sources, leading to the potential dissemination of misinformation. Moreover, since LLMs are trained on data from the Internet, they may inadvertently perpetuate biases present in the training data. This can lead to biased or offensive outputs and may collect unethical content from the web. Finally, as LLMs process user queries, there are concerns regarding user privacy and data security, especially if sensitive or personal information is involved in the queries.

Future Direction

In this survey, we comprehensively reviewed recent advancements in LLM-enhanced IR systems and discussed their limitations. Since the integration of LLMs into IR systems is still in its early stages, there are still many opportunities and challenges. In this section, we summarize the potential future directions in terms of the four modules in an IR system we just discussed, namely query rewriter, retriever, reranker, and reader. In addition, as evaluation has also emerged as an important aspect, we will also introduce the corresponding research problems that need to be addressed in the future. Another discussion about important research topics on applying LLMs to IR can be found in a recent perspective paper .

LLMs have enhanced query rewriting for both ad-hoc and conversational search scenarios. Most of the existing methods rely on prompting LLMs to generate new queries. While yielding remarkable results, the refinement of rewriting quality and the exploration of potential application scenarios require further investigation.

∙\bullet Rewriting queries according to ranking performance. A typical paradigm of prompting-based methods is providing LLMs with several ground-truth rewriting cases (optional) and the task description of query rewriting. Despite LLMs being capable of identifying potential user intents of the query , they lack awareness of the resulting retrieval quality of the rewritten query. The absence of this connection can result in rewritten queries that seem correct yet produce unsatisfactory ranking results. Although some existing studies have used reinforcement learning to adjust the query rewriting process according to generation results , a substantial realm of research remains unexplored concerning the integration of ranking results.

∙\bullet Improving query rewriting in conversational search. As yet, primary efforts have been made to improve query rewriting in ad-hoc search. In contrast, conversational search presents a more developed landscape with a broader scope for LLMs to contribute to query understanding. By incorporating historical interactive information, LLMs can adapt system responses based on user preferences, providing a more effective conversational experience. However, this potential has not been explored in depth. In addition, LLMs could also be used to simulate user behavior in conversational search scenarios, providing more training data, which are urgently needed in current research.

∙\bullet Achieving personalized query rewriting. LLMs offer valuable contributions to personalized search through their capacity to analyze user-specific data. In terms of query rewriting, with the excellent language comprehension ability of LLMs, it is possible to leverage them to build user profiles based on users’ search histories (e.g., issued queries, click-through behaviors, and dwell time). This empowers the achievement of personalized query rewriting for enhanced IR and finally benefits personalized search or personalized recommendation.

2 Retriever

Leveraging LLMs to improve retrieval models has received considerable attention, promising an enhanced understanding of queries and documents for improved ranking performance. However, despite strides in this field, several challenges and limitations still need to be investigated in the future:

∙\bullet Reducing the latency of LLM-based retrievers. LLMs, with their massive parameters and world knowledge, often entail high latency during the inferring process. This delay poses a significant challenge for practical applications of LLM-based retrievers, as search engines require in-time responses. To address this issue, promising research directions include transferring the capabilities of LLMs to smaller models, exploring quantization techniques for LLMs in IR tasks, and so on.

∙\bullet Simulating realistic queries for data augmentation. Since the high latency of LLMs usually blocks their online application for retrieval tasks, many existing studies have leveraged LLMs to augment training data, which is insensitive to inference latency. Existing methods that leverage LLMs for data augmentation often generate queries without aligning them with real user queries, leading to noise in the training data and limiting the effectiveness of retrievers. As a consequence, exploring techniques such as reinforcement learning to enable LLMs to simulate the way that real queries are issued holds the potential for improving retrieval tasks.

∙\bullet Incremental indexing for generative retrieval. As elaborated in Section 4.2.2, the emergence of LLMs has paved the way for generative retrievers to generate document identifiers for retrieval tasks. This approach encodes document indexes and knowledge into the LLM parameters. However, the static nature of LLM parameters, coupled with the expensive fine-tuning costs, poses challenges for updating document indexes in generative retrievers when new documents are added. Therefore, it is crucial to explore methods for constructing an incremental index that allows for efficient updates in LLM-based generative retrievers.

∙\bullet Supporting multi-modal search. Web pages usually contain multi-modal information, including texts, images, audios, and videos. However, existing LLM-enhanced IR systems mainly support retrieval for text-based content. A straightforward solution is to replace the backbone with multi-modal large models, such as GPT-4 . However, this undoubtedly increases the cost of deployment. A promising yet challenging direction is to combine the language understanding capability of LLMs with existing multi-modal retrieval models. By this means, LLMs can contribute their language skills in handling different types of content.

3 Reranker

In Section 5, we have discussed the recent advanced techniques of utilizing LLMs for the reranking task. Some potential future directions in reranking are discussed as follows.

∙\bullet Enhancing the online availability of LLMs. Though effective, many LLMs have a massive number of parameters, making it challenging to deploy them in online applications. Besides, many reranking methods rely on calling LLM APIs, incurring considerable costs. Consequently, devising effective approaches (such as distilling to small models) to enhance the online applicability of LLMs emerges as a research direction worth exploring.

∙\bullet Improving personalized search. Many existing LLM-based reranking methods mainly focus on the ad-hoc reranking task. However, by incorporating user-specific information, LLMs can also improve the effectiveness of the personalized reranking task. For example, by analyzing users’ search history, LLMs can construct accurate user profiles and rerank the search results accordingly, providing personalized results with higher user satisfaction.

∙\bullet Adapting to diverse ranking tasks. In addition to document reranking, there are also other ranking tasks, such as response ranking, evidence ranking, entity ranking and etc., which also belong to the universal information access system. Navigating LLMs towards adeptness in these diverse ranking tasks can be achieved through specialized methodologies, such as instruction tuning. Exploring this avenue holds promise as an intriguing and valuable research trajectory.

4 Reader

With the increasing capabilities of LLMs, the future interaction between users and IR systems will be significantly changed. Due to the powerful natural language processing and understanding capabilities of LLMs, the traditional search paradigm of providing ranking results is expected to be progressively replaced by synthesizing conclusive answering passages for user queries using the reader module. Although such strategies have already been investigated by academia and facilitated by industry as we stated in Section VII, there still exists much room for exploration.

∙\bullet Improving the reference quality for LLMs. To support answer generation, existing approaches usually directly feed the retrieved documents to the LLMs as references. However, since a document usually covers many topics, some passages in it may be irrelevant to the user queries and can introduce noise during LLMs’ generation. Therefore, it is necessary to explore techniques for extracting relevant snippets from retrieved documents, enhancing the performance of retrieval-augmented generation.

∙\bullet Improving the answer reliability of LLMs. Incorporating the retrieved references has significantly alleviated the “hallucination” problem of LLMs. However, it remains uncertain whether the LLMs refer to these supported materials during answering queries. Some studies have revealed that LLMs can still provide unfaithful answers even with additional references. Therefore, the reliability of the conclusive answers might be lower compared to the ranking results provided by traditional IR systems. It is essential to investigate the influence of these references on the generation process, thereby improving the credibility of reader-based novel IR systems.

5 Search Agent

With the outstanding performance of LLMs, the patterns of searching may completely change from traditional IR systems to autonomous search agents. In Section 7, we have discussed many existing works that utilize a static or dynamic pipeline to autonomously browse the web. These works are believed to be the pioneering works of the new searching paradigm. However, there is still plenty of room for further improvements.

∙\bullet Enhancing the Trustworthiness of LLMs. When LLMs are enabled to browse the web, it is important to ensure the validity of retrieved documents. Otherwise, the unfaithful information may increase the LLMs’ “hallucination” problem. Besides, even if the gathered information has high quality, it remains unclear whether they are really used for synthesizing responses. A potential strategy to address this issue is enabling LLMs to autonomously validate the documents they scrape. This self-validation process could incorporate mechanisms for assessing the credibility and accuracy of the information within these documents.

∙\bullet Mitigating Bias and Offensive Content in LLMs. The presence of biases and offensive content within LLM outputs is a pressing concern. This issue primarily stems from biases inherent in the training data and will be amplified by the low-quality information gathered from the web. Achieving this requires a multi-faceted approach, including improvements in training data, algorithmic adjustments, and continuous monitoring for bias and inappropriate content that LLMs collect and generate.

6 Evaluation

LLMs have attracted significant attention in the field of IR due to their strong ability in context understanding and text generation. To validate the effectiveness of LLM-enhanced IR approaches, it is crucial to develop appropriate evaluation metrics. Given the growing significance of readers as integral components of IR systems, the evaluation should consider two aspects: assessing ranking performance and evaluating generation performance.

∙\bullet Generation-oriented ranking evaluation. Traditional evaluation metrics for ranking primarily focus on comparing the retrieval results of IR models with ground-truth (relevance) labels. Typical metrics include precision, recall, mean reciprocal rank (MRR) , mean average precision (MAP), and normalized discounted cumulative gain (nDCG) . These metrics measure the alignment between ranking results and human preference on using these results. Nevertheless, these metrics may fall short in capturing a document’s role in the generation of passages or answers, as their relevance to the query alone might not adequately reflect this aspect. This effect could be leveraged as a means to evaluate the usefulness of documents more comprehensively. A formal and rigorous evaluation metric for ranking that centers on generation quality has yet to be defined.

∙\bullet Text generation evaluation. The wide application of LLMs in IR has led to a notable enhancement in their generation capability. Consequently, there is an imperative demand for novel evaluation strategies to effectively evaluate the performance of passage or answer generation. Previous evaluation metrics for text generation have several limitations, including: (1) Dependency on lexical matching: methods such as BLEU or ROUGE primarily evaluate the quality of generated outputs based on nn-gram matching. This approach cannot account for lexical diversity and contextual semantics. As a result, models may favor generating common phrases or sentence structures rather than producing creative and novel content. (2) Insensitivity to subtle differences: existing evaluation methods may be insensitive to subtle differences in generated outputs. For example, if a generated output has minor semantic differences from the reference answer but is otherwise similar, traditional methods might overlook these nuanced distinctions. (3) Lack of ability to evaluate factuality: LLMs are prone to generating “hallucination” problems . The hallucinated texts can closely resemble the oracle texts in terms of vocabulary usage, sentence structures, and patterns, while having non-factual content. Existing methods are hard to identify such problems, while the incorporation of additional knowledge sources such as knowledge bases or reference texts could potentially aid in addressing this challenge.

7 Bias

Since ChatGPT was released, LLMs have drawn much attention from both academia and industry. The wide applications of LLMs have led to a notable increase in content on the Internet that is not authored by humans but rather generated by these language models. However, as LLMs may hallucinate and generate non-factual texts, the increasing number of LLM-generated contents also brings worries that these contents may provide fictitious information for users across IR systems. More severely, researchers show that some modules in IR systems such as retriever and reranker, especially those based on neural models, may prefer LLM-generated documents, since their topics are more consistent and the perplexity of them are lower compared with human-written documents. The authors refer to this phenomenon as the “source bias” towards LLM-generated text. It is challenging but necessary to consider how to build IR systems free from this category of bias.

Conclusion

In this survey, we have conducted a thorough exploration of the transformative impact of LLMs on IR across various dimensions. We have organized existing approaches into distinct categories based on their functions: query rewriting, retrieval, reranking, and reader modules. In the domain of query rewriting, LLMs have demonstrated their effectiveness in understanding ambiguous or multi-faceted queries, enhancing the accuracy of intent identification. In the context of retrieval, LLMs have improved retrieval accuracy by enabling more nuanced matching between queries and documents, considering context as well. Within the reranking realm, LLM-enhanced models consider more fine-grained linguistic nuances when re-ordering results. The incorporation of reader modules in IR systems represents a significant step towards generating comprehensive responses instead of mere document lists. The integration of LLMs into IR systems has brought about a fundamental change in how users engage with information and knowledge. From query rewriting to retrieval, reranking, and reader modules, LLMs have enriched each aspect of the IR process with advanced linguistic comprehension, semantic representation, and context-sensitive handling. As this field continues to progress, the journey of LLMs in IR portends a future characterized by more personalized, precise, and user-centric search encounters.

This survey focuses on reviewing recent studies of applying LLMs to different IR components and using LLMs as search agents. Beyond this, a more significant problem brought by the appearance of LLMs is: is the conventional IR framework necessary in the era of LLMs? For example, traditional IR aims to return a ranking list of documents that are relevant to issued queries. However, the development of generative language models has introduced a novel paradigm: the direct generation of answers to input questions. Furthermore, according to a recent perspective paper , IR might evolve into a fundamental service for diverse systems. For example, in a multi-agent simulation system , an IR component can be used for memory recall. This implies that there will be many new challenges in future IR.

References