Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations
Xu Huang, Jianxun Lian, Yuxuan Lei, Jing Yao, Defu Lian, Xing Xie
Introduction
Recommender systems (RSs) have become an essential component of the digital landscape, playing a significant role in helping users navigate the vast array of choices available across various domains such as e-commerce and entertainment. By analyzing user preferences, historical data, and contextual information, these systems can deliver personalized recommendations that cater to individual tastes. Over the years, recommender systems have evolved from simple collaborative filtering algorithms to more advanced hybrid approaches that integrate deep learning techniques. However, as users increasingly rely on conversational interfaces for discovering and exploring products, there is a growing need to develop more sophisticated and interactive recommendation systems that can understand and respond effectively to diverse user inquiries and intents in an conversational manner.
Large language models (LLMs), such as GPT-3 (Brown et al. 2020) and PaLM (Chowdhery et al. 2022), have made significant strides in recent years, demonstrating remarkable capabilities in artificial general intelligence and revolutionizing the field of natural language processing. A variety of practical tasks can be accomplished in the manner of users conversing with AI agents such as ChatGPT https://chat.openai.com/ and Claude https://claude.ai/. With their ability to understand context, generate human-like text, and perform complex reasoning tasks, LLMs can facilitate more engaging and intuitive interactions between users and RSs, thus offering promising prospects for the next generation of RSs. By integrating LLMs into RSs, it becomes possible to provide a more natural and seamless user experience that goes beyond traditional recommendation techniques, fostering a more timely understanding of user preferences and delivering more comprehensive and persuasive suggestions.
Despite their potential, leveraging LLMs for recommender systems is not without its challenges and limitations. Firstly, while LLMs are pretrained on vast amounts of textual data from the internet, covering various domains and demonstrating impressive general world knowledge, they may fail to capture fine-grained, domain-specific behavior patterns, especially in domains with massive training data. Secondly, LLMs may struggle to understand a domain well if the domain data is private and less openly accessible on the internet. Thirdly, LLMs lack knowledge of new items released after the collection of pretraining data, and fine-tuning with up-to-date data can be prohibitively expensive. In contrast, in-domain models can naturally address these challenges. A common paradigm to overcome these limitations is to combine LLMs with in-domain models, thereby filling the gaps and producing more powerful intelligence. Notable examples include AutoGPT https://github.com/Significant-Gravitas/Auto-GPT, HuggingGPT(Shen et al. 2023), and Visual ChatGPT(Wu et al. 2023). The core idea is to utilize LLMs as the “brains” and in-domain models as “tools” that extend LLMs’ capabilities when handling domain-specific tasks.
In this paper, we connect LLMs with traditional recommendation models for interactive recommender systems. We propose InteRecAgent (Interactive Recommender Agent), a framework explicitly designed to cater to the specific requirements and nuances of recommender systems, thereby establishing a more effective connection between the LLM’s general capabilities and the specialized needs of the recommendation domain. This framework consists of three distinct sets of tools, including querying, retrieval, and ranking, which are designed to cater to the diverse needs of users’ daily inquiries. Given the typically large number of item candidates, storing item names in the tools’ input and output as observations with prompts is impractical. Therefore, we introduce a “shared candidate bus” to store intermediate states and facilitate communication between tools. To enhance the capabilities of dealing with long conversations and even lifelong conversations, we introduce a “long-term and short-term user profile” module to track the preferences and history of the user, leveraged as the input of the ranking tool to improve personalization. The “shared candidate bus” along with the “long-term and short-term user profile” constitute the advanced memory mechanisms within the InteRecAgent framework.
Regarding task planning, we employ a “plan-first execution” strategy as opposed to a step-by-step approach. This strategy not only lowers the inference costs of LLMs but can also be seamlessly integrated with the dynamic demonstration strategy to enhance the quality of plan generation. Specifically, InteRecAgent generates all the steps of tool-calling at once and strictly follows the execution plan to accomplish the task. During the conversation, InteRecAgent parses the user’s intent and retrieves a few demonstrations that are most similar to the current intent. These dynamically retrieved demonstrations help LLMs formulate a correct task execution plan. In addition, we implement a reflection strategy, wherein another LLM acts as a critic to evaluate the quality of the results and identify any errors during the task execution. If the results are unsatisfactory or errors are detected, InteRecAgent reverts to the initial state and repeats the plan-then-tool-execution process.
Employing GPT-4 as the LLM within InteRecAgent has yielded impressive results in our experiments. This naturally leads to the attractive question: is it possible to harness a smaller language model to act as the brain? To explore this, we have developed an imitation dataset featuring tool plan generations derived from interactions between InteRecAgent and a user simulator, both powered by GPT-4. Through fine-tuning the LlaMA 2 (Touvron et al. 2023b) model with this dataset, we have created RecLlama. Remarkably, RecLlama surpasses several larger models in its effectiveness as the core of a recommender agent. Our main contributions are summarized as follows:
We propose InteRecAgent, a compact LLM-based agent framework that democratizes interactive recommender systems by connecting LLMs with three distinct sets of traditional recommendation tools.
In response to the challenges posed by the application of LLM-based agents in recommendation systems, we introduce a suite of advanced modules, including shared candidate bus, long-term and short-term user profile, dynamic demonstration-augmented plan-first strategy, and a reflection strategy.
To enable small language models to serve as the brain for recommender agents, we create an imitation dataset derived from GPT-4. Leveraging this dataset, we have successfully fine-tuned a 7-billion-parameter model, which we refer to as RecLlama.
Experimental results from three public datasets demonstrate the effectiveness of InteRecAgent, with particularly significant advantages in domains that are less covered by world knowledge.
Related Work
Existing researches in conversational recommender systems (CRS) can be primarily categorized into two main areas (Gao et al. 2021): attribute-based question-answering(Zou and Kanoulas 2019; Zou, Chen, and Kanoulas 2020; Xu et al. 2021) and open-ended conversation (Li et al. 2018; Wang et al. 2022b, 2021). In attribute-based question-answering CRS, the system aims to recommend suitable items to users within as few rounds as possible. The interaction between the system and users primarily revolves around question-answering concerning desired item attributes, iteratively refining user interests. Key research challenges in this area include developing strategies for selecting queried attributes(Mirzadeh, Ricci, and Bansal 2005; Zhang et al. 2018) and addressing the exploration-exploitation trade-off(Christakopoulou, Radlinski, and Hofmann 2016; Xie et al. 2021). In open-ended conversation CRS, the system manages free-format conversational data. Initial research efforts in this area focused on leveraging pretrained language models for conversation understanding and response generation(Li et al. 2018; Penha and Hauff 2020). Subsequent studies incorporated external knowledge to enhance the performance of open-ended CRS(Chen et al. 2019; Wang, Su, and Chen 2022; Wang et al. 2022b). Nevertheless, these approaches struggle to reason with complex user inquiries and maintain seamless communication with users. The emergence of LLMs presents an opportunity to revolutionize the construction of conversational recommender systems, potentially addressing the limitations of existing approaches and enhancing the overall user experience.
2 Enhancing LLMs
The scaling-up of parameters and data has led to significant advancements in the capabilities of LLMs, including in-context learning (Brown et al. 2020; Liu et al. 2021; Rubin, Herzig, and Berant 2021), instruction following (Ouyang et al. 2022; Touvron et al. 2023a; OpenAI 2023), planning and reasoning (Wei et al. 2022; Wang et al. 2022a; Yao et al. 2022; Yang et al. 2023; Wang et al. 2023b). In recommender systems, the application of LLMs is becoming a rapidly growing trend (Liu et al. 2023a; Dai et al. 2023; Kang et al. 2023; Wang and Lim 2023).
As models show emergent intelligence, researchers have started exploring the potential to leverage LLMs as autonomous agents (Wang et al. 2023a; Zhao, Jin, and Cheng 2023), augmented with memory modules, planning ability, and tool-using capabilities. For example, (Wang et al. 2023c; Zhong et al. 2023; Liu et al. 2023b) have equipped LLMs with an external memory, empowering LLMs with growth potential. Regarding the planning, CoT (Wei et al. 2022; Kojima et al. 2022) and ReAct (Yao et al. 2022) propose to enhance planning by step-wise reasoning; ToT (Yao et al. 2023) and GoT (Besta et al. 2023) introduce multi-path reasoning to ensure consistency and correctness; Self-Refine (Madaan et al. 2023) and Reflexion (Shinn et al. 2023) lead the LLMs to reflect on errors, with the ultimate goal of improving their subsequent problem-solving success rates. To possess domain-specific skills, some works (Qin et al. 2023a) study guiding LLMs to use external tools, such as a web search engine (Nakano et al. 2021; Shuster et al. 2022), mathematical tools (Schick et al. 2023; Thoppilan et al. 2022), code interpreters (Gao et al. 2023a; Chen et al. 2022) and visual models (Wu et al. 2023; Shen et al. 2023). To the best of our knowledge, this paper is the first to explore the LLM + tools paradigm in the field of recommender systems.
Methodologies
The comprehensive framework of InteRecAgent is depicted in Figure 1. Fundamentally, LLMs function as the brain, while recommendation models serve as tools that supply domain-specific knowledge. Users engage with an LLM using natural language. The LLM interprets users’ intentions and determines whether the current conversation necessitates the assistance of tools. For instance, in a casual chit-chat, the LLM will respond based on its own knowledge; whereas for in-domain recommendations, the LLM initiates a chain of tool calls and subsequently generates a response by observing the execution results of the tools. Consequently, the quality of recommendations relies heavily on the tools, making the composition of tools a critical factor in overall performance. To ensure seamless communication between users and InteRecAgent, covering both casual conversation and item recommendations, we propose a minimum set of tools that encompass the following aspects:
(1) Information Query. During conversations, the InteRecAgent not only handles item recommendation tasks but also frequently addresses users’ inquiries. For example, within a gaming platform, users may ask questions like, “What is the release date of this game and how much does it cost?” To accommodate such queries, we include an item information query module. This module can retrieve detailed item information from the backend database using Structured Query Language (SQL) expressions.
(2) Item Retrieval. Retrieval tools aim to propose a list of item candidates that satisfy a user’s demand from the entire item pool. These tools can be compared to the retrieval stage of a recommender system, which narrows down relevant candidates to a smaller list for large-scale serving. In InteRecAgent, we consider two types of demands that a user may express in their intent: hard conditions and soft conditions. Hard conditions refer to explicit demands on items, such as “I want some popular sports games” or “Recommend me some RPG games under $100”. Soft conditions pertain to demands that cannot be explicitly expressed with discrete attributes and require the use of semantic matching models, like “I want some games similar to Call of Duty and Fortnite”. It is essential to incorporate multiple tools to address both conditions. Consequently, we utilize an SQL tool to handle hard conditions, finding candidates from the item database. For soft conditions, we employ an item-to-item tool that matches similar items based on latent embeddings.
(3) Item Ranking. Ranking tools execute a more sophisticated prediction of user preferences on the chosen candidates by leveraging user profiles. Similar to the rankers in conventional recommender systems, these tools typically employ a one-tower architecture. The selection of candidates could emerge from the output of item retrieval tools or be directly supplied by users, as in queries like “Which one is more suitable for me, item A or item B?”. Ranking tools guarantee that the recommended items are not only pertinent to the user’s immediate intent but also consonant with their broader preferences.
LLMs have the potential to handle various user inquiries when supplemented with these diverse tools. For instance, a user may ask, “I’ve played Fortnite and Call of Duty before. Now, I want to play some puzzle games with a release date after Fortnite’s. Do you have any recommendations?” In this scenario, the tool execution sequence would be “SQL Query Tool SQL Retrieval Tool Ranker Tool.” First, the release date of Fortnite is queried, then the release date and puzzle genre are interpreted as hard conditions for the SQL retrieval. Finally, Fortnite and Call of Duty are considered as the user profile for the ranking model.
Typically, the tool augmentation is implemented via ReAct (Yao et al. 2022), where LLMs generate reasoning traces, actions, and observations in an interleaved manner. We refer to this style of execution as step-by-step. Our initial implementation also employed the step-by-step approach. However, we soon observed some limitations due to various challenges. Firstly, retrieval tools may return a large number of items, resulting in an excessively long observation prompt for LLMs. Additionally, including numerous entity names in the prompt can degrade LLMs performance. Secondly, despite their powerful intelligence, LLMs may use tools incorrectly to complete tasks, such as selecting the wrong tool to call or omitting key execution steps. To tackle these challenges, we enhance the three critical components of a typical LLM-based agent, namely memory (Section 3.2), task planning (Section 3.3 and 3.4), and tool learning abilities (Section 3.5).
2 Memory Mechanism
The large number of items can pose a challenge when attempting to include items generated by tools in prompts as observations for the LLM, due to input context length limitations. Meanwhile, the input of a subsequent tool often depends on the output of preceding tools, necessitating effective communication between tools. Thus, we propose Candidate Bus, which is a separate memory to store the current item candidates, eliminating the need to append them to prompt inputs. The Candidate Bus, accessible by all tools, comprises two parts: a data bus for storing candidate items, and a tracker for recording each tool’s output.
The candidate items in the data bus are initialized to include all items at the beginning of each conversation turn by default. At the start of each tool execution, candidate items are read from the data bus, and the data bus is then refreshed with the filtered items at the end of each tool execution. This mechanism allows candidate items to flow sequentially through the various tools in a streaming manner. Notably, users may explicitly specify a set of candidate items in the conversation, such as “Which of these movies do you think is most suitable for me: [Movie List]?” In this case, the LLM will call a special tool—the memory initialization tool—to set the user-specified items as the initial candidate items.
The tracker within the memory serves to record tool execution. Each tool call record is represented as a triplet , where denotes the name of the -th tool, and are the input and output of the tool’s execution, such as the number of remaining candidates, runtime errors. The tracker’s main function is to aid the critic in making judgments within the reflection mechanism, acting as the in , as described in Section 3.4.
With the help of the Candidate Bus component, items can be transmitted in a streaming manner between various tools and continuously filtered according to conditions, presenting a funnel-like structure for the recommendation. The tracker’s records can be considered as short-term memory for further reflection. We depict an example of the memory bus in the upper of Figure 3.
User Profile
To facilitate the invocation of tools, we explicitly maintain a user profile in memory. This profile is structured as a dictionary that encapsulates three facets of user preference: “like”, “dislike”, and “expect”. The “like” and “dislike” facets reflect the user’s favorable and unfavorable tastes, respectively, whereas “expect” monitors the user’s immediate requests during the current dialogue, such as conducting a search, which is not necessarily indicative of the user’s inherent preferences. Each facet may contain content that includes item names or categories.
User profiles are synthesized by LLMs based on conversation history. To address situations where the conversation history grows excessively long, such as in lifelong learning scenarios where conversations from all days may be stored for ongoing interactions, we devise two distinct user profiles: one representing long-term memory and another for short-term memory. Should the current dialogue exceed the LLM’s input window size, we partition the dialogue, retrieve the user profile from the preceding segment, and merge it with the existing long-term memory to update the memory state. The short-term memory is consistently derived from the most recent conversations within the current prompt. When it comes to tool invocation, a comprehensive user profile is formed by the combination of both long-term and short-term memories.
3 Plan-first Execution with Dynamic Demonstrations
Rather than using the step-by-step approach, we adopt a two-phase method. In the first phase, we prompt the LLM to generate a complete tool execution plan based on the user’s intention derived from the dialogue. In the second phase, the LLM strictly adheres to the plan, calling tools in sequence while allowing them to communicate via the Candidate Bus. Concretely, the plan-first execution consists of the following two phases.
Plan: LLM accepts the user’s current input , dialogue context , descriptions of various tools , and demonstration for in-context learning. LLM formulates a tool usage plan based on user intent and preferences, providing inputs for each tool, i.e., , where consists of the tool and its input .
Execution: The tool executor invokes the tools step-by-step according to the plan and obtains outputs from each tool, i.e., . The output feedback of each tool is defined as , where only the item information from the last tool’s output serves as LLM’s observation for generating the response . The remaining information is tracked by the candidate memory bus for further reflection (see Section 3.4).
We summarize the differences between our plan-first execution strategy and step-by-step strategy in Table 1 from six aspects. Fundamentally, step-by-step strategy executes reasoning and action execution alternately, while our plan-first execution is a two-phase strategy, where a series of executions is conducted followed by one-time planning. In step-by-step strategy, the LLMs are responsible for thinking and reasoning at each step. The task entails reasoning for individual observation, resulting in-context learning being challenging due to the difficulty in crafting demonstrations comprising dynamic observations. Differently, the primary task of LLM in our plan-first execution is to make a tool utilizing plan, which could be easily guided by pairs. The foremost advantage of our plan-first execution resides in the reduction of API calls. When employing N steps to address a task, our strategy necessitates merely 2 API calls, as opposed to N+1 calls in ReAct. This leads to a decrease in latency, which is of particular importance in conversational settings.
In order to improve the planning capability of LLM, demonstrations are injected into prompts for in-context learning in the Plan phase. Each demonstration consists of a user intent and tool execution path . However, the number of demonstrations is strictly limited by the contextual length that LLM can process, which makes the quality of demonstrations of paramount importance. To address the challenge, we introduce a dynamic demonstration strategy, where only a few demonstrations that are most similar to current user intent are incorporated into the prompt. For example, if the current user input is “My game history is Call of Duty and Fortnite, please give me some recommendations”, then demonstration with user intent “I enjoyed ITEM1, ITEM2 in the past, give me some suggestions” may be retrieved as a high-quality demonstration.
4 Reflection
Despite LLM’s strong intelligence, it still exhibits occasional errors in reasoning and tool utilization (Madaan et al. 2023; Shinn et al. 2023). For example, it may violate instructions in the prompt by selecting a non-existent tool, omit or overuse some tools, or fail to prepare tool inputs in the proper format, resulting in errors in tool execution.
To reduce the occurrence of such errors, some studies have employed self-reflection (Shinn et al. 2023) mechanisms to enable LLM to have some error-correcting capabilities during decision-making. In InteRecAgent, we utilize an actor-critic reflection mechanism to enhance the agent’s robustness and the error-correcting ability. In the following part, we will formalize this self-reflection mechanism.
Assume that in the -th round, the dialogue context is and the current user input is . The actor is an LLM equipped with tools and inspired by the dynamic demonstration-augmented plan-first execution mechanism. For the user input, the actor would make a plan , obtain the tools’ output and generate the response . The critic evaluates the behavioral decisions of the actor. The execution steps of the reflection mechanism are listed as follows:
Step1: The critic evaluates the actor’s output , and under the current dialogue context and obtains the judgment .
Step2: When the judgment is positive, it indicates that the actor’s execution and response are reasonable, and the response is directly provided to the user, ending the reflection phase. When the judgment is negative, it indicates that the actor’s execution or response is unreasonable. The feedback is used as a signal to instruct the actor to rechain, which is used as the input of .
In the actor-critic reflection mechanism, the actor is responsible for the challenging plan-making task, while the critic is responsible for the relative simple evaluation task. The two agents cooperate on two different types of tasks and mutually reinforce each other through in-context interactions. This endows InteRecAgent with enhanced robustness to errors and improved error correction capabilities, culminating in more precise tool utilization and recommendations. An example of reflection is shown in the lower of Figure 3.
5 Tool Learning with Small Language Models
The default LLM served as the brain is GPT-4, chosen for its exceptional ability to follow instructions compared to other LLMs. We are intrigued by the possibility of distilling GPT-4’s proficiency in instruction-following to smaller language models (SLMs) such as the 7B-parameter Llama, aiming to reduce the costs associated with large-scale online services and to democratize our InteRecAgent framework to small and medium-sized business clients. To achieve this, we utilize GPT-4 to create a specialized dataset comprising pairs of [instructions, tool execution plans]. The “instruction” element encompasses both the system prompt and the user-agent conversation history, acting as the input to elicit a tool execution plan from the LLM; the “tool execution plan” is the output crafted by GPT-4, which serves as the target for fine-tuning Llama-7B. We denote the fine-tuned version of this model RecLlama.
To ensure the high quality of the RecLlama dataset, we employ two methods to generate data samples. The first method gathers samples from dialogues between a user simulator and a recommender agent, which is powered by GPT-4. Note that during one conversation, each exchange of user-agent produces one data sample, capturing the full range of GPT-4’s responses to the evolving context of the conversation. However, this method might not encompass a sufficiently diverse array of tool execution scenarios due to the finite number of training samples we can manage. Therefore, we complement this with a second method wherein we initially craft 30 varied dialogues designed to span a wide range of tool execution combinations. Then, for each iteration, we select three of these dialogues at random and prompt GPT-4 to generate both a conversation history and a suitable tool execution plan. This approach significantly enhances the diversity of the RecLlama dataset.
To evaluate RecLlama’s capacity for domain generalization, we limit the generation of training data to the Steam and MovieLens datasets, excluding the Beauty dataset (the details of datasets will be elaborated in Section 4.1). The final RecLlama dataset comprises 16,183 samples, with 13,525 derived from the first method and 2,658 from the second.
Experiments
Evaluating conversational recommender systems presents a challenge, as the seeker communicates their preferences and the recommendation agent provides suggestions through natural, open-ended dialogues. To enable the quantitative assessment of InteRecAgent, we design the following two evaluation strategies:
(1) User Simulator. We manually tune a role-playing prompt to facilitate GPT-4 in emulating real-world users with varying preferences. A simulated user’s preference is ascertained by injecting their historical behaviors into the role-playing prompt, leaving out the last item in their history as the target of their next interest. Following this, the simulated user engages with the recommendation agent to discover content that fits their interest. In this way, GPT-4 operates from the standpoint of the user, swiftly reacting to the recommended outcomes, thereby crafting a more natural dialogue scenario. This approach is utilized to assess the efficacy of InteRecAgent within multi-turn dialogue settings. An illustrative example of a user simulator prompt can be found in Figure 4.
The default configuration for the user simulator is set to “session-wise”. This implies that the agent will only access content within the current dialogue session, and its memory will be cleared once the user either successfully locates what they are seeking or fails to do so. The conversation turns in “session-wise” setting is usually limited, thus, the long-term memory module in InteRecAgent will not be activated. In order to assess the performance while handling “lifelong memory” (refer to Section 3.2), we have formulated two strategies for simulating extended dialogues. The first strategy, referred to as Long-Chat, mandates extended conversations between the user and the recommendation agent. This is achieved by alternately incorporating three types of chat intents within the user simulator: sharing history, detailing the target item, and participating in casual conversation. The simulator alternates between providing information (either historical or target-related) and casual chat every five rounds. During this process, if the agent mentions the target item, the conversation can be terminated and labeled as a success. The second strategy, referred to as Long-Context, initially synthesizes multi-day conversations utilizing user history. Subsequently, based on these extended dialogues, the user simulator interacts with the agent in a manner akin to the “session-wise” setting. For our method, the lengthy conversation history is loaded into the long-term memory module. However, for baseline methods, the extended conversation history will be truncated if it surpasses the maximum window size of the LLM.
(2) One-Turn Recommendation. Following the settings of traditional conversational recommender systems on ReDial (Li et al. 2018), we also adopt the one-turn recommendation strategy. Given a user’s history, we design a prompt that enables GPT-4 to generate a dialogue, thereby emulating the interaction between a user and a recommendation agent. The objective is to ascertain whether the recommendation agent can accurately suggest the ground truth item in its next response. We assess both the item retrieval task (retrieval from the entire space) and the ranking task (ranking of provided candidates). Specifically, the dialogue context is presented to the recommendation agent, accompanied by the instruction Please give me k recommendations based on the chat history for the retrieval task, and the instruction Please rank these candidate items based on the chat history for the ranking task. To ensure a fair comparison with baseline LLMs, the One-Turn Recommendation evaluation protocol employs only the “session-wise” setting, and the long-term memory module in InteRecAgent remains deactivated.
Dataset.
To compare methods across different domains, we conduct experiments using three datasets: Steamhttps://github.com/kang205/SASRec, MovieLenshttps://grouplens.org/datasets/movielens/10m and Amazon Beautyhttp://jmcauley.ucsd.edu/data/amazon/links.html. Each dataset comprises user-item interaction history data and item metadata. We apply the leave-one-out method to divide the interaction data into training, validation, and testing sets. The training of all utilized tools is performed on the training and validation sets. Due to budget constraints, we randomly sample 1000 and 500 instances from the testing set for user simulator and one-turn benchmarking respectively. For the lifelong simulator, due to the costly long conversation, we use 100 instances in evaluation.
Baselines.
As dialogue recommendation agents, we compare our methods with the following baselines:
Random: Sample k items uniformly from entire item set.
Popularity: Sample k items with item popularity as the weight.
LlaMA-2-7B-chat, LlaMA-2-13B-chat (Touvron et al. 2023b): The second version of the LlaMA model released by Meta.
Vicuna-v1.5-7B, Vicuna-v1.5-13B (Chiang et al. 2023): Open-source models fine-tuned with user-shared data from the ShareGPThttps://sharegpt.com/ based on LlaMA-2 foundation models.
Chat-Rec (Gao et al. 2023b): A recently proposed conversational recommendation agent utilizes a text-embedding tool (OpenAI text-embedding-ada-002) to retrieve candidates. It then processes the content with an LLM before responding to users. We denote the use of GPT-3.5 as the LLM in the second stage with ”Chat-Rec (3.5)” and the use of GPT-4 with ”Chat-Rec (4)”.
GPT-3.5, GPT-4 (OpenAI 2023): We access these LLMs from OpenAI by API service. The GPT-3.5 version in use is gpt-3.5-turbo-0613 and GPT-4 version is gpt-4-0613https://platform.openai.com/docs/models/.
For the LlaMA and Vicuna models, we employ the FastChat (Zheng et al. 2023) package to establish local APIs, ensuring their usage is consistent with GPT-3.5 and GPT-4.
Metrics.
Since both our method and baselines utilize LLMs to generate response, which exhibit state-of-the-art text generation capabilities, our experiments primarily compare recommendation performance of different methods. For the user simulator strategy, we employ two metrics: Hit@ and AT@, representing the success of recommending the target item within turns and the average turns (AT) required for a successful recommendation, respectively. Unsuccessful recommendations within k rounds are recorded as in calculating AT. In the one-turn strategy, we focus on the Recall@ and NDCG@ metric for retrieval and ranking task, respectively. In Recall@, the represents the retrieval of items, whereas in NDCG@, the denotes the number of candidates to be ranked.
Implementation Details.
We employ GPT-4 as the brain of the InteRecAgent for user intent parsing and tool planing. Regarding tools, we use SQL as information query tool, SQL and ItemCF (Linden, Smith, and York 2003) as hard condition and soft condition item retrieval tools, respectively, and SASRec (Kang and McAuley 2018) without position embedding as the ranking tool. SQL is implemented with SQLite integrated in pandasqlhttps://github.com/yhat/pandasql/ and retrieval and ranking models are implemented with PyTorch. The framework of InteRecAgent is implement with Python and LangChainhttps://www.langchain.com/. For dynamic demonstration selection, we employ sentence-transformershttps://huggingface.co/sentence-transformers to encode demonstrations into vectors and store them using ChromaDBhttps://www.trychroma.com/, which facilitates ANN search during runtime. Regarding hyperparameter settings, we set the number of dynamic demonstrations to 3, the maximum number of candidates for hard condition retrieval to 1000, and the threshold for soft condition retrieval cut to the top 5%.
2 Evaluation with User Simulator
Table 2 presents the results of evaluations conducted using the user simulator strategy. Our method surpasses other LLMs in terms of both hit rate and average turns across the three datasets. These results suggest that our InteRecAgent is capable of delivering more accurate and efficient recommendations in conversations compared to general LLMs. Overall, LLMs with larger parameter sizes perform better. GPT-3.5 and GPT4, with parameter sizes exceeding 100B, significantly outperform LlaMA2 and Vicuna-v1.5 13B models from the same series almost always surpass 7B models, except for LlaMA2-7B and LlaMA2-13B, which both perform extremely poorly on the Beauty dataset.
Another interesting observation is the more significant improvement in relatively private domains, such as Amazon Beauty. In comparison to gaming and movie domains, the beauty product domain is more private, featuring a larger number of items not well-covered by common world knowledge or being new. Table 2 reveals that GPT-3.5 and GPT-4 exhibit competitive performance in gaming and movie domains. However, in the Amazon Beauty domain, most LLMs suffer severe hallucination issue due to the professional, long, and complex item names, resulting in a significant drop in performance. This phenomenon highlights the necessity of recommender agents in private domains. Leveraging the text embedding retrieval tool, Chat-Rec shows superior performance compared to GPT-3.5 and GPT-4, but still falling short of the performance achieved by InteRecAgent. Chat-Rec can be seen as a simplified version of InteRecAgent, incorporating just a single tool within the agent’s framework. Consequently, Chat-Rec lacks the capability to handle multifaceted queries, such as procuring detailed information about an item or searching for items based on intricate criteria.
Lifelong conversation setting.
Table 3 and Table 4 demonstrate the performance of two lifelong memory configurations, specifically, Long-Chat and Long-Context. For Long-Chat, the recommender agent engages a maximum of 50 rounds of dialogue with the user simulator. In both configurations, InteRecAgent without long-term memory modules (denoted as “Ours” in the tables) consistently outperforms GPT-4 across all datasets, which validates the robustness of our tool-enhanced recommender agent framework. After activating the long-term memory modules, the performance gets further improved under both Long-Chat and Long-Context configurations. This confirms the necessity and effectiveness of memory on capturing user preference during lifelong interactions between the user and AI agent.
3 Evaluation with One-Turn Recommendation
In this part, we evaluate both the retrieval and ranking recommendation tasks. For the Retrieval task, we set the recommendation budget to 5 for all methods, with Recall@5 being the evaluation metric. For the Ranking task, we randomly sample 19 negative items, and together with the one positive item, they form the candidate list proactively provided by users. The evaluation metric for this task is NDCG@20. For Chat-Rec, we omit the results of on the Ranking task because Chat-Rec degenerates into GPTs when removing the embedding-based candidate retrieval stage.
The results are shown in Table 5. Based on the results, we can draw conclusions similar to those in Section 4.2. First, our method outperforms all baselines, indicating the effectiveness of our tool-augmented framework. Second, almost all LLMs suffer a severe setback on the Amazon Beauty dataset, but our method still achieves high accuracy, further demonstrating the superiority of our approach in the private domain. Notably, some LLMs underperform compared to random and popularity methods in ranking tasks, particularly in the Amazon dataset. This can be primarily attributed to LLMs not adhering to the ranking instructions, which arise due to LLMs’ uncertainty and produce out-of-scope items, especially for smaller LLMs.
4 Comparions of Different LLMs as the Brain
In previous experiments, we utilized GPT-4 as the LLM for the InteRecAgent framework. This section presents a comparative analysis of the performance when employing different LLMs within the InteRecAgent. Note that RecLlama is our finetuned 7B model introduced in Section 3.5. ToolLlaMA2-7B (Qin et al. 2023b) is another fine-tuned model designed to interact with external APIs in response to human instructions. Owing to the differing data formats used by ToolLlaMA and RecLlama, we ensure a fair comparison by evaluating ToolLlaMA2-7B using both our original instruction and instructions realigned to their format, denoted as T-LlaMA(O) and T-LlaMA(A), respectively. The outcomes are tabulated in Table 6.
Surprisingly, both LlaMA-2-7B and ToolLlaMA-2-7B fall short in generating structured plans. Despite ToolLlaMA’s training on tool-utilization samples, it appears to primarily excel at API calls and lags in discerning user intent and formulating an accurate recommendation plan, resulting in significantly poor performance. Another intriguing finding is that GPT-3.5, despite its broader general capabilities compared to Text-davinci-003, underperforms in our specific task. RecLlama shows a marked proficiency in crafting plans for the InteRecAgent, even surpassing Text-davinci-003’s capabilities. Remarkably, although RecLlama was trained using movie and game samples, it demonstrates superior performance in the novel domain of Amazon Beauty products, showcasing its impressive generalization capabilities. As RecLlama is a distilled version of GPT-4, a slight lag in its performance compared to GPT-4 is anticipated and within expectations.
5 Ablation Study
This paper introduces several key mechanisms to enhance LLM’s ability to better utilize tools. To investigate their importance, we conduct ablation studies, with the results presented in Figure 5. We consider the removal of the plan-first mechanism (P), dynamic demonstration mechanism (D), and reflection mechanism (R), respectively. Experiments are carried out using the user simulator setting, as it provides a more comprehensive evaluation, encompassing both accuracy (hit rate) and efficiency (average turn) metrics.
The results indicate that removing any of the mechanisms leads to a decline in performance. Among these mechanisms, the removal of the reflection mechanism has the most significant impact on performance, as it can correct tool input format errors and tool misuse. Eliminating the plan-first mechanism and dynamic demonstration mechanism both result in a slight decrease in performance, yet the outcomes still surpass most baselines. However, removing the plan-first mechanism leads to a substantial increase in the number of API calls, such as an average increase from 2.78 to 4.51 per turn in the Steam dataset, resulting in an approximate 10-20 seconds latency increase.
6 Case Study
To effectively visualize InteRecAgent’s performance, we present case studies in chit-chat and two domains: gaming and beauty products, as shown in Figure 6. We compare the outputs of GPT-4 and InteRecAgent for given user inputs.
In chit-chat scenario (Figure 6a), InteRecAgent retains the capabilities of GPT-4 while also possessing the added ability to query domain-specific data (such as the number of products), yielding more accurate information.
In the game domain (Figure 6b), user input conditions are complex, encompassing user history and various demands. GPT-4’s recommendations mostly align with conditions, except for a 3D game Northgard misidentified as 2D. InteRecAgent’s response adheres to user conditions, and notably, includes the subsequent game in the user’s historical sequence, RimWorld, owing to its superior ranking performance.
In the e-commerce domain (Figure 6c), GPT-4’s hallucination phenomenon intensifies, resulting in giving products not existing in Amazon platform. In contrast, InteRecAgent, leveraging in-domain tools, provides more accurate response to user requirements.
Conclusion
In this paper, we introduce InteRecAgent, a compact framework that transforms traditional recommender models into interactive systems by harnessing the power of LLMs. We identify a diverse set of fundamental tools, categorized into information query tools, retrieval tools, and ranking tools, which are dynamically interconnected to accomplish complex user inquiries within a task execution framework. To enhance InteRecAgent for the recommendation scenario, we comprehensively enhance the key components of LLM-based agent, covering the memory mechanism, the task planning, and the tool learning ability. Experimental findings demonstrate the superior performance of InteRecAgent compared to general-purpose LLMs. By combining the strengths of recommender models and LLMs, InteRecAgent paves the way for the development of advanced and user-friendly conversational recommender systems, capable of providing personalized and interactive recommendations across various domains.
References
Appendix A Dataset
To evaluate the performance of our methods, we conduct experiments on three datasets: Steam, MovieLens and Amazon Beauty. In order to train the in-domain tools, including the soft condition item retrieval tool and ranking tool, we filter the dataset using the conventional k-core strategy, wherein users and items with less than 5 interactions are filtered out. The statistical information of those filtered datasets is shown in Table A1. Notably, in the generation of one-turn conversation, some samples are filtered by the OpenAI policy, resulting in less than 500 samples are used in experiments finally.
Appendix B Prompts
In this section, we will share our prompts used in different components.
The overall task description is illustrated in Figure C1.
B.2 Tool Descriptions
We employ one SQL query tool, two item retrieval tools, one item ranking tool plus two auxiliary tools in InteRecAgent. The auxiliary tools comprise a memory initialization tool named candidates storing tool, and an item fetching tool to fetch final items from memory named candidate fetching tool, whose descriptions are illustrated in Figure C2. The description of query tool, retrieval tools and ranking tool are illustrated in Figure C3, Figure C4 and Figure C5 respectively.
B.3 Reflection
The task description of critic used in reflection mechanism is illustrated in Figure C6.
B.4 Demonstration Generation
As described in Section 3.3, we use input-first and output-fist strategies to generate various pairs as demonstrations. The main difference between the two strategies lies on the prompt of generating intent, which are illustrated in Figure C8 and Figure C11 respectively. The prompt for generating plans is illustrated in Figure C7.
B.5 User Simulator
The prompt to instruct LLM to play as a user is illustrated in Figure 4.
B.6 One-Turn Conversation Generation
One-turn recommendation comprises two tasks: retrieval and ranking. Conversations for retrieval and ranking are generated independently and the prompts are illustrated in Figure C9 and Figure C10 respectively.