Retrieval-Augmented Purifier for Robust LLM-Empowered Recommendation

Liangbo Ning, Wenqi Fan, Qing Li

Introduction

In today’s era of information explosion, recommender systems play a vital role in enhancing user experiences and influencing user decisions by filtering out irrelevant information in various applications such as streaming platforms (e.g., YouTube , TikTok ) and e-commerce (e.g., Amazon , Taobao ). Technically, most existing representative recommendation methods aim to capture collaborative signals by modeling user-item interactions . Recently, large language models (LLMs) have been widely applied in real-life scenarios due to their powerful capabilities in language comprehension and generation, and rich store of open-world knowledge . For example, as one of the most famous AI chatbots in recent years, ChatGPT has showcased human-level intelligence with impressive logical reasoning, open-ended conversation, and personalized content recommendation abilities. To fully leverage the powerful capabilities of large language models, a significant amount of research has utilized LLMs to revolutionize recommender systems for next-generation RecSys . For instance, Geng et al. propose P5, which unifies various recommendation tasks by converting user-item interactions to natural language sequences, achieving outstanding recommendation performance due to the rich textual information that can help capture complex semantics for personalization and recommendations.

Despite the remarkable success, most existing LLM-empowered RecSys still encounter a key limitation, in which they have been demonstrated to be highly vulnerable to minor perturbations in the input prompt , greatly constraining their practical applicability. Suppose that attackers might post products with enticing images and titles to attract user clicks on an e-commerce platform. Users are easily drawn to these clickbait products and interact with them, even though the content of these goods may not truly align with their preferences . Such minor perturbations (e.g., irrelevant items) can easily lead the LLM-empowered RecSys to misunderstand the user preferences by capturing the collaborative knowledge from the user’s historical interactions towards items. For example, as illustrated in Figure 1, when perturbation item "ties" is inserted into the user’s interaction sequence, the perturbed collaborative knowledge makes LLM-empowered RecSys struggle to discern whether the user is seeking men’s clothing (i.e., "suits") or women’s clothing (i.e., "dresses"), leading to inaccurate recommendation outcomes. That is due to the fact that attackers tend to add items that are irrelevant to users’ behaviors for hindering collaborative knowledge learning . In order to defend such minor perturbations for robust recommendations, one of the promising solutions is to purify malicious collaborative knowledge from the user’s historical interactions towards items in LLM-based recommender systems. In most recommender systems, collaborative graphs based on item-item co-occurrence are commonly employed as a collaborative signal to represent the relationships among items, where items frequently interacted together by different users are related (e.g., substitutable or complementary) . Following the insertion of perturbations by attackers, item co-occurrence collaborative graphs as the external knowledge source can provide valuable evidence on whether such perturbations are relevant to other items in the user’s interaction history and effectively filter out malicious collaborative signals (i.e., perturbed items), which can be achieved by retrieving subgraphs and examining the connection between the perturbation and the retrieved subgraphs.

Recently, to mitigate the problems usually caused by insufficient intrinsic knowledge of LLMs, including outdated knowledge, hallucination, and so on , retrieval-augmented generation (RAG) techniques have been proposed to expand the internal knowledge of large language models with an external database. The relevant knowledge is retrieved from the external database and employed to augment LLMs without changing the parameters of LLM backbone, achieving outstanding success for various knowledge-intensive domains such as open question answering , medicine , and finance . For example, Lewis et al. propose to utilize Wikipedia for knowledge retrieval and combine the retrieved documents with the input to augment the generation process, significantly improving the performance of LLMs for various complex tasks and mitigating the hallucination problem. In the context of RecSys, there exists a vast amount of publicly available external collaborative knowledge collected from various public platforms such as Amazon , Yelp , and Steam . Given the success of expanding the internal knowledge of LLMs through the use of external databases to enhance their capabilities in a training-free manner, along with the abundant collaborative knowledge available in the RecSys community, RAG techniques provide unprecedented opportunities to enhance the robustness of LLM-empowered RecSys with external collaborative signals. For example, as shown in Figure 1, LLM-empowered RecSys might generate an incorrect recommendation to a user who interacted with "skirt, ties, heels". To produce reliable recommendation results, a collaborative item graph (i.e., external databases) based on user-item interactions can be constructed to provide useful external collaborative knowledge for better understanding users’ preferences in LLM-based recommender systems. LLM-based RecSys can purify the noisy users’ online behaviors (i.e., perturbation "ties") by retrieving collaborative signals (i.e., subgraph) from the collaborative item graph for recommendation generation, where items "skirt" and "heels" rarely appear together with "ties" in most users’ shopping behaviors.

To effectively take advantage of external collaborative signals from item-item collaborative graph, in this paper, a novel framework RETURN is proposed as a retrieval-augmented purifier for enhancing the robustness of LLM-empowered recommender systems in a plug-and-play manner. Specifically, the users’ historical sequences within the external databases are first encoded into collaborative item graphs to capture the extensive collaborative knowledge. After that, a retrieved-augmented perturbation positioning strategy is proposed to identify potential perturbations by retrieving relevant collaborative signals from the collaborative item graphs. Then, we further cleanse the potential perturbations within the user profile by using either deletion or replacement strategies based on the external collaborative item graphs. Finally, a robust ensemble recommendation strategy is proposed to guide the LLM-empowered RecSys to generate robust recommendation results. Our major contributions are summarised as follows:

We introduce a novel strategy for denoising in LLM-empowered recommendation, in which training-free retrieval-augmented denoising strategy is proposed to leverage the collaborative signals of collaborative item graphs to purify the poisoned user profiles.

We propose a novel framework (RETURN) to enhance the robustness of LLM-empowered RecSys by harnessing collaborative signals from external databases in a plug-and-play manner. Meanwhile, a robust ensemble recommendation is proposed to cleanse user profiles multiple times and generate robust recommendations by using a decision fusion strategy.

We conduct extensive experiments on three real-world datasets to demonstrate the effectiveness of the proposed method. Comprehensive results indicate that RETURN can significantly mitigate the negative impact of the perturbations, highlighting the potential of introducing external collaborative knowledge to enhance the robustness of LLM-empowered recommender systems.

The rest of this paper is organized as follows: Section 2 reviews multiple related studies. Section 3 provides the basic definition of the research problem, and the details of the proposed RETURN are presented in Section 4. Then, we conduct comprehensive experiments to investigate the effectiveness of RETURN in Section 5. Finally, we conclude the whole work in Section 6.

Related Works

Numerous defense strategies have been devised to mitigate LLM vulnerabilities and safeguard against harmful information in LLM responses. These methods are categorized into two main classes based on whether they are employed during training or inference.

1) Defense in LLMs Training. The security of LLMs is significantly dependent on their training data, resulting in several defense strategies aimed at enhancing and purifying the training data . For example, Wenzek et al. introduced CCNet, an automated pipeline designed to efficiently extract vast amounts of high-quality monolingual datasets from the Common Crawl corpus across various languages. Beyond enhancing the quality of training data, adversarial training techniques are widely employed to guide LLMs towards appropriate behaviors by introducing adversarial perturbations into training examples to improve model robustness and performance . For example, Liu et al. introduced a general algorithm known as adversarial training for large neural language models (ALUM), aimed at enhancing the robustness of language models. ALUM enhances model resilience by regularizing the training objective via incorporating perturbations within the embedding space and focusing on maximizing adversarial loss. Wang et al. ) propose a simple yet effective adversarial training method that incorporates adversarial perturbations into the output embedding layer during model training. Li and Qiu employs a token-level accumulated perturbation vocabulary to initialize the adversarial perturbations and use a token-level normalization ball to regulate the generated perturbations for virtual adversarial training .

2) Defense in LLMs Inference. The large scale of parameters of LLMs renders their retraining or fine-tuning processes both time-consuming and computationally expensive. Therefore, training-free defense methods during inference have drawn considerable attention . For example, Kirchenbauer et al. and Jain et al. undertake extensive experiments to evaluate the effectiveness of different defense methods, such as perplexity-based detection, retokenization, and paraphrasing. Li et al. introduce an adversarial purification method that masks input texts and leverages masked language models for text reconstruction. Wei et al. and Mo et al. propose enhancing model robustness through contextual demonstrations. Wang et al. propose RMLM, aimed at countering attacks by confusing attackers and correcting adversarial contexts stemming from malicious perturbations. Helbling et al. incorporate generated content into a predefined prompt and utilize another LLM to analyze the text and assess its potential harm.

2 LLM-Empowered Recommender Systems

Currently, LLMs are widely employed in enhancing the capabilities of recommender systems due to their powerful language understanding, logical reasoning, and generation abilities. These studies can be generally divided into three categories based on the item information utilized.

1) ID-Based LLM-Empowered Recommender Systems. ID-based LLM-empowered recommender systems represent an item with a numerical index and use the item IDs for recommendations . For example, Geng et al. propose P5, which unifies various recommendation tasks by converting user-item interactions to natural language sequences. P5 introduces whole-word embedding to represent the token IDs, bridging the gap between large language models and recommender systems. Zheng et al. propose a learning-based vector quantization method for assigning meaningful item indices for items and introduce specialized tasks to facilitate the integration of collaborative semantics in LLMs, leading to an effective adaptation to recommender systems.

2) Text-Based LLM-Empowered Recommender Systems. To effectively harness the natural language understanding and generation capabilities of LLMs, text-based LLM-empowered recommender systems primarily leverage textual information such as item titles and item descriptions for recommendation . For example, Bao et al. introduce TALLRec, a novel tuning paradigm designed to tailor LLMs for recommendation tasks effectively, guides the model to assess user interest in a target item by analyzing their historical interactions that encompass textual descriptions like item titles. Du et al. propose a novel LLM-based approach for job recommendation that enhances user profiling for resume completion by extracting both explicit and implicit user characteristics based on users’ self-description and behaviors. A GANs-based method is introduced to refine the representations of low-quality resumes, and a multi-objective learning framework is utilized for job recommendations.

3) Hybrid LLM-Empowered Recommender Systems. These approaches effectively integrate both textual information and ID-based knowledge to generate recommendations . For example, Ren et al. leverage text-format knowledge from LLMs and item IDs to enhance recommendation performance, along with a novel alignment training method and an asynchronous technique to refine LLMs’ generation process for improved knowledge augmentation and accelerated training. Liao et al. propose a novel hybrid prompting approach that integrates ID-based item embedding generated by traditional RecSys with textual item features. Besides, LLaRA utilizes a projector to align traditional recommender ID embeddings with LLM input space and incorporates a curriculum learning strategy to gradually train the model to integrate behavioral knowledge from traditional sequential recommenders, thereby enhancing recommendation performance seamlessly.

3 Denoising for Traditional Recommender Systems

With the development of RecSys, a growing body of research has focused on their vulnerability to noisy data, subsequently driving the advancement of various denoising approaches to improve system robustness . For example, GraphRfi proposes an innovative end-to-end framework that integrates Graph Convolutional Networks (GCN) and neural random forests to simultaneously enhance robust recommendation accuracy and fraudster detection. By leveraging user reliability features and prediction errors of RecSys, GraphRfi effectively mitigates the impact of shilling attacks . LoRec proposes to enhance the robustness of sequential recommender systems against poisoning attacks by integrating the open-world knowledge of large language models. Through LLM-Enhanced Calibration, LoRec employs a user-wise reweighting strategy to generalize defense mechanisms beyond specific known attacks, effectively mitigating the impact of fraudsters. LLM4DASR introduces an LLM-assisted denoising framework for sequential recommendations, combining self-supervised fine-tuning with uncertainty estimation to address output quality challenges. This model-agnostic framework effectively identifies and corrects noisy interactions, enhancing recommendation performance across various models.

4 Difference between Existing Denoising Approaches and RETURN

Despite the presence of existing denoising techniques, they are fundamentally different from our approach in terms of task formulation and technical details:

1) Denoising for different phases. Existing denoising methods primarily focus on purifying the training set and ensuring accurate representation learning for RecSys to mitigate the impact of shilling attacks during training, assuming that user historical interactions during inference contain no perturbations. However, during the inference phase, users may still be attracted to clickbait items and interact with them, leading to perturbations that do not align with their true preferences. Moreover, studies have highlighted the vulnerability of LLM-empowered recommender systems during inference, where even a well-trained LLM-based RecSys frequently produces inaccurate recommendations for users affected by poisoned interactions. In other words, even after the training set has been purified, if a user inadvertently interacts with a few clickbait or disliked items during inference, LLM-empowered RecSys may still misinterpret the user’s preferences and generate unsatisfied recommendations. In this paper, we assume that the LLM-empowered RecSys is well-trained, while user interaction sequences may contain noise or adversarial perturbations during inference. In other words, RETURN is designed to address the inference-phase vulnerability of LLM-empowered RecSys and enhance their robustness, which is fundamentally different from the objective of previous denoising methods. Additionally, RETURN can be seamlessly integrated with prior denoising approaches. For instance, existing methods can be employed to cleanse the training set and train a powerful LLM-empowered RecSys, while RETURN ensures the robustness of the RecSys during the inference phase.

2) Novel purification techniques based on external collaborative signals. Existing denoising methods primarily rely on leveraging the characteristics of perturbations or the open-world knowledge of LLMs to identify perturbations within the training set, largely overlooking the potential of external collaborative knowledge. With the advancement of recommender systems, numerous publicly available datasets have been introduced to evaluate algorithm performance. These datasets offer abundant external collaborative signals that can be leveraged to purify perturbations within user historical interactions. Specifically, after attackers introduce perturbations, collaborative graphs based on item-item co-occurrence can be constructed from the external database and leveraged to assess whether these perturbations are consistent with other items in the user’s interaction history, thereby effectively filtering out malicious collaborative signals (i.e., perturbed items). Therefore, RETURN effectively extracts collaborative knowledge from external databases to purify the user historical interactions in a plug-and-play manner, providing a promising solution for enhancing the robustness of LLM-empowered RecSys.

Problem Statement

The objective of recommender systems is to capture the users’ preferences from their historical interactions, such as browsing, clicking, and purchasing. In the era of LLMs, the recommendation task is usually converted to the natural language format, consisting of user ui∈U={u1,u2,...,u∣U∣}u_{i}\in U=\{u_{1},u_{2},...,u_{|U|}\}, and user’s interaction history (also called user’s profile) Iui=[I1,I2,...,I∣Iui∣]\mathcal{I}_{u_{i}}=[I_{1},I_{2},...,I_{|\mathcal{I}_{u_{i}}|}], and a recommendation prompt P=[p1,p2,...,p∣P∣]\mathcal{P}=[p_{1},p_{2},...,p_{|\mathcal{P}|}], where pip_{i} is the textual token used to guide the RecSys Rθ\mathcal{R}_{\theta} to generate recommendations. Ii∈I={I1,I2,...,I∣I∣}I_{i}\in\mathcal{I}=\{I_{1},I_{2},...,I_{|\mathcal{I}|}\} is the interacted item from the item pool I\mathcal{I} of user uiu_{i}. Based on the above definition, a textual recommendation query can be represented as x=[P∘ui∘Iui]\textbf{x}=[\mathcal{P}\circ u_{i}\circ\mathcal{I}_{u_{i}}], where ∘\circ represents inserting the information of user uiu_{i} and the corresponding interaction list Iui\mathcal{I}_{u_{i}} into the designated position of prompt P\mathcal{P}. For example, as shown in Figure 2, after inserting the user information and item interaction sequence into the prompt P\mathcal{P}, the specific input used for recommendation can be denoted by:

where ui=[User_235]u_{i}=[User\_235] and Iui=[item_123,...,item_928]\mathcal{I}_{u_{i}}=[item\_123,...,item\_928] are the specific user and the historical interactions of user uiu_{i}, respectively. In different LLM-empowered RecSys, Iui\mathcal{I}_{u_{i}} can take various forms, such as numeric IDs or item titles for recommendations. Assume the target item is y, the performance of the LLM-empowered RecSys can be defined by:

where D\mathcal{D} evaluates the discrepancy between the generated recommendations Rθ(x)\mathcal{R}_{\theta}(\textbf{x}) and the ground truth. During training, the negative log-likelihood function could be used as the D\mathcal{D}, while during inference, the Hit Ratio or Normalized Discounted Cumulative Gain (NDCG) could be employed to evaluate the recommendation performance.

2 Vulnerabilities of LLM-Based RecSys

where x^=[P∘ui∘I^ui]\hat{\textbf{x}}=[\mathcal{P}\circ u_{i}\circ\mathcal{\hat{I}}_{u_{i}}] is the perturbed input and △\triangle constrains the magnitude the perturbations. D\mathcal{D} evaluates the discrepancy between the generated recommendations Rθ(x)\mathcal{R}_{\theta}(\textbf{x}) and the ground truth.

3 Robust LLM-based Recommendation

The primary objective of robust LLM-based recommendations is to prevent the negative impact of the perturbations contained in the users’ profiles, thereby enhancing the system’s reliability and robustness. There are mainly two approaches to achieve this goal: adversarial training-based methods and training-free methods . Adversarial training-based methods intentionally create multiple perturbed training samples to guide the RecSys in learning patterns of perturbations, thereby improving system robustness. However, these methods usually retrain or fine-tune the whole RecSys, which is extremely time-consuming due to the large number of trainable parameters in LLMs. Consequently, this paper primarily concentrates on the training-free methods, which improve the model’s robustness without introducing additional changes in model parameters. Specifically, when the users’ profiles contain adversarial perturbations, we aim to accurately identify and filter out these perturbations to ensure the appropriate recommendations for users during the inference process. Mathematically, if the input with adversarial perturbations is denoted by x^\hat{\textbf{x}}, we aim to cleanse the input for robust recommendations, formulated as follows:

Methodology

RETURN is proposed to leverage the collaborative knowledge of users within external databases to filter out adversarial perturbations, thereby enhancing the robustness of the existing LLM-based RecSys. As shown in Figure 2, RETURN mainly contains three components: Retrieval-Augmented Perturbation Positioning, Retrieval-Augmented Denoising, and Robust Ensemble Recommendation. First, we convert the user interaction sequences within the external database to collaborative item graphs to encode the collaborative knowledge without introducing additional training processes. After that, the probability of each item in the user profile being a perturbation is computed by retrieving collaborative signals from collaborative item graphs. Second, retrieval-augmented denoising filters out the potential perturbations in the user profiles using either deletion or replacement strategies based on the collaborative signals of the generated item graphs. Finally, robust ensemble recommendation purifies input query multiple times and adopts an ensemble strategy to generate the final recommendations.

2 Retrieval-Augmented Perturbation Positioning

To mitigate the negative impact of perturbations, the first crucial step is to accurately locate the perturbations from extensive interactions within user profiles. To achieve this goal, we propose to use collaborative item graphs to encode the collaborative signals from users in the external database and retrieve relevant collaborative knowledge for perturbation positioning. By encoding the user’s interaction history into collaborative item graphs, we can clearly understand the relationships between items , and such strong collaborative signals can provide evidence for subsequent denoising processes. Furthermore, this approach enables the direct use of a one-hot vector for information retrieval, eliminating the need to explicitly train a universal retriever as required by other RAG techniques, thereby improving efficiency.

Let E={UE,IE}\mathcal{E}=\{U_{\mathcal{E}},\mathcal{I}_{\mathcal{E}}\} be the external database, where UE={u1,u2,...,uE}U_{\mathcal{E}}=\{u_{1},u_{2},...,u_{\mathcal{E}}\} and IE={Iu1,Iu2,...,IuE}\mathcal{I}_{\mathcal{E}}=\{\mathcal{I}_{u_{1}},\mathcal{I}_{u_{2}},...,\mathcal{I}_{u_{\mathcal{E}}}\} denote the users and their interaction sequences, respectively. The most straightforward method of generating collaborative item graphs is to count the occurrence frequency of ii-th and jj-th items appearing together in the historical interaction of the same user. However, such a vanilla strategy overlooks the temporal relationships among items, which are significantly crucial for subsequent denoising processes. For instance, mobile phones and phone cases are usually interacted with consecutively by users, whereas mobile phones and furniture are typically not sequentially interacted with by users. During the denoising process, if a user consecutively interacts with both mobile phones and furniture, there is a high likelihood of perturbations in the user’s historical interactions. Therefore, to provide precise collaborative signals, we also consider the gap between two items and generate a set of multi-hop collaborative item graphs, which encode not only the relevance between items but also the temporal relationships of items. Given the external database E={UE,IE}\mathcal{E}=\{U_{\mathcal{E}},\mathcal{I}_{\mathcal{E}}\}, the multi-hop collaborative item graph can be represented as:

2.2 Perturbation Positioning

After encoding the external users’ collaborative knowledge into collaborative item graphs, the next step is to locate the perturbations within the input query based on the generated graphs. Specifically, if one item has never appeared together with the remaining items in the user’s historical interactions based on the collaborative item graphs derived from the majority of users’ behavior, this indicates that such item is unrelated to the other items the user has interacted with. Thus, the likelihood of this item appearing within the user’s interaction history is minimal, and the occurrence of such a low-probability event strongly implies that this item is likely introduced as a perturbation by an attacker. Therefore, to accurately locate the potential perturbations, we propose to retrieve co-occurrence frequency from the collaborative item graphs and assess the probability of each item appearing within the user’s interaction history.

Given the multi-hop collaborative item graphs Gϵ\mathcal{G}_{\epsilon} and the user’s historical interactions Iui=[I1,I2,...,I∣Iui∣]\mathcal{I}_{u_{i}}=[I_{1},I_{2},...,I_{|\mathcal{I}_{u_{i}}|}], the co-occurrence frequency for a pair of items can be defined as:

A smaller value of AiA_{i} indicates a lower probability of the current item co-occurring with other items, making it more likely to be a perturbation.

3 Retrieval-Augmented Denoising

Once the occurrence probability of each item has been computed, it is necessary to purify the input query based on such collaborative knowledge. However, directly removing numerous items that may be perturbations usually leads to the RecSys failing to capture the user’s preferences accurately since there are limited remaining interactions. To mitigate the negative impact of perturbations while maintaining the integrity of user interaction sequences, a hybrid strategy is proposed to eliminate a small subset of items that are most likely perturbations and replace the remaining potential perturbations with items that align with the user’s preferences.

If an item’s occurrence probability Ai=0A_{i}=0, it indicates that this item has never co-occurred with the other items in the user’s historical interactions based on the collaborative signals of most users from external databases. Thus, this item is highly likely a perturbation inserted by attackers due to its lack of relevance to the user’s other interaction items, and its deletion typically helps RecSys accurately capture the user’s genuine preferences. Mathematically, given the interaction history Iui\mathcal{I}_{u_{i}} of the user uiu_{i} and the occurrence probability A\mathcal{A}, we first delete the most likely perturbation items whose occurrence probability is zero, defined by:

where \mathds1(Ii,Ai)\mathds{1}(I_{i},A_{i}) represents to execute deletion operation when Ai=0A_{i}=0 and preservation otherwise.

After removing items with Ai=0A_{i}=0 that are most likely perturbations, some items usually remain in the user’s historical interactions with very low but non-zero occurrence probabilities. Simply deleting these items usually results in sparse user-item interactions, and such limited collaborative knowledge leads to cold start issues , hindering RecSys from capturing user preferences effectively. To maintain the integrity of user interaction history, a retrieval-augmented replacement strategy is proposed to replace the remaining potential perturbations with items that align with user preferences. Specifically, all items that have co-occurred with the other remaining items in the user’s interaction history are retrieved from the collaborative item graphs, and the item that shows the highest co-occurrence frequency is considered the prime candidates that best align with the current user preferences among all the retrieved items. Given the user’s historical interactions Iui=[I1,...,II∣ui∣]\mathcal{I}_{u_{i}}=[I_{1},...,I_{\mathcal{I}_{|u_{i}|}}] and the potential perturbation IiI_{i} with the low occurrence probability, the replacement operation is defined by:

4 Robust Ensemble Recommendation

By repeating the purification process multiple times on Iui\mathcal{I}_{u_{i}}, we can obtain mm cleansed user profiles, where mm is a hyperparameter. These purified prompts are individually fed into the LLM-empowered RecSys, and the results are subsequently integrated to produce the final recommendation output by using voting mechanisms, defined by:

where xˉi=[P∘ui∘Iˉui]\bar{\textbf{x}}_{i}=[\mathcal{P}\circ u_{i}\circ\mathcal{\bar{I}}_{u_{i}}] is the purified input and yˉ\bar{\textbf{y}} is the final recommendation. The pseudo-code of RETURN is shown in Algorithm 1.

Experiments

All experiments are conducted on three real-world datasets in RecSys: Movielens-1M (ML1M) , Taobao , and LastFM datasets. The ML1M dataset contains one million movie ratings collected from around 6,040 users and their interactions with around 4,000 movies, which is widely used for various recommendation tasks and evaluation of recommendation techniques. The LastFM dataset is a widely used music recommendation dataset that contains user listening histories and preferences, which is frequently used to study user preferences, understand music consumption patterns, and evaluate recommendation algorithms. The Taobao dataset comprises a massive collection of user interactions on the Taobao e-commerce platform, including browsing, searching, and purchasing activities. It consists of a million records from around 987,994 users and their interactions with around 4,162,024 items and offers valuable insights into user behavior and preferences in the online retail environment. For P5 model, all the aforementioned datasets are preprocessed following the strategies proposed by Xu et al. . For TALLRec model, it needs to divide the users’ historical sequences into users’ liked items and disliked items based on their ratings. Since LastFM and Taobao datasets lack rating information from users, we only process the ML1M dataset according to the study of Ning et al. .

1.2 Victim LLM-based Recommender Systems.

Two representative LLM-based RecSys, i.e., P5 and TALLRec, are employed as the victim models to investigate the performance of different defense techniques.

P5 is a typical ID-based LLM-empowered RecSys, which assigns each item a numerical number and converts the user-item interactions to natural language sequences for recommendations. P5 introduces several item indexing strategies, which can be employed to test the robustness of the defense methods for ID-based RecSys with different indexing strategies.

TALLRec is a representative text-based LLM-empowered recommender system, which integrates textual information (i.e., item title) into a pre-defined prompt template for recommendation. By constructing experiments based on TALLRec, we can investigate the performance of different defense methods for LLM-empowered RecSys employing textual knowledge.

1.3 Attackers.

We employ CheatAgent as the attacker to generate adversarial perturbations and insert them into the user’s historical interactions. It should be noted that CheatAgent is an evasion attack method that uses LLMs as the agent to generate high-quality perturbations for misleading the target LLM-empowered RecSys during the inference phase. Currently, there is limited research on poisoning attacks for LLM-empowered RecSys. Poisoning attacks require retraining the model, but the large parameter size of LLMs makes frequent retraining infeasible. In other words, poisoning attacks are highly time-consuming for large language models, and they are ineffective if retraining cannot be performed. Therefore, in this paper, we solely consider the evasion attack (i.e., CheatAgent) since it is a more efficient attacking method in the era of LLMs. We use CheatAgent to generate item perturbations and insert them into the user’s history interactions to test the defense performance of different methods. The primary objective of CheatAgent is to investigate the vulnerabilities of exiting LLM-empowered RecSys, and it allows the insertion of perturbations in both prompt and users’ profiles. However, during real-world applications, the attacker and users usually have no access to the prompt P\mathcal{P}, which makes the prompt attack infeasible. Therefore, in this paper, we only use CheatAgent to generate item perturbations and insert them into the user’s history interactions.

1.4 Baselines.

Several baselines are utilized to investigate the defense performance of different methods:

RD randomly deletes some items within the users’ historical sequences to filter out the adversarial perturbations.

PD computes the perplexity for each item and filters out the item with high perplexity for defense.

RPD uses an LLM to paraphrase the input prompt, which is widely used as the safeguard for LLMs.

RTD retokenizes the input prompt, which aims to break tokens apart and disrupt adversarial behaviors.

LLMSI provides a safety instruction, i.e., "Please take into account the noise present in the user’s historical interactions and filter them", along with the input prompt to guide the LLM-empowered RecSys to defense adversarial attacks by themselves.

RDE randomly deletes some items within the users’ interaction sequences and generates multiple cleansed prompts. The final recommendations are obtained by majority voting .

ICL randomly retrieves several users with different historical sequences as the demonstrations and integrates the retrieved users’ profiles with the original prompt for recommendation.

1.5 Implementation.

The proposed RETURN and all baselines are implemented by Pytorch. All victim models (i.e., P5 and TALLRec) and the attacker algorithm (i.e., CheatAgent) are implemented based on their official codes. The training and test set are constructed according to the studies of Xu et al. and Bao et al. for P5 and TALLRec, respectively. We adopt CheatAgent to generate adversarial perturbations and insert them into the benign users’ interaction history of the test set to investigate the defense performance of different methods. The magnitude of perturbations △\triangle is set to 3, consistent with the study of Ning et al. . For the proposed RETURN, we directly use the training set as the external database. During the recommendation generation process, m=10m=10 is set as default, meaning that the final ensemble recommendation is obtained based on these 10 purified prompts. For RD, we randomly delete 3 items and generate recommendations. For PD, we select the top 3 items with the highest perplexity as the perturbations and delete these items for recommendations. For RTD, we adopt the BPE-dropout to tokenize the input query to mitigate the impact of adversarial perturbations. RDE generates 10 purified prompts and integrates their recommendation outcomes as the final prediction. ICL randomly retrieves 5 users’ interaction sequences from the external database and integrates them with the original input for recommendations. All random seeds were fixed throughout the experiments, consistent with the used victim RecSys P5 and TALLRec . This ensures that the experimental results are reproducible, and therefore, we do not include variance in the reported results.

1.6 Evaluation Metrics.

For P5 model, Top-kk Hit Ratio (H@kk) and Normalized Discounted Cumulative Gain (NDCG) (N@kk) are employed to evaluate the recommendation performance. In this paper, we set k=5k=5 and k=10k=10, respectively. A-H@kk and A-N@kk represent the extent of the decrease in H@kk and N@kk after inserting adversarial perturbations into the benign prompt, which are used to measure the attack performance , formulated as:

where H^@k\widehat{\text{H}}@k and N^@k\widehat{\text{N}}@k evaluate the recommendation performance of the victim model when it is under attack. D-H@kk and D-N@kk are utilized to evaluate the performance of defense algorithms, which represent the decrease ratio in A-H@kk and A-N@kk, defined as:

where A-H~@k\widetilde{\text{A-H}}@k and A-N~@k\widetilde{\text{A-N}}@k represent the attack performance when adversarial examples are processed by defense algorithms. A greater decrease in A-H@k\text{A-H}@k and A-N@k\text{A-N}@k indicates reduced attack performance and improved performance of the defense methods. For TALLRec model, we utilize the Area Under the Receiver Operating Characteristic (AUC) to assess the recommendation performance, which is consistent with the study of Bao et al. . ASR-A and D-A are employed to evaluate the performance of the attack and defense methods, defined as:

where AUC^\widehat{\text{AUC}} and ASR-A~\widetilde{\text{ASR-A}} represent the AUC when the input contains perturbations and when the input is purified by the defense methods, respectively.

2 Defense Effectiveness

In this subsection, we investigate the defense performance of different methods. The results based on P5 with different indexing methods are summarised in Table 1 and Table 2, and the results based on TALLRec are shown in Figure 3. Benign denotes the use of the original prompt without perturbations for recommendations, and CheatAgent represents the recommendation performance under attacks. Based on these experiments, some insights are obtained as follows:

As shown in Table 1 and Table 2, the recommendation performance increases after deleting high perplexity items using PD. However, the effectiveness of this method is not robust. For instance, on the ML1M dataset, PD can significantly enhance the recommendation performance of RecSys under attacks. While on the Taobao dataset, the defense performance of PD is limited.

RPD and RTD, two common defense methods for LLMs, cannot achieve the desired performance for LLM-empowered RecSys in most cases. The reason is that LLM-empowered RecSys have captured the domain-specific knowledge of recommendations (e.g., the meaning of item IDs and item relationships) during the training process. However, the LLMs employed by RPD struggle to understand item IDs, making it challenging to effectively rewrite the input prompt. Additionally, RTD disrupts the item ID structure, which further degrades recommendation performance.

Adversarial perturbations are typically carefully crafted, so disrupting any component may reduce the attack’s effectiveness. Therefore, randomly removing a few items from the user’s interaction history (i.e., RD) can improve the robustness of the LLM-powered RecSys. Furthermore, RDE generally outperforms RD, suggesting that an ensemble strategy can further enhance system robustness.

The proposed RETURN outperforms all other baselines on three datasets and significantly improves the recommendation performance even under attacks, demonstrating the potential of introducing collaborative knowledge from external databases. For example, on the Taobao dataset, CheatAgent reduces the H@5 from 0.1420 to 0.0863. By introducing collaborative knowledge for input purification, RETURN raises the H@5 to 0.1124, nearly approaching the recommendation performance of using benign prompts, which fully demonstrates the effectiveness of RETURN.

TALLRec uses item titles to construct the input prompt, which has distinct inherent mechanisms with P5. As shown in Figure 3, the proposed RETURN also dramatically increases the AUC of TALLRec and decreases the attack performance, demonstrating the robustness of RETURN to the architecture of the LLM-empowered RecSys.

3 Model Analysis

The attack discussed in this paper mirrors a real-world phenomenon, commonly known as clickbait . Clickbait refers to the scenario in which attackers might post products with enticing images and titles to attract user clicks on an e-commerce platform. Users are easily drawn to these clickbait products and interact with them, even though the content of these goods may not truly align with their preferences . However, existing studies have demonstrated that LLM-empowered RecSys is vulnerable to minor perturbations in user historical interactions. If users are attracted by clickbait products and engage with them, minor perturbations will be introduced to their historical interactions. Such minor perturbations (e.g., irrelevant items) can easily lead the LLM-empowered RecSys to misunderstand the user preferences by capturing the collaborative knowledge from the user’s historical interactions. This leads to inaccurate recommendations, affecting user experience and engagement and consequently diminishing company profits. Therefore, enhancing the robustness of the LLM-empowered RecSys is crucial to mitigate the clickbait issue, which is a practical necessity.

During experiments, to simulate the worst-case scenario, we adopt CheatAgent , which is a powerful attacker, to insert perturbations to the user’s historical sequences. Besides, we also employ various attack methods and perturbation intensities to simulate the scenario in which the user’s historical interactions contain minor perturbations. We adopt two other methods to generate adversarial perturbations: PA adopts an LLM to generate perturbations, and RA randomly selects the items from the item pool as the perturbations.

As shown in Table 3, we can observe that the proposed defense method significantly reduces the effectiveness of various attack methods (i.e., CheatAgent, PA). This implies that even if users interact with clickbait items that trigger vulnerabilities in the recommendation system, the proposed RETURN method can effectively cleanse these malicious disturbances, ensuring the correctness of recommendations. Regarding RA, its attack capability is constrained, and it is aimed at simulating scenarios where perturbation items do not cause the RecSys to misinterpret user preferences. In this case, RETURN still improves or maintains the recommendation performance of the RecSys. This demonstrates the robustness of the proposed RETURN against different attack intensities and scenarios.

3.2 Ablation Study

Three variants RETURN-ROP, RETURN-RR, and RETURN-w/o Ens are employed for comparison: 1) RETURN-ROP randomly creates the collaborative item graphs to demonstrate the effectiveness and importance of introducing the external database. 2) RETURN-RR directly deletes all items with low occurrence probabilities. 3) RETURN-w/o Ens generates recommendations without using the ensemble strategy and only creates one purified prompt by processing a fixed number of items. The results are summarised in Table 4. RETURN-ROP generates recommendations without constructing collaborative item graphs from the external database, resulting in a significant decrease in its defense performance. This highlights the importance of introducing accurate collaborative knowledge from the external database. Since directly deleting all items with low occurrence probabilities may result in the RecSys failing to capture users’ preferences effectively, especially for users with limited interactions, there is a significant decrease in the defense performance of RETURN-RR, illustrating the importance of employing the retrieval-augmented denoising strategy. Since the number of the perturbations is unknown, RETURN-w/o Ens fixes the number of purification items. This approach usually leads to information loss if an excessive number of items are deleted, or incomplete purification if not all perturbations are eliminated, demonstrating the importance of the robust ensemble recommendation strategy.

3.3 Parameter Analysis

We investigate the sensitivity of RETURN to the hyperparameter mm. We sample varying values for mm and test the defense performance of the proposed method. and the results are illustrated in Figure 4. We observe that as mm increases, the recommendation performance and the defense capability of RETURN fluctuate within a small range, demonstrating the robustness of the proposed method to hyperparameters.

3.4 The Robustness to the Perturbation Intensity

In this subsection, we investigate the robustness of RETURN to the perturbation intensity △\triangle. We insert varying numbers of perturbations into benign users and evaluate the defense performance of the proposed method. As shown in Table 5, the proposed method significantly enhances the recommendation performance of LLM-empowered RecSys regardless of the number of perturbations inserted into the input. This is attributed to robust recommendation generation strategies that avoid introducing fixed thresholds, thereby improving the robustness of the proposed RETURN to the number of perturbations.

3.5 Impact on Benign Users

It is crucial that defense algorithms should not affect the recommendation performance of RecSys for users whose interaction histories contain no perturbations. Therefore, in this subsection, the impact of RETURN on benign users is investigated, and the results are shown in Table 6. We can observe that if the users’ profiles consist of no perturbations, RETURN can almost maintain the recommendation performance even though RETURN deletes or replaces some items. Note that the deletion or replacement operations are implemented based on the collaborative co-occurrence frequency, indicating that the selected items usually fail to align with the users’ preferences. Therefore, RETURN has little impact on the recommendation effectiveness for benign users, which demonstrates its practical applicability in enhancing the robustness of LLM-empowered RecSys.

3.6 Time Complexity

To address the concern regarding the computational overhead introduced by the RETURN framework, we conduct additional experiments to analyse the time complexity of RETURN. We measure the average time taken by the LLM-empowered RecSys to generate recommendations after incorporating different defense methods on the LastFM dataset. As shown in Tabel 7, we can observe that methods requiring minimal computational resources (e.g., RD, LLMSI, etc.) exhibit significantly shorter recommendation generation times, typically less than 0.5 seconds. However, their defense performance is notably limited. In contrast, more powerful methods, including RETURN, exhibit slightly longer recommendation generation times, with RETURN taking approximately 0.8599 seconds. This is comparable to other advanced defense methods like PD (0.7314 seconds) and RDE (0.7222 seconds), which also take around 1 second.

The results indicate that while RETURN introduces additional computational steps, such as voting operations, it does not significantly increase the overall computational burden of the RecSys. Importantly, RETURN achieves this while substantially enhancing the robustness of RecSys against perturbations. Thus, the framework strikes a balance between computational efficiency and defense effectiveness, making it a practical choice for real-world applications.

3.7 Impact of Poor Quality Data

We conduct additional experiments to investigate the impact of data quality. We introduce two variants: RETURN-A-k and RETURN-D-k, where perturbations are injected into or items are deleted from the historical interactions of users in the external database to generate collaborative item graphs. Here, kk=0.15 and kk=0.3 represent the proportion of perturbations or deletions, respectively. The results are shown in Table 8. The results demonstrate that RETURN-A-k still achieves remarkable defense performance even when perturbations are introduced into the external database. This is because the collaborative item graphs store co-occurrence frequencies, and minor perturbations do not significantly alter the overall co-occurrence distribution among items. After normalization, these perturbations have minimal impact on RETURN’s ability to cleanse user interaction data and generate accurate recommendations. Additionally, RETURN-D-k fails to achieve the desired defense performance because the lack of sufficient collaborative signals prevents it from accurately capturing relationships between items, thereby hindering its ability to identify perturbations.

These experimental results indicate that the presence of noisy data in the external database (i.e., low-quality data) does not significantly deteriorate the performance of RETURN, as the co-occurrence distribution remains relatively stable. However, insufficient data (e.g., due to deletions) can degrade RETURN’s defense effectiveness, as it relies on sufficient collaborative signals to accurately model item relationships. Therefore, while RETURN is robust to minor data quality issues, ensuring an adequate volume of data is crucial for maintaining its performance.

3.8 The Adoption of Normal Distribution

During the robust ensemble recommendation process, RETURN randomly samples an integer nn from a normal distribution, and Top-nn items with the lowest occurrence probabilities are identified from the user’s historical interactions for purification. The normal distribution is chosen because it allows for better control over the strength of perturbation filtering in RETURN. If a majority of users’ interaction histories contain significant perturbations, making it difficult for RecSys to accurately capture their preferences, the mean can be adjusted to enhance the purification strength of RETURN.

During experiments, the mean and the variance are 3.5 and 0.5, respectively. Moreover, we conducted additional experiments to demonstrate that RETURN is robust to the different values of mean and variance. The results are shown in Table 9. The performance of RETURN fluctuates within a reasonable range as the mean and variance change, demonstrating its robustness to different parameter settings. This indicates that RETURN can adapt to varying distributions while maintaining its effectiveness in generating accurate recommendations.

3.9 Impact on the Personalization of Recommendations

To evaluate the impact of RETURN on personalized recommendations, we separately analyze the recommendation results of the LLM-empowered RecSys for benign users and the results after introducing RETURN for denoising. We calculate the frequency of different items in both sets of results, computed the Jaccard similarity coefficient between the two distributions, determined the proportion of items that co-occurred, and measured the Shannon entropy of each distribution. The results are presented in Table 10. Some observations can be obtained as follows:

Jaccard Similarity (0.7605): The high Jaccard similarity coefficient indicates that the recommendation results before and after applying RETURN are highly consistent for benign users. This suggests that RETURN preserves the majority of the original recommendations, although some items are removed or replaced.

Common Items Ratio (0.8706): The proportion of items that co-occur in both the benign and RETURN-processed recommendations is 87.06%. This further demonstrates that RETURN maintains the core set of recommended items, ensuring minimal disruption to the personalized recommendations.

Shannon Entropy: The Shannon entropy values for both the benign (9.8616) and RETURN-processed (9.8292) recommendations are nearly identical. This indicates that RETURN does not significantly reduce the diversity of the recommendations, preserving the richness and variety of the suggested items.

Conclusion

In this paper, we propose a novel framework RETURN by retrieving collaborative knowledge from external databases to enhance the robustness of existing LLM-empowered RecSys in a plug-and-play manner. Specifically, the proposed RETURN first converts the user interactions within external databases into collaborative item graphs to implicitly encode the collaborative signals. Then, the potential perturbations are located by retrieving relevant knowledge from the generated graphs. To mitigate the negative impact of perturbations and maintain the integrity of user preference, a retrieval-augmented denoising strategy is introduced to purify the input user profile. Finally, a robust ensemble recommendation method is proposed to generate the final recommendations by adopting a decision fusion strategy. Comprehensive experiments on real-world datasets demonstrate the effectiveness of the proposed RETURN and highlight the potential of introducing external collaborative knowledge to enhance the robustness of LLM-empowered RecSys.

References