LLMRec: Large Language Models with Graph Augmentation for Recommendation

Wei Wei, Xubin Ren, Jiabin Tang, Qinyong Wang, Lixin Su, Suqi Cheng, Junfeng Wang, Dawei Yin, Chao Huang

Introduction

Recommender systems play a crucial role in mitigating information overload by providing online users with relevant content (Meng et al., 2023b; Wei et al., 2022). To achieve this, an effective recommender needs to have a precise understanding of user preferences, which is not limited to analyzing historical interaction patterns but also extends to incorporating rich side information associated with users and items (Zou et al., 2022).

In modern recommender systems, such as Netflix, the side information available exhibits heterogeneity, including item attributes (Ying et al., 2023), user-generated content (Fan et al., 2019; Meng et al., 2023c), and multi-modal features (Yi et al., 2022) encompassing both textual and visual aspects. This diverse content offer distinct ways to characterize user preferences. By leveraging such side information, models can obtain informative representations to personalize recommendations. However, despite significant progress, these methods often face challenges related to data scarcity and issues associated with handling side information.

Sparse Implicit Feedback Signals. Data sparsity and the cold-start problem hinder collaborative preference capturing (Wei et al., 2021b). While many efforts (e.g., NGCF (Wang et al., 2019), LightCGN (He et al., 2020)) tried powerful graph neural networks(GNNs) in collaborative filtering(CF), they face limits due to insufficient supervised signals. Some studies (Ren et al., 2023b) used contrastive learning to add self-supervised signals (e.g., SGL (Wu et al., 2021), SimGCL (Yu et al., 2022)). However, considering that real-world online platforms (e.g., Netflix, MovieLens) derive benefits from modal content, recent approaches, unlike general CF, are dedicated to incorporating side information as auxiliary for recommenders. For example, MMGCN (Wei et al., 2019) and GRCN (Wei et al., 2020) incorporate item-end content into GNNs to discover high-order content-aware relationships. LATTICE (Zhang et al., 2021b) leverages auxiliary content to conduct data augmentation by establishing i-i relationships. Recent efforts (e.g., MMSSL (Wei et al., 2023a), MICRO (Zhang et al., 2022)) address sparsity by introducing self-supervised tasks that maximize the mutual information between multiple content-augmented views. However, strategies for addressing data sparsity in recommender systems, especially in multi-modal content, can sometimes be limited. This is because the complexity and lack of side information relevance to CF can introduce distortions in the underlying patterns (Wei et al., 2020). Therefore, it becomes crucial to ensure the accurate capture of realistic user preferences when incorporating side information in CF, in order to avoid suboptimal results.

Data Quality Issues of Side Information. Recommender systems that incorporate side information often encounter significant issues that can negatively impact their performance. i) Data Noise is an important limitation faced by recommender systems utilizing side information is the issue of data noise(Tian et al., 2022), where attributes or features may lack direct relevance to user preferences. For instance, in a micro video recommender, the inclusion of irrelevant textual titles that fail to capture the key aspects of the video’s content introduces noise, adversely affecting representation learning. The inclusion of such invalid information confuse the model and lead to biased or inaccurate recommendations. ii) Data heterogeneity(Chen et al., 2023a) arises from the integration of different types of side information, each with its own unique characteristics, structures, and representations. Ignoring this heterogeneity leads to skewed distributions (Ying et al., 2023; Meng et al., 2023a). Bridging heterogeneous gap is crucial for successfully incorporating side information uniformly. iii) Data incompleteness (Ko et al., 2022; Liang et al., 2023b) occurs when side information lacks certain attributes or features. For instance, privacy concerns(Zhang et al., 2023a) may make it difficult to collect sufficient user profiles to learn their interests. Additionally, items may have incomplete textual descriptions or missing key attributes. This incompleteness impairs the model’s ability to fully capture the unique characteristics of users and items, thereby affecting the accuracy of recommendations.

Having gained insight into data sparsity and low-quality encountered by modern recommenders with auxiliary content, this work endeavors to overcome these challenges through explicit augment potential user-item interactive edges as well as enhances user/item node side information (e.g., language, genre). Inspired by the impressive natural language understanding ability of large language models (LLMs), we utilize LLMs to augment the interaction graph. Firstly, LLMRec embraces the shift from an ID-based recommendation framework to a modality-based paradigm (Li et al., 2023a; Yuan et al., 2023). It leverages large language models (LLMs) to predict user-item interactions from a natural language perspective. Unlike previous approaches that rely solely on IDs, LLMRec recognizes that valuable item-related details are often overlooked in datasets (Li et al., 2023b). Natural language representations provide a more intuitive reflection of user preferences compared to indirect ID embeddings. By incorporating LLMs, LLMRec captures the richness and context of natural language, enhancing the accuracy and effectiveness of recommendations. Secondly, to elaborate further, the low-quality and incomplete side information is enhanced by leveraging the extensive knowledge of LLMs, which brings two advantages: i) LLMs are trained on vast real-world knowledge, allowing them to understand user preferences and provide valuable completion information, even for privacy-constrained user profiles. ii) The comprehensive word library of LLMs unifies embeddings in a single vector space, bridging the gap between heterogeneous features and facilitating encoder computations. This integration prevents the dispersion of features across separate vector spaces and provide more accurate results.

Enabling LLMs as effective data augmentors for recommenders poses several technical challenges that need to be addressed:

C1: How to enable LLMs to reason over user-item interaction patterns by explicitly augmenting implicit feedback signals?

C2: How to ensure the reliability of the LLM-augmented content to avoid introducing noise that could compromise the results?

The potential of LLM-based augmentation to enhance recommenders by addressing sparsity and improving incomplete side information is undeniable. However, effectively implementing this approach requires addressing the aforementioned challenges. Hence, we have designed a novel framework LLMRec to tackle these challenges.

Solution. Our objective is to address the issue of sparse implicit feedback signals derived from user-item interactions while simultaneously improving the quality of side information. Our proposed LLMRec incorporates three LLM-based strategies for augmenting the interaction graph: i) Reinforcing user-item interaction edges, ii) Enhancing item attribute modeling, and iii) Conducting user profiling. To tackle C1 for ’i)’, we devise an LLM-based Bayesian Personalized Ranking (BPR)(Rendle et al., 2012) sampling algorithm. This algorithm uncover items that users may like or dislike based on textual content from from natural language perspective. These items are then used as positive and negative samples in the BPR training process. It is important to note that LLMs are unable to perform all-item ranking, so the selected items are chosen from a candidate item pool provided by the base recommender for each user. During the node attribute generation process (corresponding to ’ii)’ and ’iii)’), we create additional attributes for each user/item using existing text and interaction history. However, it is important to acknowledge that both the augmented edges and node features can contain noise. To address C2, our denoised data robustification mechanism comes into play by integrating noisy edge pruning and feature MAE (Tian et al., 2023a) to ensure the quality of the augmented data.

In summary, our contributions can be outlined as follows:

The LLMRec is the pioneering work that using LLMs for graph augmentation in recommender by augmenting: user-item interaction edges, ii) item node attributes, iii) user node profiles.

The proposed LLMRec addresses the scarcity of implicit feedback signals by enabling LLMs to reason explicitly about user-item interaction patterns. Additionally, it resolves the low-quality side information issue through user/item attribute generation and a denoised augmentation robustification mechanism with the noisy feedback pruning and MAE-based feature enhancement.

Our method has been extensively evaluated on real-world datasets, demonstrating its superiority over state-of-the-art baseline methods. The results highlight the effectiveness of our approach in improving recommendation accuracy and addressing sparsity issues. Furthermore, in-depth analysis and ablation studies provide valuable insights into the impact of our LLM-enhanced data augmentation strategies, further solidifying the model efficacy.

Preliminary

Recommendation with Graph Embedding. Collaborative filtering (CF) learns from sparse implicit feedback E+\mathcal{E}^{+}, with the aim of learning collaborative ID-corresponding embeddings Eu,Ei\textbf{E}_{u},\textbf{E}_{i} for recommender prediction, given user u∈Uu\in\mathcal{U} and item i∈Ii\in\mathcal{I}. Recent advanced recommenders employ GNNs to model complex high-order(Tian et al., 2023b) u-i relation by taking E+\mathcal{E}^{+} as edges of sparse interactive graph. Therefore, the CF process can be separated into two stages, bipartite graph embedding, and u-i prediction. Optimizing collaborative graph embeddings E={Eu,Ei}\textbf{E}=\{\textbf{E}_{u},\textbf{E}_{i}\} aims to maximize the posterior estimator with E+\mathcal{E}^{+}, which is formally presented below:

Here, p(E∣E+)p(\textbf{E}|\mathcal{E}^{+}) is to encode as much u-i relation from E+\mathcal{E}^{+} into Eu,Ei\textbf{E}_{u},\textbf{E}_{i} as possible for accurate u-i prediction y^u,i=eu⋅ei\hat{y}_{u,i}=\textbf{e}_{u}\cdot\textbf{e}_{i}.

Recommendation with Side Information. However, sparse interactions in E+\mathcal{E}^{+} pose a challenge for optimizing the embeddings. To handle data sparsity, many efforts introduced side information in form of node features F, by taking recommender encoder fΘf_{\Theta} as feature graph. The learning process of the fΘf_{\Theta} (including Eu,Ei\textbf{E}_{u},\textbf{E}_{i} and feature encoder) with side information F is formulated as maximizing the posterior estimator p(Θ∣F,E+)p(\Theta|\textbf{F},\mathcal{E}^{+}):

fΘf_{\Theta} will output the final representation h contain both collaborative signals from E and side information from F, i.e., h=fΘ(f,E+)\textbf{h}=f_{\Theta}(\textbf{f},\mathcal{E}^{+}).

Recommendation with Data Augmentation. Despite significant progress in incorporating side information into recommender, introducing low-quality side information may even undermine the effectiveness of sparse interactions E+\mathcal{E}^{+}. To address this, our LLMRec focuses on user-item interaction feature graph augmentation, which involves LLM-augmented u-i interactive edges EA\mathcal{E}_{\mathcal{A}}, and LLM-generated node features FA\textbf{F}_{\mathcal{A}}. The optimization target with augmented interaction feature graph is as:

The recommender fΘf_{\Theta} input union of original and augmented data, which consist of edges {E+,EA}\{\mathcal{E}^{+},\mathcal{E}_{\mathcal{A}}\} and node features {F,FA}\{\textbf{F},\textbf{F}_{\mathcal{A}}\}, and output quality representation h to predicted preference scores y^u,i\hat{y}_{u,i} by ranking the likelihood of user uu will interact with item ii.

Methodology

To conduct LLM-based augmentation, in this section, we address these questions: Q1: How to enable LLMs to predict u-i interactive edges? Q2: How to enable LLMs to generate valuable content? Q3: How to incorporate augmented contents into original graph contents? Q4: How to make model robust to the augmented data?

To directly confront the scarcity of implicit feedback, we employ LLM as a knowledge-aware sampler to sample pair-wise (Rendle et al., 2012) u-i training data from a natural language perspective. This increases potential effective supervision signals and helps gain a better understanding of user preferences by integrating contextual knowledge into the u-i interactions. Specifically, we feed each user’s historical interacted items with side information (e.g., year, genre) and an item candidates pool Cu={iu,1,iu,2,...,iu,∣Cu∣}\mathcal{C}_{u}={\{i_{u,1},i_{u,2},...,i_{u,|\mathcal{C}_{u}|}\}} into LLM. LLM then is expected to select items that user uu might be likely (iu+i^{+}_{u}) or unlikely (iu−i^{-}_{u}) to interact with from Cu\mathcal{C}_{u}. Here, we introduce Cu\mathcal{C}_{u} because LLMs can’t rank all items. Selecting items from the limited candidate set recommended by the base recommender (e.g., MMSSL (Wei et al., 2023a), MICRO (Zhang et al., 2022)), is a practical solution. These candidates Cu\mathcal{C}_{u} are hard samples with high prediction score y^ui\hat{y}_{ui} to provide potential, valuable positive samples and hard negative samples. It is worth noting that we represent each item using textual format instead of ID-corresponding indexes (Li et al., 2023b). This kind of representation offers several advantages: (1) It enables recommender to fully leverage the content in datasets, and (2) It intuitively reflects user preferences. The process of augmenting user-item interactive edges and incorporating it into the training data can be formalized as:

The utilization of LLMs-based sampler in this study to some extent alleviate noise (i.e., false positive) and non-interacted items issue (i.e., false negative) (Chen et al., 2023b; Lee et al., 2021) exist in raw implicit feedback. In this context, (i) false positive are unreliable u-i interactions, which encompass items that were not genuinely intended by the user, such as accidental clicks or instances influenced by popularity bias (Wang et al., 2021a); (ii) false negative represented by non-interacted items, which may not necessarily indicate user dispreference but are conventionally treated as negative samples (Chen et al., 2020). By taking LLMs as implicit feedback augmentor, LLMRec enables the acquisition of more meaningful and informative samples by leveraging the remarkable reasoning ability of LLMs with the support of LLMs’ knowledge. The specific analysis is supported by theoretical discussion in Sec. 3.4.1.

2. LLM-based Side Information Augmentation

Leveraging knowledge base and reasoning abilities of LLMs, we propose to summarize user profiles by utilizing users’ historical interactions and item information to overcome limitation of privacy. Additionally, the LLM-based item attributes generation aims to produce space-unified, and informative item attributes. Our LLM-based side information augmentation paradigm consists of two steps:

i) User/Item Information Refinement. Using prompts derived from the dataset’s interactions and side information, we enable LLM to generate user and item attributes that were not originally part of the dataset. Specific examples are shown in Fig. 2(b)(c).

ii) LLM-enhanced Semantic Embedding. The augmented user and item information will be encoded as features and used as input for the recommender. Using LLM as an encoder offers efficient and state-of-the-art language understanding, enabling profiling user interaction preferences and debiasing item attributes.

Formally, the LLM-based side information augmentation is as:

2.2. Side Information Incorporation (Q3)

After obtaining the augmented side information for user/item, an effective incorporation method is necessary. LLMRec includes a standard procedure: (1) Augmented Semantic Projection, (2) Collaborative Context Injection, and (3) Feature Incorporation. Let’s delve into each:

Collaborative Context Injection. To inject high-order (Wang et al., 2019) collaborative connectivity into augmented features f‾A,u\overline{\textbf{f}}_{\mathcal{A},u} and f‾A,i\overline{\textbf{f}}_{\mathcal{A},i}, LLMRec employs light weight GNNs (He et al., 2020) as the encoder.

Semantic Feature Incorporation. Instead of taking augmented features F‾A\overline{\textbf{F}}_{\mathcal{A}} as initialization of learnable vectors of recommender fΘf_{\Theta}, we opt to treat F‾A\overline{\textbf{F}}_{\mathcal{A}} as additional compositions added to the ID-corresponding embeddings (eu\textbf{e}_{u}, ei\textbf{e}_{i}). This allows flexibly adjust the influence of LLM-augmented features using scale factors and normalization. Formally, the F‾A\overline{\textbf{F}}_{\mathcal{A}}’s incorporation is presented as:

3. Training with Denoised Robustification (Q4)

In this section, we outline how LLMRec integrate augmented data into the optimization. We also introduce two quality constraint mechanisms for augmented edges and node features: i) Noisy user-item interaction pruning, and ii) MAE-based feature enhancement.

We train our recommender using the union set E∪EA\mathcal{E}\cup\mathcal{E}_{\mathcal{A}}, which includes the original training set E\mathcal{E} and the LLM-augmented set EA\mathcal{E}_{\mathcal{A}}. The objective is to optimize the BPR LBPR\mathcal{L}_{\text{BPR}} loss with increased supervisory signals E∪EA\mathcal{E}\cup\mathcal{E}_{\mathcal{A}}, aiming to enhance the recommender’s performance by leveraging the incorporated LLM-enhanced user preference:

Noise Pruning. To enhance the effectiveness of augmented data, we prune out unreliable u-i interaction noise. Technically, the largest values before minus are discarded after sorting each iteration. This helps prioritize and emphasize relevant supervisory signals while mitigating the influence of noise. Formally, the objective LBPR\mathcal{L}_{\text{BPR}} in Eq. 6 with noise pruning can be rewritten as follows:

The function SortAscend(⋅)[0:N]SortAscend(\cdot)[0:N] sorts values and selects the top-N. The retained number NN is calculated by N=(1−ω4)⋅∣E∪EA∣N=(1-\omega_{4})\cdot|\mathcal{E}\cup\mathcal{E}_{\mathcal{A}}|, where ω4\omega_{4} is a rate. This approach allows for controlled pruning of loss samples, emphasizing relevant signals while reducing noise. This can avoid the impact of unreliable gradient backpropagation, thus making optimization more stable and effective.

3.2. Enhancing Augmented Semantic Features via MAE

To mitigate the impact of noisy augmented features, we employ the Masked Autoencoders (MAE) for feature enhancement (He et al., 2022). Specifically, the masking technique is to reduce the model’s sensitivity to features, and subsequently, the feature encoders are strengthened through reconstruction objectives. Formally, we select a subset of nodes V~⊂V\widetilde{\mathcal{V}}\subset\mathcal{V} and mask their features using a mask token [MASK], denoted as f[MASK]\textbf{f}_{[MASK]} (e.g., a learnable vector or mean pooling). The mask operation can be formulated as follows:

The augmented feature after the mask operation is denoted as f‾~A\widetilde{\overline{\textbf{f}}}_{\mathcal{A}}. It is substituted as mask token f[MASK]\textbf{f}_{[MASK]} if the node is selected (V~⊂V\widetilde{\mathcal{V}}\subset\mathcal{V}), otherwise, it corresponds to the original augmented feature f‾A\overline{\textbf{f}}_{\mathcal{A}}. To strengthen the feature encoder, we introduce the feature restoration loss LFR\mathcal{L}_{FR} by comparing the masked attribute matrix f‾~A,i\widetilde{\overline{\textbf{f}}}_{\mathcal{A},i} with the original augmented feature matrix f‾A\overline{\textbf{f}}_{\mathcal{A}}, with a scaling factor γ\gamma. The restoration loss function LFR\mathcal{L}_{FR} is as follows:

The final optimization objective is the weighted sum of the noise-pruned BPR loss LBPR\mathcal{L}_{\text{BPR}} and the feature restoration (FR) loss LFR\mathcal{L}_{FR}.

4. In-Depth Analysis of our LLMRec

This section highlights challenges addressed by LLM-based augmentation in recommender systems. False negatives (non-interacted interactions) and false positives (noise) as in Fig. 3 (a) can affect data quality and result accuracy (Chen et al., 2020; Wang et al., 2021a). Non-interacted items do not necessarily imply dislike (Chen et al., 2020), and interacted one may fail to reflect real user preferences due to accidental clicks or misleading titles, etc. Mixing of unreliable data with true user preference poses a challenge in build accurate recommender. Identifying and utilizing reliable examples is key to optimizing the recommender (Chen et al., 2023b).

In theory, non-interacted and noisy interactions are used for negative y^u,i−\hat{y}_{u,i^{-}} and positive y^u,i+\hat{y}_{u,i^{+}} scores, respectively. However, their optimization directions oppose the true direction with large magnitudes, i.e., the model optimizes significantly in the wrong directions (as in Fig. 3 (b)), resulting in sensitive suboptimal results. Details. By computing the derivatives of the LBPR\mathcal{L}_{BPR} ( Eq. 10), we obtain positive gradients ∇u,i+\nabla_{u,i^{+}} = 1−σ(y^u+−)1-\sigma(\hat{y}_{u^{+-}}) and negative gradients ∇u,i−\nabla_{u,i^{-}} = σ(y^u+−)−1\sigma(\hat{y}_{u^{+-}})-1, where y^u+−=y^u,i+−y^u,i−\hat{y}_{u^{+-}}=\hat{y}_{u,i^{+}}-\hat{y}_{u,i^{-}}. Fig. 3 (b) illustrates these gradients and unveils some observations. Noisy interactions, although treated as positives, often have small values y^u,i+\hat{y}_{u,i^{+}} as false positives, resulting in large gradients ∇u,i+\nabla_{u,i^{+}}. Conversely, unobserved items, treated as negatives, tend to have relatively large values y^u,i−\hat{y}_{u,i^{-}} as false negatives, leading to small y^u+−\hat{y}_{u^{+-}} and large gradients ∇u,i−\nabla_{u,i^{-}}.

Conclusion. Wrong samples possess incorrect directions but are influential. LLM-based augmentation uses the natural language space to assist the ID vector space to provide a comprehensive reflection of user preferences. With real-world knowledge, LLMRec gets quality samples, reducing the impact of noisy and unobserved implicit feedback, improving accuracy, and speeding up convergence.

4.2. Time Complexity.

We analyze the time complexity. The projection of augmented semantic features has a time complexity of O(∣U∪I∣×dLLM×d)\mathcal{O}(|\mathcal{U}\cup\mathcal{I}|\times d_{LLM}\times d). The GNN encoder for graph-based collaborative context learning takes O(L×∣E+∣×d)\mathcal{O}(L\times|\mathcal{E}^{+}|\times d) time. The BPR loss function computation has a time complexity of O(d×∣E∪EA∣)\mathcal{O}(d\times|\mathcal{E}\cup\mathcal{E_{\mathcal{A}}}|), while the feature reconstruction loss has a time complexity of O(d×∣V~∣)\mathcal{O}(d\times|\widetilde{\mathcal{V}}|), where ∣V~∣|\widetilde{\mathcal{V}}| represents the count of masked nodes.

Evaluation

To evaluate the performance of LLMRec, we conduct experiments, aiming to address the following research questions:

RQ1: How does our LLM-enhanced recommender perform compared to the current state-of-the-art baselines?

RQ2: What is the impact of key components on the performance?

RQ3: How sensitive is the model to different parameters?

RQ3: Are the data augmentation strategies in our LLMRec applicable across different recommendation models?

RQ5: What is the computational cost associated with our devised LLM-based data augmentation schemes?

We perform experiments on publicly available datasets, i.e., Netflix and MovieLens, which include multi-modal side information. Tab. 1 presents statistical details for both the original and augmented datasets for both user and item domains. MovieLens. We utilize the MovieLens dataset derived from ML-10Mhttps://files.grouplens.org/datasets/movielens/ml-10m-README.html. Side information includes movie title, year, and genre in textual format. Visual content consists of movie posters obtained through web crawling by ourselves. Netflix. We collected its multi-model side information through web crawling. The implicit feedback and basic attribute are sourced from the Netflix Prize Datahttps://www.kaggle.com/datasets/netflix-inc/netflix-prize-data on Kaggle. For both datasets, CLIP-ViT(Radford et al., 2021) is utilized to encode visual features.

LLM-based Data Augmentation. The study employs the OpenAI package, accessed through LLMs’ APIs, for augmentation. The OpenAI Platform documentation provides detailshttps://platform.openai.com/docs/api-reference. Augmented implicit feedback is generated using the ”gpt-3.5-turbo-0613” chat completion model. Item attributes such as directors, country, and language are gathered using the same model. User profiling, based on the ”gpt-3.5-turbo-16k” model, includes age, gender, preferred genre, disliked genre, preferred directors, country, and language. Embedding is performed using the ”text-embedding-ada-002” model. The approximate cost of augmentation strategies on two datasets is 15.65 USD, 20.40 USD, and 3.12 USD, respectively.

1.2. Implementation Details

The experiments are conducted on a 24 GB Nvidia RTX 3090 GPU using PyTorch(Paszke et al., 2019) for code implementation. The AdamW optimizer(Loshchilov et al., 2017) is used for training, with different learning rate ranges of [5e−55e^{-5}, 1e−31e^{-3}] and [2.5e−42.5e^{-4}, 9.5e−49.5e^{-4}] for Netflix and MovieLens, respectively. Regarding the parameters of the LLMs, we choose the temperature from larger values {0.0, 0.6, 0.8, 1 } to control the randomness of the generated text. The value of top-p is selected from smaller values {0.0, 0.1, 0.4, 1} to encourage probable choices. The stream is set to false to ensure the completeness of responses. For more details on the parameter analysis, please refer to Section 4.4. To maintain fairness, both our method and the baselines employ a unified embedding size of 64.

1.3. Evaluation Protocols

We evaluate our approach in the top-K item recommendation task using three common metrics: Recall (R@k), Normalized Discounted Cumulative Gain (N@k), and Precision (P@k). To avoid potential biases from test sampling, we employ the all-ranking strategy(Wei et al., 2021a, 2020). We report averaged results from five independent runs, setting K to 10, 20, and 50 (reasonable for all-ranking). Statistical significance analysis is conducted by calculating pp-values against the best-performing baseline.

1.4. Baseline Description

Four distinct groups of baseline methods for thorough comparison. i) General CF Methods: MF-BPR (Rendle et al., 2012), NGCF (Wang et al., 2019) and LightGCN (He et al., 2020). ii) Methods with Side Information: VBPR (He and McAuley, 2016), MMGCN (Wei et al., 2019) and GRCN (Wei et al., 2020). iii) Data Augmentation Methods: LATTICE (Zhang et al., 2021b). iv) Self-supervised Methods: CLCRec (Wei et al., 2021b), MMSSL (Wei et al., 2023a) and MICRO (Zhang et al., 2022).

2. Performance Comparison (RQ1)

Tab. 2 compares our proposed LLMRec method with baselines.

Overall Model Superior Performance. Our LLMRec outperforms the baselines by explicitly augmenting u-i interactive edges and enhancing the quality of side information. It is worth mentioning that our model based on LATTICE’s (Zhang et al., 2021b) encoder, consisting of a ID-corresponding encoder and a feature encoder. This improvement underscores the effectiveness of our framework.

Effectiveness of Side Information Incorporation. The integration of side information significantly empowers recommenders. Methods like MMSSL (Wei et al., 2023a) and MICRO (Zhang et al., 2022) stand out for their effective utilization of multiple modalities of side information and GNNs. In contrast, approaches rely on limited content, such as VBPR (He and McAuley, 2016) using only visual features, or CF-based architectures like NGCF (Wang et al., 2019), without side information, yield significantly diminished results. This highlights the importance of valuable content, as relying solely on ID-corresponding records fails to capture the complete u-i relationships.

Inaccurate Augmentation yields Limited Benefits. Existing methods, such as LATTICE(Zhang et al., 2021b), MICRO(Zhang et al., 2022) that also utilize side information for data augmentation have shown limited improvements compared to our LLMRec. This can be attributed to two main factors: (1) The augmentation of side information with homogeneous relationships (e.g., i-i or u-u) may introduce noise, which can compromise the precise of user preferences. (2) These methods often not direct augmentation of u-i interaction data.

Advantage over SSL Approaches. Self-supervised models like, MMSSL(Wei et al., 2023a), MICRO(Zhang et al., 2022), have shown promising results in addressing sparsity through SSL signals. However, they do not surpass the performance of LLMRec, possibly because their augmented self-supervision signals may not align well with the target task of modeling u-i interactions. In contrast, we explicitly tackle the scarcity of training data by directly establishing BPR triplets.

3. Ablation and Effectiveness Analyses (RQ2)

We conduct an ablation study of our proposed LLMRec approach to validate its key components, and present the results in Table 3.

(1). w/o-u-i: Disabling the LLM-augmented implicit feedback EA\mathcal{E}_{\mathcal{A}} results in a significant decrease. This indicates that LLMRec increases the potential supervision signals by including contextual knowledge, leading to a better grasp of user preferences.

(2). w/o-u: Removing our augmentor for user profiling result in a decrease in performance, indicating that our LLM-enhanced user side information can effectively summarize useful user preference profile using historical interactions and item-end knowledge.

(3). w/o-u&i: when we remove the augmented side information for both users and items (FA,u,FA,i,1\textbf{F}_{\mathcal{A},u},\textbf{F}_{\mathcal{A},i,1}), lower recommendation accuracy is observed. This finding indicates that the LLM-based augmented side information provides valuable augmented data to the recommender system, assisting in obtaining quality and informative representations.

3.2. Impact of the Denoised Data Robustification.

w/o-prune: The removal of noise pruning results in worse performance. This suggests that the process of removing noisy implicit feedback signals helps prevent incorrect gradient descent.

w/o-QC: The performance suffer when both the limits on implicit feedback and semantic feature quality are simultaneously removed (i.e., w/o-prune + w/o-MAE). This indicates the benefits of our denoised data robustification mechanism by integrating noise pruning and semantic feature enhancement.

4. Hyperparameter Analysis (RQ3)

Temperature τ\tau of LLM: The temperature parameter τ\tau affects text randomness. Higher values (¿1.0) increase diversity and creativity, while lower values (¡0.1) result in more focus. We use τ\tau from {0,0.6,0.8,1}\{0,0.6,0.8,1\}. As shown in Table 4, increasing τ\tau initially improves most metrics, followed by a decrease.

Top-p pp of LLM: Top-p Sampling(Holtzman et al., 2019) selects tokens based on a threshold determined by the top-p parameter pp. Lower pp values prioritize likely tokens, while higher values encourage diversity. We use pp from {0,0.1,0.4,1}\{0,0.1,0.4,1\} and smaller pp values tend to yield better results, likely due to avoiding unlisted candidate selection. Higher ρ\rho values cause wasted tokens due to repeated LLM inference.

# of Candidate C\mathcal{C}: We use C\mathcal{C} to limit item candidates for LLM-based recommendation. {3,10,30}\{3,10,30\} are explored due to cost limitations, and Table 5 shows that C=10\mathcal{C}=10 yields the best results. Small values limit selection, and large values increase recommendation difficulty.

Prune Rate ω4\omega_{4}: LLMRec uses ω4\omega_{4} to control noise in augmented training data to be pruned. We set ω4\omega_{4} to {0.0, 0.2, 0.4, 0.6, 0.8} on both datasets. As shown in Fig. 4 (a), ω4=0\omega_{4}=0 yields the worst result, highlighting the need to constrain noise in implicit feedback.

4.2. Sensitivity of Recommenders to the Augmented Data.

# of Augmented Samples per Batch ∣EA∣|\mathcal{E}_{\mathcal{A}}| : LLMRec uses ω3\omega_{3} and batch size BB to control the number of augmented BPR training data samples per batch. ω3\omega_{3} is set to {0.0, 0.1, 0.2, 0.3, 0.4} on Netflix and {0.0, 0.2, 0.4, 0.6, 0.8} on MovieLens. Suboptimal results occur when ω3\omega_{3} is zero or excessively large. Increasing diversity and randomness can lead to a more robust gradient descent.

Scale ω2\omega_{2} for Incorporating Augmented Features: LLMRec uses ω2\omega_{2} to control feature magnitude, with values set to {0.0, 0.8, 1.6, 2.4, 3.2} on Netflix and {0.0, 0.1, 0.2, 0.3, 0.4} on MovieLens. Optimal results depend on the data, with suboptimal outcomes occurring when ω2\omega_{2} is too small or too large, as shown in Fig. 4 (c).

5. Model-agnostic Property (RQ4)

We conducted model-agnostic experiments on Netflix to validate the applicability of our data augmentation. Specifically, we incorporated the augmented implicit feedback EA\mathcal{E}_{\mathcal{A}} and features FA,u,FA,i\textbf{F}_{\mathcal{A},u},\textbf{F}_{\mathcal{A},i} into baselines MICRO, MMSSL, and LATTICE. As shown in Tab. 6, our LLM-based data improved the performance of all models, demonstrating their effectiveness and reusability. Some results didn’t surpass our model, maybe due to: i) the lack of a quality constraint mechanism to regulate the stability and quality of the augmented data, and ii) the absence of modeling collaborative signals in the same vector space, as mentioned in Sec. • ‣ 3.2.2.

6. Cost/Improvement Conversion Rate (RQ5)

To evaluate the cost-effectiveness of our augmentation strategies, we compute the CIR as presented in Tab. 7. The CIR is compared with the ablation of three data augmentation strategies and the best baseline from Tab. 3 and Tab. 2. The cost of the implicit feedback augmentor refers to the price of GPT-3.5 turbo 4K. The cost of side information augmentation includes completion (using GPT-3.5 turbo 4K or 16K) and embedding (using text-embedding-ada-002). We utilize the HuggingFace API tool for tokenizer and counting. The results in Tab. 7 show that ’U’ (LLM-based user profiling) is the most cost-effective strategy, and the overall investment is worthwhile.

Related Work

Content-based Recommendation. Existing recommenders have explored the use of auxiliary multi-modal side knowledge(Liang et al., 2023c, 2022), with methods like VBPR (He and McAuley, 2016) combine traditional CF with visual features, while MMGCN (Wei et al., 2019), GRCN (Wei et al., 2020) leverage GNNs to capture modality-aware higher-order collaborative signals. Recent approaches MMSSL (Wei et al., 2023a) and MICRO (Zhang et al., 2022) align modal signals with collaborative signals through contrastive SSL(Liang et al., 2023a), revealing the informative aspects of modal signals that benefit recommendations. However, the data noise, heterogeneity, and incompleteness can introduce bias. To overcome this, LLMRec explores LLM-based augmentation to improve the quality of the data.

Large Language Models (LLMs) for Recommendation. LLMs have gained attention in recommendation systems, with various efforts to use them for modeling user behavior (Wang et al., 2023; Ren et al., 2023a; Kang et al., 2023). LLMs have been employed as an inference model in diverse recommendation tasks, including rating prediction, sequential recommendation, and direct recommendation (Bao et al., 2023; Zhang et al., 2023b; Chen, 2023; Dai et al., 2023). Some efforts (Tang et al., 2023; Tian et al., 2023c) also tried to utilize LLMs to model structure relations. However, most previous methods primarily used LLMs as recommenders, abandoning the base model that has been studied for decades. We combine LLM-based data augmentation with classic CF, achieving both result assurance and enhancement concurrently.

Data Augmentation for Recommendation. Extensive research has explored data augmentation in recommendation systems (Lee et al., 2021; Huang et al., 2021). Various operations, such as permutation, deletion, swap, insertion, and duplication, have been proposed for sequential recommendation (Liu et al., 2021b; Petrov and Macdonald, 2022). Commonly used techniques include counterfactual reasoning (Wang et al., 2021b; Zhang et al., 2021a) and contrastive learning (Liu et al., 2021a). Our LLMRec use LLMs as an inference model to augment edge and enhance node features by leveraging consensus knowledge from the large model.

Conclusion

This study focuses on the design of LLM-enhanced models to address the challenges of sparse implicit feedback signals and low-quality side information by profiling user interaction preferences and debiasing item attributes. To ensure the quality of augmented data, a denoised augmentation robustification mechanism is introduced. The effectiveness of LLMRec is supported by theoretical analysis and experimental results, demonstrating its superiority over state-of-the-art recommendation techniques on benchmark datasets. Future directions for investigation include integrating causal inference into side information debiasing and exploring counterfactual factors for context-aware user preference.

References