Do LLMs Understand User Preferences? Evaluating LLMs On User Rating Prediction
Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta, Maheswaran Sathiamoorthy, Lichan Hong, Ed Chi, Derek Zhiyuan Cheng
Introduction
Large language models (LLMs) have shown an uncanny ability to handle a wide variety of tasks such as text generation (Chowdhery et al., 2022; Devlin et al., 2018; Raffel et al., 2020a; Brown et al., 2020), translation (Wang et al., 2019; Yang et al., 2020), and summarization (Liu and Lapata, 2019). The recent fine-tuning of LLMs on conversations and the use of techniques like instruction fine-tuning (Chung et al., 2022) and reinforcement learning from human feedback (RLHF) (Christiano et al., 2017) led to tremendous success to bring highly human-like chatbots (e.g. ChatGPT (OpenAI, 2022b), and Bard (Google, 2023)) to the average household. There are three key factors that directly contribute to LLMs’ versatility and effectiveness:
Knowledge from internet-scale real world information: LLMs are trained on enormous datasets of text, providing access to a wealth of real-world information. This information is converted to knowledge that can be used to answer questions, creatives writing (e.g., poems, and articles), and translate between languages.
Incredible generalization ability through effective few-shot learning: within certain context, LLMs are able to learn new tasks from an extremely small number of examples (a.k.a., few-shot learning). Strong few-shot learning capability gears up LLMs to be highly adaptable to new tasks.
Strong reasoning capability: LLMs are able to reason through a chain-of-thought process (Wei et al., 2022c, b), significantly improving their performance across many tasks (Wang et al., 2023).
Recently, there has been some early exploratory work to make use of LLMs for Search (Microsoft, 2023), Learning to Rank (Zou et al., 2021; Han et al., 2020), and Recommendation Systems (Geng et al., 2022; Cui et al., 2022; Liu et al., 2023). Specifically for recommendation systems, P5 (Geng et al., 2022) fine-tunes T5-small (60M) and T5-base(220M) (Raffel et al., 2020b), unifying both ranking, retrieval and other tasks like summary explanation into one model. M6-Rec (Cui et al., 2022) tackles the CTR prediction task by finetuning a LLM called M6 (300M) (Lin et al., 2021). Liu et al. (Liu et al., 2023) looked into whether conversational agents like ChatGPT could serve as an off-the-shelf recommender model with prompts as the interface and reported zero-shot performance on rating prediction against baselines like MF and MLP. However, there is a noticeable absence of a comprehensive study that meticulously evaluates LLMs of varying sizes and contrasts them against carefully optimized, strong baselines.
In this paper, we explore the use of off-the-shelf large language models (LLMs) for recommendation systems. We study a variety of LLMs of various sizes ranging from 250M to 540B parameters. We focus on the specific task of user rating prediction, and evaluate the performance of these LLMs under three different regimes: 1. zero-shot 2. few-shot, and 3. fine-tuning. We then carefully compare them with the state-of-the-art recommendation models on two widely adopted recommendation benchmark datasets.
We empirically study the zero-shot and few-shot performance of off-the-shelf LLMs with a wide spectrum of model sizes. We found that larger models (over 100B parameters) can provide reasonable recommendations under the cold-start scenario, achieving comparable performance to decent heuristic-based baselines.
We show that zero-shot LLMs still fall behind traditional recommender models that utilize human interaction data. Zero-shot LLMs only achieve comparable performance than two surprisingly trivial baselines that always predicts the average item or user rating. Furthermore, they significantly underperform traditional supervised recommendation models, indicating the importance of user interaction data.
Through numerous experiments that fine-tune LLMs on human interaction data, we demonstrate that fine-tuned LLMs can achieve comparable or even better performance than traditional models with only a small fraction of the training data, showing its promise in data efficiency.
Related Work
One of the earliest works that explored formulating the recommendation problem as a natural language task is (Zhang et al., 2021). They used BERT (Devlin et al., 2018) and GPT-2 (Radford et al., 2019) on the Movielens dataset (Harper and Konstan, 2015) to show that such language models perform surprisingly well, though not as good as well tuned baselines like GRU4Rec (Hidasi et al., 2015).
P5 (Geng et al., 2022) fine-tunes a popular open-sourced T5 (Raffel et al., 2020b) model, unifying both ranking, retrieval and other tasks like summary explanation into one model. M6-Rec (Cui et al., 2022) is another related work, but they tackle the CTR prediction task by finetuning a LLM called M6 (Lin et al., 2021).
Two recent works explore the use of LLMs for zero-shot prediction. ChatRec (Gao et al., 2023) handles zero-shot prediction as well as being interactive and providing explanations. (Wang and Lim, 2023) takes a three-stage prompting approach to generate next item recommendation in the Movielens dataset and achieves competitive metrics, although not being able to beat strong sequential recommender baselines such as SASRec (Kang and McAuley, 2018).
2. Large Language Models
Once people realized that scaling up sizes of data and model helps language models, there has been a series of large language models proposed and built: e.g. PaLM (Chowdhery et al., 2022), GPT-3 (Brown et al., 2020) and recent ones such as OPT (Zhang et al., 2022) and LLaMA (Touvron et al., 2023). One of the unique abilities of LLMs has been in their ability to reason about things, which is further improved by techniques such as chain-of-thought prompting (Wei et al., 2022c), self-consistency (Wang et al., 2022) and self-reflection (Shinn et al., 2023).
Another major strong capability of LLMs is instruction following that models can generalize to unseen tasks by following the given natural language instructions. Researchers have found that techniques like instruction fine-tuning (Chung et al., 2022) and RLHF (Christiano et al., 2017) can significantly improve LLMs’ capability to perform tasks given natural language descriptions that align with human’s preferences. As one of the tasks that can be described in natural language, ‘recommendation’ has become a promising new capability for LLMs. In this work, we focus on the models that have been fine-tuned to improve their instruction following capability such as ChatGPT (OpenAI, 2022b), GPT-3 (text-davinci-003 (OpenAI, 2022a)), Flan-U-PaLM and Flan-T5 (Chung et al., 2022).
Method
We study the task of user rating prediction, formulated as: Given a user , a sequence of user ’s historical interactions and an item , predict the rating that the user will give to the item , where the user historical interaction sequence is ordered by time ( is the most recent item that the user consumed), and each interaction is represented by information about the item (e.g., ID, title, metadata, etc.) that the user has consumed as well as the rating the user gave to the item.
2. Zero-shot and Few-shot LLMs for Rating Prediction
We demonstrate the zero-shot and few-shot prompts used for the rating prediction task on the MovieLens dataset in figure 2. As shown in the figure, the input prompts depict several important features represented as text, including user’s past rating history and candidate item features (title and genre). Finally, to elicit a numeric rating from the model with the rating scale, the input prompt specifies a numerical rating scale. The model response is parsed to extract the rating output from the model. However, we discovered that LLMs can be highly sensitive to the input prompts and do not always follow the provided instruction. For instance, we found that certain LLMs may offer additional reasoning or not provide a numerical rating at all. To resolve this, we performed additional prompt engineering by adding additional instructions such as "Give a single number as rating without explanation" and "Do not give reasoning" to the input prompt.
3. Fine-tuning LLMs for Rating Prediction
In traditional recommender system research, it has been widely shown that training models with human interaction data is effective and critical to improve recommender’s capability of understanding user preference.
Here, we explore training the LLMs with human interaction and study how it could improve the model performance. We focus on fine-tuning a family of LLMs, namely Flan-T5, since they are publicly available and have competitive performance on a wide range of benchmarks. The rating prediction task could be formulated into one of two tasks: (1) multi-class classification; or (2) regression, as shown in Figure 3(b).
Following (Raffel et al., 2020a; Chung et al., 2022), we formulate the rating regression task as a 5-way classification task, where we take the rating 1 to 5 as 5 classes. During training, we use the cross-entropy loss as other classification tasks, as shown below:
where is the ground-truth rating for the -th item and is the number of total training examples.
During inference, we compute the log-likelihood for the model output each class and choose the class with the largest probability as the final prediction.
Experiments
We conduct extensive experiments to answer the following research questions: RQ1: Do off-the-shelf LLMs perform well for zero-shot and few-shot recommendations? RQ2: How do LLMs compare with traditional recommenders in a fair setting RQ3: How much does model size matter for LLMs when used for recommenders? RQ4: Do LLMs converge faster than traditional recommender models?
To evaluate the user rating prediction task, we use two widely adopted benchmark datasets for evaluating model performance on recommendations. Both datasets consist of user review ratings that range from 1 to 5.
MovieLens (Harper and Konstan, 2016): We use the version MovieLens-1M that includes 1 million user ratings for movies.
Amazon-Books (Ni et al., 2019): We use the “Books” category of the Amazon Review Dataset with users’ ratings on items. We use the 5-core version that filters out users and items with less than 5 interactions.
1.2. Training / Test Split
To create the training and test sets, we follow the single-time-point split (Sun, 2022). We first filter out the ratings associated with items that don’t have metadata, then sort all user ratings in chronological order. Finally, we take the first 90% ratings as the training set and the remaining as the test set. Each training example is a tuple of , where the label is a 5 Likert-scale rating. The input features are , , and a list of item_metadata features. The statistics of the datasets are shown in the Table 1. Due to the high computation cost of zero-shot and few-shot experiments based on LLMs, we randomly sample from the test set of each dataset into 2,000 tuples as a smaller test set. For all our experiments, we report results on the sampled test set. And we truncate the user sequence to the most recent 10 interactions during training and evaluation.
1.3. Evaluation Metrcis
We use the widely adopted metrics RMSE (Root Mean Squared Error) and MAE (Mean Average Error) to measure model performance on rating prediction. Moreover, we use ROC-AUC to evaluate the model’s performance on ranking, where ratings greater than or equal to 4 are considered as positives and the rest as negatives. In this case, AUC measures whether the model ranks the positives higher than negatives.
2. Baselines and LLMs
Traditional Recommeder: We consider several traditional recommendation models as strong baselines, including 1. Matrix Factorization (MF) (Rendle et al., 2012), and 2. Multi-layer Perceptrons (MLP) (He et al., 2017). For MF and MLP, only user ID and item ID are used as input features.
Attribute and Rating-aware Sequential Rating Predictor: In our experiments, we supply the LLM with historical item metadata, such as titles and categories, along with historical ratings. However, to the best of our knowledge, there is no existing method designed for this settingThe most related works are SASRec (Kang and McAuley, 2018) and CARCA (Rashed et al., 2022), however they are designed for next item prediction instead of rating prediction, and thus not directly applicable to our case.. To ensure a fair comparison, we construct a Transformer-MLP model to efficiently process the same input information provided to the LLM.
There are three key design choices: (i) feature processing: We treat all features as sparse features, and learn their embeddings end-to-end. For example, we use one-hot encoding for genres, and create an embedding table, where the i-th row is genre i’s embedding. Similarly, we obtain bag-of-words encodings via applying a tokenizerhttps://www.tensorflow.org/text/api_docs/python/text/WhitespaceTokenizer on titles, and then look up the corresponding embedding. (ii) user modeling: for each user behavior, we use Add or Concat to aggregate all embeddings (e.g. item ID, title, genres/category, rating) into one, and then adopt bi-directional self-attention (Vaswani et al., 2017) layers with learned position embeddings to model users’ past behaviors. Similar to SASRec, we use the most recent behavior’s output embedding as the user summary; (iii) Fuse user and candidate for rating prediction: we apply a MLP on top of the user embedding along with other candidate item features to generate the final rating prediction, and optimize for minimizing MSE.
To properly tune the baseline models, we defined a hyper-parameter search space (e.g. for embedding dimension, learning rate, network size, Add or Concat aggregation, etc.), and perform more than 100 search trials using Vizier (Golovin et al., 2017), a black-box hyper-parameter optimization tool.
Heuristics: We also include three heuristic-based baselines: (1) global average rating: (2) candidate item average rating, and (3) user past average rating, meaning the model’s prediction is depending on (1) the average rating among all user-item ratings, (2) the average rating from the candidate item or (3) the user’s average rating in the past.
2.2. LLMs for Zero-shot and Few-shot Learning
: We used the LLMs listed below for zero-shot and few-shot learning. We use a temperature of 0.1 for all LLMs, as the LLM’s output in our case is simply a rating prediction. We use GPT-3 models from OpenAI (OpenAI, 2023): (i) text-davinci-003 (175B): The most capable GPT-3 model with Reinforcement Learning from Human Feedback (RLHF) (Stiennon et al., 2020); (ii) ChatGPT: the default model is gpt-3.5-turbo, fine-tuned on both human-written demonstrations and RLHF, and further optimized for conversation. Flan-U-PaLM (540B) is the largest and strongest model in (Chung et al., 2022), it applies both FLAN instruction tuning (Wei et al., 2022a) and UL2 training objective (Chung et al., 2022) on PaLM (Chowdhery et al., 2022).
2.3. LLMs for Fine-tuning
For fine-tuning methods, we use Flan-T5-Base (250M) and Flan-T5-XXL (11B) models in the experiments. We set the learning rate to 5e-5, batch size to 64, drop out rate to 0.1 and train 50k steps on all datasets.
3. Zero-Shot and Few-shot LLMs (RQ1)
As shown in Table 2, we conduct experiments on several off-the-shelf LLMs in the zero-shot setting. We observed that LLMs seem to understand the task from the prompt description, and predict reasonable ratings. LLMs outperform global average rating in most cases, and perform comparable with item or user average ratings. For example, text-davinci-003 performs slightly worse than candidate item average ratings on Movielens but outperforms on Amazon-Books. For few-shot experiments, we provide 3 examples in the prompt (3-shot). Compared against zero-shot, we found that the AUC for few-shot LLMs are improved, while there is no clear pattern in RMSE and MAE.
Furthermore, we found that both zero-shot and few-shot LLMs under-perform traditional recommendation models trained with interaction data. As shown in Table 2, GPT-3 and Flan-U-PaLM models achieve significantly lower performance compared to supervised models. The inferior performance could be due to the lack of user-item interaction data in LLMs’ pre-training, and hence they don’t have knowledge about human preference for different recommendation tasks. Moreover, recommendation tasks are highly dataset-dependent: (e.g.) the same movie can have different average ratings on different platform. Hence, without knowing the dataset-specific statistics, it’s impossible for LLM to provide a universal prediction that is suitable for every datasets.
4. LLMs vs. Traditional Recommender Models (RQ2)
Fine-tuning LLMs is an effective way to feed dataset statistics into LLMs, and we found the performance of fine-tune LLMs are much better than zero/few-shot LLMs. Also, when fine-tuning the Flan-T5-base model with the classification loss, the performance is much worse than fine-tuning with the regression loss on all three metrics. This indicates the importance of choosing the right optimizing objective for fine-tuning LLMs.
Comparing against the strongest baseline Transformer-MLP, we found fine-tuned Flan-T5-XXL has better MAE and AUC, implying fine-tune LLMs may be more suitable for ranking tasks.
5. Effect of Model Size (RQ3)
For all the LLMs we studied of different model sizes vary from 250M and 500B parameters, we were able to use zero-shot or few-shot prompts to let them output a rating prediction between 1 to 5. This shows the effectiveness of instruction tuning that enables these LLMs (Flan-T5, Flan-U-PaLM, GPT-3) to follow the prompt. We further found that only LLMs with size greater than 100B perform reasonably well on rating prediction in the zero-shot setting, as shown in Figure 1. For fine-tuning experiments, we also found that Flan-T5-XXL outperforms Flan-T5-Base on both datasets, as shown in the last two rows of Table 2.
6. Data Efficiency of LLMs (RQ4)
As LLMs have learned vast amounts of world knowledge during pre-training, while traditional recommender models are trained from scratch, we compare their convergence curves in Figure 4 to examine whether LLMs have better data efficiency. We can see that for RMSE, both methods could converge to reasonable performance with a small fraction of data. This is probably because that even average rating of all items has a relatively low RMSE, and thus as long as a model learns to predict a rating near the average rating, it could achieve reasonable performance. For AUC the trend is more clear, as simply predicting average rating results in an AUC of 0.5. We found that a small fraction of data is required for LLM to achieve good performance, while Transformer+MLP needs much more training data (at least 1 epoch) for convergence.
Conclusion
In this paper, we evaluate the effectiveness of large language models as a recommendation system for user rating prediction in three settings: 1. zero-shot; 2. few-shot; and 3. fine-tuning. Compared to traditional recommender methods, our results revealed that LLMs in zero-shot and few-shot LLMs fall behind fully supervised methods, implying the importance of incorporating the target dataset distribution into LLMs. On the other hand, fine-tuned LLMs can largely close the gap with carefully designed baselines in key metrics. LLM-based recommenders have several benefits: (i) better data efficiency; (ii) simplicity for feature processing and modeling: we only need to convert information into a prompt without manually designing feature processing strategies, embedding methods, and network architectures to handle various kind of information; (iii) potential for unlock conversational recommendation capabilities. Our work sheds light on the current status of LLM-based recommender systems, and in the future we will further look into improving the performance via methods like prompt tuning, and explore novel recommendation applications enabled by LLMs.