Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding

Jiaxi Tang, Ke Wang

introduction

Recommender systems have become a core technology in many applications. Most systems, e.g., top-NN recommendation (Hu et al. 2008)(Pan et al. 2008), recommend the items based on the user’s general preferences without paying attention to the recency of items.

For example, some user always prefer Apple’s products to Samsung’s products. General preferences represent user’s long term and static behaviors. Another type of user behaviors is sequential patterns where the next item or action more likely depends on the items or actions the user engaged recently. Sequential patterns represent the user’s short term and dynamic behaviors and come from a certain relationship between the items within a close proximity of time. For example, a user likely buys phone accessories soon after buying an iPhone, though in general the user does not buy phone accessories. In this case, the systems that consider only general preferences will miss the opportunity of recommending phone accessories after selling an iPhone since buying phone accessories is not a long term user behavior.

To model user’s sequential patterns, the work in (Rendle et al. 2010; Liu et al. 2009) considers top-NN sequential recommendation that recommends NN items that a user likely interacts with in a near future. This problem assumes a set of users U={u1,u2,⋯ ,u∣U∣}\mathcal{U}=\{u_{1},u_{2},\cdots,u_{|\mathcal{U}|}\} and a universe of items I={i1,i2,⋯ ,i∣I∣}\mathcal{I}=\{i_{1},i_{2},\cdots,i_{|\mathcal{I}|}\}. Each user uu is associated with a sequence of some items from I\mathcal{I}, Su=(S1u,⋯ ,S∣Su∣u)\mathcal{S}^{u}=(\mathcal{S}^{u}_{1},\cdots,\mathcal{S}^{u}_{|\mathcal{S}^{u}|}), where Siu∈I\mathcal{S}^{u}_{i}\in\mathcal{I}. The index tt for Stu\mathcal{S}^{u}_{t} denotes the order in which an action occurs in the sequence Su\mathcal{S}^{u}, not the absolute timestamp as in temporal recommendation like (Wu et al. 2017; Zhang et al. 2014; Koren 2010). Given all users’ sequences Su\mathcal{S}^{u}, the goal is to recommend each user a list of items that maximize her/his future needs, by considering both general preferences and sequential patterns. Unlike conventional top-NN recommendation, top-NN sequential recommendation models the user behavior as a sequence of items, instead of a set of items.

2. Limitations of Previous Work

The Markov chain based model (Rendle et al. 2010; He and McAuley 2016; Cheng et al. 2013; Wang et al. 2015a) is an early approach to top-NN sequential recommendation, where an LL-order Markov chain makes recommendations based on LL previous actions. The first-order Markov chain is an item-to-item transition matrix learnt using maximum likelihood estimation. Factorized personalized Markov chains (FPMC) (Rendle et al. 2010) proposed by Rendle et al. and its variant (Cheng et al. 2013) improved this method by factorizing this transition matrix into two latent and low-rank sub-matrices. Factorized Sequential Prediction with Item Similarity ModeLs (Fossil) (He and McAuley 2016) proposed by He et al. generalizes this method to high-order Markov chains using a weighted sum aggregation over previous items’ latent representations. However, existing approaches suffered from two major limitations:

Fail to model union-Level sequential patterns. As shown in Figure 1(a), the Markov chain models only point-level sequential patterns where each of the previous actions (blue) influences the target action (yellow) individually, instead of collectively. FPMC and Fossil fall into this taxonomy. Although Fossil (He and McAuley 2016) considers a high-order Markov chain, the overall influence is a weighted sum of previous items’ latent representations factorized from first-order Markov transition matrices. Such aggregation of point-level influences is not sufficient to model the union-level influences shown in Figure 1(b) where several previous actions, in that order, jointly influence the target action. For example, buying both milk and butter together leads to a higher probability of buying flour than buying milk or butter individually; buying both RAM and Hard Drive is a better indication of buying Operating System next than buying only one of the components.

Fail to allow skip behaviors. Existing models don’t consider skip behaviors of sequential patterns as shown in Figure 1(c), where the impact from past behaviors may skip a few steps and still have strength. For example, a tourist has check-ins sequentially at airport, hotel, restaurant, bar, and attraction. While the check-ins at the airport and hotel do not immediately precede the check-in of the attraction, they are strongly associated with the latter. On the other hand, the check-in at the restaurant or bar has little influence on the check-in of the attraction (because they do not necessarily occur). A LL-order Markov chain does not explicitly model such skip behaviors because it assumes that the LL previous steps have an influence on the immediate next step.

To provide evidences of union-level influences and skip behaviors, we mine sequential association rules (Agrawal and Srikant 1995; Han et al. 2011) of the following form from two real life data sets, MovieLens and Gowalla (see the details of these data sets in Section 4)

For a rule X→YX\rightarrow Y of the above form, the support count sup(XY)sup(XY) is the number of sequences in which XX and YY occur in order as in the rule, and the confidence, sup(XY)sup(X)\frac{sup(XY)}{sup(X)}, is the percentage of the sequences in which YY follows XX among those in which XX occurs. This rule represents the joint influence of all the items in XX on YY. By changing the right hand side to St+1u\mathcal{S}^{u}_{t+1} or St+2u\mathcal{S}^{u}_{t+2}, the rule also captures the influences with one or two step skips. Figure 2 summarizes the number of rules found versus the Markov order LL and skip steps with the minimum support count = 5 and the minimum confidence = 50% (we also tried the minimum confidence of 10%, 20%, and 30%, these trends are similar). Most rules have the orders L=2L=2 and L=3L=3 and the confidence of rules gets higher for larger LL. The figure also tells that a sizable number of rules have skip steps 1 or 2. These findings support the existence of union-level influences and skip behaviors.

3. Contributions

To address these above limitations of existing works, we propose a ConvolutionAl Sequence Embedding Recommendation Model, or Caser for short, as a solution to top-NN sequential recommendation. This model leverages the recent success of convolution filters of Convolutional Neural Network (CNN) to capture local features for image recognition (Krizhevsky et al. 2012; Karpathy et al. 2014) and natural language processing (Kim 2014). The novelty of Caser is to represent the previous LL items as an L×dL\times d matrix E\bm{E}, where dd is the number of latent dimensions and the rows preserve the order of the items. Similar to (Kim 2014), we regard this embedding matrix as the “image” of the LL items in the latent space and search for sequential patterns as local features of this “image” using various convolutional filters. Unlike image recognition, however, this “image” is not given in the input and must be learnt simultaneously with all filters.

Compared to existing methods, Caser offers several distinct advantages. (1) Caser uses horizontal and vertical convolutional filters to capture sequential patterns at point-level, union-level, and of skip behaviors. (2) Caser models both users’ general preferences and sequential patterns, and generalizes several existing state-of-the-art methods in a single unified framework. (3) Caser outperforms state-of-the-art methods for top-NN sequential recommendation on real life data sets. In the rest of the paper, we discuss further related work in Section 2, the Caser method in Section 3, and experimental studies in Section 4.

Further Related Work

Conventional recommendation methods, e.g., collaborative filtering (Sarwar et al. 2001), matrix factorization (Koren et al. 2009; Salakhutdinov and Mnih 2007), and top-NN recommendation (Hu et al. 2008)(Pan et al. 2008), are not suitable for capturing sequential patterns because they do not model the order of actions. Early works on sequential pattern mining (Agrawal and Srikant 1995; Han et al. 2011) find explicit sequential association rules based on statistical co-occurrences (Liu et al. 2009). This approach depends on the explicit representation of patterns, thus, could miss patterns in unobserved states. Also, it suffers from a potentially large search space, sensitivity to threshold settings, and a large number of rules, most being redundant.

Restricted Bolzmann Machine (RBM) (Salakhutdinov et al. 2007) is the first successful 2-layers neural network that is applied to recommendation problems. Auto-encoder framework (Sedhain et al. 2015; Wang et al. 2015b) and its variant denoising auto-encoder (Wu et al. 2016) also produce a good recommendation performance. Convolutional neural network (CNN) (Zheng et al. 2017) has been used to extract users’ preferences from their reviews. None of these works is for sequential recommendation.

Recurrent neural networks (RNN) was used for session-based recommendation (Hidasi et al. 2015; Jannach and Ludewig 2017). While RNN has shown to have an impressive capability in modeling sequences (Mikolov et al. 2010), its sequentially connected network structure may not work well under sequential recommendation setting. Because in sequential recommendation problem, not all adjacent actions have dependency relationships (e.g. a user bought i2i_{2} after i1i_{1} only because she loves i2i_{2}). Our experimental results in Section 4 verify this point: RNN-based method performs better when data sets contains considerable sequential patterns. While our proposed method doesn’t model sequential pattern as adjacent actions, it adopts convolutional filters from CNN and model sequential patterns as local features of the embeddings of previous items. This approach offers the flexibility of modeling sequential patterns at both point level and union level, and skip behaviors in a single unified framework. In fact, we will show that Caser generalizes several state-of-the-art methods.

A related but different problem is temporal recommendation (Zhang et al. 2014; Wu et al. 2017; Song et al. 2016). For example, temporal recommendation recommends coffee in the morning, instead of evening, whereas our top-NN sequential recommendation would recommend phone accessories soon after a user bought an iPhone, independently of the time. Clearly, the two problems are different and require different solutions.

Proposed Methodology

The proposed model, ConvolutionAl Sequence Embedding Recommendation (Caser), incorporates the Convolutional Neural Network (CNN) to learn sequential features, and Latent Factor Model (LFM) to learn user specific features. The goal of Caser’s network design is multi-fold: capture both user’s general preferences and sequential patterns, at both union-level and point-level, and capture skip behaviors, all in unobserved spaces. Shown in Figure 3 Caser consists of three components: Embedding Look-up, Convolutional Layers, and Fully-connected Layers. To train the CNN, for each user uu, we extract every LL successive items as input and their next TT items as the targets from the user’s sequence Su\mathcal{S}^{u}, shown on the left side of Figure 3. This is done by sliding a window of size L+TL+T over the user’s sequence, and each window generates a training instance for uu, denoted by a triplet (uu, previous LL items, next TT items).

2. Convolutional Layers

Our approach leverages the recent success of convolution filters of CNN in capturing local features for image recognition (Krizhevsky et al. 2012; Karpathy et al. 2014) and natural language processing (Kim 2014). Borrows the idea of using CNN in text classification (Kim 2014), our approach regards the L×dL\times d matrix E\bm{E} as the “image” of the previous LL items in the latent space and regard sequential patterns as local features of this “image”. This approach enables the use of convolution filters to search for sequential patterns. Figure 4 shows two “horizontal filters” that capture two union-level sequential patterns. These filters, represented as h×dh\times d matrices, have the height h=2h=2 and the full width equal to dd. They pick up signals for sequential patterns by sliding over the rows of E\bm{E}. For example, the first filter picks up the sequential pattern “(Airport, Hotel) →\rightarrow Great Wall” by having larger values in the latent dimensions where Airport and Hotel have larger values. Similarly, a “vertical filter” is a L×1L\times 1 matrix and will slide over the columns of E\bm{E}. More details are explained below. Unlike image recognition, the “image” E\bm{E} is not given because the embedding Qi\bm{Q}_{i} for all items ii must be learnt simultaneously with all filters.

where the symbol ⊙\odot denotes the inner product operator and ϕc(⋅)\phi_{c}(\cdot) is the activation function for convolutional layers. This value is the inner product between Fk\bm{F}^{k} and the sub-matrix formed by the row ii to row i−h+1i-h+1 of E\bm{E}, denoted by Ei:i+h−1\bm{E}_{i:i+h-1}. The final convolution result of Fk\bm{F}^{k} is the vector

Horizontal filters interact with every successive hh items through their embeddings E\bm{E}. Both the embeddings and the filters are learnt to minimize an objective function that encodes the prediction error of target items (more in Section 3.4). By sliding filters of various heights, a significant signal will be picked up regardless of location. Therefore, horizontal filters can be trained to capture union-level patterns with multiple union sizes.

3. Fully-connected Layers

We concatenate the outputs of the two convolutional layers and feed them into a fully-connected neural network layer to get more high-level and abstract features:

To capture user’s general preferences, we also look-up the user embedding Pu\bm{P}_{u} and concatenate the two dd-dimensional vectors, z\bm{z} and Pu\bm{P}_{u}, together and project them to an output layer with ∣I∣|\mathcal{I}| nodes, written as

4. Network Training

To train the network, we transform the values of the output layer, y(u,t)\bm{y}^{(u,t)}, to probabilities by:

where σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) is the sigmoid function. Let Cu={L+1,L+2,...,∣Su∣}\mathcal{C}^{u}=\{L+1,L+2,...,|\mathcal{S}^{u}|\} be the collection of time steps for which we would like to make predictions for user uu. The likelihood of all sequences in the dataset is:

To further capture skip behaviors, we could consider the next TT target items, Dtu={Stu,St+1u,...,St+Tu}\mathcal{D}^{u}_{t}=\{\mathcal{S}^{u}_{t},\mathcal{S}^{u}_{t+1},...,\mathcal{S}^{u}_{t+T}\}, at once by replacing the immediate next item Stu\mathcal{S}^{u}_{t} in the above equation with Dtu\mathcal{D}^{u}_{t}. Taking the negative logarithm of likelihood, we get the objective function, also known as binary cross-entropy loss:

Following previous works (Rendle et al. 2010; He and McAuley 2016; Wu et al. 2016), for each target item ii, we randomly sample several (3 in our experiments) negative instances jj in the second term.

5. Recommendation

After obtaining the trained neural network, to make recommendations for a user uu at time step tt, we take uu’s latent embedding Pu\bm{P}_{u} and extract his last LL items’ embeddings given by Eqn (2) as the neural network input. We recommend the NN items that have the highest values in the output layer y\bm{y}. The complexity for making recommendations to all users is O(∣U∣∣I∣d)O(|\mathcal{U}||\mathcal{I}|d), where the complexity of convolution operations is ignored. Note that the number of target items TT is a hyperparameter used during the model training, whereas NN is the number of items recommended after the model is trained.

6. Connection to Existing Models

We show that Caser is a generalization of several previous models.

Caser vs. MF. By discarding all convolutional layers and all bias terms, our model becomes a vanilla LFM with user embeddings as user latent factors and its associated weights as item latent factors. MF usually contains bias terms Top-NN recommendation ranks the items for each user individually, which is invariant to user bias and global bias., which is b′\bm{b^{\prime}} in our model. After discarding all convolutional layers, the resulting model is the same as MF:

Caser vs. FPMC. FPMC fuses factorized first-order Markov chain with LFM and is optimized by Bayesian personalized ranking (BPR). Although Caser uses a different optimization criterion, i.e., the cross-entropy, it is able to generalize FPMC by copying the previous item’s embedding to the hidden layer z\bm{z} and not using any bias terms:

As FPMC uses BPR as the criterion, our model is not exactly the same as FPMC. However, BPR is limited to have only 1 target and negative sample at each time step. Our cross-entropy loss does not have these limitations.

As discussed for Eqn (7), this vertical filter serves as the weighted sum of the embeddings of the LL previous items, like in Fossil, though Fossil uses Similarity Model instead of LFM and factorizes it in the same latent space as Markov model. Another difference is that Fossil uses one local weighting for each user while we use a number of global weighting through vertical filters.

Experiments

We compare Caser with state-of-the-art methods. The source code of Caser and processed data sets are available online https://github.com/graytowne/caser.

Datasets. Sequential recommendation makes sense only when the data set contains sequential patterns. To identify such data sets, we applied sequential association rule mining to several public data sets and computed their sequential intensity defined by:

The numerator is the total number of rules in the form of Eqn (1) found using a minimum threshold on support (i.e., 5) and confidence(i.e., 50%) with Markov order LL range from 1 to 5. The denominator is the total number of users. We use SISI to estimate the intensity of sequential signals in a data set.

The four data sets with their SISI are described in Table 1. MovieLens https://grouplens.org/datasets/movielens/1m/ is the widely used movie rating data. Gowalla https://snap.stanford.edu/data/loc-gowalla.html constructed by (Cho et al. 2011) and Foursquare obtained from (Yuan et al. 2014) contain implicit feedback through user-venue check-ins. Tmall, the largest B2C platform in China, is a user-purchase data obtained from IJCAI 2015 competition https://ijcai-15.org/index.php/repeat-buyers-prediction-competition, which aims to forecast repeated buyers. Following previous works (He and McAuley 2016; Rendle et al. 2009; Wu et al. 2016), we converted all numeric ratings to implicit feedback of 1. We also removed cold-start users and items of having less than nn feedbacks, as dealing with cold-start recommendation is usually treated as a separate issue in the literature (Wu et al. 2016; He et al. 2017b; He and McAuley 2016; Rendle et al. 2010). nn is 5,15,10,10 for MovieLens, Gowalla, Foursquare, and Tmall. The Amazon data previously used in (He and McAuley 2016; He et al. 2017a) was not used due to its SISI (0.0026 for ‘Office Products’ category, 0.0019 for ‘Clothing, Shoes, Jewelry’ and ’Video Games’ category), in other words, its sequential signals are much weaker than the above data sets.

Following (Liu et al. 2009; Zhao et al. 2016; Yuan et al. 2014), we hold the first 70% of actions in each user’s sequence as the training set and use the next 10% of actions as the validation set to search the optimal hyperparameter settings for all models. The remaining 20% actions in each user’s sequence are used as the test set for evaluating a model’s performance.

Evaluation Metrics. As in (Pan et al. 2008; Rendle et al. 2010; Wang et al. 2015b; Wu et al. 2016), we evaluate a model by Precision@NN, Recall@NN, and Mean Average Precision (MAP). Given a list of top NN predicted items for a user, denoted R^1:N\hat{R}_{1:N}, and the last 20% of actions in her/his sequence (i.e., denoted RR (i.e., the test set), Precision@NN and Recall@NN are computed by

We report the average of these values of all users. N∈{1,5,10}N\in\{1,5,10\}. The Average Precision (AP) is defined by

where rel(N)=1rel(N)=1 if the NN-th item in R^\hat{R} is in RR. The Mean Average Precision (MAP) is the average of AP for all users.

2. Performance Comparison

We compare our method, Caser, proposed in Section 3 with the following baselines.

POP. All items are ranked by their popularity in all users’ sequences, and the popularity is determined by the number of interactions.

BPR. Combined with Matrix Factorization model, Bayesian personalized ranking (Rendle et al. 2009) is the state-of-the-art method for non-sequential item recommendation on implicit feedback data.

FMC and FPMC. As introduced in (Rendle et al. 2010), FMC factorizes the first-order Markov transition matrix into two low-dimensional sub-matrices, and FPMC is a fusion of FMC and LFM. These are the state-of-the-art sequential recommendation methods. FPMC allows a basket of several items at each step. For our sequential recommendation problem, each basket has a single item.

Fossil. Fossil (He and McAuley 2016) models high-order Markov chains and uses Similarity Model instead of LFM for modeling general user preferences.

GRU4Rec. This is the session-based recommendation proposed by (Hidasi et al. 2015). This model uses RNN to capture sequential dependencies and make predictions.

For each method, the grid search is applied to find the optimal settings of hyperparameters using the validation set. These include latent dimensions dd from {5,10,20,30,50,100}\{5,10,20,30,50,100\}, regularization hyperparameters, and the learning rate from {1,10−1,...,10−4}\{1,10^{-1},...,10^{-4}\}. For Fossil, Caser and GRU4Rec, the Markov order LL is from {1,⋯ ,9}\{1,\cdots,9\}. For Caser itself, the height hh of horizontal filters is from {1,⋯ ,L}\{1,\cdots,L\}, the target number TT is from {1,2,3}\{1,2,3\}, the activation functions ϕa\phi_{a} and ϕc\phi_{c} are from \{\emph{identity},\emph{sigmoid},\emph{tanh},\emph{relu}\}. For each height hh, the number of horizontal filters is from {4,8,16,32,64}\{4,8,16,32,64\}. The number of vertical filters is from {1,2,4,8,16}\{1,2,4,8,16\}. We report the result of each method under its optimal hyperparameter settings.

The best results of the six baselines and Caser are summarized in Table 2. The best performer on each row is highlighted in bold face. The last column is the improvement of Caser relative to the best baseline, defined as Caser−baselinebaseline\frac{Caser-baseline}{baseline}. Except for MovieLens, Caser improved the best baseline on all NN tested by a large margin w.r.t. the three metrics. Among the baseline methods, the sequential recommenders (e.g., FPMC and Fossil) usually outperform non-sequential recommenders (i.e., BPR) on all data sets, suggesting the importance of considering sequential information. FPMC and Fossil outperform FMC on all data sets, suggesting the effectiveness of personalization. On MovieLens, GRU4Rec achieved a performance close to Caser’s, but got a much worse performance on the other three data sets. In fact, MovieLens has more sequential signals than the other three data sets, thus, the RNN-based GRU4Rec could perform well on MovieLens but can easily get biased on training sets of the other three data sets despite the use of regularization and dropout as described in (Hidasi et al. 2015). In addition, GRU4Rec’s recommendation is session-based, instead of personalized, which enlarge the generalization error to some extent.

In the following studies, we examine the impact of the hyperparameters d,L,Td,L,T one at a time by holding the remaining hyperparameters at their optimal settings. We focus on MAP as it is an overall performance indicator and consistent with other metrics.

Figure 5 shows MAP for various dd while keeping the other optimal hyperparameters unchanged. On the denser MovieLens, a larger dd does not always lead to a better model performance. A model achieves its best performance when dd is chosen properly and gets worse for a larger dd because of over-fitting. But for the other three sparser data sets, each model requires more latent dimensions to achieve their best results. For all data sets, Caser beats the strongest baseline performance by using a relatively small number of latent dimensions.

2.2. Influence of Markov Order LL and Target Number TT

We vary LL to explore how much of Fossil, GRU4Rec and Caser can gain from high-order information while keeping other optimal hyperparameters unchanged. Caser-1, Caser-2, and Caser-3 denote Caser with the target number TT at 1, 2, 3 to study the effect of skip behaviors. The results are shown in Figure 6. On the dense MovieLens, Caser best utilizes the extra information provided by a larger LL and Caser-3 performs the best, suggesting the benefits of skip steps. However, for the sparser data sets, all models do not consistently benefit from a larger LL. This is reasonable, because for a sparse data set, a higher order Markov chain tends to introduce both extra information and more noises. In most cases, Caser-2 slightly outperforms the other models on these three data sets.

2.3. Analysis of Caser Components

3. Network Visualization

We have a closer look at some trained networks and prediction. Figure 7 shows the values of four vertical convolutional filters after training Caser on MovieLens with L=9L=9. In the micro perspective, the four filters are trained to be diverse, but in the macro perspective, they follow an ascending trend from past positions to recent positions. With each vertical filter serving as a way of weighting the embeddings of previous actions (see the related discussion in Section 3), this trend indicates that Caser puts more emphasis on recent actions, demonstrating a major difference from the conventional top-NN recommendation.

To see the effectiveness of horizontal filters, Figure 8(a) shows top N=3N=3 ranked movies recommended by Caser, i.e., R^1\hat{R}_{1} (Mad Max), R^2\hat{R}_{2} (Star War), R^3\hat{R}_{3} (Star Trek) in that order, for a user with L=5L=5 previous movies, i.e., S1S_{1} (13th Warrior), S2S_{2} (American Beauty), S3S_{3} (Star Trek), S4S_{4} (Star Trek III), and S5S_{5} (Star Trek IV). R^3\hat{R}_{3} is the ground truth (i.e., the next movie in the user sequence). Note that R^1\hat{R}_{1} and R^2\hat{R}_{2} are quite similar to R^3\hat{R}_{3}, i.e., all being action and science fiction movies, so are also recommended to the user. Figure 8(b) shows the new rank of R^3\hat{R}_{3} after masking some of the LL previous movies by setting their item embeddings to zeros in the trained network. Masking S1S_{1} and S2S_{2} actually increases the rank of R^3\hat{R}_{3} to 2 (from 3); in fact, S1S_{1} and S2S_{2} are history or romance movies and act like noises for recommending R^3\hat{R}_{3}. Masking each of S3S_{3}, S4S_{4} and S5S_{5} decreases the rank of R^3\hat{R}_{3} because these movies are in the same category as R^3\hat{R}_{3}. The most decrease occurs after masking S3S_{3}, S4S_{4} and S5S_{5} all together. This study clearly indicates that our model correctly captures the dependence of R^3\hat{R}_{3} on the related {S3,S4,S5}\{S_{3},S_{4},S_{5}\} as a union-level sequential feature for recommending R^3\hat{R}_{3}.

Conclusion

Caser is a novel solution to top-NN sequential recommendation by modeling recent actions as an “image” among time and latent dimensions and learning sequential patterns using convolutional filters. This approach provides a unified and flexible network structure for capturing many important features of sequential recommendation, i.e., point-level and union-level sequential patterns, skip behaviors, and long term user preferences. Our experiments and case studies on public real life data sets suggested that Caser outperforms the state-of-the-art methods for top-NN sequential recommendation.

Acknowledgement

The work of the second author is partially supported by a Discovery Grant from Natural Sciences and Engineering Research Council of Canada.

References