ASAP: A Chinese Review Dataset Towards Aspect Category Sentiment Analysis and Rating Prediction
Jiahao Bu, Lei Ren, Shuang Zheng, Yang Yang, Jingang Wang, Fuzheng Zhang, Wei Wu
Introduction
With the rapid development of e-commerce, massive user reviews available on e-commerce platforms are becoming valuable resources for both customers and merchants. Aspect-based sentiment analysis(ABSA) on user reviews is a fundamental and challenging task which attracts interests from both academia and industries (Hu and Liu, 2004; Ganu et al., 2009; Jo and Oh, 2011; Kiritchenko et al., 2014). According to whether the aspect terms are explicitly mentioned in texts, ABSA can be further classified into aspect term sentiment analysis (ATSA) and aspect category sentiment analysis (ACSA), we focus on the latter which is more widely used in industries. Specifically, given a review ”Although the fish is delicious, the waiter is horrible!”, the ACSA task aims to infer the sentiment polarity over aspect category food is positive while the opinion over the aspect category service is negative.
The user interfaces of e-commerce platforms are more intelligent than ever before with the help of ACSA techniques. For example, Figure 1 presents the detail page of a coffee shop on a popular e-commerce platform in China. The upper aspect-based sentiment text-boxes display the aspect categories (e.g., food, sanitation) mentioned frequently in user reviews and the aggregated sentiment polarities on these aspect categories (the orange ones represent positive and the blue ones represent negative). Customers can focus on corresponding reviews effectively by clicking the aspect-based sentiment text-boxes they care about (e.g., the orange filled text-box “卫生条件好” (good sanitation)). Our user survey based on valid questionnaires demonstrates that customers agree that the aspect-based sentiment text-boxes are helpful to their decision-making on restaurant choices. Besides, the merchants can keep track of their cuisines and service qualities with the help of the aspect-based sentiment text-boxes. Most Chinese e-commerce platforms such as Taobaohttps://www.taobao.com/, Dianpinghttps://www.dianping.com/, and Koubeihttps://www.koubei.com/ deploy the similar user interfaces to improve user experience.
Users also publish their overall -star scale ratings together with reviews. Figure 1 displays a sample of -star rating to the coffee shop. In comparison to fine-grained aspect sentiment, the overall review rating is usually a coarse-grained synthesis of the opinions on multiple aspects. Rating prediction(RP) (Jin et al., 2016; Li et al., 2018; Wu et al., 2019a) which aims to predict the “seeing stars” of reviews also has wide applications. For example, to promise the aspect-based sentiment text-boxes accurate, unreliable reviews should be removed before ACSA algorithms are performed. Given a piece of user review, we can predict a rating for it based on the overall sentiment polarity underlying the text. We assume the predicted rating of the review should be consistent with its ground-truth rating as long as the review is reliable. If the predicted rating and the user rating of a review disagree with each other explicitly, the reliability of the review is doubtful. Figure 2 demonstrates an example review of low-reliability. In summary, RP can help merchants to detect unreliable reviews.
Therefore, both ACSA and RP are of great importance for business intelligence in e-commerce, and they are highly correlated and complementary. ACSA focuses on predicting its underlying sentiment polarities on different aspect categories, while RP focuses on predicting the user’s overall feelings from the review content. We reckon these two tasks are highly correlated and better performance could be achieved by considering them jointly.
As far as we know, current public datasets are constructed for ACSA and RP separately, which limits further joint explorations of ACSA and RP. To address the problem and advance the related researches, this paper presents a large-scale Chinese restaurant review dataset for Aspect category Sentiment Analysis and rating Prediction, denotes as ASAP for short. All the reviews in ASAP are collected from the aforementioned e-commerce platform. There are restaurant reviews attached with -star scale ratings. Each review is manually annotated according to its sentiment polarities towards fine-grained aspect categories. To the best of our knowledge, ASAP is the largest Chinese large-scale review dataset towards both ACSA and RP tasks.
We implement several state-of-the-art (SOTA) baselines for ACSA and RP and evaluate their performance on ASAP. To make a fair comparison, we also perform ACSA experiments on a widely used SemEval-2014 restaurant review dataset (Pontiki et al., 2014). Since BERT (Devlin et al., 2018) has achieved great success in several natural language understanding tasks including sentiment analysis (Xu et al., 2019; Sun et al., 2019; Jiang et al., 2019), we propose a joint model that employs the fine-to-coarse semantic capability of BERT. Our joint model outperforms the competing baselines on both tasks.
Our main contributions can be summarized as follows. (1) We present a large-scale Chinese review dataset towards aspect category sentiment analysis and rating prediction, named as ASAP, including as many as real-world restaurant reviews annotated from pre-defined aspect categories. Our dataset has been released at https://github.com/Meituan-Dianping/asap. (2) We explore the performance of widely used models for ACSA and RP on ASAP. (3) We propose a joint learning model for ACSA and RP tasks. Our model achieves the best results both on ASAP and SemEval Restaurant datasets.
Related Work and Datasets
Aspect Category Sentiment Analysis. ACSA (Zhou et al., 2015; Movahedi et al., 2019; Ruder et al., 2016; Hu et al., 2018) aims to predict sentiment polarities on all aspect categories mentioned in the text. The series of SemEval datasets consisting of user reviews from e-commerce websites have been widely used and pushed forward related research (Wang et al., 2016; Ma et al., 2017; Xu et al., 2019; Sun et al., 2019; Jiang et al., 2019). The SemEval-2014 task-4 dataset (SE-ABSA14) (Pontiki et al., 2014) is composed of laptop and restaurant reviews. The restaurant subset includes aspect categories (i.e., Food, Service, Price, Ambience and Anecdotes/Miscellaneous) and polarity labels (i.e., Positive, Negative, Conflict and Neutral). The laptop subset is not suitable for ACSA. The SemEval-2015 task-12 dataset (SE-ABSA15) (Pontiki et al., 2015) builds upon SE-ABSA14 and defines its aspect category as a combination of an entity type and an attribute type(e.g., Food#Style_Options). The SemEval-2016 task-5 dataset (SE-ABSA16) (Pontiki et al., 2016) extends SE-ABSA15 to new domains and new languages other than English. MAMS (Jiang et al., 2019) tailors SE-ABSA14 to make it more challenging, in which each sentence contains at least two aspects with different sentiment polarities.
Compared with the prosperity of English resources, high-quality Chinese datasets are not rich enough. “ChnSentiCorp” (Tan and Zhang, 2008), “IT168TEST” (Zagibalov and Carroll, 2008), “Weibo”http://tcci.ccf.org.cn/conference/2014/pages/page04_dg.html, “CTB” (Li et al., 2014) are popular Chinese datasets for general sentiment analysis. However, aspect category information is not annotated in these datasets. Zhao et al. (2014) presents two Chinese ABSA datasets for consumer electronics (mobile phones and cameras). Nevertheless, the two datasets only contain documents ( sentences), in which each sentence only mentions one aspect category at most. BDCIhttps://www.datafountain.cn/competitions/310 automobile opinion mining and sentiment analysis dataset (Dai et al., 2019) contains user reviews in automobile industry with pre-defined categories. Peng et al. (2017) summarizes available Chinese ABSA datasets. While most of them are constructed through rule-based or machine learning-based approaches, which inevitably introduce additional noise into the datasets. Our ASAP excels above Chinese datasets both on quantity and quality.
Rating Prediction. Rating prediction (RP) aims to predict the “seeing stars” of reviews, which represent the overall ratings of reviews. In comparison to fine-grained aspect sentiment, the overall review rating is usually a coarse-grained synthesis of the opinions on multiple aspects. Ganu et al. (2009); Li et al. (2011); Chen et al. (2018) form this task as a text classification or regression problem. Considering the importance of opinions on multiple aspects in reviews, recent years have seen numerous work (Jin et al., 2016; Cheng et al., 2018; Li et al., 2018; Wu et al., 2019a) utilizing the information of the aspects to improve the rating prediction performance. This trending also inspires the motivation of ASAP.
Most RP datasets are crawled from real-world review websites and created for RP specifically. Amazon Product Review English dataset (McAuley and Leskovec, 2013) containing product reviews and metadata from Amazon has been widely used for RP (Cheng et al., 2018; McAuley and Leskovec, 2013). Another popular English dataset comes from Yelp Dataset Challenge 2017http://www.yelp.com/dataset_challenge/, which includes reviews of local businesses in metropolitan areas across countries. Openricehttps://www.openrice.com is a Chinese RP dataset composed of reviews. Both the English and Chinese datasets don’t annotate fine-grained aspect category sentiment polarities.
Dataset Collection and Analysis
We collect reviews from one of the most popular O2O e-commerce platforms in China, which allows users to publish coarse-grained star ratings and writing fine-grained reviews to restaurants (or places of interest) they have visited. In the reviews, users comment on multiple aspects either explicitly or implicitly, including ambience,price, food, service, and so on.
First, we retrieve a large volume of user reviews from popular restaurants holding more than user reviews randomly. Then, pre-processing steps are performed to promise the ethics, quality, and reliability of the reviews. (1) User information (e.g., user-ids, usernames, avatars, and post-times) are removed due to privacy considerations. (2) Short reviews with less than Chinese characters, as well as lengthy reviews with more than Chinese characters are filtered out. (3) If the ratio of non-Chinese characters within a review is over %, the review is discarded. (4) To detect the low-quality reviews (e.g., advertising texts), we build a BERT-based classifier with an accuracy of in a leave-out test-set. The reviews detected as low-quality by the classifier are discarded too.
2 Aspect Categories
Since the reviews already hold users’ star ratings, this section mainly introduces our annotation details for ACSA. In SE-ABSA14 restaurant dataset (denoted as Restaurant for simplicity), there are coarse-grained aspect categories, including food, service, price, ambience and miscellaneous. After an in-depth analysis of the collected reviews, we find the aspect categories mentioned by users are rather diverse and fine-grained. Take the text “…The restaurant holds a high-end decoration but is quite noisy since a wedding ceremony was being held in the main hall… (…环境看起来很高大上的样子,但是因为主厅在举办婚礼非常混乱,感觉特别吵…)” in Table 3 for example, the reviewer actually expresses opposite sentiment polarities on two fine-grained aspect categories related to ambience. The restaurant’s decoration is very high-end (Positive), while it’s very noisy due to an ongoing ceremony (Negative). Therefore, we summarize the frequently mentioned aspects and refine the coarse-grained categories into fine-grained categories. We replace miscellaneous with location since we find users usually review the restaurants’ location (e.g., whether the restaurant is easy to reach by public transportation.). We denote the aspect category as the form of “Coarse-grained Category#Fine-grained Categoty”, such as “Food#Taste” and “Ambience#Decoration”. The full list of aspect categories and definitions are listed in Table 1.
3 Annotation Guidelines & Process
Bearing in mind the pre-defined aspects, assessors are asked to annotate sentiment polarities towards the mentioned aspect categories of each review. Given a review, when an aspect category is mentioned within the review either explicitly and implicitly, the sentiment polarity over the aspect category is labeled as (Positive), (Neutral) or (Negative) as shown in Table 3.
We hire vendor assessors, project managers, and expert reviewer to perform annotations. Each assessor needs to attend a training to ensure their intact understanding of the annotation guidelines. Three rounds of annotation are conducted sequentially. First, we randomly split the whole dataset into groups, and every group is assigned to assessors to annotate independently. Second, each group is split into subsets according to the annotation results, denoted as Sub-Agree and Sub-Disagree. Sub-Agree comprises the data examples with agreement annotation, and Sub-Disagree comprises the data examples with disagreement annotation. Sub-Agree will be reviewed by assessors from other groups. The controversial examples during the review are considered as difficult cases. Sub-Disagree will be reviewed by the project managers independently and then discuss to reach an agreement annotation. The examples that could not be addressed after discussions are also considered as difficult cases. Third, for each group, the difficult examples from two subsets are delivered to the expert reviewer to make a final decision. More details of difficult cases and annotation guidelines during annotation are demonstrated in Table 2.
Finally, ASAP corpus consists of pieces of real-world user reviews, and we split it into a training set (), a validation set () and a test set () randomly. Table 3 presents an example review of ASAP and corresponding annotations on the aspect categories.
4 Dataset Analysis
Figure 3 presents the distribution of aspect categories in ASAP. Because ASAP concentrates on the domain of restaurant, % reviews mention Food#Taste as expected. Users also pay great attention to aspect categories such as Service#Hospitality, Price#Level and Ambience#Decoration. The distribution proves the advantages of ASAP, as users’ fine-grained preferences could reflect the pros and cons of restaurants more precisely.
The statistics of ASAP are presented in Table 4. We also include a tailored SE-ABSA14 Restaurant dataset for reference. Please note that we remove the reviews holding aspect categories with sentiment polarity of “conflict” from the original Restaurant dataset.
Compared with Restaurant, ASAP excels in the quantities of training instances, which supports the exploration of recent data-intensive deep neural models. ASAP is a review-level dataset, while Restaurant is a sentence-level dataset. The average length of reviews in ASAP is much longer, thus the reviews tend to contain richer aspect information. In ASAP, the reviews contain aspect categories in average, which is times of Restaurant. Both review-level ACSA and RP are more challenging than their sentence-level counterparts. Take the review in Table 3 for example, the review contains several sentiment polarities towards multiple aspect categories. In addition to aspect category sentiment annotations, ASAP also includes overall user ratings for reviews. With the help of ASAP, ACSA and RP can be further optimized either separately or jointly.
Methodology
We use to denote the collection of user review corpus in the training data. Given a review which consists of a series of words: , ACSA aims to predict the sentiment polarity of review with respect to the mentioned aspect category , . denotes the length of review . is the number of pre-defined aspect categories (i.e., in this paper). Suppose there are mentioned aspect categories in . We define a mask vector to indicate the occurrence of aspect categories. When the aspect category is mentioned in , , otherwise . So we have . In terms of RP, it aims to predict the -star rating score of , which represents the overall rating of the given review .
2 Joint Model
Given a user review, ACSA focuses on predicting its underlying sentiment polarities on different aspect categories, while RP focuses on predicting the user’s overall feelings from the review content. We reckon these two tasks are highly correlated and better performance could be achieved by considering them jointly.
The advent of BERT has established the success of the “pre-training and then fine-tuning” paradigm for NLP tasks. BERT-based models have achieved impressive results in ACSA (Xu et al., 2019; Sun et al., 2019; Jiang et al., 2019). Review rating prediction can be deemed as a single-sentence classification (regression) task, which could also be addressed with BERT. Therefore, we propose a joint learning model to address ACSA and RP in a multi-task learning manner. Our joint model employs the fine-to-coarse semantic representation capability of the BERT encoder. Figure 4 illustrates the framework of our joint model.
If the aspect category is not mentioned in , is set as a random value. The serves as a gate function, which filters out the random and ensures only the mentioned aspect categories can participate in the calculation of the loss function.
Hence the RP loss for a given review is defined as follows,
The final loss of our joint model becomes as follows.
Experiments
We perform an extensive set of experiments to evaluate the performance of our joint model on ASAP and Restaurant (Pontiki et al., 2014). Ablation studies are also conducted to probe the interactive influence between ACSA and RP.
Baseline Models We implement several ACSA baselines for comparison. According to the different structures of their encoders, these models are classified into Non-BERT based models or BERT-based models. Non-BERT based models include TextCNN (Kim, 2014), BiLSTM+Attn (Zhou et al., 2016), ATAE-LSTM (Wang et al., 2016) and CapsNet (Sabour et al., 2017). BERT-based models include vanilla BERT (Devlin et al., 2018), QA-BERT (Sun et al., 2019) and CapsNet-BERT (Jiang et al., 2019).
Implementation Details of Experimental Models In terms of non-BERT-based models, we initialize their inputs with pre-trained embeddings. For Chinese ASAP, we utilize Jiebahttps://github.com/fxsjy/jieba to segment Chinese texts and adopt Tencent Chinese word embeddings Song et al. (2018) composed of words. For English Restaurant, we adopt -dimensional word embeddings pre-trained by Glove Pennington et al. (2014).
In terms of BERT-based models, we adopt the -layer Google BERT Basehttps://github.com/google-research/bert to encode the inputs.
The batch sizes are set as and for non-BERT-based models and BERT-based models respectively. Adam optimizer Kingma and Ba (2014) is employed with and . The maximum sequence length is set as . The number of epochs is set as . The learning rates are set as and for non-BERT-based models and BERT-based models respectively. All the models are trained on a single NVIDIA Tesla G V Volta GPU.
Evaluation Metrics Following the settings of Restaurant, we adopt Macro-F1 and Accuracy (Acc) as evaluation metrics.
Experimental Results & Analysis We report the performance of aforementioned models on ASAP and Restaurant in Table 5. Generally, BERT-based models outperform Non-BERT based models on both datasets. The two variants of our joint model perform better than vanilla-BERT, QA-BERT, and CapsNet-BERT, which proves the advantages of our joint learning model. Given a user review, vanilla-BERT, QA-BERT, and CapsNet-BERT treat the pre-defined aspect categories independently, while our joint model combines them together with a multi-task learning framework. On one hand, the encoder-sharing setting enables knowledge transferring among different aspect categories. On the other hand, our joint model is more efficient than other competitors, especially when the number of aspect categories is large. The ablation of RP (i.e., joint model(w/o RP)) still outperforms all other baselines. The introduction of RP to ACSA brings marginal improvement. This is reasonable considering that the essential objective of RP is to estimate the overall sentiment polarity instead of fine-grained sentiment polarities.
We visualize the attention weights produced by our joint model on the example of Table 3 in Figure 5. Since different aspect category information is dispersed across the review of , we add an attention-pooling layer Wang et al. (2016) to aggregate the related token embeddings dynamically for every aspect category. The attention-pooling layer helps the model focus on the tokens most related to the target aspect categories. Figure 5 visualizes attention weights of given aspect categories. The intensity of the color represents the magnitude of attention weight, which means the relatedness of tokens to the given aspect category. It’s obvious that our joint model focus on the tokens most related to the aspect categories across the review of .
2 Rating Prediction
We compare several RP models on ASAP, including TextCNN (Kim, 2014), BiLSTM+Attn (Zhou et al., 2016) and ARP (Wu et al., 2019b). The data pre-processing and implementation details are identical with ACSA experiments.
Evaluation Metrics. We adopt Mean Absolute Error (MAE) and Accuracy (by mapping the predicted rating score to the nearest category) as evaluation metrics.
Experimental Results & Analysis The experimental results of comparative RP models are illustrated in Table 6.
Our joint model which combines ACSA and RP outperforms other models considerably. On one hand, the performance improvement is expected since our joint model is built upon BERT. On the other hand, the ablation of ACSA (i.e., joint model(w/o ACSA)) brings performance degradation of RP on both metrics. We can conclude that the fine-grained aspect category sentiment prediction of the review indeed helps the model predict its overall rating more accurately.
This section conducts preliminary experiments to evaluate classical ACSA and RP models on our proposed ASAP dataset. We believe there still exists much room for improvements to both tasks, and we will leave them for future work.
Conclusion
This paper presents ASAP, a large-scale Chinese restaurant review dataset towards aspect category sentiment analysis (ACSA) and rating prediction (RP). ASAP consists of restaurant user reviews with star ratings from a leading e-commerce platform in China. Each review is manually annotated according to its sentiment polarities on fine-grained aspect categories. Besides evaluations of ACSA and RP models on ASAP separately, we also propose a joint model to address ACSA and RP synthetically, which outperforms other state-of-the-art baselines considerably. we hope the release of ASAP could push forward related researches and applications.