Informative and Controllable Opinion Summarization

Reinald Kim Amplayo, Mirella Lapata

Introduction

The proliferation of opinions expressed in online reviews, blogs, and social media has created a pressing need for automated systems which enable customers and companies to make informed decisions without having to absorb large amounts of opinionated text. Opinion summarization is the task of automatically generating summaries for a set of opinions about a specific target Conrad et al. (2009). Figure 1 shows various reviews about the movie “Coach Carter” and example summaries generated by humans and automatic systems.

The vast majority of previous work Hu and Liu (2004) views opinion summarization as the final stage of a three-step process involving: (1) aspect extraction (i.e., finding features pertaining to the target of interest, such as battery life or sound quality); (2) sentiment prediction (i.e., determining the sentiment of the extracted aspects); and (3) summary generation (i.e., presenting the identified opinions to the user). Textual summaries are created following mostly extractive methods which select representative segments (usually sentences) from the source text Popescu and Etzioni (2005); Blair-Goldensohn et al. (2008); Lerman et al. (2009). Despite being less popular, abstractive approaches seem more appropriate for the task at hand as they attempt to generate summaries which are maximally informative and minimally redundant without simply rearranging passages from the original opinions Ganesan et al. (2010); Carenini et al. (2013); Gerani et al. (2014).

General-purpose summarization approaches have recently shown promising results with end-to-end models which are data-driven and take advantage of the success of sequence-to-sequence neural network architectures. Most approaches Rush et al. (2015); See et al. (2017) encode documents and then decode the learned representations into an abstractive summary, often by attending to the source input Bahdanau et al. (2014) and copying words from it Vinyals et al. (2015). Under this modeling paradigm, it is no longer necessary to identify aspects and their sentiment for the opinion summarization task, as these are learned indirectly from training data (i.e., sets of opinions and their corresponding summaries). These models are usually tested on domains where the input is either one document or a small set of documents.

However, the number of input reviews for each target entity tends to be very large (150 for the example in Figure 1). It is therefore practically unfeasible to train a model in an end-to-end fashion, given the memory limitations of modern hardware. As a result, current approaches Wang and Ling (2016); Liu et al. (2018); Liu and Lapata (2019) sacrifice end-to-end elegance in favor of a two-stage framework which we call Extract-Abstract (EA): an extractive model first selects a subset of opinions and an abstractive model then generates the summary while conditioning on the extracted subset (see Figure 2(a)). The extractive pass unfortunately has two drawbacks. Firstly, on account of having access to only a small subset of reviews, the summaries can be less informative and inaccurate, as shown in Figure 1. And secondly, user preferences cannot be easily taken into account (e.g., a user may wish to obtain a summary focusing on the acting or plot of a movie as opposed to a general-purpose summary) since more specialized information might have been removed.

In this paper, we propose Condense-Abstract (CA), an alternative two-stage framework which enables the use of all input reviews when generating the summary (see Figure 2(b)). The Condense model first represents the input reviews as encodings, aiming to condense their meaning and distill information relating to sentiment and various aspects of the target being reviewed. The Abstract model then fuses these condensed representations into one aggregate encoding and generates an opinion summary from it. We implement a simple yet effective instantiation of the CA framework, using a vanilla autoencoder as the Condense model, and a decoder with attention and copy mechanisms as the Abstract model. We also introduce a zero-shot customization technique allowing users to control important aspects of the generated summary at test time. Our approach enables controllable generation while leveraging the full spectrum of opinions available for a specific target.

We perform experiments on a dataset consisting of movie reviews and opinion summaries elicited from the Rotten Tomatoes website (Wang and Ling, 2016; see Figure 1). Our proposed approach outperforms state-of-the-art models by a large margin using automatic metrics and in a judgment elicitation study. We also verify that our zero-shot customization technique can effectively generate need-specific summaries.

Related Work

Most opinion summarization models follow extractive methods (see Kim et al., 2011 and Angelidis and Lapata, 2018 for overviews), with the exception of a few systems which are able to generate novel words and phrases not featured in the source text. Ganesan et al. (2010) propose a graph-based framework for generating concise opinion summaries, while Gerani et al. (2014) represent reviews as discourse trees which they aggregate to a global graph to generate a summary. Other work Carenini et al. (2013); Mukherjee and Joshi (2013) takes the distribution of opinions and their aspects into account so as to generate more readable summaries. Di Fabbrizio et al. (2014) present a hybrid system which uses extractive techniques to select salient quotes from the input reviews and embeds them into an abstractive summary to provide evidence for positive or negative opinions.

More recent work has seen the effective application of sequence-to-sequence models Sutskever et al. (2014); Bahdanau et al. (2014) to various abstractive summarization tasks including headline generation Rush et al. (2015), single- See et al. (2017); Nallapati et al. (2016), and multi-document summarization Wang and Ling (2016); Liu et al. (2018); Liu and Lapata (2019). Closest to our approach is the work of Wang and Ling (2016) who generate opinion summaries following a two-stage process which first selects/extracts reviews bearing pertinent information, and then generates the summary by conditioning on these reviews. More recent models Chu and Liu (2019); Bražinskas et al. (2020); Amplayo and Lapata (2020) perform opinion summarization in an unsupervised way. However, these are mostly done on toy datasets Chu and Liu (2019), typically with a small number of reviews per target entity.

Our proposed framework works better on real-world datasets with a large number of reviews, since it eliminates the need to rely only on pre-selected salient reviews which we argue leads to information loss and subsequently less customizable generation. Instead, our model first condenses the source reviews into multiple dense vectors which serve as input to a decoder to generate an abstractive summary. Beyond producing more informative summaries, we demonstrate that our approach also allows to customize them. Recent conditional generation models have focused on controlling various aspects of the output such as politeness Sennrich et al. (2016), length Kikuchi et al. (2016), content Fan et al. (2018), or style Ficler and Goldberg (2017). In contrast, our zero-shot customization technique requires neither training examples of documents and corresponding (customized) summaries nor specialized pre-processing to encode which tokens in the input might give rise to customization.

Condense-Abstract Framework

We propose an alternative to the Extract-Abstract (EA) approach which enables the use of all input reviews when generating the summary. Figure 2(b) illustrates our proposed Condense-Abstract (CA) framework. In lieu of an integrated encoder-decoder, we generate summaries using two separate models. The Condense model returns review encodings for NN input reviews, while the Abstract model uses these encodings to create an abstractive summary. This two-step approach has two advantages for multi-document summarization. Firstly, CA-based models are more space-efficient, since the set of NN reviews is not treated as one large instance but as NN separate instances when training the Condense model. And secondly, it is possible to generate maximally informative and customizable summaries targeting specific aspects of the input since the Abstract model operates over the encodings of all available reviews.

In the following subsections, we explain how we instantiate a model using the CA framework, which we call CondaSum, with an LSTM-based We use LSTMs as our text encoder instead of other popular alternatives, such as Transformers Vaswani et al. (2017), since LSTMs work better on autoencoder architectures, as shown in the literature Liu et al. (2019); Zhang et al. (2020), as well as during our preliminary experiments. vanilla autoencoder (Condense model) and a decoder with attention and copy mechanisms (Abstract model).

Let D\mathcal{D} denote a cluster of NN reviews about a specific target (e.g., a movie or product). For each review X={w1,w2,...,wM}∈DX=\{w_{1},w_{2},...,w_{M}\}\in\mathcal{D}, the Condense model learns an encoding dd, and word-level encodings h1,h2,...,hMh_{1},h_{2},...,h_{M}. We employ a Bidirectional Long Short Term Memory (BiLSTM) encoder Hochreiter and Schmidhuber (1997) as our Condense model:

where h→i\overrightarrow{h}_{i} and h←i\overleftarrow{h}_{i} are forward and backward hidden states of the BiLSTM at timestep ii, and ; denotes concatenation.

Training is performed with a reconstruction objective. We use a separate LSTM as the decoder where the first hidden state z0z_{0} is set to dd. Words wt′w^{\prime}_{t} are generated using a softmax classifier:

The auto-encoder is trained with a maximum likelihood loss:

Once training has taken place, we use the Condense model to obtain NN pairs of review encodings {di}\{d_{i}\} and word-level encodings {hi,1,hi,2,...,hi,M}\{h_{i,1},h_{i,2},...,h_{i,M}\}, 1≤i≤N1\leq i\leq N as representations for the reviews in D\mathcal{D}.

2 The Abstract Model

The Abstract model first fuses the multiple encodings obtained from the Condense stage and then generates a summary using a decoder.

We aggregate NN pairs of review encodings {di}\{d_{i}\} and word-level encodings {hi,1,hi,2,...,hi,M}\{h_{i,1},h_{i,2},...,h_{i,M}\}, 1≤i≤N1\leq i\leq N into a single pair of review encoding d′d^{\prime} and word-level encodings h1′,h2′,...,hV′h^{\prime}_{1},h^{\prime}_{2},...,h^{\prime}_{V}, where VV is the number of total unique tokens in the input.

We also fuse word-level encodings, since the same words may appear in multiple reviews. To do this, we simply average all encodings of the same word, if multiple tokens of the word exist:

where VwjV_{w_{j}} is the number of tokens for word wjw_{j} in the input.

Decoder

The decoder generates summaries conditioned on the fused review encoding d′d^{\prime} and word-level encodings h1′,h2′,...,hV′h^{\prime}_{1},h^{\prime}_{2},...,h^{\prime}_{V}. We use a simple LSTM decoder enhanced with attention Bahdanau et al. (2014) and copy mechanisms Vinyals et al. (2015). We set the first hidden state s0s_{0} to d′d^{\prime}, and run an LSTM to calculate the current hidden state using the previous hidden state st−1s_{t-1} and word yt−1′y^{\prime}_{t-1} at time step tt:

At each time step tt, we use an attention mechanism over word-level encodings to output the attention weight vector ata_{t} and context vector ctc_{t}:

Finally, we employ a copy mechanism over the input words to output the final word probability p(yt′)p(y^{\prime}_{t}) as a weighted sum over the generation probability pg(yt′)p_{g}(y^{\prime}_{t}) and the copy probability pc(yt′)p_{c}(y^{\prime}_{t}):

where WW, vv, and bb are learned parameters, and tt is the current timestep.

Salience-biased Extracts

The model presented so far has no explicit mechanism to encourage salience among reviews. We direct the decoder towards salient reviews by incorporating information from an extractive step. Specifically, we use BertCent, a centroid-based Radev et al. (2000) document extraction method that obtains document representations by resorting to BERT Devlin et al. (2019).

BertCent can be simply described as follows. Firstly, given a review, we obtain its encoding as the average of its token encodings obtained from BERT. We then take the average of the review encodings and treat it as the centroid of the input reviews, which approximately represents the information that is considered salient. We select the top kk reviews whose encodings are the nearest neigbors to the centroid. The selected reviews are concatenated into a long sequence and encoded using a separate BiLSTM whose output serves as input to an LSTM decoder. This decoder generates a salience-biased hidden state rtr_{t}. We then update hidden state sts_{t} in Equation (10) as st=[st;rt]s_{t}=[s_{t};r_{t}].

Using these extracts, we still take all input reviews into account, while acknowledging that some might be more descriptive than others. This module is a key component to generating general-purpose opinion summaries, where a set of aspects is deemed more salient than others (e.g., in general, people care more about the plot rather than the special effects of a movie). However, this extractive module may hurt the customizability of the model (e.g., generating need-specific summaries, details explained in Section 3.3), which we show in our experiments in Section 5.

Training

We use two objective functions to train the Abstract model. Firstly, we use a maximum likelihood loss to optimize the generation probability distribution p(yt′)p(y^{\prime}_{t}) based on gold summaries Y={y1,y2,...,yL}Y=\{y_{1},y_{2},...,y_{L}\} provided at training time:

Secondly, we propose a way to introduce supervision and guide the attention pooling weights WpW_{p} in Equation (7) when fusing the review encodings. Our motivation is that the resulting fused encoding d′d^{\prime} should be roughly equivalent to the encoding of summary yy, which can be calculated as z=Condense(y)z=\text{{Condense}}(y). Specifically, we use a hinge loss that maximizes the inner product between d′d^{\prime} and zz and simultaneously minimizes the inner product between d′d^{\prime} and nin_{i}, where nin_{i} is the encoding of one of five randomly sampled negative summaries:

The final objective is then the sum of both loss functions:

3 Zero-shot Customization

At test time, we can either generate a general-purpose summary or a need-specific summary. To generate the former, we run the trained model as is and use beam search to find the sequence of words with the highest cumulative probability. To generate the latter, we employ the following simple technique that revises the query vector dˉ\bar{d} in Equation (6).

More concretely, in the movie review domain, users might wish to obtain a summary that focuses on a specific sentiment (positive or negative) or aspect (e.g., acting, plot, etc.) of a movie. In a different domain, users might care about the price of a product, its comfort, and so on. Since these summaries are not available at training time, we undertake such customization without requiring access to need-specific summaries. Instead, at test time, we assume access to background reviews to represent the user need. For example, if we wish to generate a positive summary, our method requires a set of reviews with positive sentiment. This is an easy and practical way to approximately provide the model some background on how sentiment is communicated in a review.

We use these background reviews conveying a user need xx (e.g., acting, plot, positive or negative sentiment) in the multi-source fusion module to attend more to input reviews related to xx. Let CxC_{x} denote the set of background reviews. We obtain a new query vector d^=∑c=1∣Cx∣dc/∣Cx∣\hat{d}=\sum_{c=1}^{|C_{x}|}d_{c}/|C_{x}|, where dcd_{c} is the encoding of the cc’th review in CxC_{x}, calculated using the Condense model. This simple change allows the model to focus on input reviews with semantics similar to the user’s need as conveyed by the background reviews CxC_{x}. The new query vector d^\hat{d} is used instead of dˉ\bar{d} to obtain review encoding d′d^{\prime} (see Equation (6)).

Experimental Setup

We performed experiments on the Rotten Tomatoes datasethttp://www.ccs.neu.edu/home/luwang/publications.html provided in Wang and Ling (2016). It contains 3,731 movies; for each movie we are given a large set of reviews written by professional critics and users and a gold-standard consensus summary written by an editor (see an example in Figure 1). We report the dataset statistics in Table 1. Following previous work Wang and Ling (2016), we used a generic label for movie titles during training which we replace with the original titles during inference.

Training Configuration

For all experiments, our model used word embeddings with 128 dimensions, pretrained using GloVe Pennington et al. (2014). We set the dimensions of all hidden vectors to 256 and the batch size to 8. For decoding summaries, we use a length-normalized beam search with beam size of 5. We applied dropout Srivastava et al. (2014) at a rate of 0.5. The model was trained using the Adam optimizer Kingma and Ba (2015) with default parameters and l2l_{2} constraint Hinton et al. (2012) of 2. We performed early stopping based on model performance on the development set. Our model is implemented in PyTorchOur code can be downloaded from xxx.yyy.zzz..

Comparison Systems

We compare our approach against two types of methods: one-pass methods and methods that use the EA framework. One-pass methods include (a) LexRank Erkan and Radev (2004), a PageRank-like summarization algorithm which generates a summary by selecting the nn most salient units, until the length of the target summary is reached; (b) Opinosis Ganesan et al. (2010), a graph-based abstractive summarizer that generates concise summaries of highly redundant opinions; (c) SummaRunner Nallapati et al. (2017), a supervised neural extractive model where each review is classified as to whether it should be part of the summary or not; and (d) BertCent, a centroid-based method discussed in Section 3.2 that selects k=1k=1 review nearest to the centroid.

EA-based methods include (g) Regress+S2S Wang and Ling (2016), an instantiation of the EA framework where a ridge regression model with hand-engineered features implements the Extract model, while an attention-based sequence-to-sequence neural network is the Abstract model; (h) BertCent+S2S, our implementation of an EA-based system which uses BertCent instead of Regress as the Extract model; and (i) BertCent+PtGen, the same model as (h) but enhanced with a copy mechanism Vinyals et al. (2015). For all extractive steps, we set k=5k=5, which is tuned on the development set.

Results

We considered two evaluation metrics which are also reported in Wang and Ling (2016): METEOR Denkowski and Lavie (2014), a recall-oriented metric that rewards matching stems, synonyms, and paraphrases, and ROUGE-SU4 Lin (2004) which is calculated as the recall of unigrams and skip-bigrams up to four words. We also report F1-scores for ROUGE-1/2/L Lin (2004). Unigram and bigram overlap (ROUGE-1 and ROUGE-2) are a proxy for assessing informativenes while the longest common subsequence (ROUGE-L) measures fluency.

Our results are presented in Table 2. Among one-pass systems, the extractive model BertCent performs the best; despite being unsupervised and extractive, it benefits from the ability of large neural language models to learn general-purpose representations. When used in EA-based systems, BertCent also improves the system performance, where BertCent+PtGen performs the best. Interestingly, BertCent performs better than BertCent+PtGen in terms of METEOR and ROUGE-SU4, while the latter performs better in terms of ROUGE-1/2/L. Our CA-based model CondaSum outperforms all other models across all metrics, showing that exploiting information about all reviews helps in improving performance.

We present in Table 3 various ablation studies, which assess the contribution of different model components. Results confirm that our multi-source fusion method and the fusion loss improve performance. Morevoer, using BertCent for the salient-biased extractive step is better than no extractive step or using SummaRunner, which is a weaker extractive model. Both multi-source fusion and salient-biased extracts help create better general-purpose summaries; the former learns which reviews to focus on while the latter explicitly selects the most important ones.

Human Evaluation

In addition to automatic evaluation, we also assessed system output by eliciting human judgments. Participants compared summaries produced from the best extractive baseline (BertCent), the best EA system (BertCent+PtGen), and our model CondaSum, respectively. As an upper bound, we also included Gold standard summaries.

The study was conducted on the Amazon Mechanical Turk platform using Best-Worst Scaling (BWS; Louviere et al., 2015), a less labor-intensive alternative to paired comparisons that has been shown to produce more reliable results than rating scales Kiritchenko and Mohammad (2017). Specifically, participants were shown the movie title and basic background information (i.e., synopsis, release year, genre, director, and cast). They were also presented with three system summaries and asked to select the best and worst among them according to three criteria: Informativeness (i.e., does the summary convey opinions about specific aspects of the movie in a concise manner?), Correctness (i.e., is the information in the summary factually accurate and corresponding to the information given about the movie?), and Grammaticality (i.e., is the summary fluent and grammatical?). Examples of summaries are shown in Figure 1 and more can be found in the Appendix. We randomly selected 50 movies from the test set and compared all possible combinations of summary triples for each movie. We collected three judgments for each comparison. The order of summaries and movies was randomized per participant.

The scores are computed as the percentage of times it was chosen as best minus the percentage of times it was selected as worst. The scores range from -1 (worst) to 1 (best) and are shown in Table 4. Perhaps unsurprisingly, the human-generated gold summaries were considered best, whereas our model CondaSum was ranked second, indicating that humans find its output more informative, correct, and grammatical compared to other systems. BertCent was ranked third followed by BertCent+PtGen. We inspected the summaries produced by the latter system and found they were factually incorrect bearing little correspondence to the movie (examples shown in the Appendix), possibly due to the huge information loss at the extraction stage.

Customizing Summaries

We further assessed the ability of CA systems to generate customized summaries at test time. We evaluate CondaSum models with and without the salience-biased extractive step. The latter model biases summary generation towards the kk most salient extracted opinions using an additional extractive module which may discard information relevant to the user’s need. We thus expect this model to be less effective for customization than CondaSum which makes no assumptions regarding which summaries to consider.

In this experiment, we assume users may wish to control the output summaries in four ways focusing on acting- and plot-related aspects of a movie review, as well as its sentiment, which may be positive or negative. Let Cust(xx) be the zero-shot customization technique discussed in the Section 3.3, where xx is an information need (i.e., acting, plot, positive, or negative). We sampled a set of background reviews CxC_{x} (∣Cx∣|C_{x}|=1,000) from a corpus of 1 million reviews covering 7,500 movies from the Rotten Tomatoes website, made available in Ficler and Goldberg (2017). The reviews contain sentiment labels provided by their authors and heuristically classified aspect labels. We then ran Cust(xx) using both the CondaSum models. We show in Figure 3 customized summaries generated by the models.

To determine which system is better at customization, we again conducted a judgment elicitation study on Amazon Mechanical Turk. Participants read a summary which was created by a general-purpose system or its customized variant. They were then asked to decide if the summary is generic or focuses on a specific aspect (plot or acting) and expresses positive, negative, or neutral sentiment. We selected 50 movies (from the test set) which had mixed reviews and collected judgments from three different participants per summary. The summaries were presented in random order per participant.

Table 5 shows what participants thought of summaries produced by non-customized systems (see column No) and systems which had customization switched on (see column Yes). Overall, we observe that CondaSum without the extractive step is able to customize summaries to a great extent. In all cases, crowdworkers perceive a significant increase in the proportion of aspect xx when using Cust(xx). CondaSum with the extractive step is unable to generate need-specific summaries, showing no discernible difference between generic and customized summaries. This indicates that the use of an extractive module, which is one of the main components of EA-based approaches, limits the flexibility of the abstractive model to customize summaries based on a user need.

Conclusions

We introduced the Condense-Abstract (CA) framework for opinion summarization which eliminates the need to rely only on a small subset of extracted reviews and allows the use of all reviews to generate maximally informative summaries. We presented CondaSum, an instantiation of this framework and showed in both automatic and human-based evaluation that it is superior to purely extractive models and abstractive models that include an extractive pre-selection stage. We also showed that when an extractive step is not used, our zero-shot customization technique is able to generate need-specific summaries at test time. In the future, we plan to apply the CA framework to other multi-document summarization tasks.

References

Appendix A Appendices

We performed ablation studies on CondaSum to four different versions: (a) using a mean document fusion instead of our multi-source fusion module, (b) without using a fusion loss, (c) without using salience-biased extracts, and (d) using outputs from SummaRunner Nallapati et al. (2017) as salience-biased extracts. Table 6 shows the ROUGE-1/2/L F1-scores of our model and various versions thereof. The final model consistently performs better on all metrics.

A.2 Amazon Mechanical Turk Human Evaluation Experiments

We conducted three different human evaluation experiments: the best-worst scaling evaluation, the aspect-specific (acting vs. plot) customization evaluation, and the sentiment-specific (positive vs. negative) customization evaluation. To lessen the burden on the annotators and consequently gather more accurate responses, we conducted three separate Amazon Mechanical Turk (AMT) experiments. For all experiments, we ensure turkers have an approval rate of 98% (or above) with at least 1,000 tasks approved. Furthermore, turkers should be a (self-reported) native English speakers from one of the following countries: Australia, Canada, Ireland, New Zealand, United Kingdom, and United States. We discuss specific configurations for each experiment in the next paragraphs.

Each Human Intelligence Task (HIT) consists of five questions a turker must answer to receive payment. Each question includes the title of the movie, and corresponding basic background information: synopsis, release year, genre, director, and actors (see examples in Figures 4 and 5). The summaries shown are randomly shuffled and are labeled A, B, or C. Turkers are then asked which to select the best or worst summary amongst A, B, and C, according to informativeness (i.e., does the summary convey opinions about specific aspects of the movie in a concise manner?), correctness (i.e., is the information in the summary factually accurate and corresponding to the information given about the movie?), and grammaticality (i.e., is the summary fluent and grammatical?). The criteria and their definitions are shown to turkers to guide them while they select their answers.

Aspect-Specific Customization Evaluation

Similar to the best-worst scaling template, each HIT also consists of five questions a turker must answer to receive payment. Each question also includes the movie title and its basic background information. Turkers are given a summary which they have to read. They then choose the most appropriate answer from the following choices: (a) mentions neither acting nor plot, (b) mentions acting, (c) mentions plot, or (d) mentions both acting and plot.

Sentiment-Specific Customization Evaluation

For this evaluation, the template is similar with that of the aspect-specific customization evaluation, but the choices are different. The turkers instead are given three choices: (a) neutral sentiment, (b) positive sentiment, and (c) negative sentiment.

A.3 Example Summaries

Finally, we show more example system summaries generated by SummaRunner, SummaRunner+PtGen, CondaSum, and CondaSum+Salient, together with the Gold summary in Figures 4–5. Figure 4 additionally shows summaries that are customized based on the plot or acting aspect of a movie, while Figure 5 shows customized summaries according to positive or negative sentiment. The examples show similar trends with the examples in the main paper.