Unsupervised Opinion Summarization with Noising and Denoising

Reinald Kim Amplayo, Mirella Lapata

Introduction

The proliferation of massive numbers of online product, service, and merchant reviews has provided strong impetus to develop systems that perform opinion mining automatically Pang and Lee (2008). The vast majority of previous work Hu and Liu (2006) breaks down the problem of opinion aggregation and summarization into three inter-related tasks involving aspect extraction Mukherjee and Liu (2012), sentiment identification Pang et al. (2002); Pang and Lee (2004), and summary creation based on extractive Radev et al. (2000); Lu et al. (2009) or abstractive methods Ganesan et al. (2010); Carenini et al. (2013); Gerani et al. (2014); Di Fabbrizio et al. (2014). Although potentially more challenging, abstractive approaches seem more appropriate for generating informative and concise summaries, e.g., by performing various rewrite operations (e.g., deletion of words or phrases and insertion of new ones) which go beyond simply copying and rearranging passages from the original opinions.

Abstractive summarization has enjoyed renewed interest in recent years thanks to the availability of large-scale datasets Sandhaus (2008); Hermann et al. (2015); Grusky et al. (2018); Liu et al. (2018); Fabbri et al. (2019) which have driven the development of neural architectures for summarizing single and multiple documents. Several approaches See et al. (2017); Celikyilmaz et al. (2018); Paulus et al. (2018); Gehrmann et al. (2018); Liu et al. (2018); Perez-Beltrachini et al. (2019); Liu and Lapata (2019); Wang and Ling (2016) have shown promising results with sequence-to-sequence models that encode one or several source documents and then decode the learned representations into an abstractive summary.

The supervised training of high-capacity models on large datasets containing hundreds of thousands of document-summary pairs is critical to the recent success of deep learning techniques for abstractive summarization. Unfortunately, in most domains (other than news) such training data is not available and cannot be easily sourced. For instance, manually writing opinion summaries is practically impossible since an annotator must read all available reviews for a given product or service which can be prohibitively many. Moreover, different types of products impose different restrictions on the summaries which might vary in terms of length, or the types of aspects being mentioned, rendering the application of transfer learning techniques Pan and Yang (2010) problematic.

Motivated by these issues, Chu and Liu (2019) consider an unsupervised learning setting where there are only documents (product or business reviews) available without corresponding summaries. They propose an end-to-end neural model to perform abstractive summarization based on (a) an autoencoder that learns representations for each review and (b) a summarization module which takes the aggregate encoding of reviews as input and learns to generate a summary which is semantically similar to the source documents. Due to the absence of ground truth summaries, the model is not trained to reconstruct the aggregate encoding of reviews, but rather it only learns to reconstruct the encoding of individual reviews. As a result, it may not be able to generate meaningful text when the number of reviews is large. Furthermore, autoencoders are constrained to use simple decoders lacking attention Bahdanau et al. (2014) and copy Vinyals et al. (2015) mechanisms which have proven useful in the supervised setting leading to the generation of informative and detailed summaries. Problematically, a powerful decoder might be detrimental to the reconstruction objective, learning to express arbitrary distributions of the output sequence while ignoring the encoded input Kingma and Welling (2014); Bowman et al. (2016).

In this paper, we enable the use of supervised techniques for unsupervised summarization. Specifically, we automatically generate a synthetic training dataset from a corpus of product reviews, and use this dataset to train a more powerful neural model with supervised learning. The synthetic data is created by selecting a review from the corpus, pretending it is a summary, generating multiple noisy versions thereof and treating these as pseudo-reviews. The latter are obtained with two noise generation functions targeting textual units of different granularity: segment noising introduces noise at the word- and phrase-level, while document noising replaces a review with a semantically similar one. We use the synthetic data to train a neural model that learns to denoise the pseudo-reviews and generate the summary. This is motivated by how humans write opinion summaries, where denoising can be seen as removing diverging information. Our proposed model consists of a multi-source encoder and a decoder equipped with an attention mechanism. Additionally, we introduce three modules: (a) explicit denoising guides how the model removes noise from the input encodings, (b) partial copy enables to copy information from the source reviews only when necessary, and (c) a discriminator helps the decoder generate topically consistent text.

We perform experiments on two review datasets representing different domains (movies vs businesses) and summarization requirements (short vs longer summaries). Results based on automatic and human evaluation show that our method outperforms previous unsupervised summarization models, including the state-of-the-art abstractive system of Chu and Liu (2019) and is on the same par with a state-of-the-art supervised model Wang and Ling (2016) trained on a small sample of (genuine) review-summary pairs.

Related Work

Most previous work on unsupervised opinion summarization has focused on extractive approaches Carenini et al. (2006); Ku et al. (2006); Paul et al. (2010); Angelidis and Lapata (2018) where a clustering model groups opinions of the same aspect, and a sentence extraction model identifies text representative of each cluster. Ganesan et al. (2010) propose a graph-based abstractive framework for generating concise opinion summaries, while Di Fabbrizio et al. (2014) use an extractive system to first select salient sentences and then generate an abstractive summary based on hand-written templates Carenini and Moore (2006).

As mentioned earlier, we follow the setting of Chu and Liu (2019) in assuming that we have access to reviews but no gold-standard summaries. Their model learns to generate opinion summaries by reconstructing a canonical review of the average encoding of input reviews. Our proposed method is also abstractive and neural-based, but eschews the use of an autoencoder in favor of supervised sequence-to-sequence learning through the creation of a synthetic training dataset. Concurrently with our work, Bražinskas et al. (2019) use a hierarchical variational autoencoder to learn a latent code of the summary. While they also use randomly sampled reviews for supervised training, our dataset construction method is more principled making use of linguistically motivated noise functions.

Our work relates to denoising autoencoders (DAEs; Vincent et al., 2008), which have been effectively used as unsupervised methods for various NLP tasks. Earlier approaches have shown that DAEs can be used to learn high-level text representations for domain adaptation Glorot et al. (2011) and multimodal representations of textual and visual input Silberer and Lapata (2014). Recent work has applied DAEs to text generation tasks, specifically to data-to-text generation Freitag and Roy (2018) and extractive sentence compression Fevry and Phang (2018). Our model differs from these approaches in two respects. Firstly, while previous work has adopted trivial noising methods such as randomly adding or removing words Fevry and Phang (2018) and randomly corrupting encodings Silberer and Lapata (2014), our noise generators are more linguistically informed and suitable for the opinion summarization task. Secondly, while in Freitag and Roy (2018) the decoder is limited to vanilla RNNs, our noising method enables the use of more complex architectures, enhanced with attention and copy mechanisms, which are known to improve the performance of summarization systems Rush et al. (2015); See et al. (2017).

Modeling Approach

We sample a review as a candidate summary and generate noisy versions thereof, using two functions: (a) segment noising adds noise at the token and chunk level, and (b) document noising adds noise at the text level. The noise functions are illustrated in Figure 1.

Segment Noising

Segment-level noise involves token- and chunk-level alterations. Token-level alterations are performed by replacing tokens in yy with probability pRp^{\mathcal{R}}. Specifically, we replace token wjw_{j} in yy, by sampling token wj′w^{\prime}_{j} from the BiLM predicted word distribution (see in Figure 1). We use nucleus sampling Holtzman et al. (2019), which samples from a rescaled distribution of words with probability higher than a threshold pNp^{\mathcal{N}}, instead of the original distribution. This has been shown to yield better samples in comparison to top-kk sampling, mitigating the problem of text degeneration Holtzman et al. (2019).

Document Noising

Given candidate summary y={w1,...,wL}y=\{w_{1},...,w_{L}\}, we also create another set of document-level noisy versions X(d)={x1(d),...,xN(d)}\mathbf{X}^{(d)}=\{x^{(d)}_{1},...,x^{(d)}_{N}\}. Instead of manipulating parts of the summary, we altogether replace it with a similar review from the corpus and treat it as a noisy version. Specifically, we select NN reviews that are most similar to yy and discuss the same product. To measure similarity, we use IDF-weighted ROUGE-1 F1 Lin (2004), where we calculate the lexical overlap between the review and the candidate summary, weighted by token importance:

where xx is a review in the corpus, 1(⋅)1(\cdot) is an indicator function, and P, R, and F1 are the ROUGE-1 precision, recall, and F1, respectively. The reviews with the highest F1 are selected as noisy versions of yy, resulting in the noisy set X(d)\mathbf{X}^{(d)} (see Figure 1).

2 Summarization via Denoising

We summarize (aka denoise) the input X\mathbf{X} with our model which we call DenoiseSum, illustrated in Figure 2. A multi-source encoder produces an encoding for each pseudo-review. The encodings are further corrected via an explicit denoising module, and then fused into an aggregate encoding for each type of noise. Finally, the fused encodings are passed to a decoder with a partial copy mechanism to generate the summary yy.

For each pseudo-review xj∈Xx_{j}\in\mathbf{X} where xj={w1,...,wL}x_{j}=\{w_{1},...,w_{L}\} and wkw_{k} is the kkth token in xjx_{j}, we obtain contextualized token encodings {hk}\{h_{k}\} and an overall review encoding djd_{j} with a BiLSTM encoder Hochreiter and Schmidhuber (1997):

where h→k\overrightarrow{h}_{k} and h←k\overleftarrow{h}_{k} are forward and backward hidden states of the BiLSTM at timestep kk, and ; denotes concatenation (see module (a) in Figure 2).

Explicit Denoising

The model should be able to remove noise from the encodings before decoding the text. While previous methods Vincent et al. (2008); Freitag and Roy (2018) implicitly assign the denoising task to the encoder, we propose an explicit denoising component (see module (b) in Figure 2). Specifically, we create a correction vector cj(c)c^{(c)}_{j} for each pseudo-review dj(c)d^{(c)}_{j} which resulted from the application of segment noise. cj(c)c^{(c)}_{j} represents the adjustment needed to denoise each dimension of dj(c)d^{(c)}_{j} and is used to create d^j(c)\hat{d}^{(c)}_{j}, a denoised encoding of dj(c)d^{(c)}_{j}:

Noise-Specific Fusion

For each type of noise (segment and document), we create a noise-specific aggregate encoding by fusing the denoised encodings into one (see module (c) in Figure 2). Given {d^j(c)}\{\hat{d}^{(c)}_{j}\}, the set of denoised encodings corresponding to segment noisy inputs, we create aggregate encoding s0(c)s^{(c)}_{0}:

where αj\alpha_{j} is a gate vector with the same dimensionality as the denoised encodings. Analogously, we obtain s0(d)s^{(d)}_{0} from the denoised encodings {d^j(d)}\{\hat{d}^{(d)}_{j}\} corresponding to document noisy inputs.

Decoder with Partial Copy

Our decoder generates a summary given encodings s0(c)s^{(c)}_{0} and s0(d)s^{(d)}_{0} as input. An advantage of our method is its ability to incorporate techniques used in supervised models, such as attention Bahdanau et al. (2014) and copy Vinyals et al. (2015). Pseudo-reviews created using segment noising include various chunk permutations, which could result to ungrammatical and incoherent text. Using a copy mechanism on these texts may hurt the fluency of the output. We therefore allow copy on document noisy inputs only (see module (d) in Figure 2).

We use two LSTM decoders for the aggregate encodings, one equipped with attention and copy mechanisms, and one without copy mechanism. We then combine the results of these decoders using a learned gate. Specifically, token wtw_{t} at timestep tt is predicted as:

where sts_{t} and p(wt)p(w_{t}) are the hidden state and predicted token distribution at timestep tt, and σ(⋅)\sigma(\cdot) is the sigmoid function.

3 Training and Inference

We use a maximum likelihood loss to optimize the generation probability distribution based on summary y={w1,...,wL}y=\{w_{1},...,w_{L}\} from our synthetic dataset:

The decoder depends on Lgen\mathcal{L}_{gen} to generate meaningful, denoised outputs. As this is a rather indirect way to optimize our denoising module, we additionally use a discriminative loss providing direct supervision. The discriminator operates at the output of the fusion module and predicts the category distribution p(z)p(z) of the output summary yy (see module (e) in Figure 2). The type of categories varies across domains. For movies, categories can be information about their genre (e.g., drama, comedy), while for businesses their specific type (e.g., restaurant, beauty parlor). This information is often included in reviews but we assume otherwise and use an LDA topic model Blei et al. (2003) to infer p(z)p(z) (we present experiments with human labeled and automatically induced categories in Section 5). An MLP classifier takes as input aggregate encodings s(c)s^{(c)} and s(d)s^{(d)} and infers q(z)q(z). The discriminator is trained by calculating the KL divergence between predicted and actual category distributions q(z)q(z) and p(z)p(z):

The final objective is the sum of both loss functions:

At test time, we are given genuine reviews X\mathbf{X} as input instead of the synthetic ones. We generate a summary by treating X\mathbf{X} as X(c)\mathbf{X}^{(c)} and X(d)\mathbf{X}^{(d)}, i.e., the outcome of segment and document noising.

Experimental Setup

We performed experiments on two datasets which represent different domains and summary types. The Rotten Tomatoes datasethttp://www.ccs.neu.edu/home/luwang/data.html Wang and Ling (2016) contains a large set of reviews for various movies written by critics. Each set of reviews has a gold-standard consensus summary written by an editor. We follow the partition of Wang and Ling (2016) but do not use ground truth summaries during training to simulate our unsupervised setting. The Yelp datasethttps://github.com/sosuperic/MeanSum in Chu and Liu (2019) includes a large training corpus of reviews without gold-standard summaries. The latter are provided for the development and test set and were generated by an Amazon Mechanical Turker. We follow the splits introduced in their work. A comparison between the two datasets is provided in Table 1. As can be seen, Rotten Tomatoes summaries are generally short, while Yelp reviews are three times longer. Interestingly, there are a lot more reviews to summarize in Rotten Tomatoes (approximately 100 reviews) while input reviews in Yelp are considerably less (i.e., 8 reviews).

Implementation

To create the synthetic dataset, we sample candidate summaries using the following constraints: (1) the number of non-alphanumeric symbols must be less than 3, (2) there must be no first-person singular pronouns (not used for Yelp), and (3) the number of tokens must be between 20 to 30 (50 to 90 for Yelp). We set pRp^{\mathcal{R}} to 0.8 and 0.4 for token and chunk noise, and pNp^{\mathcal{N}} to 0.9. For each review-summary pair, the number of reviews NN is sampled from the Gaussian distribution N(μ,σ2)\mathcal{N}(\mu,\sigma^{2}) where μ\mu and σ\sigma are the mean and standard deviation of the number of reviews in the development set. We created 25k (Rotten Tomatoes) and 100k (Yelp) pseudo-reviews for our synthetic datasets (see Table 1).

We set the dimensions of the word embeddings to 300, the vocabulary size to 50k, the hidden dimensions to 256, the batch size to 8, and dropout Srivastava et al. (2014) to 0.1. For our discriminator, we employed an LDA topic model trained on the review corpus, with 50 (Rotten Tomatoes) and 100 (Yelp) topics (tuned on the development set). The LSTM weights were pretrained with a language modeling objective, using the corpus as training data. For Yelp, we additionally trained a coverage mechanism See et al. (2017) in a separate training phase to avoid repetition. We used the Adam optimizer Kingma and Ba (2015) with a learning rate of 0.001 and l2l_{2} constraint of 3. At test time, summaries were generated using length normalized beam search with a beam size of 5. We performed early stopping based on the performance of the model on the development set. Our model was trained on a single GeForce GTX 1080 Ti GPU and is implemented using PyTorch.Our code can be downloaded from https://github.com/rktamplayo/DenoiseSum.

Comparison Systems

We compared DenoiseSum to several unsupervised extractive and abstractive methods. Extractive approaches include (a) LexRank Erkan and Radev (2004), an algorithm similar to PageRank that generates summaries by selecting the most salient sentences, (b) Word2Vec Rossiello et al. (2017), a centroid-based method which represents the input as IDF-weighted word embeddings and selects as summary the review closest to the centroid, and (c) SentiNeuron, which is similar to Word2vec but uses a language model called Sentiment Neuron Radford et al. (2017) as input representation. As an upper bound, Oracle selects as summary the review which maximizes the ROUGE-1/2/L F1 score against the gold summary.

Abstractive methods include (d) Opinosis Ganesan et al. (2010), a graph-based summarizer that generates concise summaries of highly redundant opinions, and (e) MeanSum Chu and Liu (2019), a neural model that generates a summary by reconstructing text from aggregate encodings of reviews. Finally, for Rotten Tomatoes, we also compared with the state-of-the-art supervised model proposed in Amplayo and Lapata (2019) which used the original training split. Examples of system summaries are shown in the Appendix.

Results

Our results on Rotten Tomatoes are shown in Table 3. Following previous work Wang and Ling (2016); Amplayo and Lapata (2019) we report five metrics: METEOR Denkowski and Lavie (2014), a recall-oriented metric that rewards matching stems, synonyms, and paraphrases; ROUGE-SU4 Lin (2004), the recall of unigrams and skip-bigrams of up to four words; and the F1-score of ROUGE-1/2/L, which respectively measures word-overlap, bigram-overlap, and the longest common subsequence between system and reference summaries. Results on Yelp are given in Table 3 where we compare systems using ROUGE-1/2/L F1, following Chu and Liu (2019).

As can be seen, DenoiseSum outperforms all competing models on both datasets. When compared to MeanSum, the difference in performance is especially large on Rotten Tomatoes, where we see a 4.01 improvement in ROUGE-L. We believe this is because MeanSum does not learn to reconstruct encodings of aggregated inputs, and as a result it is unable to produce meaningful summaries when the number of input reviews is large, as is the case for Rotten Tomatoes. In fact, the best extractive model, SentiNeuron, slightly outperforms MeanSum on this dataset across metrics with the exception of ROUGE-L. When compared to the best supervised system, DenoiseSum performs comparably on several metrics, specifically METEOR and ROUGE-1, however there is still a gap on ROUGE-2, showing the limitations of systems trained without gold-standard summaries.

Table 4 presents various ablation studies on Rotten Tomatoes (RT) and Yelp which assess the contribution of different model components. Our experiments confirm that increasing the size of the synthetic data improves performance, and that both segment and document noising are useful. We also show that explicit denoising, partial copy, and the discriminator help achieve best results. Finally, human-labeled categories (instead of LDA topics) decrease model performance, which suggests that more useful labels can be approximated by automatic means.

Human Evaluation

We also conducted two judgment elicitation studies using the Amazon Mechanical Turk (AMT) crowdsourcing platform. The first study assessed the quality of the summaries using Best-Worst Scaling (BWS; Louviere et al., 2015), a less labor-intensive alternative to paired comparisons that has been shown to produce more reliable results than rating scales Kiritchenko and Mohammad (2017). Specifically, participants were shown the movie/business name, some basic background information, and a gold-standard summary. They were also presented with three system summaries, produced by SentiNeuron (best extractive model), MeanSum (most related unsupervised model), and DenoiseSum.

Participants were asked to select the best and worst among system summaries taking into account how much they deviated from the ground truth summary in terms of: Informativeness (i.e., does the summary present opinions about specific aspects of the movie/business in a concise manner?), Coherence (i.e., is the summary easy to read and does it follow a natural ordering of facts?), and Grammaticality (i.e., is the summary fluent and grammatical?). We randomly selected 50 instances from the test set. We collected five judgments for each comparison. The order of summaries was randomized per participant. A rating per system was computed as the percentage of times it was chosen as best minus the percentage of times it was selected as worst. Results are reported in Table 5, where Inf, Coh, and Gram are shorthands for Informativeness, Coherence, and Grammaticality. DenoiseSum was ranked best in terms of informativeness and coherence, while the extractive system SentiNeuron was ranked best on grammaticality. This is not entirely surprising since extractive summaries written by humans are by definition grammatical.

Our second study examined the veridicality of the generated summaries, namely whether the facts mentioned in them are indeed discussed in the input reviews. Participants were shown reviews and the corresponding summary and were asked to verify for each summary sentence whether it was fully supported by the reviews, partially supported, or not at all supported. We performed this experiment on Yelp only since the number of reviews is small and participants could read them all in a timely fashion. We used the same 50 instances as in our first study and collected five judgments per instance. Participants assessed the summaries produced by MeanSum and DenoiseSum. We also included Gold-standard summaries as an upper bound but no output from an extractive system as it by default contains facts mentioned in the reviews.

Table 5 reports the percentage of fully (FullSupp), partially (PartSupp), and un-supported (NoSupp) sentences. Gold summaries display the highest percentage of fully supported sentences (63.3%), followed by DenoiseSum (55.1%), and MeanSum (41.7%). These results are encouraging, indicating that our model hallucinates to a lesser extent compared to MeanSum.

Conclusions

We consider an unsupervised learning setting for opinion summarization where there are only reviews available without corresponding summaries. Our key insight is to enable the use of supervised techniques by creating synthetic review-summary pairs using noise generation methods. Our summarization model, DenoiseSum, introduces explicit denoising, partial copy, and discrimination modules which improve overall summary quality, outperforming competitive systems by a wide margin. In the future, we would like to model aspects and sentiment more explicitly as well as apply some of the techniques presented here to unsupervised single-document summarization.

Acknowledgments

We thank the anonymous reviewers for their feedback. We gratefully acknowledge the support of the European Research Council (Lapata, award number 681760). The first author is supported by a Google PhD Fellowship.

References

Appendix A Appendix

Algorithm 1 shows how segment noising (i.e., token- and chunk-level alterations) is applied step-by-step (see Section 3.1). Segment noising assumes we have access to a language model LMLM that is able to output token-level predictions given neighboring tokens, and a syntactic chunker SCSC that is able to return shallow parses (i.e., chunks with corresponding syntactic labels). Function TokenAlter takes candidate summary yy as input and generates token-level alterations. The noisy summary is then passed on to function ChunkAlter to create chunk-level alterations. Sequentially executing both functions produces noisy version X(c)\mathbf{X}^{(c)}.

A.2 Ablation Studies

We performed ablation studies on DenoiseSum by comparing it to versions (a) using less synthetic training data, (b) using one kind of noising function, and (c) with missing one module (explicit denoising, partial copy, or discriminator). We also compared our model with a version that uses human-labeled categories, instead of induced topic distribution, as the ground-truth category distribution p(z)p(z). For Rotten Tomatoes, we used movie genres (e.g., comedy, drama) as categories, while for Yelp, we use business types (e.g., restaurant, beauty parlor) as categories. In total, there are 21 movie genres and 898 business types.The original versions of the datasets from Wang and Ling (2016) and Chu and Liu (2019) do not contain this information. We share our version of these datasets here: https://github.com/rktamplayo/DenoiseSum.

Table 6 shows the ROUGE-1/2/L F1-scores of our model and various versions thereof. The final model consistently performs better on all metrics when compared to versions with less synthetic data (second block), versions with only one type of noise (second block), and versions with a module removed (third block). When using human-labeled categories, we see a slight improvement in ROUGE-2 on Rotten Tomatoes, however the model performs substantially worse on other metrics. We believe there are several reasons for this. Firstly, human-labeled categories, at least the ones available, are not fine-grained enough to capture various aspects mentioned in the reviews and their sentiment (e.g., did the actors perform well? was the plot convoluted?). Secondly, the number of business types available on Yelp is very large (i.e., 898 types), which makes the discriminator loss Ldisc\mathcal{L}_{disc} hard to optimize. This explains the relatively larger decrease in performance on Yelp.

A.3 Example Noisy Versions

Figure 3 shows example noisy versions of a candidate summary using both segment and document noising methods. Although segment noising yields texts which may not be entirely comprehensible to humans, a few segments contain understandable content that could be perceived as diverging information and as such should not be included in the summary (e.g., “some can’t laugh hard” in S.1 of Rotten Tomatoes). Similarly, noisy versions generated by document noising include content that is somewhat related but not critical for generating the summary (e.g., “remains just as relevant now as it did in 1989” in D.3 of Rotten Tomatoes). These examples show that our noise functions are not entirely random, in contrast to previous trivial and un-informed approaches Fevry and Phang (2018); Silberer and Lapata (2014).

A.4 Human Evaluation on Amazon Mechanical Turk

We conducted three different experiments using the Amazon Mechanical Turk platform: the best-worst scaling evaluations for Rotten Tomatoes and Yelp, and the summary veridicality experiment on the Yelp dataset. For all experiments, we made sure crowdworkers had an approval rate of 98% (or above) with at least 1000 tasks approved. Furthermore, turkers were (self reported) native English speakers from one of the following countries: Australia, Canada, Ireland, New Zealand, United Kingdom, and United States. We discuss further specifications for each experiment in the next paragraphs.

For the best-worst scaling experiments, we created different templates for each dataset. For Rotten Tomatoes, each Human Intelligence Task (HIT) included the title of the movie and basic background information: synopsis, release date, genre, director, and actors (see the examples in Figure 4). We also included the human-written gold-standard summary (highlighted in blue), emphasizing that the AMT workers must use it as a reference. System summaries were randomly shuffled and labeled A, B, or C. Turkers were then asked which to select the best or worst summary, according to informativeness (i.e., does the summary present opinions about specific aspects of the movie in a concise manner?), coherence (i.e., is the summary easy to read and does it follow a natural ordering of facts?), and grammaticality (i.e., is the summary fluent and grammatical?). The criteria and their definitions were shown to crowdworkers. In total, 92 turkers participated in the study annotating a total of 200 HITs.

For Yelp, each HIT also showed the name of the business and basic information such as location and the type of service provided (see the Figure 5. Again, we showed the gold-standard summary highlighted in blue, and randomly shuffled system summaries (A, B, and C). Crowdworkers selected the best/worst summary according to informativeness, coherence, and grammaticality. In total, 94 turkers participated in the study annotating a total of 200 HITs.

Summary Veridicality

In this experiment we only used the Yelp dataset since the number of reviews is small and participants could read them in a timely manner. Each HIT presented the business name, location, and type of service together with eight reviews and a summary produced by one of the following systems: MeanSum, DenoiseSum, and Gold-standard summaries (see the example in Figure 5). Turkers were asked to verify whether the facts mentioned in the summary were true (summaries were at most ten sentences long). Specifically they had to decide for each summary sentence whether it was fully supported, partially supported, or not supported by the reviews. In total, 79 turkers participated in this study annotating a total of 600 HITs.

A.5 Example Summaries

We show example summaries produced by three systems: SentiNeuron, MeanSum, and our model DenoiseSum, as well as the Gold-standard summary in Figure 4 (for Rotten Tomatoes) and Figure 5 (for Yelp). The extractive model SentiNeuron tends to select reviews that are longer and more verbose. Summaries generated by MeanSum on Rotten Tomatoes are mostly gibberish, which we argue is due to the model being unable to handle the large number of input reviews in this dataset. Overall, DenoiseSum produces the best summaries among the three systems.