Self-Supervised and Controlled Multi-Document Opinion Summarization
Hady Elsahar, Maximin Coavoux, Matthias Gallé, Jos Rozen
Introduction
Recent progress in unsupervised methods has created breakthroughs in natural language processing applications, such as machine translation (Artetxe et al., 2018; Lample et al., 2018). Those have been mostly based on a bootstrapping approach, which consists in iteratively alternating between two representations, and optimizing a reconstruction loss. Machine translation is the most successful of those applications, but other applications include Question-Answering (Lewis et al., 2019) and parsing (Drozdov et al., 2019). While similar ideas have been applied as well for video summarization (Yuan et al., 2019), such a bootstrapping approach seems less suited for summarization, because of the inherent information loss when going from the full text to the summarized one. Existing unsupervised approaches for summarization therefore relied mostly on extractive graph-based systems (Mihalcea & Tarau, 2004). Only recently have there been proposals for unsupervised abstractive summarization, using auto-encoders (Chu & Liu, 2019; Bražinskas et al., 2019). However, these set-ups are quite complex, requiring a combination of loss functions (Chu & Liu, 2019) or hierarchical latent variables (Bražinskas et al., 2019) to ensure that the generated summaries remain on-topic.
In this paper, we investigate a self-supervised approach for multi-document opinion summarization. In this setting, there are multiple opinions (reviews), one entity (products, venues, movies, etc) and the goal is to extract a short summary of those opinions. Our approach is based on self-supervision and does not require any gold summaries. We train a supervised model on examples artificially created by selecting (i) one review that will act as a target summary and (ii) a subset of reviews of the same entity that acts as a document collection.
Neural models have a known problem of hallucination (Rohrbach et al., 2018), which can be utmost misleading in natural language generation tasks as the fluency of those models often distract from the wrong facts stated in the generated text. To reduce this effect, we propose to use control tokens (Fan et al., 2017; Keskar et al., 2019). Control tokens are discrete variables that are used to condition the generation. Different from previous work, our goal is not to allow users to control the generated text, but instead to steer the generated text to produce an output which is consistent with the input documents to be summarized.
Our main contributions are therefore three-fold:
performing multi-document summarization by modelling it as a self-supervised problem where one document acts as the summary of a subset. We carefully select those two, and link the resulting formulation to a recently proposed theoretical framework (Peyrard, 2019) (Sect. 3);
using control tokens to steer the model towards consistency, increasing relevance of the generated summary (Sect. 4);
an application of multi-input transformer model (Libovický et al., 2018) to summarization. This model encodes each input independently, and at decoding time applies parallel attention to each encoded input (Sect. 5).
Our experimental results (Sect. 6 and 7) show that our proposed approach outperforms existing models on two datasets: Yelp reviews on venues (Chu & Liu, 2019) and Rotten Tomatoes movie reviews (Wang & Ling, 2016). We focus the human evaluation on the faithfulness of the generated reviews and they confirm that the generated summaries are more factually correct than the compared baseline.
Related Work
Unsupervised Multi-Document summarization methods encompass both extractive and abstractive approaches. Extractive summarization consists in selecting a few sentences from the input documents to form the output summary. Radev et al. (2004) proposed to rank sentences according to their relevance to the whole input, representing sentences as tfidf bags of words and the input as the centroid vector of its sentences. Recent refinements of this approach include using distributed word representations (Rossiello et al., 2017) or ranking whole summaries instead of individual sentences (Gholipour Ghalandari, 2017). Graph-based methods, such as LexRank (Erkan & Radev, 2004) or TextRank (Mihalcea & Tarau, 2004; Zheng & Lapata, 2019), work by constructing a graph whose nodes are the sentences from the input documents and whose edges indicate a high word overlap between two sentences. Then, they use the PageRank algorithm to extract the sentences with the highest centrality. In contrast to these methods, we focus on abstractive summarization methods.
Abstractive methods for summarization are in principle able to generate new words and sentences that do not occur in the input documents and therefore produce more fluent text. Non-neural abstractive methods (Ganesan et al., 2010; Nayeem et al., 2018) are also graph-based, but construct graphs whose nodes are word types and edges indicate the immediate precedence relationship between two instantiations of the word type in a sentence. The summary is extracted by finding salient paths in the graph.
Recently, a few approaches for neural unsupervised abstractive summarization have been proposed. Chu & Liu (2019, MeanSum) introduced a summarization system based on a review autoencoder. At inference time, MeanSum encodes every review for a product to a vector, computes the centroid of reviews’ vectors and uses this centroid to seed the decoder and generate a summary. However, averaging representations of statements that are sometimes contradictory tends to confuse the decoder, and to lead it to rely on only language modeling for generating the output summary, thus ignoring the input signal. To deal with this limitation, Coavoux et al. (2019) proposed to add a clustering step to identify similar reviews and to generate one sentence per such found cluster: the averaging step only targets similar reviews. Contemporaneous to this work, Bražinskas et al. (2019) proposed to solve the problem of unsupervised summarization of reviews through an auto-encoder with latent variables. Their proposed way of solving the problem of hallucinating content from other categories is to use one latent variable per product, and let the decoder access all the reviews of a product. Compared to it, we argue that our self-supervised setting is simpler as it relies on training with standard cross-entropy. In addition, the use of Transformer (as opposed to GRU in their case) makes it possible to apply separate attentions to each input.
A similar idea was very recently proposed for pretraining summarization models. Zhang et al. (2019) masks out full sentences from a document, and trains a model that predicts those sentence from the surrounding text. Our self-supervision training mechanism can be seen as a multi-document version of that.
West et al. (2019) introduced a self-supervised system for sentence compression: they design an unsupervised extractive system and use it to generate data to train a supervised neural sentence compressor. However, their two-level system works at the level of single sentences whereas our end-to-end approach summarizes sets of reviews with multiple sentences.
Controlled Generation
We rely on controlled natural language generation to steer the generation away from hallucinations. Controllable text generation has been previously investigated to apply global constraints on text generation. Previous work proposed fine-tuning NLG models to provide control. To allow back-propagation through the discrete sampling process of text generation several proposals have used policy gradient methods, most notably REINFORCE (Williams, 1992) for applications such as machine translation (Ranzato et al., 2016; Wu et al., 2018), image-to-text generation (Liu et al., 2017), dialogue generation (Li et al., 2016b) and visual question answering (Yi et al., 2018). Other work has relied on continuous approximation methods, most notably Gumbel-Softmax (Jang et al., 2017) such as in Chu & Liu (2019); Yang et al. (2018).
Other methods of control applied control only at inference time. Weighted decoding, introduced by Holtzman et al. (2018), was shown to be challenging, and to often lead to sacrificing fluency and coherence (See et al., 2019). Constrained beam search (Anderson et al., 2017; Hokamp & Liu, 2017; Post & Vilar, 2018) is slower, requires in practice very large beam sizes, and does not enable soft constraints. Finally, updating the decoder hidden states (Chen et al., 2018; Dathathri et al., 2019) requires an extra training step.
The first introduction of control codes to neural generation models has been an early form of copy mechanism to overcome the rare word problem (Luong et al., 2015; ElSahar et al., 2018) and has recently shown a wide adoption - due to its simplicity and effectiveness - to steer large scale language models toward desired traits such as general aspects (Keskar et al., 2019) or structured fields (Zellers et al., 2019).
Previous work for controlling large-scale language models has relied on a predefined set of bag of control tokens, collected either manually (Keskar et al., 2019) or from dictionaries (Dathathri et al., 2019), which can lead to low domain coverage. Regularized classification models have intrinsic feature selection capabilities (Ng, 2004), that have been exploited before for lexicon generation from sentiment classifiers (Nabil et al., 2014; ElSahar & El-Beltagy, 2015). These approaches generate more relevant lexicons than traditional topic models such as LDA (Blei et al., 2003). In this work, to automatically generate bag of control tokens we follow the same approach which does only rely on the category meta-data provided with the reviews. In the absence of such meta-data information several approaches have relied instead creating seed lexicons using unsupervised or weakly supervised aspect extractors (He et al., 2017; Angelidis & Lapata, 2018)
Hierarchical encoding.
In order to allow a neural summarizer to read several sections Cohan et al. (2018) proposes a hierarchical LSTM that works at two level. Most similar to our proposed method, Liu & Lapata (2019) extends a Transformer network to read several ranked paragraphs as input, avoiding a retrieve-then-read pipeline. In multi-document summarization the paragraphs are not ranked but independent. This entails a significant change model-wise, and we propose to encode each review independently (avoiding their inter-paragraph self-attention) and only adapt the decoder-encoder attention.
Self-Supervision
In order to create our training dataset we assume that a review for an entity (venue or product) can serve as a summary for a set of other similar reviews . This simple intuition allows us to create training points in a very similar way to what the model will experience at inference time. However, there are two issues with this approach. First, the potential set of training points is too large to explore exhaustively. Given the set of all reviews the total number of possible input-output pairs is . Second, the assumption that any review is fit to serve as a summary for any set of other reviews is obviously not true, and might yield a very noisy training dataset.
To solve the combinatorial explosion, we limit the size of to , and from a given , we look for a set of good reviews , for which serves as a good summary. Fixing also simplifies training, and enables comparison with previous work where the number of input reviews is fixed (Chu & Liu, 2019; Bražinskas et al., 2019). Both and all members of are reviews of the same entity.
Having fixed, we now search for reviews for which is a relevant review:
Note that fixing first the target summaries turns traditional approaches upside down. In particular, a recently proposed theoretical model of importance in summarization (Peyrard, 2019) defines the importance of a summary based on three aspects: (i) minimum redundancy, (ii) maximum relevance with the input document, and (iii) maximum informativeness. In that line of work is considered fixed: redundancy and informativeness are not dependent on and can therefore be ignored when is fixed. In this setting Peyrard (2019) reduces then to Eq. 1
Then, we sort the data-points according to the value of the relevance (). Depending on the desired size of the target dataset, we keep the top- pairs for training. Limiting inherently increases informativeness, since it limits the creation of training examples where input and outputs are repetitive similar reviews that might be very prominent on corpora level (e.g. “Great restaurant.”). In addition to simplicity, this method enables a fast implementation using state-of-the-art nearest neighbour search libraries (Pedregosa et al., 2011b). For all our experiments we defined sim to be the cosine similarity over a tf-idf bag-of-word representation (Ramos et al., 2003).
Controlling Hallucinations
Hallucinations are pieces of generated text that bear no relationship to the text they were conditioned on. They are likely to happen in our self-supervised setting, due to the noise from the construction of training instances. This might happen, for instance, if the synthetically created training data contains a variety of contradictory signals, or because certain types of review are overly present (e.g. “great movie”). The model might default to those very frequent patterns if it finds itself in a unfrequent state during decoding time.
To alleviate the problem of hallucinations, we propose to use control tokens that represent desired traits of the output text to steer the generated text towards more input-coherent summaries.
These control tokens are inferred from each review, and used as prompts at inference time. We use two types of codes as follows:
1) Metadata control tokens. Those are special tokens that are associated with each input review, and are the capitalized control tokens in Fig. 1. We use two types of metadata that represent (i) the review polarity, a numerical value denoting the average sentiment score of the input reviews; (ii) and categorical tokens representing the type of the entity of the review (e.g. Deli, Beauty&Spa, Furniture Stores). In the case of the unavailability of meta-data labels for all reviews (as in Rotten-Tomatoes dataset), we infer control tokens with the same process, but using categories predicted by a trained classifiers on labeled examples from the same domain.
During training, each review is enriched with its tailored control codes. In particular, the reviews acting as summary also contain them, and by construction those are -grams present in the text. At inference time – when the target side and hence its control codes are not available – we select the most repeated control tokens from the input side and feed them as a prefix to the decoder before the start of generation. There is clearly a risk that the model just learns to copy the control codes it has seen somewhere in the text. We check whether this is the case in Sect. 7.
Multi-source Transformer Model
Previous work for multi-document summarization (Chu & Liu, 2019) built multi-source input representations through a simple mean over the last hidden states of the encoder. An intrinsic limitation of this method is that the full set of reviews is represented as a single vector. This aggregation might cause information distortion especially when some input reviews are expected to have conflicted opinions in between. Standard transformer models (Vaswani et al., 2017) consider only a single input to the decoder part of the model. Aggregating all input reviews into a single input (Junczys-Dowmunt, 2019) with special tokens to represent document boundaries might be slow and impractical due the complexity of the self-attention mechanism. We therefore experiment with several input combination strategies of the transformer cross-attention (Libovický et al., 2018). Parallel. At each cross-attention head, the decoder set of queries attend to each of the encoded inputs separately from which the set of keys () and values () are generated and then the yielded context is averaged and followed by a residual connection from the previous decoder layer. This corresponds to box (C) in Fig. 1.
Mean. We also propose a simpler input combination strategy, which is less computationally demanding. It does not apply the cross-attention with each encoder separately. Instead, the set of keys and values coming from each input encoder are aggregated using the average at each absolute position. Afterwards the decoder set of queries attend to this aggregated set of keys and values. This combination can be seen as a more efficient variation of the flat combination strategy (Libovický et al., 2018) with mean instead of concatenation. Fig. 2 depicts this strategy, which replaces box (C) in Fig. 1.
In Sect. 7, we compare both approaches through an ablation study, focusing on summary quality as well as empirical training times.
Experimental Setup
All our models are implemented with PyTorch (Paszke et al., 2019) and Fairseq (Ott et al., 2019) libraries, as well as scikit-learn (Pedregosa et al., 2011a) for the classifiers used either for inferring control tokens or for evaluation. For all our models we use sentence piece (Kudo & Richardson, 2018) as a tokenizer with a vocabulary size of 32 000. We use the same hyperparameters as the Transformer Big model described by Vaswani et al. (2017) (, , , ). We optimize them with a Nesterov accelerated SGD optimizer with a learning rate of . We train all models for a total of 80 000 steps across epochs, with linear warm-up for the first 8 000 steps. We select the best model checkpoint based on perplexity on the validation set. All models were trained on one machine with 4 NVIDIA V100 GPUs, the longest model took hours to train. For inference, we use a beam size of 35. We discard hypotheses that contains twice the same trigram. We limit generation of each summary to a maximum budget of 150 tokens for each summary for Yelp, as was done by Chu & Liu (2019), and a budget of 50 tokens for Rotten Tomatoes. We set a similar budget for all other extractive baselines in the experiments. Finally, we use length normalization (Wu et al., 2016) with length penalty 1.2 to account for the model’s bias towards shorter sequences.
Datasets
We evaluate our proposal on two English datasets: Yelphttps://www.yelp.com/dataset/challenge (Chu & Liu, 2019) and Rotten Tomatoes (Wang & Ling, 2016). The Yelp dataset contains reviews of businesses (approximately one million reviews for around 40k venues). As described in Section 3, for each venue, we select the best reviews to use as target summaries: either the top- (with ) or the top- (with ) reviews, whichever is smaller. For each target summary thus selected, we then take its most similar reviews (cosine similarity) to form its input. We obtain around 340k training examples, representing 22.5k venues. The Rotten Tomatoes dataset was constructed by (Wang & Ling, 2016) from the movie review website rottentomatoes.com. We use the same process as for Yelp, but use and . We construct around 170k training examples, representing 3.7k movies. Details about dataset sizes and splits are in the Appendix.
Evaluation Metrics
We evaluate summary systems with the classical ROUGE-F-{1,2,L} metrics (Lin, 2004).For Yelp we use the MeanSum (Chu & Liu, 2019) implementation to keep results comparable while for RottenTomatoes we use py-rouge package pypi.python.org/pypi/pyrouge/0.1.3 We also report BERT-score (Zhang et al., 2020), a metric that uses pre-trained BERT (Devlin et al., 2019) to compute the semantic similarity between a candidate summary and the gold summary. and - () scores (Li et al., 2016a) are the percentage of distinct -grams in the generated text on the summary level or the corpora level respectively. Dist- is an indicator of repetitiveness within a single summary while Distc- indicates the diversity of different generations. Finally, as done by Chu & Liu (2019), we use a classifier to check whether the sentiment of the summary is consistent with the sentiment of input reviews (Sentiment Acc. in Table 1).We use a 3-class classification: negative (1 or 2 star), neutral (3), positive (4 and 5)). As a result, the numbers are not comparable with those reported Chu & Liu (2019). We extend this method to check whether the correct product category can also be inferred from the summary, we report the micro F-score of the multi-label category classifier.
Baselines and Other Systems
We compare our system to three unsupervised baselines. TextRank (Mihalcea & Tarau, 2004) and LexRank (Radev et al., 2004) are extractive systems based on the PageRank algorithm. Opinosis (Ganesan et al., 2010) is an abstractive graph-based system. We use openly available Python implementations for TextRankhttps://github.com/summanlp/textrank (Barrios et al., 2016) and LexRank.https://github.com/crabcamp/lexrank We use the default parameters of the implementations. For Opinosis, we use the official Java implementationhttps://github.com/kavgan/opinosis-summarization with default hyperparameters.Except for the redundancy parameter which was set to one, since the default led to many empty outputs.
We also compare our systems with more recent neural unsupervised summarization systems. For the Yelp dataset, we rerun the released pretrained version of MeanSumhttps://github.com/sosuperic/MeanSum/ (Chu & Liu, 2019).
Evaluation Results
Table 1 contains the automatic evaluation metrics with respect to reference summaries. The proposed multi-input self-supervised model with control codes perform consistently better in the Yelp dataset across the benchmarked models, inlcuding the recent neural unuspervised models of MeanSum and H-VAE. Note that because of the concurrent nature of the Bražinskas et al. (2019) paper, the H-VAE model is not available and we report the numbers from their paper.While the ROUGE implementation might be different, the numbers of the common baselines are very close. For MeanSum we re-run their provided checkpoint and run evaluation through the same pipeline. The BERTScore (Zhang et al., 2020) differences are closer and seem to favour neural models.
With the RottenTomatoes dataset we only benchmarked the graph-based unsupervised methods, since the released pretrained MeanSum model does not cover the domain of movie reviews. We attribute the lower score in sentiment accuracy to the fact that the “summaries” in RottenTomatoes are critical reviews, written in a very different style than the original reviews.
Table 2 contains reference-less evaluation, analyzing the number of distinct -grams (an indicator of repetitiveness) on the summary level and corpora level. On the summary level our model outperforms all the baselines, meaning, our model is capable of generating more rich and less repetitive summaries. On the level of all generations our model generates text with more diversity than MeanSum. In general however extractive models tend to have more diversity on the corpus level as they directly copy from each input separately, while abstractive models tend to learn repetitive patterns present in the training set.
Fig. 3 shows summaries generated by different models from the same input. We notice that our model learned to copy aspects of the input documents such as restaurant names “Capricotti’s” and menu items “the Bobbie”, this is possibly attributed to the cross-attention mechanism in our proposed model. More examples are provided in the supplementary material Appendix.
Human Evaluation
Existing natural language generation systems are known to generate very fluent language, that looks very natural to native speakers. On the other side, current neural models are known to generate factually incorrect data, something which was less of a concern in pre-neural methods but also much harder to detect. As mentioned by Kryscinski et al. (2019): “Neither of the methods explicitly examines the factual consistency of summaries, leaving this important dimension unchecked.” Inspired by Falke et al. (2019) we decided to focus the human evaluation on those aspects of the summarization evaluation in which existing models risk failing the most, the one of faithfulness.
We annotated 94 summaries through a crowd-sourcing platform, comparing 3 systems (Gold, MeanSum and ours). Workers were asked if “the summary contains correct information given the original reviews”. In total we had 282 tasks () and each task was labeled by 3 annotators and paid 15 / hour) and restricted to experienced, English-speaking workers. A full description of the campaign, including the filtering of the annotations, is detailed in Appendix.
The results in Table 4 show that 92.6% of the generated summaries of our system are considered factually correct (compare with 95.7% for the gold summaries), as opposed to 79.7% of MeanSum.
Ablation
We analyzed the impact of our proposed variations of the basic self-supervised setting in Table 3. Removing control codes degrades significantly – as expected – sentiment and category classification of the produced summary . It also impacts greatly on the ROUGE score. Changing the decoder-encoder attention from parallel to mean (see Sect. 5) also degrades ROUGE. The difference of this attention change without control codes is smaller but – surprisingly – in the different direction.
Control Codes
The previous ablation study shows the importance of the control codes in the quality of the final summaries. In order to see how rigidly the model follows those control codes we devise the following experiment to see if the tokens used as control codes are forced to appear in the output text, independent of the input text.
For this, we sample reviews (for venues from the Yelp validation set). For each input example, we randomly sample control tokens (inferred control codes, see Sect 4) from the tokens occurring in the review. We refer to these as correct control tokens. We run the decoder using these control tokens as prompt and count the proportion of them that also occurs in the generated summary. For comparison, we repeat the same experiment but sampling instead control tokens that do not occur in the input text. We refer to these as incorrect control tokens.
To minimize the possibility of conditioning on control tokens that might show up naturally in the generated text, for both settings, we repeat the process times per input example (resulting in with correct control tokens as prefix and using incorrect). We report in Fig. 5 the proportion of fed control codes that are generated by the model in both cases. We observe that the model tends to comply with the correct control tokens that occur in the input documents (eg: of the summaries contain more than of the control tokens), but tends to ignore the control tokens when they do not occur in the input. Fig. 4 shows a set of generated examples for the same input when the model is conditioned on different control tokens.
Conclusion
Neural methods have shown great promises for unsupervised multi-document abstractive summarization, overcoming the lack of fluency of extractive models. However, those models are often complex to train and more importantly tend to generate incorrect statements; characteristics which are exacerbated in the unsupervised setting. Our proposed models aim to overcome those problems by proposing a simple training mechanism relying on a self-supervised formulation. In addition to our use of multi-input transformers and control codes, we show that the resulting summaries are better (as measured by ROUGE and other automatic measures), and produce more faithful summaries (as measured by human evaluation). The use of control codes makes it easy to extend for other multi-document summarization use-cases.
While the generated reviews are more factual than those generated by other models, we want to stress that inaccuracies can still appear and that special care should be taken if such methods are to be deployed. In particular, the models learn the conjugations from the input, which is mostly in first persons. Such summaries might be misleading as it could lend to believe that an actual human wrote those. We recommend strongly that any use of such algorithms to be accompanied by a clear disclaimer on its true nature.
References
Appendix A Generated Examples
Fig. 6,7 include a set of samples generated from our model and baselines, full generations our model can be downloaded https://www.dropbox.com/s/w6eqviy5fnda11f/hypos_and_refs.zip?dl=0
Appendix B Inferred Control Tokens
Fig. 8 shows examples of the top inferred tokens for some categories in the Yelp dataset, those tokens have been inferred using our proposed method in this work.
Appendix C Human Evaluation Campaign
We used Amazon Mechanical Turk to ask 3 “workers” to assess if 282 summaries produced by 3 systems (94 from each: ours, gold from human experts and Meansum) aligned correctly with sets of 8 reviews. Workers had to read the reviews, the summary and answer the question: “does the summary contain correct information given the original reviews?” Instructions specified to “assess the faithfulness of the summary with respect to [the] set of reviews,” specifically to “verify that the summary [did] not contain factually incorrect or self-contradicting statements that could not be inferred from what [was] provided in the original reviews.” Using Mechanical Turk qualification criteria, we asked for the workers: (1) to be located in the United States, Canada or United Kingdom; (2) to have a HIT approval rate higher than 98; (3) to have more than 1000 HITs approved.
We did an internal run to estimate the time needed per individual assignment –each Human Intelligence Task, or HIT, an annotation in our case, was assigned to 3 workers. We followed it by a short pilot to validate the average 2 minutes we had estimated. This is important to establish the rate to pay: 2 minutes translate into 30 potential assignments per hour, we picked 15 hourly wage. Beyond the timing, the pilot was also used as a dry run for the full campaign. Average time to Answer and the theoritical hourly wage is available in Table 6
By using shuffled gold summaries, hence written for another set of reviews, we included 21 badly aligned “negatives.” Workers who answered yes for these obvious no were filtered out as “dubious” from the results: all their answers were discarded. After filtering out the “negatives” HITs and the ones from “dubious” answers, we were left with 446 annotations. We further discarded all annotations made in less than a minute to keep 377 realistic answers.
Finally we looked for full agreement at the HIT level and kept only the ones with either 0 yes or 0 no, with varying numbers, from 1 to 3, of the alternatives after the filtering of the ”dubious” and ”unrealistic” answers. Not surprisingly, as we focused on alignment, Gold summaries scored best but ours scored nicely, with a very low number of misaligned summaries:
Assessing the alignment of summaries to a set of reviews is not an easy task. We decided to discard all answers from the ”dubious” workers who erred on our ”negatives” summaries to be on the safe side. Mechanical Turk reports the time taken for an assignment, their averages is an interesting metric to look at, especially the way it evolves along our filterings —we translated it to the associated theoretical hourly wages, alas all under the $15 we initially targeted.