Self-Supervised Multimodal Opinion Summarization

Jinbae Im, Moonki Kim, Hoyeop Lee, Hyunsouk Cho, Sehee Chung

Introduction

Opinion summarization is the task of automatically generating summaries from multiple documents containing users’ thoughts on businesses or products. This summarization of users’ opinions can provide information that helps other users with their decision-making on consumption. Unlike conventional single-document or multiple-document summarization, where we can obtain the prevalent annotated summaries (Nallapati et al., 2016; See et al., 2017; Paulus et al., 2018; Liu et al., 2018; Liu and Lapata, 2019; Perez-Beltrachini et al., 2019), opinion summarization is challenging; it is difficult to find summarized opinions of users. Accordingly, studies used an unsupervised approach for opinion summarization (Ku et al., 2006; Paul et al., 2010; Carenini et al., 2013; Ganesan et al., 2010; Gerani et al., 2014). Recent studies (Bražinskas and Titov, 2020; Amplayo and Lapata, 2020; Elsahar et al., 2021) used a self-supervised learning framework that creates a synthetic pair of source reviews and a pseudo summary by sampling a review text from a training corpus and considering it as a pseudo summary, as in Figure 1(a).

Users’ opinions are based on their perception of a specific entity and perceptions originate from various characteristics of the entity; therefore, opinion summarization can use such characteristics. For instance, Yelp provides users food or menu images and various metadata about restaurants, as in Figure 1(b). This non-text information influences the review text generation process of users (Truong and Lauw, 2019). Therefore, using this additional information can help in opinion summarization, especially under unsupervised settings (Su et al., 2019; Huang et al., 2020). Furthermore, the training process of generating a review text (a pseudo summary) based on the images and metadata for self-supervised learning is consistent with the actual process of writing a review text by a user.

This study proposes a self-supervised multimodal opinion summarization framework called MultimodalSum by extending the existing self-supervised opinion summarization framework, as shown in Figure 1. Our framework receives source reviews, images, and a table on the specific business or product as input and generates a pseudo summary as output. Note that images and the table are not aligned with an individual review in the framework, but they correspond to the specific entity. We adopt the encoder–decoder framework and build multiple encoders representing each input modality. However, a fundamental challenge lies in the heterogeneous data of various modalities (Baltrušaitis et al., 2018).

To address this challenge, we propose a multimodal training pipeline. The pipeline regards the text modality as a pivot modality. Therefore, we pretrain the text modality encoder and decoder for a specific business or product via the self-supervised opinion summarization framework. Subsequently, we pretrain modality encoders for images and a table to generate review texts belonging to the same business or product using the pretrained text decoder. When pretraining the non-text modality encoders, the pretrained text decoder is frozen so that the image and table modality encoders obtain homogeneous representations with the pretrained text encoder. Finally, after pretraining input modalities, we train the entire model in an end-to-end manner to combine multimodal information.

Our contributions can be summarized as follows:

this study is the first work on self-supervised multimodal opinion summarization;

we propose a multimodal training pipeline to resolve the heterogeneity between input modalities;

we verify the effectiveness of our model framework and model training pipeline through various experiments on Yelp and Amazon datasets.

Related Work

Generally, opinion summarization has been conducted in an unsupervised manner, which can be divided into extractive and abstractive approaches. The extractive approach selects the most meaningful texts from input opinion documents, and the abstractive approach generates summarized texts that are not shown in the input documents. Most previous works on unsupervised opinion summarization have focused on extractive approaches. Clustering-based approaches (Carenini et al., 2006; Ku et al., 2006; Paul et al., 2010; Angelidis and Lapata, 2018) were used to cluster opinions regarding the same aspect and extract the text representing each cluster. Graph-based approaches (Erkan and Radev, 2004; Mihalcea and Tarau, 2004; Zheng and Lapata, 2019) were used to construct a graph—where nodes were sentences, and edges were similarities between sentences—and extract the sentences based on their centrality.

Although some abstractive approaches were not based on neural networks (Ganesan et al., 2010; Gerani et al., 2014; Di Fabbrizio et al., 2014), neural network-based approaches have been gaining attention recently. Chu and Liu (2019) generated an abstractive summary from a denoising autoencoder-based model. More recent abstractive approaches have focused on self-supervised learning. Bražinskas and Titov (2020) randomly selected NN review texts for each entity and constructed NN synthetic pairs by sequentially regarding one review text as a pseudo summary and the others as source reviews. Amplayo and Lapata (2020) sampled a review text as a pseudo summary and generated various noisy versions of it as source reviews. Elsahar et al. (2021) selected review texts similar to the sampled pseudo summary as source reviews, based on TF-IDF cosine similarity. We construct synthetic pairs based on Bražinskas and Titov (2020) and extend the self-supervised opinion summarization to a multimodal version.

Multimodal text summarization has been mainly studied in a supervised manner. Text summaries were created by using other modality data as additional input (Li et al., 2018, 2020a), and some studies provided not only a text summary but also other modality information as output (Zhu et al., 2018; Chen and Zhuge, 2018; Zhu et al., 2020; Li et al., 2020b; Fu et al., 2020). Furthermore, most studies summarized a single sentence or document. Although Li et al. (2020a) summarized multiple documents, they used non-subjective documents. Our study is the first unsupervised multimodal text summarization work that summarizes multiple subjective documents.

Problem Formulation

The goal of the self-supervised multimodal opinion summarization is to generate a pseudo summary from multimodal data. Following existing self-supervised opinion summarization studies, we consider a review text selected from an entire review corpus as a pseudo summary. We extend the formulation of Bražinskas and Titov (2020) to a multimodal version. Let RR = {r1,r2,...,rN}\{r_{1},r_{2},...,r_{N}\} denote the set of reviews about an entity (e.g., a business or product). Each review, rjr_{j}, consists of review text, djd_{j}, and review rating, sjs_{j}, that represents the overall sentiment of the review text. We denote images uploaded by a user or provided by a company for the entity as II = {i1,i2,...,iM}\{i_{1},i_{2},...,i_{M}\} and a table containing abundant metadata about the entity as TT. Here, TT consists of several fields, and each field contains its own name and value. We set jj-th review text djd_{j} as the pseudo summary and let it be generated from R−jR_{-j}, II, and TT, where R−j={r1,...,rj−1,rj+1,...,rN}R_{-j}=\{r_{1},...,r_{j-1},r_{j+1},...,r_{N}\} denotes source reviews. To help the model summarize what stands out overall in the review corpus, we calculate the loss for all NN cases of selecting djd_{j} from RR, and train the model using the average loss. During testing, we generate a summary from RR, II, and TT.

Model Framework

The proposed model framework, MultimodalSum, is designed with an encoder–decoder structure, as in Figure 1(b). To address the heterogeneity of three input modalities, we configure each modality encoder to effectively process data in each modality. We set a text decoder to generate summary text by synthesizing encoded representations from the three modality encoders. Details are described in the following subsections.

2 Image Encoder

3 Table Encoder

where nn and vv denote eTe_{T}-dimensional representations of field name and value, respectively, and Wf∈R2eT×eTW_{f}\in R^{2e_{T}\times e_{T}}, bf∈ReTb_{f}\in R^{e_{T}} are parameters. By stacking lTl_{T} field representations, we obtain F∈R1×lT×eTF\in R^{1\times l_{T}\times e_{T}}. The additional linear weights Wtable∈ReT×eDW_{table}\in R^{e_{T}\times e_{D}} play the same role as in the image encoder, and htable∈R1×lT×eDh_{table}\in R^{1\times l_{T}\times e_{D}}.

Model Training Pipeline

To effectively train the model framework, we set a model training pipeline, which consists of three steps, as in Figure 2. The first step is text modality pretraining, in which a model learns unsupervised summarization capabilities using only text modality data. Next, during the pretraining for other modalities, an encoder for each modality is trained using the text modality decoder learned in the previous step as a pivot. The main purpose of this step is that other modalities have representations whose distribution is similar to that of the text modality. In the last step, the entire model framework is trained using all the modality data. Details of each step can be found in the next subsections.

The limitation of the self-supervised opinion summarization is that training and inference tasks are different. The model learns a review generation task using a review text as a pseudo summary; however, the model needs to perform a summary generation task at inference. To close this gap, we use a rating deviation between the source reviews and the target as an additional input feature of the text decoder, inspired by Bražinskas et al. (2020). We define the average ratings of the source reviews minus the rating of the target as the rating deviation: sdj=∑i≠jNsi/(N−1)−sj{sd}_{j}=\sum_{i\neq j}^{N}{s_{i}}/(N-1)-s_{j}. We use sdj{sd}_{j} to help generate a pseudo summary djd_{j} during training and set it as to generate a summary with average semantic of input reviews during inference. To reflect the rating deviation, we modify the way in which a Transformer creates input embeddings, as in Figure 3. We create deviation embeddings with the same dimensionality as token embeddings and add sdj{sd}_{j} ×\times deviation embeddings to the token embeddings in the same way as positional embeddings.

Our methods to close the gap between training and inference tasks do not require additional modeling or training in comparison with previous works. We achieve noising and denoising effects by simply using rating deviation embeddings without variational inference in Bražinskas and Titov (2020). Furthermore, the information that the rating deviation is plays the role of an input prompt for inference, without the need to train a separate classifier for selecting control tokens to be used as input prompts (Elsahar et al., 2021).

2 Other Modalities Pretraining

3 Training for Multiple Modalities

Experimental Setup

To evaluate the effectiveness of the model framework and training pipeline on datasets with different domains and characteristics, we performed experiments on two review datasets: Yelp Dataset Challengehttps://www.yelp.com/dataset and Amazon product reviews (He and McAuley, 2016). The Yelp dataset provides reviews based on personal experiences for a specific business. It also provides numerous images (e.g., food and drinks) uploaded by the users. Note that the maximum number of images, MM, was set to 1010 based on the 90th90^{th} percentile. In addition, the dataset contains abundant metadata of businesses according to the characteristics of each business. On the contrary, the Amazon dataset provides reviews with more objective and specific details about a particular product. It contains a single image provided by the supplier, and provides relatively limited metadata for the product. For evaluation, we used the data used in previous research (Chu and Liu, 2019; Bražinskas and Titov, 2020). The data were generated by Amazon Mechanical Turk workers who summarized 88 input review texts. Therefore, we set NN to 99 so that a pseudo summary is generated from 88 source reviews during training. For the Amazon dataset, 33 summaries are given per product. Simple data statistics are shown in Table 1, and other details can be found in Appendix A.1.

2 Experimental Details

All the modelsOur code is available at https://bit.ly/3bR4yod were implemented with PyTorch (Paszke et al., 2019), and we used the Transformers library from Hugging Face (Wolf et al., 2020) as the backbone skeleton. Our text encoder and decoder were initialized using BART-Large and further pretrained using the training review corpus with the same objective as BART. eDe_{D}, eIe_{I}, and eTe_{T} were all set to 1,024. We trained the entire models using the Adam optimizer (Kingma and Ba, 2014) with a linear learning rate decay on NVIDIA V100s. We decayed the model weights with 0.10.1. For each training pipeline, we set different batch sizes, epochs, learning rates, and warmup steps according to the amount of learning required at each step. We used label smoothing with 0.10.1 and set the maximum norm of gradients as 11 for other modalities pretraining and multiple-modalities training. During testing, we used beam search with early stopping and discarded hypotheses that contain twice the same trigram. Different beam size, length penalty, and max length were set for Yelp and Amazon. The best hyperparameter values and other details are described in Appendix A.2.

3 Comparison Models

We compared our model to extractive and abstractive opinion summarization models. For extractive models, we used some simple baseline models (Bražinskas and Titov, 2020). Clustroid selects one review that gets the highest ROUGE-L score with the other reviews of an entity. Lead constructs a summary by extracting and concatenating the lead sentences from all review texts of an entity. Random simply selects one random review from an entity. LexRank (Erkan and Radev, 2004) is an extractive model that selects the most salient sentences based on graph centrality.

For abstractive models, we used non-neural and neural models. Opinosis (Ganesan et al., 2010) is a non-neural model that uses a graph-based summarizer based on token-level redundancy. MeanSum (Chu and Liu, 2019) is a neural model that is based on a denoising-autoencoder and generates a summary from mean representations of source reviews. We also used three self-supervised abstractive models. DenoiseSum (Amplayo and Lapata, 2020) generates a summary by denoising source reviews. Copycat (Bražinskas and Titov, 2020) uses a hierarchical variational autoencoder model and generates a summary from mean latent codes of the source reviews. Self & Control (Elsahar et al., 2021) generates a summary from Transformer models and uses some control tokens as additional inputs to the text decoder.

Results

We evaluated our model framework and model training pipeline. In particular, we evaluated the summarization quality compared to other baseline models in terms of automatic and human evaluation, and conducted ablation studies.

To evaluate the summarization quality, we used two automatic measures: ROUGE-{1,2,L} (Lin, 2004) and BERT-score (Zhang et al., 2020). The former is a token-level measure for comparing 11, 22, and adaptive L-gram matching tokens, and the latter is a document-level measure using pretrained BERT (Devlin et al., 2019). Contrary to ROUGE-score, which is based on exact matching between n-gram words, BERT-score is based on the semantic similarity between word embeddings that reflect the context of the document through BERT. It is approved that BERT-score is more robust to adversarial examples and correlates better with human judgments compared to other measures for machine translation and image captioning. We hypothesize that BERT-score is strong in opinion summarization as well, and BERT-score would complement ROUGE-score.

The results for opinion summarization on two datasets are shown in Table 2. MultimodalSum showed superior results compared with extractive and abstractive baselines for both token-level and document-level measures. From the results, we conclude that the multimodal framework outperformed the unimodal framework for unsupervised opinion summarization. In particular, our model achieved state-of-the-art results on the Amazon dataset and outperformed the comparable model by a large margin in the R-L representing the ROUGE scores on the Yelp dataset. Although Self & Control showed high R-2 score, we attributed their score to the inferred NN-gram control tokens used as additional inputs to the text decoder.

Sample summaries on the Yelp dataset are shown in Table 3. They were generated from source reviews on Baby Cakes bakery. Copycat misused “sweet tooth” and generated “lemon mernigue pie” that was not mentioned in the source reviews. Self & Control generated a summary about a buffet by totally misunderstanding one sentence from source reviews: “If you love the desserts in Studio B Buffet in the M Hotel but don’t want to wait in the massive buffet line or even eat in the buffet, Baby Cakes in the M Hotel is really nice fix.” Furthermore, “Matos Buffet” is a non-existent word. On the contrary, MultimodalSum generated a good summary with a rich description of chocolate croissants. Although “chocolate chip cookie” was not found in the source reviews, our model generated it from cookie images. Note that the term can be found in other reviews that were not used as source reviews. Additional sample summaries on two datasets are shown in Appendix A.5.

1.2 Human Evaluation

To evaluate the quality of summarization based on human criteria, we conducted a user study. We assessed the quality of summaries using Best-Worst Scaling (BWS; Louviere et al. (2015)). BWS is known to produce more reliable results than raking scales (Kiritchenko and Mohammad, 2017) and is widely used in self-supervised opinion summarization studies. We recruited 1010 NLP experts and asked each participant to choose one best and one worst summary from four summaries for three criteria. For each participant’s response, the best model received +11, the worst model received -11, and the rest of the models received scores. The final scores were obtained by averaging the scores of all the responses from all participants.

For Overall criterion, Self & Control, Copycat, MultimodalSum, and gold summaries scored -0.5270.527, -0.1130.113, +0.2600.260, and +0.3800.380 on the Yelp dataset, respectively. MultimodalSum showed superior performance in human evaluation as well as automatic evaluation. We note that human judgments correlate better with BERT-score than ROUGE-score. Self & Control achieved a very low human evaluation score despite its high ROUGE-score in automatic evaluation. We analyzed the summaries of Self & Control, and we found several flaws such as redundant words, ungrammatical expressions, and factual hallucinations. It generated a non-existent word by combining several subwords. It was particularly noticeable when a proper noun was generated. Furthermore, Self & Control generated an implausible sentence by copying some words from source reviews. From the results, we conclude that both automatic evaluation and human evaluation performances should be supported to be a good summarization model and BERT-score can complement ROUGE-score in automatic evaluation. Details on human evaluation and full results can be found in Appendix A.3.

1.3 Effects of Multimodality

To analyze the effects of multimodal data on opinion summarization, we analyzed the multimodal gate. Since the multimodal gate is a eDe_{D}-dimensional vector, we averaged it by a scalar value. Furthermore, as multimodal gates exist for each layer of the text decoder, we averaged them to measure the overall influence of a table or images when generating each token in the decoder. An example of aggregated multimodal gates is shown in Figure 4. It shows the table and images used for generating a summary text, and the multimodal gates for a part of the generated summary are expressed as heatmaps. As we intended, table and image information was selectively used to generate a specific word in the summary. The aggregated value of the table was relatively high for generating “Red Lobster”, which is the name of the restaurants. It was relatively high for images, when generating “food” that is depicted in two images. Another characteristic of the result is that aggregated values of the table were higher than those of the image: mean values for the table and image in the entire test data were 0.103 and 0.045, respectively. This implies that table information is more used when creating a summary, and this observation is valid in that the table contains a large amount of metadata. Note that the values displayed on the heatmaps are small by and large, as they were aggregated from eDe_{D}-dimensional vector.

2 Ablation Studies

For ablation studies, we analyzed the effectiveness of our model framework and model training pipeline in Table 4. To analyze the model framework, we first compared the summarization quality with four versions of unimodal model framework, as in the first block of Table 4. BART denotes the model framework in Figure 1(a), whose weights are the weights of BART-Large. It represents the lower bound of our model framework without any training. BART-Review denotes the model framework whose weights are from further pretrained BART using the entire training review corpus. UnimodalSum refers to the results of the text modality pretraining, and we classified it into two frameworks according to the use of the rating deviation.

Surprisingly, using only BART achieved comparable or better results than many extractive and abstractive baselines in Table 2. Furthermore, further pretraining using the review corpus brought performance improvements. Qualitatively, BART with further pretraining generated more diverse words and rich expressions from the review corpus. This proved our assumption that denoising autoencoder-based pretraining helps in self-supervised multimodal opinion summarization. Based on the BART-Review, UnimodalSum achieved superior results. Furthermore, the use of rating deviation improved the quality of summarization. We conclude that learning to generate reviews based on wide ranges of rating deviations including during training helps to generate a better summary of the average semantics of the input reviews.

To analyze the effect of other modalities in our model framework, we compared the summarization quality with three versions of multimodal model frameworks, as in the second block of Table 4. We removed the image or table modality from MultimodalSum to analyze the contribution of each modality. Results showed that both modalities improved the summarization quality compared with UnimodalSum, and they brought additional improvements when used altogether. This indicates that using non-text information helps in self-supervised opinion summarization. As expected, the utility of the table modality was higher than that of the image modality. The image modality contains detailed information not revealed in the table modality (e.g., appearance of food, inside/outside mood of business, design of product, and color/texture of product). However, the information is unorganized to the extent that the utility of the image modality depends on the capacity of the image encoder to extract unorganized information. Although MultimodalSum used a representative image encoder because our study is the first work on multimodal opinion summarization, we expect that the utility of the image modality will be greater if unorganized information can be extracted effectively from the image using advanced image encoders.

For analyzing the model training pipeline, we removed text modality or/and other modalities pretraining from the pipeline. By removing each of them, the performance of MultimodalSum declined, and removing all of the pretraining steps caused an additional performance drop. Although MultimodalSum without other modalities pretraining has the capability of text summarization, it showed low summarization performance at the beginning of the training due to the heterogeneity of the three modality representations. However, MultimodalSum without text modality pretraining, whose image and table encoders were pretrained using BART-Review as a pivot, showed stable performance from the beginning, but the performance did not improve significantly. From the results, we conclude that both text modality and other modalities pretraining help the training of multimodal framework. For the other modalities pretraining, we conducted a further analysis in the Appendix A.4.

Conclusions

We proposed the first self-supervised multimodal opinion summarization framework. Our framework can reflect text, images, and metadata together as an extension of the existing self-supervised opinion summarization framework. To resolve the heterogeneity of multimodal data, we also proposed a multimodal training pipeline. We verified the effectiveness of our multimodal framework and training pipeline with various experiments on real review datasets. Self-supervised multimodal opinion summarization can be used in various ways in the future, such as providing a multimodal summary or enabling a multimodal retrieval. By retrieving reviews related to a specific image or metadata, controlled opinion summarization will be possible.

Acknowledgments

We thank the anonymous reviewers for their insightful comments and suggestions.

References

Appendix A Appendix

We selected businesses and products with a minimum of 10 reviews and popular entities above the 90th90^{th} percentile were removed. The minimum and maximum length of the words were set as 3535 and 100100 for Yelp, and 4545 and 7070 for Amazon, respectively. We set the maximum number of tokens as 128128 using the BART tokenizer for training, and we did not limit the maximum tokens for inference. For the Amazon dataset, we selected 4 categories: Electronics; Clothing, Shoes and Jewelry; Home and Kitchen; Health and Personal Care. As Yelp dataset contains unlimited number of images for each entity, we did not use images for popular entities above the 90th90^{th} percentile. On the other hand, Amazon dataset contains a single image for each entity. Therefore, we did not use images only when meaningless images such as non-image icon or update icon were used or the image links had expired.

For Yelp dataset, we selected name, ratings, categories, hours, and attributes among the metadata. We used the hours of each day of the week as seven fields and used all metadata contained in attributes as each field. For some attributes (‘Ambience’, ‘BusinessParking’, ‘GoodForMeal’) that have subordinate attributes, we used each subordinate attribute. Among the fields, we selected 4747 fields used by at least 10%10\% of the entities. We set the maximum number of categories as 66 based on the 90th90^{th} percentile, and averaged the representations of each category. For ratings, we converted it to binary notation consisting of 44 digits (22,21,20,2−12^{2},2^{1},2^{0},2^{-1}). For hours, we considered (open hour, close hour) as a 22-dimensional vector, and conducted KK-means clustering. We selected four clusters based on silhouette score: (16.5,23.216.5,23.2), (8.7,17.18.7,17.1), (6.4,236.4,23), and (10.6,22.610.6,22.6). Based on the clusters, we converted hours into a categorical type.

For Amazon dataset, we selected six fields: name, price, brand, categories, ratings, and description. We set the maximum number of categories as 33 based on the 90th90^{th} percentile, and averaged the representations of each category. Furthermore, as each category consists of hierarchies with a maximum of 8 depths, we averaged the representations of hierarchies to get each category representation. For price and ratings, we converted them to binary notation consisting of 1111 and 44 digits, respectively, after rounding them to the nearest 0.50.5 to contain digit for 2−12^{-1}. As some descriptions consist of many tokens, we set the maximum number of tokens as 128128. We regarded each token in description as each field, so we got total 5+1285+128 fields.

A.2 Experimental Details

Our image encoder is based on ResNet101. ResNet101 is composed of 1 convolution layer, 4 convolution layer blocks, and 1 fully connected layer block. Among them, 4 convolution layer blocks play an important role in analyzing image. Through each convolution layer block, the size of the image feature map is reduced to 1/4, but it gets high-level features. To maintain the ability to extract low-level features of the image, we set the model weights up to the second convolution layer block not to be trained further. We only used up to the third convolution layer block to increase the resolution of feature maps without using too high-level features for image classification. In this way, lIl_{I} was set to 14×1414\times 14 and eIe_{I} was set to 1,024.

To use the knowledge of text modality in table encoder, we obtained field name embeddings by summing the BART token embeddings for the tokens contained in the field name. Because various data types can be used for field value, we used different processing methods for each data type. Nominal values were handled in the same way as the field name. Binary and ordinal values were processed by replacing them with nominal values of corresponding meanings: ‘true’ and ‘false’ were used for binary values, and ‘cheap’, ‘average’, ‘expensive’, and ‘very expensive’ were used for ‘RestaurantsPriceRange’. Numerical values were converted to binary notation, and we obtained the representations by summing embeddings corresponding to the place, where the place value is 1. For other categorical values, we simply trained embeddings corresponding to each category.

We set each hyperparameter value different for each step in the model training pipeline, as in Table 5. We set the batch size according to the memory usage and set other values according to the amount of learning required. Hyperparameter ranges for epochs and lr (learning rate) were and [1e-03, 1e-04, 5e-05, 1e-05, 5e-06], respectively, and optimized values were chosen from validation loss in one trial. For summary generation at test time, we set different hyperparameter values for each dataset. Beam size, length penalty, and max length were set to 44, 0.970.97, and 105105 for Yelp and 22, 0.90.9, and 8080 for Amazon, respectively. Note that max length was set first to prevent incomplete termination and length penalty was determined based on the ROUGE scores on validation dataset. The number of training parameters for text, image, and table modality pretraining are 406.3406.3M, 27.127.1M, and 3.23.2M, respectively, and that for multimodal training is 486.9486.9M. Run time for text modality pretraining was 1616h on 44 GPUs, and it took 4141h and 4343h on 22 GPUs for image and table modality training, respectively. For final multimodal training, it took 1414h on 88 GPUs.

A.3 Human Evaluation

For human evaluation, we randomly selected 3030 entities from Yelp test data, and used three criteria: Grmmaticality (the summary should be fluent and grammatical), Coherence (the summary should be well structured and well organized), and Overall (based on your own criteria, select the best and the worst summary of the reviews). Results for three criteria are shown in Table 6. Self & Control achieved very poor performance for all criteria due to its flaws that were not revealed in the automatic evaluation. Surprisingly, MultimodalSum outperformed gold summaries for two criteria; however, its overall performance lagged behind Gold. As our model was initialized from BART-Large that had been pretrained using large corpus and further pretrained using training review corpus, it may have generated fluent and coherent summaries. It seems that our model lagged behind Gold in Overall due to various criteria other than those two. The fact that Gold scored lower than Copycat in Grammaticality may seem inconsistent with the result from Bražinskas and Titov (2020). However, we assumed that this result was due to a combination of the four models in relative evaluation. The ranking for Copycat and Gold may have changed in absolute evaluation.

A.4 Analysis on Other Modalities Pretraining

To analyze the various models for the other modalities pretraining, we evaluated the performance of the reference review generation task that generates corresponding reviews from images or a table. For evaluation, we used the data that were not used for training data: we left 10%10\% of the data for Yelp and 5%5\% for Amazon. We chose two comparison models: Untrained and Triplet. Untrained denotes the model that image encoder or table encoder keeps untrained. This option indicates the lower bound containing only the effect of the text decoder. Triplet denotes the triplet-based metric-learning model, based on Lee et al. (2018) and Vo and Hays (2016). For triplet (images or a table, reviews of positive entity, reviews of negative entities), we trained the image or table encoder based on the pretrained text encoder, by placing the image or table encoded representations close to the positive reviews representations and far from the negative reviews representations. Note that pretrained text encoder was not trained further.

Results on the other modalities pretraining are shown in Table 7. For each model, the pretrained decoder generated a review from image or table encoded representations. We measured the average ROUGE scores between the generated review and NN reference reviews. The first finding was that results of table outperformed those of image. It indicates that table has more helpful information for generating reference review. The second finding was that our method based on the text decoder outperformed the Triplet based on the text encoder. Especially, Triplet achieved very poor performance for image because it is hard to match MM images to NN reference reviews for metric learning. On the contrary, our method achieved much better performance by pivoting the text decoder. Triplet showed good performance on table because it is relatively easy to match 11 table to NN reference reviews; however, our method outperformed it. We conclude that our method lets the image and table encoder get proper representations to generate reference reviews regardless of the number of inputs.

A.5 Example Summaries

Table 8, 9 show sample summaries generated from our model and baseline models on Yelp and Amazon datasets. Full summaries from our model are available at https://bit.ly/3bR4yod.