Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space
Chunyuan Li, Xiang Gao, Yuan Li, Baolin Peng, Xiujun Li, Yizhe Zhang, Jianfeng Gao
Introduction
Pre-trained language models (PLMs) have substantially advanced the state-of-the-art across a variety of natural language processing (NLP) tasks Peters et al. (2018); Devlin et al. (2019); Yang et al. (2019); Radford et al. (2019); Liu et al. (2019); Keskar et al. (2019); Shoeybi et al. (2019). PLMs are often trained to predict words based on their context on massive text data, and the learned models can be fine-tuned to adapt to various downstream tasks.
PLMs can generally play two different roles: a generic encoder such as BERT Devlin et al. (2019) to provide contextualized representations for language understanding tasks, and a powerful decoder such as GPT-2 Radford et al. (2019) to generate text sequences in an auto-regressive manner. In a bid to combine language understanding and generation tasks in one unified framework, several model variants have been proposed, including UniLM Dong et al. (2019), BART Lewis et al. (2019), and T5 Raffel et al. (2019). Although significant performance improvement has been reported on a wide range of NLP tasks, these models lack of explicit modeling of structures in a compact latent space, rendering it difficult to control language generation/representation from an abstract level.
Variational Autoencoders (VAEs) Kingma and Welling (2013); Rezende et al. (2014) provide a tractable method to train latent-variable generative models. In NLP, latent variables may assume the role of higher-level sentence representations, which govern a lower-level word-by-word generation process, thus facilitating controlled text generation Bowman et al. (2016); Hu et al. (2017). By representing sentences in a low-dimensional latent space, VAEs allow easy manipulation of sentences using the corresponding compact vector representations, such as feature regularization specified by prior distributions, and guided sentence generation with interpretable vector operators. Despite the attractive theoretical strengths, the current language VAEs are often built with shallow network architectures, such as two-layer LSTMs Hochreiter and Schmidhuber (1997). This limits the model’s capacity and leads to sub-optimal performance.
In this paper, we propose Optimus, the first large-scale pre-trained deep latent variable models for natural language. Optimus is pre-trained using the sentence-level (variational) auto-encoder objectives on large text corpus. This leads to a universal latent space to organize sentences (hence named Optimus). Optimus enjoys several favorable properties: It combines the strengths of VAE, BERT and GPT, and supports both natural language understanding and generation tasks. Comparing to BERT, Optimus learns a more structured semantic space due to the use of the prior distribution in training. As a result, the language representations learned by Optimus are more universal / general in that they can be more easily adapted to a new domain/task. Different from GPT-2, which generates human-like text but may lack effective means of controlling its high-level semantics (such as tense, topics, sentiment), Optimus can be easily deployed for guided text generation. The effectiveness of Optimus has been demonstrated with extensive experiments on language modeling, dialog response generation, text style transfer and low-resource language understanding. It achieves lower perplexity than GPT-2 on standard benchmarks, produces strong performance on guided text generation, and improves BERT on feature-based language understanding tasks. The code and pre-trained models are released on Githubhttps://github.com/ChunyuanLI/Optimus.
Along the way to build the first big VAE language model, there are several technical contributions/implications that are novel: Latent vector injection: this work demonstrates two schemes to discuss how to effectively inject conditioning vectors into GPT-2 without re-training it. The design idea to combine BERT/GPT-2 serves as a practical recipe to inspire people to integrate and reuse existing PLMs for larger and complex models. Pre-training on massive datasets itself is an effective approach to reduce KL vanishing, as demonstrated by the state of-the-art performance on four VAE language modeling datasets. The proof of VAE objective from the lens of IB, showing that VAE is a principled approach to balance the compactness and usability of learned representations. Improved performance on several language tasks shows the importance and necessity of pre-training a latent space.
Related Work
Large-scale Transformer-based PLMs have recently achieved state-of-the-art performance on various natural language understanding and generation tasks Devlin et al. (2019); Yang et al. (2019); Radford et al. (2019); Liu et al. (2019); Keskar et al. (2019). Prior to Transformer-based PLMs, non-generative methods have seen some early success in pre-training sequence models for supervised downstream tasks including standard sequence auto-encoders Dai and Le (2015); Li et al. (2015), skip-thought models Kiros et al. (2015) and paragraph vector models Le and Mikolov (2014) etc. However, all of these models do not generally learn a smooth, interpretable feature space for sentence encoding, or generating novel sentences. In this work, we aim to fill the gap to learn such a universal latent space in the field of Transformer-based PLMs.
Language VAEs have inspired new applications in NLP, via exploiting many interesting properties of the model’s latent space Bowman et al. (2016); Kim et al. (2018b). Its modeling capacity and empirical performance is somewhat limited, partially due to the KL vanishing issue described in Section 4.3. Several attempts have been made to alleviate this issue, including different KL annealing/thresholding schemes Bowman et al. (2016); Fu et al. (2019); Higgins et al. (2017); Li et al. (2019), decoder architectures Yang et al. (2017); Dieng et al. (2018), auxiliary loss Zhao et al. (2017), semi-amortized inference Kim et al. (2018a), aggressive encoder training schedule He et al. (2019), batch normalized inference Zhu et al. (2020) and flexible posterior Fang et al. (2019). Subramanian et al. (2018) have shown some promise that general encoder can benefit language generation. Transformers Vaswani et al. (2017) are recently considered in VAEs for classification Gururangan et al. (2019) and storytelling Wang and Wan (2019). Pre-training VAEs has been recently considered in conditional text generation to amortize the training of decoders and to allow easy adaptation in new generation tasks Duan et al. (2019).
All these efforts utilize simple LSTM Hochreiter and Schmidhuber (1997) and shallow Transformer Vaswani et al. (2017) architectures, thus with limited capacity. Our paper is the first big VAE model at the same scale of recent PLMs such as BERT and GPT-2. More importantly, we show that pre-training a meaningful latent space on a large text corpus can largely reduce the KL vanishing issue, and lead to new state-of-the-art performance.
Background on NLMs & GPT-2
To generate a text sequence of length , , neural language models (NLM) Mikolov et al. (2010) generate every token conditioned on the previous word tokens:
where indicates all tokens before , and is the model parameter. In NLMs, each one-step-ahead conditional in (1) is modeled by an expressive family of neural networks, and is typically trained via maximum likelihood estimate (MLE). Perhaps the most well-known NLM instance is GPT-2 Radford et al. (2019), which employs Transformers Vaswani et al. (2017) for each conditional, and is learned on a huge amount of OpenWeb text corpus. GPT-2 has shown surprisingly realistic text generation results, and low perplexity on several benchmarks. GPT-3 Brown et al. (2020) was recently proposed to further scale up NLMs to 175 billion parameters, showing impressive results on few-shot learning on multiple language tasks.
However, the only source of variation in NLMs, GPT2 and GPT3 is modeled in the conditionals at every step: the text generation process only depends on previous word tokens, and there is limited capacity for the generation to be guided by the higher-level structures that are likely presented in natural language, such as tense, topics or sentiment.
Pre-trained Latent Space Modeling
To facilitate high-level guidance in sentence generation, Optimus organizes sentences in a universal latent (or semantic) space, via pre-training on large text corpora. Each sample in this space can be interpreted as outlines of the corresponding sentences, guiding the language generation process performed in the symbolic space Subramanian et al. (2018). This naturally fits within the learning paradigm of latent variable models such as VAEs Kingma and Welling (2013); Bowman et al. (2016), where the latent representations capture the high-level semantics/patterns. It consists of two parts, generation and inference, enabling a bidirectional mapping between the latent space and symbolic space.
The generative model (decoder) draws a latent vector from the continuous latent space with prior , and generates the text sequence from a conditional distribution ; is typically assumed a multivariate Gaussian, and represents the neural network parameters. The following auto-regressive decoding process is usually used:
Intuitively, VAE provides a “hierachical” generation procedure: determines the high-level semantics, followed by (2) to produce the output sentences with low-level syntactic and lexical details. This contrasts with (1) in the explicit dependency on .
Similar to GPT-2, parameters are typically learned by maximizing the marginal log likelihood . However, this marginal term is intractable to compute for many decoder choices. Thus, variational inference is considered, and the true posterior is approximated via the variational distribution is (often known as the inference model or encoder), implemented via a -parameterized neural network. It yields the evidence lower bound objective (ELBO):
Typically, is modeled as a Gaussian distribution, and the re-parametrization trick is used for efficient learning Kingma and Welling (2013).
There is an alternative interpretation of the ELBO: the VAE objective can be viewed as a regularized version of the autoencoder (AE) Goodfellow et al. (2016). It is thus natural to extend the negative of in (3) by introducing a hyper-parameter to control the strength of regularization:
where is the reconstruction error (or negative log-likelihood (NLL)), and is a KL regularizer. The cost function provides a unified perspective for understanding various autoencoder variants and training methods. We consider two types of latent space with the following objectives:
AE. Only is considered (), while the Gaussian sampling in remains. In other words, the regularization is removed, and a point-estimate is likely to be learned to represent the text sequence’s latent feature. Note our reconstruction is on sentence-level, while other PLMs Devlin et al. (2019); Yang et al. (2019) employ masked LM loss, performing token-level reconstruction.
VAE. The full VAE objective is considered (). It tends to learn a smooth latent space due to .
From an information theory perspective, information bottleneck (IB) provides a principled approach to find the trade-off between predictive power and complexity (compactness) when summarizing observed data in learned representations. We show that our Optimus pre-training objectives effectively practice the IB principle as follows.
The objective in (4) shows the -VAE loss for one single sentence . The training objective over the dataset can be written as:
2 Model Architectures
The model architecture of Optimus is composed of multi-layer Transformer-based encoder and decoder, based on the original implementation described in Vaswani et al. (2017). The overall architecture is illustrated in Figure 1. To leverage the expressiveness power of existing PLMs, we initialize our encoder and decoder with weights of BERT and GPT-2 , respectively. This procedure is seamless, as all of these models are trained in a self-supervised/unsupervised manner.
We denote the number of layers (i.e., Transformer blocks) as , the hidden size as , and the number of self-attention heads as . Specifically, we consider BERT (L=12, H=768, A=12, Total Parameters=110M) and GPT-2 (L=12, H=768, A=12, Total Parameters=117M). We hope that our approach can provide a practical recipe to inspire future work to integrate larger pre-trained encoder and decoder for higher performance models.
Two technical questions remain, when pre-training Optimus from BERT & GPT-2: How to represent sentences, since the two PLMs employ different tokenization schemes? How to adapt a pre-trained GPT-2 to arbitrary conditional input without re-training the model again? Controllable GPT-2 models have been studied in Keskar et al. (2019); Zellers et al. (2019); Peng et al. (2020a, b) when prescribed control codes/tokens are provided, but it is still unknown how to ground GPT-2 to arbitrary conditional inputs.
In BERT, WordPiece Embeddings (WPE) is used for tokenization (vocabulary size is 28996 for the cased version). In GPT-2, the modified Byte Pair Encoding (BPE) Radford et al. (2019) is used for tokenization (vocabulary size is 50260). A given token is represented as , by summing the corresponding token, position and segment embeddings Optimus does not require segment embeddings, but we remain it due to BERT initialization.. For a sentence, we present it in both types of tokenization: the input of encoder is WPE, and the output of decoder is BPE to compute the reconstruction loss.
We study their empirical performance in Section B.1 of Appendix, and observe that Memory is significantly more effective than Embedding, and the integration of both schemes yields slightly better results. We hypothesize that the reason why Memory is superior is because it allows the decoder to attend the latent information at every layer of the network directly, while the Embedding method only allows the decoder to see the latent information at the input and output layer. In our experiments, we use the integration scheme by default. In summary, the encoder parameters , and decoder parameters .
3 Learning Procedures
We train the model parameters using two objectives: AE and VAE, discussed in Section 4.1. Pre-training AE using (5) is straightforward. However, pre-training VAE can be challenging due to the notorious KL vanishing issue Bowman et al. (2016), where an encoder that produces posteriors almost identical to the Gaussian prior for all sentences (rather than a more interesting posterior); and a decoder that completely ignores in (2), and a learned model that reduces to a simpler NLM.
To reduce this issue, we follow the intuition that if the encoder is providing useful information from the beginning of decoder training, the decoder is more likely to make use of Fu et al. (2019); He et al. (2019). Specifically, we use the cyclical schedule to anneal for 10 periods Fu et al. (2019). Within one period, there are three consecutive stages: Training AE () for 0.5 proportion, annealing from 0 to 1 for 0.25 proportion, and fixing for 0.25 proportion. When , we use the KL thresholding scheme Li et al. (2019); Kingma et al. (2016), and replace the KL term in (6) with a hinge loss term that maxes each component of the original KL with a constant :
Here, denotes the th dimension of . Using the thresholding objective causes learning to give up driving down KL for dimensions of that are already beneath the target compression rate.
The pre-training procedure largely follows the existing literature on language model pre-training. We use English Wikipedia to pre-train our AE and VAE objectives. As our main interest is to model sentences (rather than text sequences of a fixed length), we pre-process Wikipedia with maximum sentences length 64. It leads to 1990K sentences, which accounts 96.45% Wikipedia sentences used in BERT. More data pre-processing details are in Section B.2 of Appendix.
Experimental Results
We consider to apply the pre-trained Optimus models to three types of downstream tasks: language modeling, where Optimus is compared with SoTA VAE methods and GPT-2. Guided language generation, where Optimus shows its unique advantage in producing controllable sentences in contrast to GPT-2. Low-resource language understanding, where the learned structured latent features can be used for fast adaptation in new tasks.
Fine-tuning LM on new datasets is straightforward. We load the pre-trained Optimus, and update the model with one additional scheduling cycle for one epoch. The semantic latent vectors are first pre-trained off-the-shelf, and then easily leveraged to train the decoder on downstream datasets. From this perspective, our pre-training can be viewed as an effective approach to reduce KL vanishing.
We consider four datasets: the Penn Treebank () Marcus et al. (1993), Bowman et al. (2015), , and corpora Yang et al. (2017); He et al. (2019).
There are two types of metrics to evaluate language VAEs. Generation capability: we use perplexity (PPL). Note that NLM and GPT-2 has exactly PPL, while VAEs does not. Following He et al. (2019), we use the importance weighted bound in Burda et al. (2015) to approximate , and report PPL. Representation learning capability: Active units (AU) of and its Mutual Information (MI) with . We report the full results with ELBO, KL and Reconstruction in Appendix, but note that higher ELBO does not necessarily yield better language modeling.
GPT-2. A large-scale LM trained on OpoenWebText Radford et al. (2019). We load the pre-trained GPT-2 weights, and refine the model for 1 epoch on the new datasets. Annealing. is gradually annealed from 0 to 1. This annealing procedure can be used once (M.A.) Bowman et al. (2016) or multiple times (C.A.) Fu et al. (2019). Aggressive Training He et al. (2019). Training the encoder multiple times per decoder update. AE-FB Li et al. (2019). Training AE, and then VAE using the KL thresholding in (9), the results on are reported as a good trade-off.
The results are shown in Table 1. Various values are used, we observe a trade-off between language modeling and representation learning, controlled by . Compared with existing VAE methods, Optimus achieve significantly lower perplexity, and higher MI/AU. This indicates that our pre-training method is an effective approach to reduce KL vanishing issue and training VAEs, especially given the fact that we only fine-tune on these datasets for one epoch. Optimus achieves lower perplexity compared with GPT-2 on three out of four datasets. Intuitively, this is because the model can leverage the prior language knowledge encoded in . This gap is larger, when the sentences in the dataset exhibit common regularities, such as , where the prior plays a more important/effective role in this scenario. Though the form of our model is simple, Optimus shows stronger empirical performance than sophisticated models that are particularly designed for long-text, such as hVAE in Shen et al. (2019). For example, the KL and PPL of Optimus (15.09 and 22.79) are much better than hVAE (6.8 and 45.8) on Yelp dataset. This verifies the importance of pre-training a latent space. The full experimental results are shown in Table 8, 9, 10 and 11 of Appendix.
2 Guided Language Generation
Different from the traditional NLMs or GPT-2, VAEs learns bidirectional mappings between the latent and symbolic space. It enables high-level sentence editing as arithmetic latent vector operations, and thus allows guided language generation. The reason that Optimus supports arithmetic operations are two-fold: (1) Pre-training on large datasets with large networks allows all sentences to be densely and faithfully represented in the latent space. (2) The continuity property of neural nets and KL regularization of VAE encourage latent vectors with similar semantics are smoothly organized together.
This is demonstrated with two simple schemes to manipulate pre-trained latent spaces: sentence transfer and interpolation, with results in Table 2 and Table 3, respectively. Details and more results are shown in Appendix. They showcase that Optimus enables new ways that one can play with language generation using pre-trained models, compared with GPT-2 that can only fulfill text sequences with given prompts. A website demohttp://aka.ms/optimus is released to the public to interact with the model, exhibiting the power of latent-vector-based controllable text generation. We demonstrate more sophisticated ways to manipulate pre-trained latent spaces in three real applications as follows.
The open-domain dialog response generation task is considered: generating responses given a dialog history . Following Gao et al. (2019a), we embed the history and response in a joint latent space as and , respectively. A fusion regularization is used to match the responses to the context. We consider Li et al. (2017c) used in Gu et al. (2019), which has 13,118 daily conversations. Each utterance is processed as the response of previous 10 context utterances from both speakers. The baseline methods are described in Appendix. We measure the performance using Bleu Chen and Cherry (2014), and compute the precision, recall and F1 in Table 4. Optimus shows higher Bleu scores than all existing baselines.
Following StyleFusion Gao et al. (2019b), we consider generating responses for in the style of Holmes. The comparison is shown in Table 5. In addition to Bleu, we use neural and N-gram classifier scores to evaluate the accuracy of the generated responses that belong to the desired style. Optimus achieves better performance on all metrics.
The short dataset collected in Shen et al. (2017) is used. It contains 444K training sentences, and we use separated datasets of 10K sentences for validation/testing, respectively. The goal is to generate text reviews given the positive/negative sentiment. We fine-tune Optimus using the VAE objective on the dataset, then freeze backbone weights. A conditional GAN Mirza and Osindero (2014) is trained on the fixed latent space. The generation process is to first produce a latent vector based on a given label using conditional GAN, then generate sentences conditioned on using the decoder. The baselines are described in Appendix. G-score computes the geometric mean of Accuracy and Bleu, measuring the comprehensive quality of both content and style. Self-Bleu measures the diversity of the generated sentences. The results are shown in Table 6, Optimus achieves the best performance on all metrics. This verifies the importance of learning a smooth and meaningful latent space. The conditional generated sentences are shown in Appendix.
3 Low-resource Language Understanding
A varying number of training samples are randomly chosen, ranging from 1 to 10K per class. 10 trials are used when the number of available training samples are small, each is trained in 100 training epochs. The results are shown in Figure 3. When pre-trained models are used to provide sentence embeddings, the proposed Optimus consistently outperforms BERT. It demonstrates that the latent structure learned by Optimus is more separated, and helps generalize better. When the entire network is fine-tuned, Optimus can adapt faster than BERT, when the available number of training samples is small. The two methods perform quite similarly when more training data is provided. This is because the pre-trained backbone network size is much larger than the classifier, where the performance is dominated by the backbone networks.
We use tSNE Maaten and Hinton (2008) to visualize the learned feature on a 2D map. The validation set of Yelp is used to extract the latent features. Compared with BERT, Optimus learns a smoother space and more structured latent patterns, which explains why Optimus can yield better classification performance and faster adaptation.
We further consider the GLUE benchmark Wang et al. (2019), which consists of nine datasets for general language understanding. Following the finetuning schedule in Devlin et al. (2019), we use learning rate and train the model for 3 epochs. We select the best performance among different runs. We show the results on the validation set in Table 7. With the feature-based scheme, Optimus yields higher performance than BERT, especially on the large datasets such as MNLI, QQP and QNLI. When the full models are fine-tuned, the two methods perform quite similarly.
In summary, the scenarios that Optimus fit the low-resource settings are two-fold: (1) The required computing resource is low: the feature-based approach only updates the classifier, whose computing requirement is much lower than full-model fine-tuning; (2) The number of required labelled data is low: when labelled data is rare, Optimus adapts better. The results confirm that Optimus can maintain and exploit the structures learned in pre-training, and presents a more general representation that can be adapted to new tasks more easily than BERT – feature-based adaption is much faster and easier to perform than fine-tuning.
Discussion
We present Optimus, a large-scale pre-trained deep latent variable model for natural language. It introduces a smooth and universal latent space, by combining the advantages of VAEs, BERT and GPT-2 in one model. Experimental results on a wide range of tasks and datasets have demonstrated the strong performance of Optimus, including new state-of-the-art for language VAEs.
There are several limitations in current Optimus. First, our pre-trained language VAE is still under-trained due to limited compute resource, as the training reconstruction loss can still decrease. One may further train the models with higher latent dimension and longer time to fully release the power of pre-trained latent spaces. Second, the current model can only control sentences of moderate length. One future direction is to consider more sophisticated mechanisms to gain stronger control-ability over longer sentences while maintaining the compactness of latent representations.
While deep generative models (DGMs) such as VAEs are theoretically attractive due to its principle nature, it is now rarely used by practitioners in the modern pre-trained language modeling era where BERT/GPT dominate with strong empirical performance. That’s why this paper makes a timely contribution to making DGMs practical for NLP. We hope that this paper will help renew interest in DGMs for this purpose. Hence, we deliberately keep a simple model, believing that the first pre-trained big VAE model itself and its implications are novel: it helps the community to recognize the importance of DGMs in the pre-training era, and revisit DGMs to make it more practical. Indeed, Optimus is uniquely positioned to learn a smooth latent space to organize sentences, which can enable guided language generation compared with GPT-2, and yield better generalization in low-resource language understanding tasks than BERT.
Acknowledgments
The authors gratefully acknowledge Jason Yosinski, Changyou Chen, Yang Zhao and Le Fang for helpful discussion. Additional thanks go to the entire Project Philly team inside Microsoft, who provided us the computing platform for our research. The implementation in our experiments depends on open source GitHub repositories; we acknowledge all the authors who made their code public, which tremendously accelerates our project progress.
References
Appendix A Information Bottleneck and VAEs
Tishby et al. (2000) presented the Information Bottleneck (IB) method via solving the Lagrange relaxation of the optimization problem:
In the following, we first show that the KL and reconstruction terms of VAE are the bounds of MI, respectively. Further, we put the bounds together, and show that VAE objective can optimize IB.
Following Makhzani et al. (2016), we refer to as the aggregated posterior. This marginal distribution captures the aggregated over the entire dataset. The KL term (6) in can be decomposed into two refined terms Chen et al. (2018); Hoffman and Johnson (2016):
where is the mutual information (MI) measured by . Higher MI can lead to a higher correlation between the latent variable and data variable, and encourages a reduction in the degree of KL vanishing. The marginal KL is represented by , and it measures the fitness of the aggregated posterior to the prior distribution.
The reconstruction term in (5) provides a lower bound for MI measured by , based on Corollary 3 in Li et al. (2017a):
When scheduled with , the training objective over the dataset can be written as:
This recovers IB principle in (10). When , we have the AE variant of our Optimus, the model fully focuses on maximizing the MI to recover sentence from the latent space. As increases, the model gradually transits towards fitting the aggregated latent codes to the given prior, leading the VAE variant of our Optimus.
Appendix B Pre-training Details
We compare three different schemes to inject latent vector into GPT2 in Figure 5:
Mem. Latent vector is used as additional memory token for GPT2 to attend.
Emb. Latent vector is used as additional embedding to add into other embeddings.
Mem+Emb. The integration of the above two schemes.
On both Yelp and PTB datasets, 5 training epochs are considered. Yelp generally has longer sentences than PTB. The encoder is initialized with BERT, and decoder is initialized with GPT-2. Lower reconstruction error per word indicates a more effective approach to pass the information flow from encoder to decoder. We see that it is significantly more efficient to use as a memory vector for GPT-2 to attend, than as the additional embedding. The combined scheme yields slightly better performance in the late stage of training. In the paper, we use the combined scheme in default.
B.2 Wikipedia Dataset
We illustrate the statistics of Wikipedia dataset in Figure 6. Since we focus on modeling natural sentences (rather than text sequences of a fixed length as in GPT-2 Radford et al. (2019)) in a latent space, we pre-process Wikipedia into a set of natural sentences, with maximum sequence length as 64. This leads to 1990K sentences, which is 96.45% of entire Wikipedia dataset.
Appendix C Experiment Details
In addition to generating high-quality sentences as in the traditional language models that only, VAEs also aim to learn a good posterior distribution in the latent space. The language modeling performance is evaluated with ELBO, perplexity (PPL) or importance weighted perplexity He et al. (2019), which provides a tighter bound to . Higher ELBO and lower PPL indicate the model fits the observed sentences better. The pre-training takes around 50 hours for one epoch on eight V100 DGX2 GPU’s.
ELBO: The sum of KL divergence and reconstruction loss.
Perplexity. , where is the number of words. For latent variable models, we use a lower bound on the marginal log-likelihood , as follows from Jensen’s Inequality and the fact that the average importance weights are an unbiased estimator of :
where .
More importantly, we are interested in the learned , which is evaluated using the following three metrics:
MI: The mutual information ;
The full experimental results on shown in Table 8, 9, 10 and 11.
C.2 Dialog response generation
We interpolate samples between the context and response as , where . We fix the first 11 layers of encoder, and fine-tune from last layer to : . An additional network path is introduced from the 11th layer of encoder to to represent context. The fine-tuning objective is:
where is the same with fusion term in Gao et al. (2019a), and .
We benchmark representative baselines and state-of-the-art approaches, including: Seq2Seq: a generalized sequence-to-sequence model with hierarchical RNN encoder Serban et al. (2016); SeqGAN: a GAN based model for sequence generation Li et al. (2017b); CVAE baseline Zhao et al. (2017); Dialogue WAE, a conditional Wasserstein auto-encoder for response generation Gu et al. (2019); : A hierarchical VAE model Serban et al. (2017). VHCR: a hierarchical VAE model with conversation modeling Park et al. (2018). iVAE: An implicit VAE model augmented with mutual information regularizer Fang et al. (2019). The full comparison in shown in Table 12.
In this task, the additional sentences are used to bias the generated response towards the reference style. The biased response representation is , where and is the latent representation of . The corresponding loss for the biased target is , which is added into for training.
Two type of Accuracy are reported, based on text sequence (i.e., neural) and its N-gram information. The accuracy is assessed by an oracle classifier to correctly predict whether generated response belongs the style-reference dataset.
C.2.1 Label-Conditional Text Generation
The goal of this task is to generate sentences conditioned on a given label. We consider a two-stage algorithm to adapt Optimus for this task. First, we fine-tune a VAE language model on the downstream dataset, and freeze the model parameters. In another word, the latent space is fixed. Second, we build a conditional GAN for the latent space. Let’s denote the latent vectors for ground-trurh sentences as . We build a generator to produce , where is the random noise, and is the label. A discriminator is trained simultaneously to distinguish and . The learning objectives for conditional GAN is:
To make the model work effectively, it is key to learn a smooth and meaningful latent space of target sentences. The text generation procedure conditioned on label is:
This mimics the process to produce the outlines of the sentences using conditional GAN, and fill in details using the decoder. We show some generated sentences in Table 20.
We compare with three baselines: (1) Ctrl-Gen Hu et al. (2017); We use their released code to reproduce the results. (2) ARAE Zhao et al. (2018) proposes to learn an auto-encoder first, and then train a GAN to produce the latent vectors. (3) NN-Outlines Subramanian et al. (2018) proposes the use of a general purpose encoder for text generation, and we implement it using BERT. Note that our two-stage fine-tuning scheme borrows the ideas from ARAE and NN-Outlines. The key difference is that we employ our pre-trained Optimus model, and work on a better latent space.
We consider three metrics: (1) Bleu for sentence quality, (2) Accuracy for conditional generation capability. The accuracy is assessed by an oracle classifier to correctly predict the attributes that generated sentences are conditioned on. (3) G-score is reported as the geometric mean of Accuracy and Bleu. This is the most important metric, as it evaluates the overall performance. For label-conditional text generation, Bleu of each generated sentence is computed by comparing with all sentences in the test set, as there are no source sentences. We further report Self-Bleu Zhu et al. (2018) to evaluate the diversity of generated sentences.
C.3 Latent space interpolation & arithmetic operation
The universal latent space learned by Optimus supports arithmetic operations. Given source sentence and target , the goal is to re-write the input sentence as output in analogy to the transition from to . We first encode into the latent vectors , respectively, then apply the arithmetic operator , and generate conditioned on . One example is shown in Table 2. Interestingly, we observe consistent style transfer from to , to analogize the relation from to . For example, the subject is revised from singular to plural forms, the topic changes from daily-life to sport. In another word, Optimus supports sentence arithmetic operator at the semantic level. More latent vector arithmetic operation examples are shown in Table 17, 18, 19.
One favorable property of VAEs is to provide a smooth space that captures sentence semantics. We demonstrate linear interpolating between latent vectors. We take two sentences and , and use their posterior mean as the latent features and , respectively. We interpolate a path with increased from 0 to 1 by a step size of 0.1. Table 3 shows generated sentences using greedy decoding conditioned on . The interpolated sentences exhibit smooth semantic evolution. More interpolation examples are shown in Appendix. Note that we have observed smooth & meaningful interpolation results for almost arbitrary input sentences pairs. This demonstrates the promise that Optimus learns a universal latent space. More latent space interpolation examples are shown in Table 13, 14, 15.
While Optimus shows the potentials of latent-vector-based controllable language generations, it has several limitations: (1) The compactness of latent vectors restricts the amount of encoded information, thus the model has difficulties in representing with long or complex sentences. This can be improved with more sophisticated design of latent space. (2) The model generates repeated interpolated sentences when intrinsic language variations are limited. (3) When doing interpolation, though the model knows the basic trend of numbers, it does not fully understand how to count numbers; For example, it jumps from one to five, then to twenty, instead of outputting the smoothly changing numbers such as one, five, ten, fifth, twenty.
For more user interaction with Optimus, we have released a demo website to allows users to input sentences, and the system will provide controllable generated sentences with arithmetic or interpolating operations.
C.4 Ablation study on VAE & AE objectives
We compare the interpolation examples in Table 16, and generally observe that VAE can produce smoother sentences interpolation results than AE. We compare the two pre-training objectives on the GLUE benchmark using the feature-based approach. The results are shown in Table 7. We see that both objectives outperform than BERT on large datasets, and VAE objective performs better than AE objective. This verifies the effectiveness of smooth regularization on the latent space for the classification performance.