Deconvolutional Paragraph Representation Learning

Yizhe Zhang, Dinghan Shen, Guoyin Wang, Zhe Gan, Ricardo Henao, Lawrence Carin

Introduction

A central task in natural language processing is to learn representations (features) for sentences or multi-sentence paragraphs. These representations are typically a required first step toward more applied tasks, such as sentiment analysis , machine translation , dialogue systems and text summarization . An approach for learning sentence representations from data is to leverage an encoder-decoder framework . In a standard autoencoding setup, a vector representation is first encoded from an embedding of an input sequence, then decoded to the original domain to reconstruct the input sequence. Recent advances in Recurrent Neural Networks (RNNs) , especially Long Short-Term Memory (LSTM) and variants , have achieved great success in numerous tasks that heavily rely on sentence-representation learning.

RNN-based methods typically model sentences recursively as a generative Markov process with hidden units, where the one-step-ahead word from an input sentence is generated by conditioning on previous words and hidden units, via emission and transition operators modeled as neural networks. In principle, the neural representations of input sequences aim to encapsulate sufficient information about their structure, to subsequently recover the original sentences via decoding. However, due to the recursive nature of the RNN, challenges exist for RNN-based strategies to fully encode a sentence into a vector representation. Typically, during training, the RNN generates words in sequence conditioning on previous ground-truth words, i.e., teacher forcing training , rather than decoding the whole sentence solely from the encoded representation vector. This teacher forcing strategy has proven important because it forces the output sequence of the RNN to stay close to the ground-truth sequence. However, allowing the decoder to access ground truth information when reconstructing the sequence weakens the encoder’s ability to produce self-contained representations, that carry enough information to steer the decoder through the decoding process without additional guidance. Aiming to solve this problem, proposed a scheduled sampling approach during training, which gradually shifts from learning via both latent representation and ground-truth signals to solely use the encoded latent representation. Unfortunately, showed that scheduled sampling is a fundamentally inconsistent training strategy, in that it produces largely unstable results in practice. As a result, training may fail to converge on occasion.

During inference, for which ground-truth sentences are not available, words ahead can only be generated by conditioning on previously generated words through the representation vector. Consequently, decoding error compounds proportional to the length of the sequence. This means that generated sentences quickly deviate from the ground-truth once an error has been made, and as the sentence progresses. This phenomenon was coined exposure bias in .

We propose a simple yet powerful purely convolutional framework for learning sentence representations. Conveniently, without RNNs in our framework, issues connected to teacher forcing training and exposure bias are not relevant. The proposed approach uses a Convolutional Neural Network (CNN) as encoder and a deconvolutional (i.e., transposed convolutional) neural network as decoder. To the best of our knowledge, the proposed framework is the first to force the encoded latent representation to capture information from the entire sentence via a multi-layer CNN specification, to achieve high reconstruction quality without leveraging RNN-based decoders. Our multi-layer CNN allows representation vectors to abstract information from the entire sentence, irrespective of order or length, making it an appealing choice for tasks involving long sentences or paragraphs. Further, since our framework does not involve recursive encoding or decoding, it can be very efficiently parallelized using convolution-specific Graphical Process Unit (GPU) primitives, yielding significant computational savings compared to RNN-based models.

Convolutional Auto-encoding for Text Modeling

After this first convolutional layer, we apply the convolution operation to the feature map, C(1){\bf C}^{(1)}, using the same filter size, hh, with this repeated in sequence for L−1L-1 layers. Each time, the length along the spatial coordinate is reduced to T(l+1)=⌊(T(l)−h)/r(l)+1⌋T^{(l+1)}=\lfloor(T^{(l)}-h)/r^{(l)}+1\rfloor, where r(l)r^{(l)} is the stride length, T(l)T^{(l)} is the spatial length, ll denotes the ll-th layer and ⌊⋅⌋\lfloor\cdot\rfloor is the floor function. For the final layer, LL, the feature map C(L−1){\bf C}^{(L-1)} is fed into a fully-connected layer, to produce the latent representation h{\boldsymbol{h}}. Implementation-wise, we use a convolutional layer with filter size equals to T(L−1)T^{(L-1)} (regardless of hh), which is equivalent to a fully-connected layer; this implementation trick has been also utilized in . This last layer summarizes all remaining spatial coordinates, T(L−1)T^{(L-1)}, into scalar features that encapsulate sentence sub-structures throughout the entire sentence characterized by filters, {Wc(i,l)}\{{{\bf W}}_{c}^{(i,l)}\} for i=1,…,pli=1,\ldots,p_{l} and l=1,…,Ll=1,\ldots,L, where Wc(i,l){{\bf W}}_{c}^{(i,l)} denotes filter ii for layer ll. This also implies that the extracted feature is of fixed-dimensionality, independent of the length of the input sentence.

Having pLp_{L} filters on the last layer, results in pLp_{L}-dimensional representation vector, h=C(L){\boldsymbol{h}}={\bf C}^{(L)}, for the input sentence. For example, in Figure 1, the encoder consists of L=3L=3 layers, which for a sentence of length T=60T=60, embedding dimension k=300k=300, stride lengths {r(1),r(2),r(3)}={2,2,1}\{r^{(1)},r^{(2)},r^{(3)}\}=\{2,2,1\}, filter sizes h={5,5,12}h=\{5,5,12\} and number of filters {p1,p2,p3}={300,600,500}\{p_{1},p_{2},p_{3}\}=\{300,600,500\}, results in intermediate feature maps, C(1){\bf C}^{(1)} and C(2){\bf C}^{(2)} of sizes {28×300,12×600}\{28\times 300,12\times 600\}, respectively. The last feature map of size 1×5001\times 500, corresponds to latent representation vector, h{\boldsymbol{h}}.

Conceptually, filters from the lower layers capture primitive sentence information (hh-grams, analogous to edges in images), while higher level filters capture more sophisticated linguistic features, such as semantic and syntactic structures (analogous to image elements). Such a bottom-up architecture models sentences by hierarchically stacking text segments (hh-grams) as building blocks for representation vector, h{\boldsymbol{h}}. This is similar in spirit to modeling linguistic grammar formalisms via concrete syntax trees , however, we do not pre-specify a tree structure based on some syntactic structure (i.e., English language), but rather abstract it from data via a multi-layer convolutional network.

2 Deconvolutional decoder

Denoting w^t\hat{w}^{t} as the tt-th word in reconstructed sentence s^\hat{s}, the probability of w^t\hat{w}^{t} to be word vv is specified as

3 Model learning

The objective of the convolutional autoencoder described above can be written as the word-wise log-likelihood for all sentences s∈Ds\in\mathcal{D}, i.e.,

where D\mathcal{D} denotes the set of observed sentences. The simple, maximum-likelihood objective in (2) is optimized via stochastic gradient descent. Details of the implementation are provided in the experiments. Note that (2) differs from prior related work in two ways: ii) use pooling and un-pooling operators, while we use convolution/deconvolution with stride; and iiii) more importantly, do not use a cosine similarity reconstruction as in (1), but a RNN-based decoder. A further discussion of related work is provided in Section 3. We could use pooling and un-pooling instead of striding (a particular case of deterministic pooling/un-pooling), however, in early experiments (not shown) we did not observe significant performance gains, while convolution/deconvolution operations with stride are considerably more efficient in terms of memory footprint. Compared to a standard LSTM-based RNN sequence autoencoders with roughly the same number of parameters, computations in our case are considerably faster (see experiments) using single NVIDIA TITAN X GPU. This is due to the high parallelization efficiency of CNNs via cuDNN primitives .

The proposed framework can be seen as a complementary building block for natural language modeling. Contrary to the standard LSTM-based decoder, the deconvolutional decoder imposes in general a less strict sequence dependency compared to RNN architectures. Specifically, generating a word from an RNN requires a vector of hidden units that recursively accumulate information from the entire sentence in an order-preserving manner (long-term dependencies are heavily down-weighted), while for a deconvolutional decoder, the generation only depends on a representation vector that encapsulates information from throughout the sentence without a pre-specified ordering structure. As a result, for language generation tasks, a RNN decoder will usually generate more coherent text, when compared to a deconvolutional decoder. On the contrary, a deconvolutional decoder is better at accounting for distant dependencies in long sentences, which can be very beneficial in feature extraction for classification and text summarization tasks.

4 Semi-supervised classification and summarization

Identifying related topics or sentiments, and abstracting (short) summaries from user generated content such as blogs or product reviews, has recently received significant interest . In many practical scenarios, unlabeled data are abundant, however, there are not many practical cases where the potential of such unlabeled data is fully realized. Motivated by this opportunity, here we seek to complement scarcer but more valuable labeled data, to improve the generalization ability of supervised models. By ingesting unlabeled data, the model can learn to abstract latent representations that capture the semantic meaning of all available sentences irrespective of whether or not they are labeled. This can be done prior to the supervised model training, as a two-step process. Recently, RNN-based methods exploiting this idea have been widely utilized and have achieved state-of-the-art performance in many tasks . Alternatively, one can learn the autoencoder and classifier jointly, by specifying a classification model whose input is the latent representation, h{\boldsymbol{h}}; see for instance .

In the case of product reviews, for example, each review may contain hundreds of words. This poses challenges when training RNN-based sequence encoders, in the sense that the RNN has to abstract information on-the-fly as it moves through the sentence, which often leads to loss of information, particularly in long sentences . Furthermore, the decoding process uses ground-truth information during training, thus the learned representation may not necessarily keep all information from the input text that is necessary for proper reconstruction, summarization or classification.

We consider applying our convolutional autoencoding framework to semi-supervised learning from long-sentences and paragraphs. Instead of pre-training a fully unsupervised model as in , we cast the semi-supervised task as a multi-task learning problem similar to , i.e., we simultaneously train a sequence autoencoder and a supervised model. In principle, by using this joint training strategy, the learned paragraph embedding vector will preserve both reconstruction and classification ability. Specifically, we consider the following objective:

where α>0\alpha>0 is an annealing parameter balancing the relative importance of supervised and unsupervised loss; Dl\mathcal{D}_{l} and Du\mathcal{D}_{u} denote the set of labeled and unlabeled data, respectively. The first term in (3) is the sequence autoencoder loss in (2) for the dd-th sequence. Lsup(⋅)\mathcal{L}^{\rm sup}(\cdot) is the supervision loss for the dd-th sequence (labeled only). The classifier function, f(⋅)f(\cdot), that attempts to reconstruct ydy_{d} from hd{\boldsymbol{h}}_{d} can be either a Multi-Layer Perceptron (MLP) in classification tasks, or a CNN/RNN in text summarization tasks. For the latter, we are interested in a purely convolutional specification, however, we also consider an RNN for comparison. For classification, we use a standard cross-entropy loss, and for text summarization we use either (2) for the CNN or the standard LSTM loss for the RNN.

In practice, we adopt a scheduled annealing strategy for α\alpha as in , rather than fixing it a priori as in . During training, \eqrefeq:semi\eqref{eq:semi} gradually transits from focusing solely on the unsupervised sequence autoencoder to the supervised task, by annealing α\alpha from 1 to a small positive value αmin\alpha_{\rm min}. We set αmin=0.01\alpha_{\rm min}=0.01 in the experiments. The motivation for this annealing strategy is to first focus on abstracting paragraph features, then to selectively refine learned features that are most informative to the supervised task.

Related Work

Previous work has considered leveraging CNNs as encoders for various natural language processing tasks . Typically, CNN-based encoder architectures apply a single convolution layer followed by a pooling layer, which essentially acts as a detector of specific classes of hh-grams, given a convolution filter window of size hh. The deep architecture in our framework will, in principle, enable the high-level layers to capture more sophisticated language features. We use convolutions with stride rather than pooling operators, e.g., max-pooling, for spatial downsampling following , where it is argued that fully convolutional architectures are able to learn their own spatial downsampling. Further, uses a 29-layer CNN for text classification. Our CNN encoder is considerably simpler in structure (convolutions with stride and no more than 4 layers) while still achieving good performance.

Language decoders other than RNNs are less well studied. Recently, proposed a hybrid model by coupling a convolutional-deconvolutional network with an RNN, where the RNN acts as decoder and the deconvolutional model as a bridge between the encoder (convolutional network) and decoder. Additionally, considered CNN variants, such as pixelCNN , for text generation. Nevertheless, to achieve good empirical results, these methods still require the sentences to be generated sequentially, conditioning on the ground truth historical information, akin to RNN-based decoders, thus still suffering from the exposure bias.

Other efforts have been made to improve embeddings from long paragraphs using unsupervised approaches . The paragraph vector learns a fixed length vector by concatenating it with a word2vec embedding of history sequence to predict future words. The hierarchical neural autoencoder builds a hierarchical attentive RNN, then it uses paragraph-level hidden units of that RNN as embedding. Our work differs from these approaches in that we force the sequence to be fully restored from the latent representation, without aid from any history information.

Previous methods have considered leveraging unlabeled data for semi-supervised sequence classification tasks. Typically, RNN-based methods consider either ii) training a sequence-to-sequence RNN autoencoder, or a RNN classifier that is robust to adversarial perturbation, as initialization for the encoder in the supervised model ; or, iiii) learning latent representation via a sequence-to-sequence RNN autoencoder, and then using them as inputs to a classifier that also takes features extracted from a CNN as inputs . For summarization tasks, has considered a semi-supervised approach based on support vector machines, however, so far, research on semi-supervised text summarization using deep models is scarce.

Experiments

For all the experiments, we use a 3-layer convolutional encoder followed by a 3-layer deconvolutional decoder (recall implementation details for the top layer). Filter size, stride and word embedding are set to h=5h=5, rl=2r^{l}=2, for l=1,…,3l=1,\ldots,3 and k=300k=300, respectively. The dimension of the latent representation vector varies for each experiment, thus is reported separately.

For notational convenience, we denote our convolutional-deconvolutional autoencoder as CNN-DCNN. In most comparisons, we also considered two standard autoencoders as baselines: aa) CNN-LSTM: CNN encoder coupled with LSTM decoder; and bb) LSTM-LSTM: LSTM encoder with LSTM decoder. An LSTM-DCNN configuration is not included because it yields similar performance to CNN-DCNN while being more computationally expensive. The complete experimental setup and baseline details is provided in the Supplementary Material (SM). CNN-DCNN has the least number of parameters. For example, using 500 as the dimension of h{\boldsymbol{h}} results in about 9, 13, 15 million total trainable parameters for CNN-DCNN, CNN-LSTM and LSTM-LSTM, respectively.

Paragraph reconstruction

We first investigate the performance of the proposed autoencoder in terms of learning representations that can preserve paragraph information. We adopt evaluation criteria from , i.e., ROUGE score and BLEU score , to measure the closeness of the reconstructed paragraph (model output) to the input paragraph. Briefly, ROUGE and BLEU scores measures the nn-gram recall and precision between the model outputs and the (ground-truth) references. We use BLEU-4, ROUGE-1, 2 in our evaluation, in alignment with . In addition to the CNN-LSTM and LSTM-LSTM autoencoder, we also compared with the hierarchical LSTM autoencoder . The comparison is performed on the Hotel Reviews datasets, following the experimental setup from , i.e., we only keep reviews with sentence length ranging from 50 to 250 words, resulting in 348,544 training data samples and 39,023 testing data samples. For all comparisons, we set the dimension of the latent representation to h=500{\boldsymbol{h}}=500.

From Table 1, we see that for long paragraphs, the LSTM decoder in CNN-LSTM and LSTM-LSTM suffers from heavy exposure bias issues. We further evaluate the performance of each model with different paragraph lengths. As shown in Figure 2 and Table 2, on this task CNN-DCNN demonstrates a clear advantage, meanwhile, as the length of the sentence increases, the comparative advantage becomes more substantial. For LSTM-based methods, the quality of the reconstruction deteriorates quickly as sequences get longer. In constrast, the reconstruction quality of CNN-DCNN is stable and consistent regardless of sentence length. Furthermore, the computational cost, evaluated as wall-clock, is significantly lower in CNN-DCNN. Roughly, CNN-LSTM is 3 times slower than CNN-DCNN, and LSTM-LSTM is 5 times slower on a single GPU. Details are reported in the SM.

Character-level and word-level correction

This task seeks to evaluate whether the deconvolutional decoder can overcome exposure bias, which severely limits LSTM-based decoders. We consider a denoising autoencoder where the input is tweaked slightly with certain modifications, while the model attempts to denoise (correct) the unknown modification, thus recover the original sentence.

For character-level correction, we consider the Yahoo! Answer dataset . The dataset description and setup for word-level correction is provided in the SM. We follow the experimental setup in for word-level and character-level spelling correction (see details in the SM). We considered substituting each word/character with a different one at random with probability η\eta, with η=0.30\eta=0.30. For character-level analysis, we first map all characters into a 40 dimensional embedding vector, with the network structure for word- and character-level models kept the same.

We employ Character Error Rate (CER) and Word Error Rate (WER) for evaluation. The WER/CER measure the ratio of Levenshtein distance (a.k.a., edit distance) between model predictions and the ground-truth, and the total length of sequence. Conceptually, lower WER/CER indicates better performance. We use LSTM-LSTM and CNN-LSTM denoising autoencoders for comparison. The architecture for the word-level baseline models is the same as in the previous experiment. For character-level correction, we set dimension of h{\boldsymbol{h}} to 900. We also compare to actor-critic training , following their experimental guidelines (see details in the SM).

As shown in Figure 3 and Table 3, we observed CNN-DCNN achieves both lower CER and faster convergence. Further, CNN-DCNN delivers stable denoising performance irrespective of the noise location within the sentence, as seen in Figure 4. For CNN-DCNN, even when an error is detected but not exactly corrected (darker colors in Figure 4 indicate higher uncertainty), denoising with future words is not effected, while for CNN-LSTM and LSTM-LSTM the error gradually accumulates with longer sequences, as expected.

For word-level correction, we consider word substitutions only, and mixed perturbations from three kinds: substitution, deletion and insertion. Generally, CNN-DCNN outperforms CNN-LSTM and LSTM-LSTM, and is faster. We provide experimental details and comparative results in the SM.

Semi-supervised sequence classification & summarization

We investigate whether our CNN-DCNN framework can improve upon supervised natural language tasks that leverage features learned from paragraphs. In principle, a good unsupervised feature extractor will improve the generalization ability in a semi-supervised learning setting. We evaluate our approach on three popular natural language tasks: sentiment analysis, paragraph topic prediction and text summarization. The first two tasks are essentially sequence classification, while summarization involves both language comprehension and language generation.

We consider three large-scale document classification datasets: DBPedia, Yahoo! Answers and Yelp Review Polarity . The partition of training, validation and test sets for all datasets follows the settings from . The detailed summary statistics of all datasets are shown in the SM. To demonstrate the advantage of incorporating the reconstruction objective into the training of text classifiers, we further evaluate our model with different amounts of labeled data (0.1%, 0.15%, 0.25%, 1%, 10% and 100%, respectively), and the whole training set as unlabeled data.

For our purely supervised baseline model (supervised CNN), we use the same convolutional encoder architecture described above, with a 500-dimensional latent representation dimension, followed by a MLP classifier with one hidden layer of 300 hidden units. The dropout rate is set to 50%. Word embeddings are initialized at random.

As shown in Table 4, the joint training strategy consistently and significantly outperforms the purely supervised strategy across datasets, even when all labels are available. We hypothesize that during the early phase of training, when reconstruction is emphasized, features from text fragments can be readily learned. As the training proceeds, the most discriminative text fragment features are selected. Further, the subset of features that are responsible for both reconstruction and discrimination presumably encapsulate longer dependency structure, compared to the features using a purely supervised strategy. Figure 5 demonstrates the behavior of our model in a semi-supervised setting on Yelp Review dataset. The results for Yahoo! Answer and DBpedia are provided in the SM.

For summarization, we used a dataset composed of 58,000 abstract-title pairs, from arXiv. Abstract-title pairs are selected if the length of the title and abstract do not exceed 50 and 500 words, respectively. We partitioned the training, validation and test sets into 55000, 2000, 1000 pairs each.

We train a sequence-to-sequence model to generate the title given the abstract, using a randomly selected subset of paired data with proportion σ=(5%,10%,50%,100%)\sigma=(5\%,10\%,50\%,100\%). For every value of σ\sigma, we considered both purely supervised summarization using just abstract-title pairs, and semi-supervised summarization, by leveraging additional abstracts without titles. We compared LSTM and deconvolutional network as the decoder for generating titles for σ=100%\sigma=100\%.

Table 5 summarizes quantitative results using ROUGE-L (longest common subsequence) . In general, the additional abstracts without titles improve the generalization ability on the test set. Interestingly, even when σ=100%\sigma=100\% (all titles are observed), the joint training objective still yields a better performance than using Lsup\mathcal{L}^{sup} alone. Presumably, since the joint training objective requires the latent representation to be capable of reconstructing the input paragraph, in addition to generating a title, the learned representation may better capture the entire structure (meaning) of the paragraph. We also empirically observed that titles generated under the joint training objective are more likely to use the words appearing in the corresponding paragraph (i.e., more extractive), while the the titles generated using the purely supervised objective Lsup\mathcal{L}^{sup}, tend to use wording more freely, thus more abstractive. One possible explanation is that, for the joint training strategy, since the reconstructed paragraph and title are all generated from latent representation h{\boldsymbol{h}}, the text fragments that are used for reconstructing the input paragraph are more likely to be leveraged when “building” the title, thus the title bears more resemblance to the input paragraph.

As expected, the titles produced by a deconvolutional decoder are less coherent than an LSTM decoder. Presumably, since each paragraph can be summarized with multiple plausible titles, the deconvolutional decoder may have trouble when positioning text segments. We provide discussions and titles generated under different setups in the SM. Designing a framework which takes the best of these two worlds, LSTM for generation and CNN for decoding, will be an interesting future direction.

Conclusion

We proposed a general framework for text modeling using purely convolutional and deconvolutional operations. The proposed method is free of sequential conditional generation, avoiding issues associated with exposure bias and teacher forcing training. Our approach enables the model to fully encapsulate a paragraph into a latent representation vector, which can be decompressed to reconstruct the original input sequence. Empirically, the proposed approach achieved excellent long paragraph reconstruction quality and outperforms existing algorithms on spelling correction, and semi-supervised sequence classification and summarization, with largely reduced computational cost.

References