Neural Variational Inference for Text Processing

Yishu Miao, Lei Yu, Phil Blunsom

Introduction

Probabilistic generative models underpin many successful applications within the field of natural language processing (NLP). Their popularity stems from their ability to use unlabelled data effectively, to incorporate abundant linguistic features, and to learn interpretable dependencies among data. However these successes are tempered by the fact that as the structure of such generative models becomes deeper and more complex, true Bayesian inference becomes intractable due to the high dimensional integrals required. Markov chain Monte Carlo (MCMC) (Neal, 1993; Andrieu et al., 2003) and variational inference (Jordan et al., 1999; Attias, 2000; Beal, 2003) are the standard approaches for approximating these integrals. However the computational cost of the former results in impractical training for the large and deep neural networks which are now fashionable, and the latter is conventionally confined due to the underestimation of posterior variance. The lack of effective and efficient inference methods hinders our ability to create highly expressive models of text, especially in the situation where the model is non-conjugate.

This paper introduces a neural variational framework for generative models of text, inspired by the variational auto-encoder (Rezende et al., 2014; Kingma & Welling, 2014). The principle idea is to build an inference network, implemented by a deep neural network conditioned on text, to approximate the intractable distributions over the latent variables. Instead of providing an analytic approximation, as in traditional variational Bayes, neural variational inference learns to model the posterior probability, thus endowing the model with strong generalisation abilities. Due to the flexibility of deep neural networks, the inference network is capable of learning complicated non-linear distributions and processing structured inputs such as word sequences. Inference networks can be designed as, but not restricted to, multilayer perceptrons (MLP), convolutional neural networks (CNN), and recurrent neural networks (RNN), approaches which are rarely used in conventional generative models. By using the reparameterisation method (Rezende et al., 2014; Kingma & Welling, 2014), the inference network is trained through back-propagating unbiased and low variance gradients w.r.t. the latent variables. Within this framework, we propose a Neural Variational Document Model (NVDM) for document modelling and a Neural Answer Selection Model (NASM) for question answering, a task that selects the sentences that correctly answer a factoid question from a set of candidate sentences.

The NVDM (Figure 2) is an unsupervised generative model of text which aims to extract a continuous semantic latent variable for each document. This model can be interpreted as a variational auto-encoder: an MLP encoder (inference network) compresses the bag-of-words document representation into a continuous latent distribution, and a softmax decoder (generative model) reconstructs the document by generating the words independently. A primary feature of NVDM is that each word is generated directly from a dense continuous document representation instead of the more common binary semantic vector (Hinton & Salakhutdinov, 2009; Larochelle & Lauly, 2012; Srivastava et al., 2013; Mnih & Gregor, 2014). Our experiments demonstrate that our neural document model achieves the state-of-the-art perplexities on the 20NewsGroups and RCV1-v2.

The NASM (Figure 2) is a supervised conditional model which imbues LSTMs (Hochreiter & Schmidhuber, 1997) with a latent stochastic attention mechanism to model the semantics of question-answer pairs and predict their relatedness. The attention model is designed to focus on the phrases of an answer that are strongly connected to the question semantics and is modelled by a latent distribution. This mechanism allows the model to deal with the ambiguity inherent in the task and learns pair-specific representations that are more effective at predicting answer matches, rather than independent embeddings of question and answer sentences. Bayesian inference provides a natural safeguard against overfitting, especially as the training sets available for this task are small. The experiments show that the LSTM with a latent stochastic attention mechanism learns an effective attention model and outperforms both previously published results, and our own strong non-stochastic attention baselines.

In summary, we demonstrate the effectiveness of neural variational inference for text processing on two diverse tasks. These models are simple, expressive and can be trained efficiently with the highly scalable stochastic gradient back-propagation. Our neural variational framework is suitable for both unsupervised and supervised learning tasks, and can be generalised to incorporate any type of neural networks.

Neural Variational Inference Framework

Latent variable modelling is popular in many NLP problems, but it is non-trivial to carry out effective and efficient inference for models with complex and deep structure. In this section we introduce a generic neural variational inference framework that we apply to both the unsupervised NVDM and supervised NASM in the follow sections.

We define a generative model with a latent variable h\boldsymbol{h}, which can be considered as the stochastic units in deep neural networks. We designate the observed parent and child nodes of h\boldsymbol{h} as x\boldsymbol{x} and y\boldsymbol{y} respectively. Hence, the joint distribution of the generative model is pθ(x,y)=∑hpθ(y∣h)pθ(h∣x)p(x)p_{\theta}(\boldsymbol{x},\boldsymbol{y})=\sum_{\boldsymbol{h}}p_{\theta}(\boldsymbol{y}|\boldsymbol{h})p_{\theta}(\boldsymbol{h}|\boldsymbol{x})p(\boldsymbol{x}), and the variational lower bound L\mathcal{L} is derived as:

Construct vector representations of the observed variables: u=fx(x)\boldsymbol{u}=f_{x}(\boldsymbol{x}), v=fy(y)\boldsymbol{v}=f_{y}(\boldsymbol{y}).

Assemble a joint representation: π=g(u,v)\boldsymbol{\pi}=g(\boldsymbol{u},\boldsymbol{v}).

Parameterise the variational distribution over the latent variable: μ=l1(π),log⁡σ=l2(π)\boldsymbol{\mu}=l_{1}(\boldsymbol{\pi}),\log\boldsymbol{\sigma}=l_{2}(\boldsymbol{\pi}).

fx(⋅)f_{x}(\cdot) and fy(⋅)f_{y}(\cdot) can be any type of deep neural networks that are suitable for the observed data; g(⋅)g(\cdot) is an MLP that concatenates the vector representations of the conditioning variables; l(⋅)l(\cdot) is a linear transformation which outputs the parameters of the Gaussian distribution. By sampling from the variational distribution, h∼qϕ(h∣x,y)\boldsymbol{h}\sim q_{\phi}(\boldsymbol{h}|\boldsymbol{\boldsymbol{x}},\boldsymbol{y}), we are able to carry out stochastic back-propagation to optimise the lower bound (Eq. 1).

During training, the model parameters θ\theta together with the inference network parameters ϕ\phi are updated by stochastic back-propagation based on the samples h\boldsymbol{h} drawn from qϕ(h∣x,y)q_{\phi}(\boldsymbol{h}|\boldsymbol{\boldsymbol{x}},\boldsymbol{y}). For the gradients w.r.t. θ\theta, we have the form:

For the gradients w.r.t. ϕ\phi we reparameterise h=μ+σ⋅ϵ\boldsymbol{h}=\boldsymbol{\mu}+\boldsymbol{\sigma}\cdot\boldsymbol{\epsilon} and sample ϵ(l)∼N(0,I)\boldsymbol{\epsilon}^{(l)}\sim\mathcal{N}(0,\boldsymbol{I}) to reduce the variance in stochastic estimation (Rezende et al., 2014; Kingma & Welling, 2014). The update of ϕ\phi can be carried out by back-propagating the gradients w.r.t. μ\boldsymbol{\mu} and σ\boldsymbol{\sigma}:

It is worth mentioning that unsupervised learning is a special case of the neural variational framework where h\boldsymbol{h} has no parent node x\boldsymbol{x}. In that case h\boldsymbol{h} is directly drawn from the prior p(h)p(\boldsymbol{h}) instead of the conditional distribution pθ(h∣x)p_{\theta}(\boldsymbol{h}|\boldsymbol{x}), and s(h)=log⁡pθ(y∣h)pθ(h)−log⁡qϕ(h∣y)s(\boldsymbol{h})=\log p_{\theta}(\boldsymbol{y}|\boldsymbol{h})p_{\theta}(\boldsymbol{h})-\log q_{\phi}(\boldsymbol{h}|\boldsymbol{y}).

Here we only discuss the scenario where the latent variables are continuous and the parameterised diagonal Gaussian is employed as the variational distribution. However the framework is also suitable for discrete units, and the only modification needed is to replace the Gaussian with a multinomial parameterised by the outputs of a softmax function. Though the reparameterisation trick for continuous variables is not applicable for this case, a policy gradient approach (Mnih & Gregor, 2014) can help to alleviate the high variance problem during stochastic estimation. (Kingma et al., 2014) proposed a variational inference framework for semi-supervised learning, but the prior distribution over the hidden variable p(h)p(\boldsymbol{h}) remains as the standard Gaussian prior, while we apply a conditional parameterised Gaussian distribution, which is jointly learned with the variational distribution.

Neural Variational Document Model

As an unsupervised generative model, we could interpret NVDM as a variational autoencoder: an MLP encoder q(h∣X)q(\boldsymbol{h}|\boldsymbol{X}) compresses document representations into continuous hidden vectors (X→h\boldsymbol{X}\to\boldsymbol{h}); a softmax decoder p(X∣h)=∏i=1Np(xi∣h)p(\boldsymbol{X}|\boldsymbol{h})=\prod_{i=1}^{N}p(\boldsymbol{x}_{i}|\boldsymbol{h}) reconstructs the documents by independently generating the words (h→{xi}\boldsymbol{h}\to\{\boldsymbol{x}_{i}\}). To maximise the log-likelihood log⁡∑hp(X∣h)p(h)\log\sum\nolimits_{\boldsymbol{h}}p(\boldsymbol{X}|\boldsymbol{h})p(\boldsymbol{h}) of documents, we derive the lower bound:

where NN is the number of words in the document and p(h)p(\boldsymbol{h}) is a Gaussian prior for h\boldsymbol{h}. Here, we consider NN is observed for all the documents. The conditional probability over words pθ(xi∣h)p_{\theta}(\boldsymbol{x}_{i}|\boldsymbol{h}) (decoder) is modelled by multinomial logistic regression and shared across documents:

As there is no supervision information for the latent semantics, h\boldsymbol{h}, the posterior approximation qϕ(h∣X)q_{\phi}(\boldsymbol{h}|\boldsymbol{X}) is only conditioned on the current document X\boldsymbol{X}. The inference network qϕ(h∣X)=N(h∣μ(X),diag(σ2(X)))q_{\phi}(\boldsymbol{h}|\boldsymbol{X})=\mathcal{N}(\boldsymbol{h}|\boldsymbol{\mu}(\boldsymbol{X}),diag(\boldsymbol{\sigma}^{2}(\boldsymbol{X}))) is modelled as:

For each document X\boldsymbol{X}, the neural network generates its own parameters μ\boldsymbol{\mu} and σ\boldsymbol{\sigma} that parameterise the latent distribution over document semantics h\boldsymbol{h}. Based on the samples h∼qϕ(h∣X)\boldsymbol{h}\sim q_{\phi}(\boldsymbol{h}|\boldsymbol{X}), the lower bound (Eq. 5) can be optimised by back-propagating the stochastic gradients w.r.t. θ\theta and ϕ\phi.

Since p(h)p(\boldsymbol{h}) is a standard Gaussian prior, the Gaussian KL-Divergence DKL⁡[qϕ(h∣X)∥p(h)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{X})\|p(\boldsymbol{h})] can be computed analytically to further lower the variance of the gradients. Moreover, it also acts as a regulariser for updating the parameters of the inference network qϕ(h∣X)q_{\phi}(\boldsymbol{h}|\boldsymbol{X}).

Neural Answer Selection Model

Answer sentence selection is a question answering paradigm where a model must identify the correct sentences answering a factual question from a set of candidate sentences. Assume a question q\boldsymbol{q} is associated with a set of answer sentences {a1,a2,...,an}\{\boldsymbol{a}_{1},\boldsymbol{a}_{2},...,\boldsymbol{a}_{n}\}, together with their judgements {y1,y2,...,yn}\{\boldsymbol{y}_{1},\boldsymbol{y}_{2},...,\boldsymbol{y}_{n}\}, where ym=1\boldsymbol{y}_{m}=1 if the answer am\boldsymbol{a}_{m} is correct and ym=0\boldsymbol{y}_{m}=0 otherwise. This is a classification task where we treat each training data point as a triple (q,a,y)(\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}) while predicting y\boldsymbol{y} for the unlabelled question-answer pair (q,a)(\boldsymbol{q},\boldsymbol{a}).

The Neural Answer Selection Model (Figure 2) is a supervised model that learns the question and answer representations and predicts their relatedness. It employs two different LSTMs to embed raw question inputs q\boldsymbol{q} and answer inputs a\boldsymbol{a}. Let sq(j)\boldsymbol{s}_{q}(j) and sa(i)\boldsymbol{s}_{a}(i) be the state outputs of the two LSTMs, and ii, jj be the positions of the states. Conventionally, the last state outputs sq(∣q∣)\boldsymbol{s}_{q}(|\boldsymbol{q}|) and sa(∣a∣)\boldsymbol{s}_{a}(|\boldsymbol{a}|), as the independent question and answer representations, can be used for relatedness prediction. In NASM, however, we aim to learn pair-specific representations through a latent attention mechanism, which is more effective for pair relatedness prediction.

In this model, the conditional distribution pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) is:

For each question q\boldsymbol{q}, the neural network generates the corresponding parameters μ\boldsymbol{\mu} and σ\boldsymbol{\sigma} that parameterise the latent distribution over question semantics h\boldsymbol{h}. Following Bahdanau et al. (2015), the attention model is defined as:

where α(i)\alpha(i) is the normalised attention score at answer token ii, and the context vector c(a,h)\boldsymbol{c}(\boldsymbol{a},\boldsymbol{h}) is the weighted sum of all the state outputs sa(i)\boldsymbol{s}_{a}(i). We adopt zq(q),za(a,h)\boldsymbol{z}_{q}(\boldsymbol{q}),\boldsymbol{z}_{a}(\boldsymbol{a},\boldsymbol{h}) as the question and answer representations for predicting their relatedness y\boldsymbol{y}. zq(q)\boldsymbol{z}_{q}(\boldsymbol{q}) is a deterministic vector that is equal to sq(∣q∣)\boldsymbol{s}_{q}(|\boldsymbol{q}|), while za(a,h)\boldsymbol{z}_{a}(\boldsymbol{a},\boldsymbol{h}) is a combination of the sequence output sa(∣a∣)\boldsymbol{s}_{a}(|\boldsymbol{\boldsymbol{a}}|) and the context vector c(a,h)\boldsymbol{c}(\boldsymbol{a},\boldsymbol{h}) (Eq. 14). For the prediction of pair relatedness y\boldsymbol{y}, we model the conditional probability distribution pθ(y∣zq,za)p_{\theta}(\boldsymbol{y}|\boldsymbol{z}_{q},\boldsymbol{z}_{a}) by sigmoid function:

To maximise the log-likelihood log⁡p(y∣q,a)\log p(\boldsymbol{y}|\boldsymbol{q},\boldsymbol{a}) we use the variational lower bound:

where q\boldsymbol{q} and a\boldsymbol{a} are also modelled by LSTMsIn this case, the LSTMs for q\boldsymbol{q} and a\boldsymbol{a} are shared by the inference network and the generative model, but there is no restriction on using different LSTMs in the inference network., and the relatedness label y\boldsymbol{y} is modelled by a simple linear transformation into the vector sy\boldsymbol{s}_{y}. According to the joint representation πϕ\boldsymbol{\pi}_{\phi}, we then generate the parameters μϕ\boldsymbol{\mu}_{\phi} and σϕ\boldsymbol{\sigma}_{\phi}, which parameterise the variational distribution over the question semantics h\boldsymbol{h}. To emphasise, though both pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) and qϕ(h∣q,a,y)q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}) are modelled as parameterised Gaussian distributions, qϕ(h∣q,a,y)q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}) as an approximation only functions during inference by producing samples to compute the stochastic gradients, while pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) is the generative distribution that generates the samples for predicting the question-answer relatedness y\boldsymbol{y}.

Based on the samples h∼qϕ(h∣q,a,y)\boldsymbol{h}\sim q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}), we use SGVB to optimise the lower bound (Eq.16). The model parameters θ\theta and the inference network parameters ϕ\phi are updated jointly using their stochastic gradients. In this case, similar to the NVDM, the Gaussian KL divergence DKL⁡[qϕ(h∣q,a,y))∥pθ(h∣q)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}))\|p_{\theta}(\boldsymbol{h}|\boldsymbol{q})] can be analytically computed during training process.

Experiments

We experiment with NVDM on two standard news corpora: the 20NewsGroupshttp://qwone.com/ jason/20Newsgroups and the Reuters RCV1-v2http://trec.nist.gov/data/reuters/reuters.html. The former is a collection of newsgroup documents, consisting of 11,314 training and 7,531 test articles. The latter is a large collection from Reuters newswire stories with 794,414 training and 10,000 test cases. The vocabulary size of these two datasets are set as 2,000 and 10,000.

To make a direct comparison with the prior work we follow the same preprocessing procedure and setup as Hinton & Salakhutdinov (2009), Larochelle & Lauly (2012), Srivastava et al. (2013), and Mnih & Gregor (2014). We train NVDM models with 50 and 200 dimensional document representations respectively. For the inference network, we use an MLP (Eq. 8) with 2 layers and 500 dimension rectifier linear units, which converts document representations into embeddings. During training we carry out stochastic estimation by taking one sample for estimating the stochastic gradients, while in prediction we use 20 samples for predicting document perplexity. The model is trained by Adam (Kingma & Ba, 2015) and tuned by hold-out validation perplexity. We alternately optimise the generative model and the inference network by fixing the parameters of one while updating the parameters of the other.

2 Experiments on Document Modelling

Table 1a presents the test document perplexity. The first column lists the models, and the second column shows the dimension of latent variables used in the experiments. The final two columns present the perplexity achieved by each topic model on the 20NewsGroups and RCV1-v2 datasets. In document modelling, perplexity is computed by exp(−1D∑nNd1Ndlog⁡p(Xd))exp(-\frac{1}{D}\sum_{n}^{N_{d}}\frac{1}{N_{d}}\log p(\boldsymbol{X}_{d})), where DD is the number of documents, NdN_{d} represents the length of the ddth document and log⁡p(X)=log⁡∫p(X∣h)p(h)dh\log p(\boldsymbol{X})=\log\int p(\boldsymbol{X}|\boldsymbol{h})p(\boldsymbol{h})d\boldsymbol{h} is the log probability of the words in the document. Since log⁡p(X)\log p(\boldsymbol{X}) is intractable in the NVDM, we use the variational lower bound (which is an upper bound on perplexity) to compute the perplexity following Mnih & Gregor (2014).

While all the baseline models listed in Table 1a apply discrete latent variables, here NVDM employs a continuous stochastic document representation. The experimental results indicate that NVDM achieves the best performance on both datasets. For the experiments on RCV1-v2 dataset, the NVDM with latent variable of 50 dimension performs even better than the fDARN with 200 dimension. It demonstrates that our document model with continuous latent variables has higher expressiveness and better generalisation ability. Table 1b compares the 5 nearest words selected according to the semantic vector learned from NVDM and docNADE.

In addition to the perplexities, we also qualitatively evaluate the semantic information learned by NVDM on the 20NewsGroups dataset with latent variables of 50 dimension. We assume each dimension in the latent space represents a topic that corresponds to a specific semantic meaning. Table 2 presents 5 randomly selected topics with 10 words that have the strongest positive connection with the topic. Based on the words in each column, we can deduce their corresponding topics as: Space, Religion, Encryption, Sport and Policy. Although the model does not impose independent interpretability on the latent representation dimensions, we still see that the NVDM learns locally interpretable structure.

3 Dataset & Setup for Answer Sentence Selection

We experiment on two answer selection datasets, the QASent and the WikiQA datasets. QASent (Wang et al., 2007) is created from the TREC QA track, and the WikiQA (Yang et al., 2015) is constructed from Wikipedia, which is less noisy and less biased towards lexical overlapYang et al. (2015) provide detailed explanation of the differences between the two datasets.. Table 4 summarises the statistics of the two datasets.

In order to investigate the effectiveness of our NASM model we also implemented two strong baseline models — a vanilla LSTM model (LSTM) and an LSTM model with a deterministic attention mechanism (LSTM+Att). The former directly applies the QA matching function (Eq. 15) on the independent question and answer representations which are the last state outputs sq(∣q∣)\boldsymbol{s}_{q}(|\boldsymbol{q}|) and sa(∣a∣)\boldsymbol{s}_{a}(|\boldsymbol{a}|) from the question and answer LSTM models. The latter adds an attention model to learn pair-specific representation for prediction on the basis of the vanilla LSTM. Moreover, LSTM+Att is the deterministic counterpart of NASM, which has the same neural network architecture as NASM. The only difference is that it replaces the stochastic units h\boldsymbol{h} with deterministic ones, and no inference network is required to carry out stochastic estimation. Following previous work, for each of our models we also add a lexical overlap feature by combining a co-occurrence word count feature with the probability generated from the neural model. MAP and MRR are adopted as the evaluation metrics for this task.

To facilitate direct comparison with previous work we follow the same experimental setup as Yu et al. (2014) and Severyn (2015). The word embeddings (K=50)(K=50) are obtained by running the word2vec tool (Mikolov et al., 2013) on the English Wikipedia dump and the AQUAINThttps://catalog.ldc.upenn.edu/LDC2002T31 corpus. We use LSTMs with 33 layers and 5050 hidden units, and apply 40%40\% dropout after the embedding layer. For the construction of the inference network, we use an MLP (Eq. 10) with 2 layers and tanh units of 50 dimension, and an MLP (Eq. 17) with 2 layers and tanh units of 150 dimension for modelling the joint representation. During training we carry out stochastic estimation by taking one sample for computing the gradients, while in prediction we use 20 samples to calculate the expectation of the lower bound. Figure 4 presents the standard deviation of NASM’s MAP scores while using different numbers of samples. Considering the trade-off between computational cost and variance, we chose 20 samples for prediction in all the experiments. The models are trained using Adam (Kingma & Ba, 2015), with hyperparameters selected by optimising the MAP score on the development set.

4 Experiments on Answer Sentence Selection

Table 4 compares the results of our models with current state-of-the-art models on both answer selection datasets. On the QASent dataset, our vanilla LSTM model outperforms the deep CNN As stated in (Yih et al., 2013) that the evaluation scripts used by previous work are noisy — 4 out of 72 questions in the test set are treated answered incorrectly. This makes the MAP and MRR scores ∼4%\sim 4\% lower than the true scores. Since Severyn (2015) and Wang & Ittycheriah (2015) use a cleaned-up evaluation scripts, we apply the original noisy scripts to re-evaluate their outputs in order to make the results directly comparable with previous work. model by approximately 7%7\% on MAP and 6%6\% on MRR. The LSTM+Att performs slightly better than the vanilla LSTM model, and our NASM improves the results further. Since the QASent dataset is biased towards lexical overlapping features, after combining with a co-occurrence word count feature, our best model NASM outperforms all the previous models, including both neural network based models and classifiers with a set of hand-crafted features (e.g. LCLR). Similarly, on the WikiQA dataset, all of our models outperform the previous distributional models by a large margin. By including a word count feature, our models improve further and achieve the state-of-the-art. Notably, on both datasets, our two LSTM-based models have set strong baselines and NASM works even better, which demonstrates the effectiveness of introducing stochastic units to model question semantics in this answer sentence selection task.

In Figure 4, we compare the effectiveness of the latent attention mechanism (NASM) and its deterministic counterpart (LSTM+Att) by visualising the attention scores on the answer sentences. For most of the negative answer sentences, neither of the two attention models can attend to reasonable words that are beneficial for predicting relatedness. But for the correct answer sentences, such as the ones in Figure 4, both attention models are able to capture crucial information by attending to different parts of the sentence based on the question semantics. Interestingly, compared to the deterministic counterpart LSTM+Att, our NASM assigns higher attention scores on the prominent words that are relevant to the question, which forms a more peaked distribution and in turn helps the model achieve better performance.

In order to have an intuitive observation on the latent distributions, we present Hinton diagrams of their log standard deviation parameters (Figure 4). In a Hinton diagram, the size of a square is proportional to a value’s magnitude, and the colour (black/white) indicates its sign (positive/negative). In this case, we visualise the parameters of 50 conditional distributions pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) with the questions selected from 5 different groups, which start with ‘how’, ‘what’, ‘who’, ‘when’ and ‘where’. All the log standard deviations are initialised as zero before training. According to Figure 4, we can see that the questions starting with ‘how’ have more white areas, which indicates higher variances or more uncertainties are in these dimensions. By contrast, the questions starting with ‘what’ have black squares in almost every dimension. Intuitively, it is more difficult to understand and answer the questions starting with ‘how’ than the others, while the ‘what’ questions commonly have explicit words indicating the possible answers. To validate this, we compute the stratified MAP scores based on different question type. The MAP of ’how’ questions is 0.524 which is the lowest among the five groups. Hence empirically, ’how’ questions are harder to ’understand and answer’.

Discussion

As shown in the experiments, neural variational inference brings consistent improvements on the performance of both NLP tasks. The basic intuition is that the latent distributions grant the ability to sum over all the possibilities in terms of semantics. From the perspective of optimisation, one of the most important reasons is that Bayesian learning guards against overfitting.

According to Eq. 5 in NVDM, since we adopt p(h)p(\boldsymbol{h}) as a standard Gaussian prior, the KL divergence term DKL⁡[qϕ(h∣X)∥p(h)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{X})\|p(\boldsymbol{h})] can be analytically computed as 12(K−∥μ∥2−∥σ∥2+log⁡∣diag⁡(σ2)∣)\frac{1}{2}(K-\|\boldsymbol{\mu}\|^{2}-\|\boldsymbol{\sigma}\|^{2}+\log|\operatorname{diag}(\boldsymbol{\sigma}^{2})|). It is not difficult to find that it actually acts as L2 regulariser when we update the μ\boldsymbol{\mu}. Similarly, in NASM (Eq. 16), we also have the KL divergence term DKL⁡[qϕ(h∣q,a,y))∥pθ(h∣q)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}))\|p_{\theta}(\boldsymbol{h}|\boldsymbol{q})]. Different from NVDM, it attempts to minimise the distance between qϕ(h∣q,a,y))q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y})) and pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) that are both conditional distributions. Because pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) as well as qϕ(h∣q,a,y))q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y})) are learned during training, the two distributions are mutually restrained while being updated. Therefore, NVDM simply penalises the large μ\boldsymbol{\mu} and encourages qϕ(h∣X)q_{\phi}(\boldsymbol{h}|\boldsymbol{X}) to approach the prior p(h)p(\boldsymbol{h}) for every document X\boldsymbol{X}, but in NASM, pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}) acts like a moving baseline distribution which regularises the update of qϕ(h∣q,a,y))q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y})) for every different conditions. In practice, we carry out early stopping by observing the prediction performance on development dataset for the question answer selection task. Using the same learning rate and neural network structure, LSTM+Att reaches optimal performance and starts to overfit on training dataset generally at the 2020th iteration, while NASM starts to overfit around the 3535th iteration.

More interestingly, in the question answer selection experiments, NASM learns more peaked attention scores than its deterministic counterpart LSTM+Att. For the update process of LSTM+Att, we find there exists a relatively big variance in the gradients w.r.t. question semantics (LSTM+Att applies deterministic sq(∣q∣)\boldsymbol{s}_{q}(|\boldsymbol{q}|) while NASM applies stochastic h\boldsymbol{h}). This is because the training dataset is small and contains many negative answer sentences that brings no benefit but noise to the learning of the attention model. In contrast, for the update process of NASM, we observe more stable gradients w.r.t. the parameters of latent distributions. The optimisation of the lower bound on one hand maximises the conditional log-likelihood (that the deterministic counterpart cares about) and on the other hand minimises the KL-divergence (that regularises the gradients). Hence, each update of the lower bound actually keeps the gradients w.r.t. μ\boldsymbol{\mu} from swinging heavily. Besides, since the values of σ\boldsymbol{\sigma} are not very significant in this case, the distribution of attention scores mainly depends on μ\boldsymbol{\mu}. Therefore, the learning of the attention model benefits from the regularisation as well, and it explains the fact that NASM learns more peaked attention scores which in turn helps achieve a better prediction performance.

Since the computations of NVDM and NASM can be parallelised on GPU and only one sample is required during training process, it is very efficient to carry out the neural variational inference. Moreover, for both NVDM and NASM, all the parameters are updated by back-propagation. Thus, the increased computation time for the stochastic units only comes from the added parameters of the inference network.

Related Work

Training an inference network to approximate the variational distribution was first proposed in the context of Helmholtz machines (Hinton & Zemel, 1994; Hinton et al., 1995; Dayan & Hinton, 1996), but applications of these directed generative models come up against the problem of establishing low variance gradient estimators. Recent advances in neural variational inference mitigate this problem by reparameterising the continuous random variables (Rezende et al., 2014; Kingma & Welling, 2014), using control variates (Mnih & Gregor, 2014) or approximating the posterior with importance sampling (Bornschein & Bengio, 2015). The instantiations of these ideas (Gregor et al., 2015; Kingma et al., 2014; Ba et al., 2015) have demonstrated strong performance on the tasks of image processing. The recent variants of generative auto-encoder (Louizos et al., 2015; Makhzani et al., 2015) are also very competitive. Tang & Salakhutdinov (2013) applies the similar idea of introducing stochastic units for expression classification, but its inference is carried out by Monte Carlo EM algorithm with the reliance on importance sampling, which is less efficient and lack of scalability.

Another class of neural generative models make use of the autoregressive assumption (Larochelle & Murray, 2011; Uria et al., 2014; Germain et al., 2015; Gregor et al., 2014). Applications of these models on document modelling achieve significant improvements on generating documents, compared to conventional probabilistic topic models (Hofmann, 1999; Blei et al., 2003) and also the RBMs (Hinton & Salakhutdinov, 2009; Srivastava et al., 2013). While these models that use binary semantic vectors, our NVDM employs dense continuous document representations which are both expressive and easy to train. The semantic word vector model (Maas et al., 2011) also employs a continuous semantic vector to generate words, but the model is trained by MAP inference which does not permit the calculation of the posterior distribution. A very similar idea to NVDM is Bowman et al. (2015), which employs VAE to generate sentences from a continuous space.

Apart from the work mentioned above, there is other interesting work on question answering with deep neural networks. One of the popular streams is mapping factoid questions with answer triples in the knowledge base (Bordes et al., 2014a, b; Yih et al., 2014). Moreover, Weston et al. (2015); Sukhbaatar et al. (2015); Kumar et al. (2015) further exploit memory networks, where long-term memories act as dynamic knowledge bases. Another attention-based model (Hermann et al., 2015) applies the attentive network to help read and comprehend for long articles.

Conclusion

This paper introduced a deep neural variational inference framework for generative models of text. We experimented on two diverse tasks, document modelling and question answer selection tasks to demonstrate the effectiveness of this framework, where in both cases our models achieve state of the art performance. Apart from the promising results, our model also has the advantages of (1) simple, expressive, and efficient when training with the SGVB algorithm; (2) suitable for both unsupervised and supervised learning tasks; and (3) capable of generalising to incorporate any type of neural network.

References

Appendix A t-SNE Visualisation of Document Representations

Appendix B Details of the Deep Neural Network Structures

(1) Inference Network qϕ(h∣X)q_{\phi}(\boldsymbol{h}|\boldsymbol{X}):

(2) Generative Model pθ(X∣h)p_{\theta}(\boldsymbol{X}|\boldsymbol{h}):

(3) KL Divergence DKL⁡[qϕ(h∣X)∣∣p(h)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{X})||p(\boldsymbol{h})]:

The variational lower bound to be optimised:

B.2 Neural Answer Selection Model

(1) Inference Network qϕ(h∣q,a,y)q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y}):

pθ(h∣q)p_{\theta}(\boldsymbol{h}|\boldsymbol{q}):

pθ(y∣q,a,h)p_{\theta}(\boldsymbol{y}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{h}):

(3) KL Divergence DKL⁡[qϕ(h∣q,a,y)∣∣pθ(h∣q)]D_{\operatorname{KL}}[q_{\phi}(\boldsymbol{h}|\boldsymbol{q},\boldsymbol{a},\boldsymbol{y})||p_{\theta}(\boldsymbol{h}|\boldsymbol{q})]:

The variational lower bound to be optimised:

Appendix C Computational Complexity

The computational complexity of NVDM for a training document is Cϕ+Cθ=O(LK2+KSV)C_{\phi}+C_{\theta}=O(LK^{2}+KSV). Here, Cϕ=O(LK2)C_{\phi}=O(LK^{2}) represents the cost for the inference network to generate a sample, where LL is the number of the layers in the inference network and KK is the average dimension of these layers. Besides, Cθ=O(KSV)C_{\theta}=O(KSV) is the cost of reconstructing the document from a sample, where SS is the average length of the documents and VV represents the volume of words applied in this document model, which is conventionally much lager than KK.

The computational complexity of NASM for a training question-answer pair is Cϕ+Cθ=O((L+S)K2+SW)C_{\phi}+C_{\theta}=O((L+S)K^{2}+SW). The inference network needs Cϕ=2SW+2K+LK2=O(LK2+SW)C_{\phi}=2SW+2K+LK^{2}=O(LK^{2}+SW). It takes 2SW+2K2SW+2K to produce the joint representation for a question-answer pair and its label, where WW is the total number of parameters of an LSTM and SS is the average length of the sentences. Based on the joint representation, an MLP spends LK2LK^{2} to generate a sample, where LL is the number of layers and KK represents the average dimension. The generative model requires Cθ=2SW+LK2+SK2+5K2+2K2=O((L+S)K2+SW)C_{\theta}=2SW+LK^{2}+SK^{2}+5K^{2}+2K^{2}=O((L+S)K^{2}+SW). Similarly, it costs 2SW+LK22SW+LK^{2} to construct the generative latent distribution , where 2SW2SW can be saved if the LSTMs are shared by the inference network and the generative model. Besides, the attention model takes SK2+5K2SK^{2}+5K^{2} and the relatedness prediction takes the last 2K22K^{2}.

Since the computations of NVDM and NASM can be parallelised in GPU and only one sample is required during training process, it is very efficient to carry out the neural variational inference. As NVDM is an instantiation of variational auto-encoder, its computational complexity is the same as the deterministic auto-encoder. In addition, the computational complexity of LSTM+Att, the deterministic counterpart of NASM, is also O((L+S)K2+SW)O((L+S)K^{2}+SW). There is only O(LK2)O(LK^{2}) time increase by introducing an inference network for NASM when compared to LSTM+Att.