A Hierarchical Latent Structure for Variational Conversation Modeling

Yookoon Park, Jaemin Cho, Gunhee Kim

Introduction

Conversation modeling has been a long interest of natural language research. Recent approaches for data-driven conversation modeling mostly build upon recurrent neural networks (RNNs) (Vinyals and Le, 2015; Sordoni et al., 2015b; Shang et al., 2015; Li et al., 2017; Serban et al., 2016). Serban et al. (2016) use a hierarchical RNN structure to model the context of conversation. Serban et al. (2017) further exploit an utterance latent variable in the hierarchical RNNs by incorporating the variational autoencoder (VAE) framework Kingma and Welling (2014); Rezende et al. (2014).

VAEs enable us to train a latent variable model for natural language modeling, which grants us several advantages. First, latent variables can learn an interpretable holistic representation, such as topics, tones, or high-level syntactic properties. Second, latent variables can model inherently abundant variability of natural language by encoding its global and long-term structure, which is hard to be captured by shallow generative processes (e.g. vanilla RNNs) where the only source of stochasticity comes from the sampling of output words.

In spite of such appealing properties of latent variable models for natural language modeling, VAEs suffer from the notorious degeneration problem (Bowman et al., 2016; Chen et al., 2017) that occurs when a VAE is combined with a powerful decoder such as autoregressive RNNs. This issue makes VAEs ignore latent variables, and eventually behave as vanilla RNNs. Chen et al. (2017) also note this degeneration issue by showing that a VAE with a RNN decoder prefers to model the data using its decoding distribution rather than using latent variables, from bits-back coding perspective. To resolve this issue, several heuristics have been proposed to weaken the decoder, enforcing the models to use latent variables. For example, Bowman et al. (2016) propose some heuristics, including KL annealing and word drop regularization. However, these heuristics cannot be a complete solution; for example, we observe that they fail to prevent the degeneracy in VHRED (Serban et al., 2017), a conditional VAE model equipped with hierarchical RNNs for conversation modeling.

The objective of this work is to propose a novel VAE model that significantly alleviates the degeneration problem. Our analysis reveals that the causes of the degeneracy are two-fold. First, the hierarchical structure of autoregressive RNNs is powerful enough to predict a sequence of utterances without the need of latent variables, even with the word drop regularization. Second, we newly discover that the conditional VAE structure where an utterance is generated conditioned on context, i.e. a previous sequence of utterances, induces severe data sparsity. Even with a large-scale training corpus, there only exist very few target utterances when conditioned on the context. Hence, the hierarchical RNNs can easily memorize the context-to-utterance relations without relying on latent variables.

We propose a novel model named Variational Hierarchical Conversation RNN (VHCR), which involves two novel features to alleviate this problem. First, we introduce a global conversational latent variable along with local utterance latent variables to build a hierarchical latent structure. Second, we propose a new regularization technique called utterance drop. We show that our hierarchical latent structure is not only crucial for facilitating the use of latent variables in conversation modeling, but also delivers several additional advantages, including gaining control over the global context in which the conversation takes place.

(1) We reveal that the existing conditional VAE model with hierarchical RNNs for conversation modeling (e.g. (Serban et al., 2017)) still suffers from the degeneration problem, and this problem is caused by data sparsity per context that arises from the conditional VAE structure, as well as the use of powerful hierarchical RNN decoders.

(2) We propose a novel variational hierarchical conversation RNN (VHCR), which has two distinctive features: a hierarchical latent structure and a new regularization of utterance drop. To the best of our knowledge, our VHCR is the first VAE conversation model that exploits the hierarchical latent structure.

(3) With evaluations on two benchmark datasets of Cornell Movie Dialog (Danescu-Niculescu-Mizil and Lee, 2011) and Ubuntu Dialog Corpus (Lowe et al., 2015), we show that our model improves the conversation performance in multiple metrics over state-of-the-art methods, including HRED (Serban et al., 2016), and VHRED (Serban et al., 2017) with existing degeneracy solutions such as the word drop (Bowman et al., 2016), and the bag-of-words loss (Zhao et al., 2017).

Related Work

Conversation Modeling. One popular approach for conversation modeling is to use RNN-based encoders and decoders, such as (Vinyals and Le, 2015; Sordoni et al., 2015b; Shang et al., 2015). Hierarchical recurrent encoder-decoder (HRED) models (Sordoni et al., 2015a; Serban et al., 2016, 2017) consist of utterance encoder and decoder, and a context RNN which runs over utterance representations to model long-term temporal structure of conversation.

Recently, latent variable models such as VAEs have been adopted in language modeling (Bowman et al., 2016; Zhang et al., 2016; Serban et al., 2017). The VHRED model (Serban et al., 2017) integrates the VAE with the HRED to model Twitter and Ubuntu IRC conversations by introducing an utterance latent variable. This makes a conditional VAE where the generation process is conditioned on the context of conversation. Zhao et al. (2017) further make use of discourse act labels to capture the diversity of conversations.

Degeneracy of Variational Autoencoders. For sequence modeling, VAEs are often merged with the RNN encoder-decoder structure (Bowman et al., 2016; Serban et al., 2017; Zhao et al., 2017) where the encoder predicts the posterior distribution of a latent variable z\bm{\mathbf{z}}, and the decoder models the output distributions conditioned on z\bm{\mathbf{z}}. However, Bowman et al. (2016) report that a VAE with a RNN decoder easily degenerates; that is, it learns to ignore the latent variable z\bm{\mathbf{z}} and falls back to a vanilla RNN. They propose two techniques to alleviate this issue: KL annealing and word drop. Chen et al. (2017) interpret this degeneracy in the context of bits-back coding and show that a VAE equipped with autoregressive models such as RNNs often ignores the latent variable to minimize the code length needed for describing data. They propose to constrain the decoder to selectively encode the information of interest in the latent variable. However, their empirical results are limited to an image domain. Zhao et al. (2017) use an auxiliary bag-of-words loss on the latent variable to force the model to use z\bm{\mathbf{z}}. That is, they train an auxiliary network that predicts bag-of-words representation of the target utterance based on z\bm{\mathbf{z}}. Yet this loss works in an opposite direction to the original objective of VAEs that minimizes the minimum description length. Thus, it may be in danger of forcibly moving the information that is better modeled in the decoder to the latent variable.

Approach

We assume that the training set consists of NN i.i.d samples of conversations {c1,c2,...,cN}\{\bm{\mathbf{c}}_{1},\bm{\mathbf{c}}_{2},...,\bm{\mathbf{c}}_{N}\} where each ci\bm{\mathbf{c}}_{i} is a sequence of utterances (i.e. sentences) {xi1,xi2,...,xini}\{\bm{\mathbf{x}}_{i1},\bm{\mathbf{x}}_{i2},...,\bm{\mathbf{x}}_{in_{i}}\}. Our objective is to learn the parameters of a generative network θ\bm{\mathbf{\theta}} using Maximum Likelihood Estimation (MLE):

We first briefly review the VAE, and explain the degeneracy issue before presenting our model.

We follow the notion of Kingma and Welling (2014). A datapoint x\bm{\mathbf{x}} is generated from a latent variable z\bm{\mathbf{z}}, which is sampled from some prior distribution p(z)p(\bm{\mathbf{z}}), typically a standard Gaussian distribution N(z∣0,I)\mathcal{N}(\bm{\mathbf{z}}|\bm{\mathbf{0}},\bm{\mathbf{I}}). We assume parametric families for conditional distribution pθ(x∣z)p_{\bm{\mathbf{\theta}}}(\bm{\mathbf{x}}|\bm{\mathbf{z}}). Since it is intractable to compute the log-marginal likelihood log⁡pθ(x)\log p_{\bm{\mathbf{\theta}}}(\bm{\mathbf{x}}), we approximate the intractable true posterior pθ(z∣x)p_{\bm{\mathbf{\theta}}}(\bm{\mathbf{z}}|\bm{\mathbf{x}}) with a recognition model qϕ(z∣x)q_{\bm{\mathbf{\phi}}}(\bm{\mathbf{z}}|\bm{\mathbf{x}}) to maximize the variational lower-bound:

Eq. 2 is decomposed into two terms: KL divergence term and reconstruction term. Here, KL divergence measures the amount of information encoded in the latent variable z\bm{\mathbf{z}}. In the extreme where KL divergence is zero, the model completely ignores z\bm{\mathbf{z}}, i.e. it degenerates. The expectation term can be stochastically approximated by sampling z\bm{\mathbf{z}} from the variational posterior qϕ(z∣x)q_{\bm{\mathbf{\phi}}}(\bm{\mathbf{z}}|\bm{\mathbf{x}}). The gradients to the recognition model can be efficiently estimated using the reparameterization trick (Kingma and Welling, 2014).

2 VHRED

Serban et al. (2017) propose Variational Hierarchical Recurrent Encoder Decoder (VHRED) model for conversation modeling. It integrates an utterance latent variable ztutt\bm{\mathbf{z}}^{\text{utt}}_{t} into the HRED structure (Sordoni et al., 2015a) which consists of three RNN components: encoder RNN, context RNN, and decoder RNN. Given a previous sequence of utterances x1,...xt−1\bm{\mathbf{x}}_{1},...\bm{\mathbf{x}}_{t-1} in a conversation, the VHRED generates the next utterance xt\bm{\mathbf{x}}_{t} as:

At time step tt, the encoder RNN fθencf^{\text{enc}}_{\bm{\mathbf{\theta}}} takes the previous utterance xt−1\bm{\mathbf{x}}_{t-1} and produces an encoder vector ht−1enc\bm{\mathbf{h}}^{\text{enc}}_{t-1} (Eq. 3). The context RNN fθcxtf^{\text{cxt}}_{\bm{\mathbf{\theta}}} models the context of the conversation by updating its hidden states using the encoder vector (Eq. 4). The context htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} defines the conditional prior pθ(ztutt∣x<t)p_{\bm{\mathbf{\theta}}}(\bm{\mathbf{z}}^{\text{utt}}_{t}|\bm{\mathbf{x}}_{<t}), which is a factorized Gaussian distribution whose mean μt\bm{\mathbf{\mu}}_{t} and diagonal variance σt\bm{\mathbf{\sigma}}_{t} are given by feed-forward neural networks (Eq. 5-7). Finally the decoder RNN fθdecf^{\text{dec}}_{\bm{\mathbf{\theta}}} generates the utterance xt\bm{\mathbf{x}}_{t}, conditioned on the context vector htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} and the latent variable ztutt\bm{\mathbf{z}}^{\text{utt}}_{t} (Eq. 8). We make two important notes: (1) the context RNN can be viewed as a high-level decoder, and together with the decoder RNN, they comprise a hierarchical RNN decoder. (2) VHRED follows a conditional VAE structure where each utterance xt\bm{\mathbf{x}}_{t} is generated conditioned on the context htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} (Eq. 5-8).

The variational posterior is a factorized Gaussian distribution where the mean and the diagonal variance are predicted from the target utterance and the context as follows:

3 The Degeneration Problem

A known problem of a VAE that incorporates an autoregressive RNN decoder is the degeneracy that ignores the latent variable z\bm{\mathbf{z}}. In other words, the KL divergence term in Eq. 2 goes to zero and the decoder fails to learn any dependency between the latent variable and the data. Eventually, the model behaves as a vanilla RNN. This problem is first reported in the sentence VAE (Bowman et al., 2016), in which following two heuristics are proposed to alleviate the problem by weakening the decoder.

First, the KL annealing scales the KL divergence term of Eq. 2 using a KL multiplier λ\lambda, which gradually increases from 0 to 1 during training:

This helps the optimization process to avoid local optima of zero KL divergence in early training. Second, the word drop regularization randomly replaces some conditioned-on word tokens in the RNN decoder with the generic unknown word token (UNK) during training. Normally, the RNN decoder predicts each next word in an autoregressive manner, conditioned on the previous sequence of ground truth (GT) words. By randomly replacing a GT word with an UNK token, the word drop regularization weakens the autoregressive power of the decoder and forces it to rely on the latent variable to predict the next word. The word drop probability is normally set to 0.25, since using a higher probability may degrade the model performance (Bowman et al., 2016).

However, we observe that these tricks do not solve the degeneracy for the VHRED in conversation modeling. An example in Fig. 1 shows that the VHRED learns to ignore the utterance latent variable as the KL divergence term falls to zero.

4 Empirical Observation on Degeneracy

The decoder RNN of the VHRED in Eq. 8 conditions on two information sources: deterministic htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} and stochastic zutt\bm{\mathbf{z}}^{\text{utt}}. In order to check whether the presence of deterministic source htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} causes the degeneration, we drop the deterministic htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} and condition the decoder only on the stochastic utterance latent variable zutt\bm{\mathbf{z}}^{\text{utt}}:

While this model achieves higher values of KL divergence than original VHRED, as training proceeds it again degenerates with the KL divergence term reaching zero (Fig. 2).

This empirical observation implies that the fundamental reason behind the degeneration may originate from combination of two factors: (1) strong expressive power of the hierarchical RNN decoder and (2) training data sparsity caused by the conditional VAE structure. The VHRED is trained to predict a next target utterance xt\bm{\mathbf{x}}_{t} conditioned on the context htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} which encodes information about previous utterances {x1,…,xt−1}\{\bm{\mathbf{x}}_{1},\dots,\bm{\mathbf{x}}_{t-1}\}. However, conditioning on the context makes the range of training target xt\bm{\mathbf{x}}_{t} very sparse; even in a large-scale conversation corpus such as Ubuntu Dialog (Lowe et al., 2015), there exist one or very few target utterances per context. Therefore, hierarchical RNNs, given their autoregressive power, can easily overfit to training data without using the latent variable. Consequently, the VHRED will not encode any information in the latent variable, i.e. it degenerates. It explains why the word drop fails to prevent the degeneracy in the VHRED. The word drop only regularizes the decoder RNN; however, the context RNN is also powerful enough to predict a next utterance in a given context even with the weakened decoder RNN. Indeed we observe that using a larger word drop probability such as 0.5 or 0.75 only slows down, but fails to stop the KL divergence from vanishing.

5 Variational Hierarchical Conversation RNN (VHCR)

As discussed, we argue that the two main causes of degeneration are i) the expressiveness of the hierarchical RNN decoders, and ii) the conditional VAE structure that induces data sparsity. This finding hints us that in order to train a non-degenerate latent variable model, we need to design a model that provides an appropriate way to regularize the hierarchical RNN decoders and alleviate data sparsity per context. At the same time, the model should be capable of modeling complex structure of conversation. Based on these insights, we propose a novel VAE structure named Variational Hierarchical Conversation RNN (VHCR), whose graphical model is illustrated in Fig. 3. Below we first describe the model, and discuss its unique features.

We introduce a global conversation latent variable zconv\bm{\mathbf{z}}^{\text{conv}} which is responsible for generating a sequence of utterances of a conversation c={x1,…,xn}\bm{\mathbf{c}}=\{\bm{\mathbf{x}}_{1},\dots,\bm{\mathbf{x}}_{n}\}:

Overall, the VHCR builds upon the hierarchical RNNs, following the VHRED (Serban et al., 2017). One key update is to form a hierarchical latent structure, by using the global latent variable zconv\bm{\mathbf{z}}^{\text{conv}} per conversation, along with local the latent variable ztutt\bm{\mathbf{z}}_{t}^{\text{utt}} injected at each utterance (Fig. 3):

For inference of zconv\bm{\mathbf{z}}^{\text{conv}}, we use a bi-directional RNN denoted by fconvf^{\text{conv}}, which runs over the utterance vectors generated by the encoder RNN:

The posteriors for local variables ztutt\bm{\mathbf{z}}^{\text{utt}}_{t} are then conditioned on zconv\bm{\mathbf{z}}^{\text{conv}}:

Our solution of VHCR to the degeneration problem is based on two ideas. The first idea is to build a hierarchical latent structure of zconv\bm{\mathbf{z}}^{\text{conv}} for a conversation and ztutt\bm{\mathbf{z}}^{\text{utt}}_{t} for each utterance. As zconv\bm{\mathbf{z}}^{\text{conv}} is independent of the conditional structure, it does not suffer from the data sparsity problem. However, the expressive power of hierarchical RNN decoders makes the model still prone to ignore latent variables zconv\bm{\mathbf{z}}^{\text{conv}} and ztutt\bm{\mathbf{z}}^{\text{utt}}_{t}. Therefore, our second idea is to apply an utterance drop regularization to effectively regularize the hierarchical RNNs, in order to facilitate the use of latent variables. That is, at each time step, the utterance encoder vector htenc\bm{\mathbf{h}}^{\text{enc}}_{t} is randomly replaced with a generic unknown vector hunk\bm{\mathbf{h}}^{\text{unk}} with a probability pp. This regularization weakens the autoregressive power of hierarchical RNNs and as well alleviates the data sparsity problem, since it induces noise into the context vector htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t} which conditions the decoder RNN. The difference with the word drop (Bowman et al., 2016) is that our utterance drop depresses the hierarchical RNN decoders as a whole, while the word drop only weakens the lower-level decoder RNNs. Fig. 4 confirms that with the utterance drop with a probability of 0.250.25, the VHCR effectively learns to use latent variables, achieving a significant degree of KL divergence.

6 Effectiveness of Hierarchical Latent Structure

Is the hierarchical latent structure of the VHCR crucial for effective utilization of latent variables? We investigate this question by applying the utterance drop on the VHRED which lacks any hierarchical latent structure. We observe that the KL divergence still vanishes (Fig. 4), even though the utterance drop injects considerable noise in the context htcxt\bm{\mathbf{h}}^{\text{cxt}}_{t}. We argue that the utterance drop weakens the context RNN, thus it consequently fail to predict a reasonable prior distribution for zutt\bm{\mathbf{z}}^{\text{utt}} (Eq. 5-7). If the prior is far away from the region of zutt\bm{\mathbf{z}}^{\text{utt}} that can generate a correct target utterance, encoding information about the target in the variational posterior will incur a large KL divergence penalty. If the penalty outweighs the gain of the reconstruction term in Eq. 2, then the model would learn to ignore zutt\bm{\mathbf{z}}^{\text{utt}}, in order to maximize the variational lower-bound in Eq. 2.

On the other hand, the global variable zconv\bm{\mathbf{z}}^{\text{conv}} allows the VHCR to predict a reasonable prior for local variable ztutt\bm{\mathbf{z}}^{\text{utt}}_{t} even in the presence of the utterance drop regularization. That is, zconv\bm{\mathbf{z}}^{\text{conv}} can act as a guide for zutt\bm{\mathbf{z}}^{\text{utt}} by encoding the information for local variables. This reduces the KL divergence penalty induced by encoding information in zutt\bm{\mathbf{z}}^{\text{utt}} to an affordable degree at the cost of KL divergence caused by using zconv\bm{\mathbf{z}}^{\text{conv}}. This trade-off is indeed a fundamental strength of hierarchical models that provide parsimonious representation; if there exists any shared information among the local variables, it is coded in the global latent variable reducing the code length by effectively reusing the information. The remaining local variability is handled properly by the decoding distribution and local latent variables.

The global variable zconv\bm{\mathbf{z}}^{\text{conv}} provides other benefits by representing a latent global structure of a conversation, such as a topic, a length, and a tone of the conversation. Moreover, it allows us to control such global properties, which is impossible for models without hierarchical latent structure.

Results

We first describe our experimental setting, such as datasets and baselines (section 4.1). We then report quantitative comparisons using three different metrics (section 4.2–4.4). Finally, we present qualitative analyses, including several utterance control tasks that are enabled by the hierarchal latent structure of our VHCR (section 4.5). We defer implementation details and additional experiment results to the appendix.

Datasets. We evaluate the performance of conversation generation using two benchmark datasets: 1) Cornell Movie Dialog Corpus (Danescu-Niculescu-Mizil and Lee, 2011), containing 220,579 conversations from 617 movies. 2) Ubuntu Dialog Corpus (Lowe et al., 2015), containing about 1 million multi-turn conversations from Ubuntu IRC channels. In both datasets, we truncate utterances longer than 30 words.

Baselines. We compare our approach with four baselines. They are combinations of two state-of-the-art models of conversation generation with different solutions to the degeneracy. (i) Hierarchical recurrent encoder-decoder (HRED) (Serban et al., 2016), (ii) Variational HRED (VHRED) (Serban et al., 2017), (iii) VHRED with the word drop (Bowman et al., 2016), and (iv) VHRED with the bag-of-words (bow) loss (Zhao et al., 2017).

Performance Measures. Automatic evaluation of conversational systems is still a challenging problem (Liu et al., 2016). Based on literature, we report three quantitative metrics: i) the negative log-likelihood (the variational bound for variational models), ii) embedding-based metrics Serban et al. (2017), and iii) human evaluation via Amazon Mechanical Turk (AMT).

2 Results of Negative Log-likelihood

Table 1(b) summarizes the per-word negative log-likelihood (NLL) evaluated on the test sets of two datasets. For variational models, we instead present the variational bound of the negative log-likelihood in Eq. 2, which consists of the reconstruction error term and the KL divergence term. The KL divergence term can measure how much each model utilizes the latent variables.

We observe that the NLL is the lowest by the HRED. Variational models show higher NLLs, because they are regularized methods that are forced to rely more on latent variables. Independent of NLL values, we later show that the latent variable models often show better generalization performance in terms of embedding-based metrics and human evaluation. In the VHRED, the KL divergence term gradually vanishes even with the word drop regularization; thus, early stopping is necessary to obtain a meaningful KL divergence. The VHRED with the bag-of-words loss (bow) achieves the highest KL divergence, however, at the cost of high NLL values. That is, the variational lower-bound minimizes the minimum description length, to which the bow loss works in an opposite direction by forcing latent variables to encode bag-of-words representation of utterances. Our VHCR achieves stable KL divergence without any auxiliary objective, and the NLL is lower than the VHRED + bow model.

Table 2 summarizes how global and latent variable are used in the VHCR. We observe that VHCR encodes a significant amount of information in the global variable zconv\bm{\mathbf{z}}^{\text{conv}} as well as in the local variable zutt\bm{\mathbf{z}}^{\text{utt}}, indicating that the VHCR successfully exploits its hierarchical latent structure.

3 Results of Embedding-Based Metrics

The embedding-based metrics (Serban et al., 2017; Rus and Lintean, 2012) measure the textual similarity between the words in the model response and the ground truth. We represent words using Word2Vec embeddings trained on the Google News Corpushttps://code.google.com/archive/p/word2vec/.. The average metric projects each utterance to a vector by taking the mean over word embeddings in the utterance, and computes the cosine similarity between the model response vector and the ground truth vector. The extrema metric is similar to the average metric, only except that it takes the extremum of each dimension, instead of the mean. The greedy metric first finds the best non-exclusive word alignment between the model response and the ground truth, and then computes the mean over the cosine similarity between the aligned words.

Table 3(b) compares the different methods with three embedding-based metrics. Each model generates a single response (1-turn) or consecutive three responses (3-turn) for a given context. For 3-turn cases, we report the average of metrics measured for three turns. We use the greedy decoding for all the models.

Our VHCR achieves the best results in most metrics. The HRED is the worst on the Cornell Movie dataset, but outperforms the VHRED and VHRED + bow on the Ubuntu Dialog dataset. Although the VHRED + bow shows the highest KL divergence, its performance is similar to that of VHRED, and worse than that of the VHCR model. It suggests that a higher KL divergence does not necessarily lead to better performance; it is more important for the models to balance the modeling powers of the decoder and the latent variables. The VHCR uses a more sophisticated hierarchical latent structure, which better reflects the structure of natural language conversations.

4 Results of Human Evaluation

Table 4 reports human evaluation results via Amazon Mechanical Turk (AMT). The VHCR outperforms the baselines in both datasets; yet the performance improvement in Cornell Movie Dialog are less significant compared to that of Ubuntu. We empirically find that Cornell Movie dataset is small in size, but very diverse and complex in content and style, and the models often fail to generate sensible responses for the context. The performance gap with the HRED is the smallest, suggesting that the VAE models without hierarchical latent structure have overfitted to Cornell Movie dataset.

5 Qualitative Analyses

Comparison of Predicted Responses. Table 5 compares the generated responses of algorithms. Overall, the VHCR creates more consistent responses within the context of a given conversation. This is supposedly due to the global latent variable zconv\bm{\mathbf{z}}^{\text{conv}} that provides a more direct and effective way to handle the global context of a conversation. The context RNN of the baseline models can handle long-term context to some extent, but not as much as the VHCR.

Interpolation on zconv\bm{\mathbf{z}}^{\text{conv}}. We present examples of one advantage by the hierarchical latent structure of the VHCR, which cannot be done by the other existing models. Table 6 shows how the generated responses vary according to the interpolation on zconv\bm{\mathbf{z}}^{\text{conv}}. We randomly sample two zconv\bm{\mathbf{z}}^{\text{conv}} from a standard Gaussian prior as references (i.e. the top and the bottom row of Table 6), and interpolate points between them. We generate 3-turn conversations conditioned on given zconv\bm{\mathbf{z}}^{\text{conv}}. We see that zconv\bm{\mathbf{z}}^{\text{conv}} controls the overall tone and content of conversations; for example, the tone of the response is friendly in the first sample, but gradually becomes hostile as zconv\bm{\mathbf{z}}^{\text{conv}} changes.

Generation on a Fixed zconv\bm{\mathbf{z}}^{\text{conv}}. We also study how fixing a global conversation latent variable zconv\bm{\mathbf{z}}^{\text{conv}} affects the conversation generation. Table 7 shows an example, where we randomly fix a reference zconv\bm{\mathbf{z}}^{\text{conv}} from the prior, and generate multiple examples of 3-turn conversation using randomly sampled local variables zutt\bm{\mathbf{z}}^{\text{utt}}. We observe that zconv\bm{\mathbf{z}}^{\text{conv}} heavily affects the form of the first utterance; in the examples, the first utterances all start with a “where” phrase. At the same time, responses show variations according to different local variables zutt\bm{\mathbf{z}}^{\text{utt}}. These examples show that the hierarchical latent structure of VHCR allows both global and fine-grained control over generated conversations.

Discussion

We introduced the variational hierarchical conversation RNN (VHCR) for conversation modeling. We noted that the degeneration problem in existing VAE models such as the VHRED is persistent, and proposed a hierarchical latent variable model with the utterance drop regularization. Our VHCR obtained higher and more stable KL divergences than various versions of VHRED models without using any auxiliary objective. The empirical results showed that the VHCR better reflected the structure of natural conversations, and outperformed previous models. Moreover, the hierarchical latent structure allowed both global and fine-grained control over the conversation generation.

Acknowledgments

This work was supported by Kakao and Kakao Brain corporations, and Creative-Pioneering Researchers Program through Seoul National University. Gunhee Kim is the corresponding author.

References

Appendix A Data Processing

In both datasets, we truncate utterances longer than 30 words. Tokenization and text preprocessing is carried out using Spacy https://spacy.io/.

As Cornell Movie Dialog does not provide a separate test set, we randomly choose 80% of the conversations in Cornell Movie Dialog as training set. The remaining 20% is evenly split into validation set and test set.

Appendix B Implementation Details

We use Pytorch Framework http://pytorch.org/ for our implementations. We plan to release our code public.

We build a dictionary with the vocabulary size of 20,000, and further remove words with frequency less than five. We set the word embedding dimension to 500. We adopt Gated Recurrent Unit (GRU) (Cho et al., 2014) in our model and all baseline models, as we observe no improvement of LSTMs (Hochreiter and Schmidhuber, 1997) over GRUs in our experiments. We use one-layer GRU with the hidden dimension of 1,000 (2,000 for bi-directional GRU) for our RNN decoders. Two-layer MLPs with hidden layer size 1000 parameterizes the distribution of latent variables. All latent variables have a dimension of 100. We apply dropout ratio of 0.2 during training. Batch size is 80 for Cornell Movie Dialog, and 40 for Ubuntu Dialog. For optimization, we use Adam (Kingma and Ba, 2014) with a learning rate of 0.0001 with gradient clipping. We adopt early stopping by monitoring the performance on the validation set. We apply the KL annealing to all variational models, where the KL multiplier λ\lambda gradually increases from 0 to 1 over 15,000 steps on Cornell Movie Dialog and over 250,000 steps on Ubuntu Dialog. For both the word drop and the utterance drop, we use drop probability of 0.25.

Appendix C Experimental Results

Table 8 – 11 shows additional sample generation results.

Appendix D Human Evaluation

We perform human evaluation study on Amazon Mechanical Turk (AMT). We first filter out contexts that contain generic unknown word (unk) token from the test set. Using these contexts, we generate model response samples. Samples that contain less than 4 tokens are removed. The order of the samples and the order of model responses are randomly shuffled.

Evaluation procedure is as follows: given a context and two model responses, a Turker decides which response is more appropriate in the given context. In the case where the Turker thinks that two responses are about equally good or bad or does not understand the context, we ask the Turker to choose “tie”. We randomly select 100 samples to build a batch for a human intelligence test (HIT). For each pair of models, we perform 3 HITs on AMT and each HIT is evaluated by 5 unique humans. In total we obtain 9000 preferences in 90 HITs.