Latent Intention Dialogue Models

Tsung-Hsien Wen, Yishu Miao, Phil Blunsom, Steve Young

Introduction

Recurrent neural networks (RNNs) have shown impressive results in modeling generation tasks that have a sequential structured output form, such as machine translation (SutskeverVL14; BahdanauCB14), caption generation (KarpathyF14; xu2015icml), and natural language generation (wensclstm15; Kiddon2016). These discriminative models are trained to learn only a conditional output distribution over strings and despite the sophisticated architectures and conditioning mechanisms used to ensure salience, they are not able to model the underlying actions needed to generate natural dialogues. As a consequence, these sequence-to-sequence models are limited in their ability to exhibit the intrinsic variability and stochasticity of natural dialogue. For example both goal-oriented dialogue systems (wenN2N17; bordes16n2n) and sequence-to-sequence learning chatbots (VinyalsL15; ShangLL15; SerbanSBCP15) struggle to generate diverse yet causal responses (LiGBGD15; serban2016lv). In addition, there is often insufficient training data for goal-oriented dialogues which results in over-fitting which prevents deterministic models from learning effective and scalable interactions. In this paper, we propose a latent variable model – Latent Intention Dialogue Model (LIDM) – for learning the complex distribution of communicative intentions in goal-oriented dialogues. Here, the latent variable representing dialogue intention can be considered as the autonomous decision-making center of a dialogue agent for composing appropriate machine responses.

Recent advances in neural variational inference (icml2014Kingma; mnih14NVIL) have sparked a series of latent variable models applied to NLP (BowmanVVDJB15; serban2016lv; miaoTVAE16; cao17eacl). For models with continuous latent variables, the reparameterisation trick (icml2014Kingma) is commonly used to build an unbiased and low-variance gradient estimator for updating the models. However, since a continuous latent space is hard to interpret, the major benefits of these models are the stochasticity and the regularisation brought by the latent variable. In contrast, models with discrete latent variables are able to not only produce interpretable latent distributions but also provide a principled framework for semi-supervised learning (NIPS2014_5352). This is critical for NLP tasks, especially where additional supervision and external knowledge can be utilized for bootstrapping (FaruquiDJDHS15; miao16latentLangauage; kocisky16). However, variational inference with discrete latent variables is relatively difficult due to the problem of high variance during sampling. Hence we introduce baselines, as in the REINFORCE (Williams1992) algorithm, to mitigate the high variance problem, and carry out efficient neural variational inference (mnih14NVIL) for the latent variable model.

In the LIDM, the latent intention is inferred from user input utterances. Based on the dialogue context, the agent draws a sample as the intention which then guides the natural language response generation. Firstly, in the framework of neural variational inference (mnih14NVIL), we construct an inference network to approximate the posterior distribution over the latent intention. Then, by sampling the intentions for each response, we are able to directly learn a basic intention distribution on a human-human dialogue corpus by optimising the variational lower bound. To further reduce the variance, we utilize a labeled subset of the corpus in which the labels of intentions are automatically generated by clustering. Then, the latent intention distribution can be learned in a semi-supervised fashion, where the learning signals are either from the direct supervision (labeled set) or the variational lower bound (unlabeled set).

From the perspective of reinforcement learning, the latent intention distribution can be interpreted as the intrinsic policy that reflects human decision-making under a particular conversational scenario. Based on the initial policy (latent intention distribution) learnt from the semi-supervised variational inference framework, the model can refine its strategy easily against alternative objectives using policy gradient-based reinforcement learning. This is somewhat analagous to the training process used in AlphaGo (silver2016mastering) for the game of Go. Based on LIDM, we show that different learning paradigms can be brought together under the same framework to bootstrap the development of a dialogue agent (liEMNLP20162; NIPS2016_6264).

In summary, the contribution of this paper is two-fold: firstly, we show that the neural variational inference framework is able to discover discrete, interpretable intentions from data to form the decision-making basis of a dialogue agent; secondly, the agent is capable of revising its conversational strategy based on an external reward within the same framework. This is important because it provides a stepping stone towards building an autonomous dialogue agent that can continuously improve itself through interaction with users. The experimental results demonstrate the effectiveness of our latent intention model which achieves state-of-the-art performance on both automatic corpus-based evaluation and human evaluation.

Latent Intention Dialogue Model for Goal-oriented Dialogue

Goal-oriented dialogueLike most of the goal-oriented dialogue research, we focus on information seek type dialogues. (6407655) aims at building models that can help users to complete certain tasks via natural language interaction. Given a user input utterance utu_{t} at turn tt and a knowledge base (KB), the model needs to parse the input into actionable commands QQ and access the KB to search for useful information in order to answer the query. Based on the search result, the model needs to summarise its findings and reply with an appropriate response mtm_{t} in natural language.

2 Inference

Parameters Θ2\Theta_{2} in the generative network, however, are updated by minimising the KL divergence,

Then the parameters Φ\Phi are updated by,

and the gradient w.r.t. Φ\Phi can be rewritten as

3 Semi-Supervision

Despite the steps described above for reducing the variance, there remain two major difficulties in learning latent intentions in a completely unsupervised manner: (1) the high variance of the inference network prevents it from generating sensible intention samples in the early stages of training, and (2) the overly strong discriminative power of the LSTM language model is prone to the disconnection phenomenon between the LSTM decoder and the rest of the components whereby the decoder learns to ignore the samples and focuses solely on optimising the language model. To ensure more stable training and prevent disconnection, a semi-supervised learning technique is introduced.

The final joint objective function can then be written as L′=αL1+L2\mathcal{L^{\prime}}=\alpha\mathcal{L}_{1}+\mathcal{L}_{2}, where α\alpha controls the trade-off between the supervised and unsupervised examples.

4 Reinforcement Learning

Experiments

We explored the properties of the LIDM modelWill be available at https://github.com/shawnwun/NNDIAL using the CamRest676 corpushttps://www.repository.cam.ac.uk/handle/1810/260970 collected by Wen et al (wenN2N17), in which the task of the system is to assist users to find a restaurant in the Cambridge, UK area. The corpus was collected based on a modified Wizard of Oz (Kelley84) online data collection. Workers were recruited on Amazon Mechanical Turk and asked to complete a task by carrying out a conversation, alternating roles between a user and a wizard. There are three informable slots (food, pricerange, area) that users can use to constrain the search and six requestable slots (address, phone, postcode plus the three informable slots) that the user can ask a value for once a restaurant has been offered. There are 676 dialogues in the dataset (including both finished and unfinished dialogues) and approximately 2750 conversational turns in total. The database contains 99 unique restaurants.

To make a direct comparison with prior work we follow the same experimental setup as in Wen et al (wencond16; wenN2N17). The corpus was partitioned into training, validation, and test sets in the ratio 3:1:1. The LSTM hidden layer sizes were set to 50, and the vocabulary size is around 500 after pre-processing, to remove rare words and words that can be delexicalised2. All the system components were trained jointly by fixing the pre-trained belief trackers and the discrete database operator with the model’s latent intention size II set to 50, 70, and 100, respectively. The trade-off constants λ\lambda and α\alpha were both set to 0.1. To produce self-labeled response clusters for semi-supervised learning of the intentions, we firstly removed function words from all the responses and clustered them according to their content words. We then assigned the responses in the ii-th frequent cluster to the ii-th latent dimension as its supervised set. This results in about 35% (I=50I=50) to 43% (I=100I=100) labeled responses across the whole dataset. An example of the resulting seed set is shown in Table 1. During inference we carried out stochastic estimation by taking one sample for estimating the stochastic gradients. The model is trained by Adam (KingmaB14) and tuned (early stopping, hyper-parameters) on the held-out validation set. We alternately optimised the generative model and the inference network by fixing the parameters of one while updating the parameters of the other.

During reinforcement fine-tuning, we generated a sentence mtm_{t} from the model to replace the ground truth mt^\hat{m_{t}} at each turn and define an immediate reward as whether mtm_{t} can improve the dialogue success (SuVGKMWY15) relative to mt^\hat{m_{t}}, plus the sentence BLEU score (export217163),

where the constant η\eta was set to 0.5. We fine-tuned the model parameters using RL for only 3 epochs. During testing, we greedily selected the most probable intention and applied beam search with the beamwidth set to 10 when decoding the response. The decoding criterion was the average log-probability of tokens in the response. We then evaluated our model on task success rate (SuVGKMWY15) and BLEU score (papineni2002bleu) as in Wen et al (wencond16; wenN2N17) in which the model is used to predict each system response in the held-out test set.

2 Experiments on Goal-oriented Dialogue

Table 2 presents the results of the corpus-based evaluation. The Ground Truth block shows the two metrics when we compute them on the human-authored responses. This sets a gold standard for the task. In the Published Models block, the results for the three baseline models were borrowed from Wen et al (wencond16), they are: (1) the vanilla neural dialogue model (NDM), (2) NDM plus an attention mechanism on the belief trackers, and (3) the attentive NDM with self-supervised sub-task neurons. The results of the LIDM model with and without RL fine-tuning are shown in the LIDM Models and the LIDM Models + RL blocks, respectively. As can be seen, the initial policy learned by fitting the latent intention to the underlying data distribution yielded reasonably good results on BLEU but did not perform well on task success when compared to their deterministic counterparts (block 2 v.s. 3). This may be due to the fact that the variational lower bound of the dataset was optimised rather than task success during variational inference. However, once RL was applied to optimise the success rate as part of the reward function (Equation 21) during the fine-tuning phase, the resulting LIDM+RL models outperformed the three baselines in terms of task success without significantly sacrificing BLEU (block 2 v.s. 4)Note that both NDM+Att+SS and LIDM use self-supervised information.

In order to assess the human perceived performance, we evaluated the three models (1) NDM, (2) LIDM, and (3) LIDM+RL by recruiting paid subjects on Amazon Mechanical Turk. Each judge was asked to follow a task and carried out a conversation with the machine. At the end of each conversation the judges were asked to rate and compare the model’s performance. We assessed the subjective success rate, the perceived comprehension ability and the naturalness of responses on a scale of 1 to 5. For each model, we collected 200 dialogues and averaged the scores. During human evaluation, we sampled from the top-5 intentions of the LIDM models and decoded a response based on the sample. The result is shown in Table 3. One interesting fact to note is that although the LIDM did not perform well on the corpus-based task success metric, the human judges rated its subjective success almost indistinguishably from the others. This discrepancy between the two experiments arises mainly from a flaw in the corpus-based success metric in that it favors greedy policies because the user side behaviours are fixed rather than interactionalThe system tries to provide as much information as possible in the early turns, in case the fixed user side behaviours a few turns later do not fit the scenario the system originally planned.. Despite the fact that LIDMs are considered only marginally better than NDM on subjective success, the LIDMs do outperform NDM on both comprehension and naturalness scores. This is because the proposed LIDM models can better capture multiple modes in the communicative intention and thereby respond more naturally by sampling from the latent intention variable.