Exploiting Persona Information for Diverse Generation of Conversational Responses
Haoyu Song, Wei-Nan Zhang, Yiming Cui, Dong Wang, Ting Liu
Introduction
Building human-like dialogue agents to pass the Turing Test Turing (1950) has been a long-term goal of artificial intelligence. Open-domain conversations need to be diversified Li et al. (2015) and sustained Li et al. (2016b), as the primary goal of an open-domain chatbot is to establish a connection with the user and to accompany the user over a long period of time Shum et al. (2018). Because of the vast quantity of rational responses for an open-domain conversational message, sequence-to-sequence (Seq2Seq) model Sutskever et al. (2014) has been widely used for conversation modeling Vinyals and Le (2015); Shang et al. (2015); Serban et al. (2016). Seq2Seq model enables the incorporation of rich context when mapping between consecutive dialogue turns and is trained by predicting one target response in a given dialogue context using the maximum-likelihood estimation (MLE) objective function.
Despite the successful application of Seq2Seq in dialogue modeling, it is still quite impossible for current dialogue agents to pass the Turing Test, while one of the reasons is the lack of a coherent persona Vinyals and Le (2015). Endowing an open-domain chatbot with persona is challenging, yet impactful to deliver more realistic and engaging conversations. First, presenting a coherent persona is crucial for a chatbot to gain confidence and trust from users Li et al. (2016a); Qian et al. (2017). Second, grounded on predefined persona, a chatbot can be trained to both ask and answer questions about personal topics, thus making users more engaged Zhang et al. (2018). Due to these reasons, there has been a growing interest in modeling persona within the Seq2Seq model.
However, human dialogue has a unique property that many other learning tasks do not: given a dialogue context, there may exist more than one valid responseDenoted as one-to-many property. Zhao et al. (2017); Rajendran et al. (2018). As illustrated in Table 1, the potential candidate responses cover different aspects and are non-repetitive. Such responses can not only provide more candidates for different dialogue strategies Yu et al. (2016) but also enrich the diversity. Nevertheless, the one-to-many property is not taken into account in the present persona dialogue models. Although recent Seq2Seq-based persona models were successful, they still suffer from the low diversity problem Jiang and de Rijke (2018) brought by Seq2Seq. They tend to generate repetitive and non-specific responses like “I don’t know” Li et al. (2015, 2016b), regardless of the conversational context.
In this paper, we follow the definition of persona in Zhang et al. Zhang et al. (2018) (as exemplified in Table 1), and work on the explicit textual persona. This is because the unstructured persona data is easier to collect and more authentic For example, 80% of posts on Twitter are about personal emotional state, thoughts or activities Naaman et al. (2010).. We propose a memory-augmented architecture to exploit the explicit persona texts and incorporate a latent variable that can capture response variations under the conditional variational autoencoder model, rather than learn to predict only one response using MLE. To generate diverse persona-based responses, our model draws samples from the latent variable and combines two different decoding strategies to generate words or copy words directly from persona texts. We evaluate our model on the ConvAI2 persona-chat dataset and yield promising results in (i) incorporating persona information into responses and (ii) generating diverse and engaging responses. In summary, the contributions of this paper are three-fold: First, we propose to address the one-to-many property in end-to-end model. To the best of our knowledge, this is the first work that jointly models persona text and response diversity. Second, we present a novel end-to-end model that incorporates persona texts into diverse responses (Section 3). Third, experimental results show that our model can generate persona-based responses as well as deliver more diverse and more engaging responses than baselines (Section 4).
Related Work
This work is related to the recent advancements in persona-based dialogue models and the efforts of improving diversity in dialogue generation.
In recent years, the dialogue research community has shown a growing interest in the modeling of personality. Li et al. Li et al. (2016a) first incorporated implicit persona in dialogue generation using user embeddings, which projects each user into a dense vector. Kottur et al. Kottur et al. (2017) proposed a neural dialogue model which conditioned on both speakers and context history. However, these two models rely heavily on speaker-tagged dialogue data, which is more expensive and sparser. To present a coherent personality, Qian et al. Qian et al. (2017) defined several profile key-value pairs, including name, gender, age, location, etc., and explicitly expressed a profile value in the response. Recent works brought new persona models as well as high-quality data. Zhang et al. Zhang et al. (2018) contributed a persona-chat dataset, and they further proposed two generative models, persona-Seq2Seq and Generative Profile Memory Network, to incorporate persona information into responses. Yavuz et al. Yavuz et al. (2018) explored the use of copy mechanism in persona-based dialogue models. Our work follows the line initiated by Zhang et al. Zhang et al. (2018) but makes a step further by generating diverse persona-based responses.
On the other hand, large-scale dialogue generation still suffers from the tendency of generating dull and generic responses. The efforts to tackle the low diversity problem fall into two major categories. The first focuses on the objective function of the dialogue generation model, and they argue that the MLE objective is unable to approximate the real-world goal of the conversation Li et al. (2015, 2016b). The other line of research tries to augmenting the encoder of the generative model with richer information Xing et al. (2016); Zhao et al. (2017). As one of these improvements, Zhao et al. Zhao et al. (2017) proposed a dialogue model adapted from conditional variational autoencoders Yan et al. (2016). The dialogue CVAE learns a latent variable that captures discourse-level variations and generates diverse responses by drawing samples from the learned distribution.
We also borrow ideas from Sukhbaatar et al. Sukhbaatar et al. (2015) in the modeling of external memory. The proposed Persona-CVAE can be seen as an extension of dialogue CVAE, with several main differences: (i) the combination with external persona memory, (ii) the decoding strategies, and (iii) the ability to deliver persona-based responses.
Persona-CVAE
In this section, we describe the components of the proposed Persona-CVAE model in detail. As illustrated in Figure 1 (b), the conditional graphical model shows the core idea of proposed Persona-CVAE. We assume that the generation of depends on , and . y relies on x, p and z. The main objective is to learn the conditional probability and , with parameters and . The variational parameters are learned jointly with the generative model parameters , as described in Section 3.2.
The task can be formally defined as: given an input message and a set of persona texts , the goal is to generate diverse responses , based on both input message and persona texts.
The diverse in responses means that the utterances in are non-repetitive, cover different aspects and differ in words.
2 Persona based CVAE
The proposed model uses four random variables to represent each dyadic conversation: the dialogue context , the target response , the set of persona texts and a latent variable . is used to capture the latent distribution over the valid responses. Further, is composed of both input messages and responses in dialogue history and can be denoted as , where ( is a single word). Similarly, consists of several unstructured persona texts: , where .
Figure 2 demonstrates an overview of our model. The sentence encoder is a bidirectional recurrent neural network (BRNN) Schuster and Paliwal (1997) and is used to encode a single sentence (e.g., the target response). We concatenate the last hidden state of both forward and backward RNN to capture the semantic information from both sides, formulized as , where is the sequence length. For the context encoder, we use a hierarchical encoding strategy. First, each sentence in context is encoded by the sentence encoder to get a latent representation . Then a single layer forward RNN is used to encode the sentence representations into a final state .
Persona differs from dialogue context in two main aspects: (i) it keeps unchanged throughout the conversation, and (ii) it is only the unilateral information (here is for the chatbot), while two sides share the same dialogue context. Therefore, we conjecture that modeling persona differently from dialogue context, with a memory augmented architecture, will be beneficial for the model’s performance. The details will be discussed in Section 3.3.
Then we define the conditional distribution over the above random variables. As mentioned, we define the generation process as a conditional distribution and our objective is to approximate and with deep neural networks. Following the definition in Zhao et al. Zhao et al. (2017), we refer to as the response decoder and as the prior network. Further, to approximate the true posterior distribution, we refer to as the recognition network.
CVAE is trained to maximize the conditional log-likelihood, but this involves an intractable marginalization over the latent variable. Previous works Sohn et al. (2015); Yan et al. (2016) have shown that CVAE can be efficiently trained with the Stochastic Gradient Variational Bayes (SGVB) Kingma and Welling (2013) by maximizing the variational lower bound of the conditional log-likelihood. As proposed in Zhao et al. Zhao et al. (2017), we also assume the latent variable follows multivariate Gaussian distribution with a diagonal covariance matrix. And the variational lower bound of Persona-CVAE can be written as:
Since we assume latent variable follows isotropic Gaussian distribution, the recognition network and the prior network . We sample either from in training or in testing. To make the sampling operation differentiable, we use the reparametrization trick Kingma and Welling (2013) and we have:
We project the concatenation of , and to a vector, to serve as the initial state of response decoder. Then the response decoder predicts words under the Decoding Strategies (Section 3.4.2).
3 Persona Memory
The persona memory module has two roles: first is to encode the persona texts into a dense vector (persona memory) and sencond is to select a persona to be expressed in the response. Notice that persona is not always needed in the response, and thus model selects persona from .
The first role of the persona memory module is defined by multi-hops attention over persona texts and dialogue context. We convert the persona texts into memory vectors of dimension , computed by embedding each in a continuous space using an embedding matrix of size . We use the embedded dialogue context (with the same dimensions ) as an input. We compute the match between context vector and each memory vector in the embedding space by:
where . Meanwhile, each persona has an output vector (by another embedding matrix ). A single hop output is computed by a sum over the output vector , weighted by the probability from equation (13):
And -hops attention are stacked in the following way:
where . For the embedding matrix, we use an adjacent strategy, i.e. . Finally, the is used as our persona memoryIn our experiments, performs better than or , but there is no significant performance change when or ..
The second role of the persona memory module is to decide which persona should be expressed in the generated response. This is implemented as:
where is a weight matrix, and is the persona memory. is the sampled latent variable, as described in Section 3.2. The latent variable captures some information that is not presented in the dialogue context and is beneficial for selecting the optimal persona. As aforementioned, in some cases, there is no need to incorporate persona information, so we need an MLP here to make the decision. The model selects a persona with the maximal probability, i.e. , where . Then the words from selected persona are further used in response decoder.
4 Decoding Strategy
We propose two decoding strategies (general purpose and special cases) to better express persona information.
The soft decoding strategy assumes that at each decoding step , there is a distribution over the two types {persona, other}. Similarly, according to the selected persona words, the entire vocabulary set is also divided into two parts, i.e. {persona words, other words}.
At time step , the response decoder first estimates word generation probabilities over the two vocabulary sets independently, denoted as and :
and then computes the type distributions:
where is the hidden state of decoder RNN cell. The final probability of generating a word is a mixture of type-specific generation distributions where the coefficients are type probabilities:
4.2 Force Decoding Strategy (FDS)
We propose force decoding strategy from the observation that persona texts can directly serve as a valid response under some circumstances. In the decoding process, if the last words of the partially decoded sequence are the same as the first words of selected persona, the input words to decoder RNN cell will come from the selected persona words successively in next few steps:
where is from the selected persona. At time step , is also the output word. The decoder uses FDS only once in a decoding process. Though this strategy is simple, we found it works well in some situations.
5 Training and Optimization
The Persona-CVAE is trained to maximizing the variational lower bound of the conditional log-likelihood in Formula (3.2). Persona memory module and type distribution are trained by two standard cross-entropy loss function.
As addressed in Bowman et al. Bowman et al. (2015), the KL annealing trick is necessary for the training of RNN-based CVAE. Besides, another essential technique for the model to work is the bag-of-word loss Zhao et al. (2017). Finally, the Persona-CVAE model is trained with the sum of these losses and optimized through backpropagation.
Experiment
We perform experiments on the recently released ConvAI2 benchmark dataset, which is an extended version (with a new test set) of persona-chat dataset Zhang et al. (2018). The conversations are obtained from crowdworkers who were randomly paired and asked to act the part of a given persona.
This dataset contains 164,356 utterances in over 10,981 dialogues and has a set of 1,155 personas, each consisting of at least four profile texts. The testing set contains 1,016 dialogues and 200 never seen before personas. We set aside 800 dialogues together with its profile texts from the training set for validation. The final data have 9,181/800/1,016 dialogues for train/validate/test.
To train the persona memory module, we label each utterance with its corresponding persona.We first compute word inverse document frequency: , where is from the GloVe index via Zipf’s law Zhang et al. (2018). If the utterance has the highest tf-idf similarity with a profile sentence and the similarity is higher than a threshold, then we label the utterance with this profile sentence. Otherwise, the utterance’s persona label is none. We also label the position that shares a word with profile sentence to learn the type distribution.
2 Baselines
We compared the proposed Persona-CVAE model with five state-of-the-art generative baseline models. The five models fall into two categories, i.e., persona-free model and persona-based model: Seq2Seq: a general persona-free model with context attention mechanism Shang et al. (2015). Dialogue CVAE: a persona-free model that uses a latent variable to learn a distribution over potential conversational intents and improves the discourse-level diversity of responses Zhao et al. (2017). Persona-Seq2Seq: a persona-based model that prepends persona texts to the input sequence , i.e., ,where denotes concatenation Zhang et al. (2018). Generative Profile Memory Network (GPMN): a persona-based generative model that encodes each profile sentence as individual memory representations in a memory network Zhang et al. (2018). Oracle Seq2Seq+Copy (Oracle. Copy): a persona-based model with copy mechanism. This model is similar to Persona-Seq2Seq, but it selects the persona text whose tokens have the highest unigram tf-idf similarity to the ground truth response (the Oracle). It achieves the best results in a variety of copy-based persona models Yavuz et al. (2018). The aim of this experiment is to have a better comparison with the copy-based persona models, although using the ground truth for persona selection is not fair.
Among these baselines, we fine-tuned Seq2Seq, Dialogue CVAE and Oracle Seq2Seq+Copy on the ConvAI2 persona-chat dataset, while the other two methods are using the latest released models from Zhang et al. Zhang et al. (2018) in ParlAI. Moreover, to make different models comparable, we generate responses from the Seq2Seq based models by sampling from the softmax. For the latent variable based models, we sample times from the latent to generate responses.
3 Experimental Settings
In our experimentsCode available at: https://github.com/vsharecodes/percvae ., the RNN is two-layer GRU with a 500-dimensional hidden state. The dimension of word embedding is set to 300, and thus the persona memory size is also 300. The vocabulary size is limited to 20,000. The latent variable size is set to 100. KL annealing steps are set to 10,000. We train the model with a minibatch size of 32 and use Adam optimizer with an initial learning rate of 0.001. All parameters are initialized by sampling from a uniform distribution.
4 Automatic Evaluation
Automatically evaluating an open-domain generative dialogue model is still an unsolved research challenge. Although metrics such as BLEU and perplexity have been used for dialogue quality evaluation Vinyals and Le (2015); Serban et al. (2016), recent work has found that BLEU shows very weak correlation with human judgment Liu et al. (2016). Further, perplexity is not computed on a per-response basis. Since the goal of the proposed model is not to predict only one target response, but rather diverse persona-based responses, we do not employ BLEU or perplexity for evaluation. Following the one-to-many property, we use two metrics to evaluate how well our model enhancing diversity and incorporating persona information: Distinct-K (Dtinct-K): this metric calculates the number of distinct k-grams in generated responses and is scaled by the total number of generated tokens to avoid favoring long responses Li et al. (2015). The Distinct-1 and Distinct-2 are thus the token ratios for unigrams and bigrams. This metric is an indicator of word-level diversity for generated responses. Persona Coverage (P. Cover): we propose this metric to evaluate how well persona information is expressed. We assume that there are predefined profile sentences and model generates hypothesis responses . The response-level persona coverage is defined as:
where is the counts of shared words weighted by words’ idf (as described in Section 4.1). Assume that the words’ set shared by and are and the inverse document frequency for is , then we have:
and denotes the number of words in .
The final score is averaged over the entire test dataset, and we report the performance in Table 2. As can be seen, even if =1, our model still outperforms all the baselines in diversity but is inferior to the Oracle model in the Persona Coverage. When =5 and 10, the proposed model obtains the best performance in all metrics compared with all the baseline models. Due to the number of persona texts in each dialogue (generally 4 or 5), generating too many responses leads to a small decline in the Persona Coverage.
The results show that under the different number of , our model always generates diverse responses as well as effectively incorporates persona information into multiple responses. To better evaluate the diversity of generated responses, the following experiments are all based on =5.
5 Human Evaluation
We also explore two settings for human evaluation. In both settings, we present persona texts, input message, as well as generated responses.
In the first setting, we employ judges to evaluate a random sample of 100 items per model according to three metrics, based on a 1/0 scoring schema, similar as Qian et al. Qian et al. (2017): Engagingness (Engage.): a response is engaging only when it is appropriate, interesting and easy to answer. This is the overall score of generated responses. Variety: this metric measures the variety of generated responses. Score 1 indicates that there is a significant difference in the linguistic patterns and wordings and score 0 otherwise. Persona Detection: we measure the model’s ability to incorporate persona information by displaying two possible persona texts: one is the true persona that the model used and another is randomly sampled from the rest of persona texts. Then the judges will determine which is more likely to be the persona used by model according to generated responses.
We calculated the Fleiss’ kappa to measure inter-rater consistency. Fleiss’ kappa for Engagingness, Variety and Persona Detection is 0.5149, 0.8281 and 0.6656, indicating “moderate agreement”, “almost perfect” and “substantial agreement” respectively.
Results in Table 3 support: First, Our model performs better than all baselines in all subjective metrics. Second, Our model successfully incorporates persona information. Third, Our model delivers diverse and engaging responses.
In addition, we performed a Preference Test in three persona-based models. In this setting, the judges are presented with pairwise items and are asked to decide which of the two responses is of higher quality. And a partial ordering relation about the overall quality is given: best response quality response variety persona coverage. The judges first need to do a re-rank in mind and decide which model offers a better top response. If no difference then considering variety in responses, otherwise whether persona text is expressed. As shown in Table 4, the proposed model is significantly (2-tailed t-test, ) preferred than the other two baselines. In the pairwise setting, the diverse persona-based responses are more attractive to users.
We present some generated examples in Table 5.
6 Ablation Tests
In order to investigate the influence of different decoding strategies, we also conducted ablation tests where one or two decoding strategies were removed from Persona-CVAE. The results are shown in Table 6. As we can see, after removing the SDS, the Persona Coverage decreases the most, indicating the SDS contributes to a higher persona information coverage. However, SDS also shows a negative effect on Distinct-1. This may be because the SDS encourages the use of words from profile sentences, which leads to low-frequency words more difficult to appear.
Conclusion and Future Work
In this paper, we focus on the diverse generation of conversational responses based on chatbot’s persona and propose a memory-augmented architecture named Persona-CVAE. We conduct experiments on the ConvAI2 dataset. Experimental results show that our model outperforms the state-of-the-art methods, especially in terms of diversity and persona integration. For future work, we will explore modeling the persona information of users in open-domain conversations.
Acknowledgments
The paper is supported by the National Natural Science Foundation of China under Grant No.61772153. We also thank the anonymous reviewers for their helpful comments.