Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances

Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, Jie Zhou

Introduction

Recent intelligent open-domain chatbots Adiwardana et al. (2020); Bao et al. (2020); Smith et al. (2020) have made substantial progress thanks to the rapid development of the large-scale pre-training approaches Devlin et al. (2019); Radford et al. (2019); Brown et al. (2020) and the large amount of conversational data Dinan et al. (2019); Baumgartner et al. (2020); Smith et al. (2020). However, effectively modeling the dialogue history in large-scale dialogue pre-training is still challenging.

Most of the previous work on dialogue history modeling mainly fall into two groups. One group of works generally concatenate the dialogue history as the model input and predict the response Zhang et al. (2020); Smith et al. (2020); Bao et al. (2020), named as flat pattern, which is commonly adopted in the large-scale pre-training. However, Sankar et al. (2019) demonstrate that flat concatenation is likely to ignore the conversational dynamics across utterances in the dialogue history. Another group of works employ hierarchical modeling to encode the dialogue history Serban et al. (2016b); Shan et al. (2020); Gu et al. (2020), in which the utterances are separately encoded and then fed into an utterance-level encoder. These approaches lack the history information when encoding each individual utterance, while the history information is essential for understanding dialogue utterances. Thus, all the aforementioned methods are deficient in modeling the dynamic information in the dialogue history.

In this work, inspired by the human cognitive process that humans always consider the goal or influence of the next response before they continue the conversation Brown-Schmidt and Konopka (2015), we propose the DialoFlow to model the dynamic information flow in the dialogue history by addressing the semantic influence brought about by each utterance. As shown in Figure 1, we define the dense representation of the dialogue history at different utterances as the contexts (gray dot line) and the context transformation as the semantic influence brought by each utterance. In particular, our DialoFlow constructs the process of the utterance-level history context flow. Correspondingly, the semantic influence of each utterance can be measured by the difference between two adjacent contexts, which will be further used to guide the current response generation.

Practically, we first employ a transformer to encode the whole conversation to get the dense context representation. Then we design a uni-directional Flow module to capture the context flow on the utterance level, and design three training objectives to model the context flow and measure the semantic influence brought about by each utterance: 1) Context Flow Modeling, which aims to capture the context flow schema. 2) Semantic Influence Modeling, which targets to measure the predicted semantic influence. 3) Response Generation Modeling, which is to generate the response under the guidance of the predicted semantic influence. Furthermore, to demonstrate the effect of modeling dynamic information flow in the dialogue understanding, we propose the Flow score based on the DialoFlow, an automatic reference-free evaluation metric for interactive dialogue evaluation by measuring the semantic influence perplexity.

We pre-train the proposed DialoFlow on the large-scale Reddit comments and conduct experiments on dialogue generation and interactive dialogue quality evaluation. For dialogue generation, DialoFlow achieves significant improvements on the Reddit multi-reference dataset and the DailyDialog dataset compared to the baseline DialoGPT Zhang et al. (2020). For interactive dialogue quality evaluation, our proposed Flow score obtains an impressively high chatbot-level correlation (r=0.9r=0.9) with human ratings on 2200 human-bot dialogues from 11 chatbots.

Our contributions are summarized as follows:

We propose the DialoFlow, a new paradigm to construct the dynamic information flow in the dialogue history by addressing the semantic influence brought about by each utterance. Besides, we design an automatic reference-free evaluation metric Flow score based on the pre-trained DialoFlow for interactive dialogue quality evaluation.

The experimental results illustrate that DialoFlow achieves significant improvements on dialogue generation compared to the DialoGPT, and Flow score shows impressively high chatbot-level correlation (r=0.9r=0.9) with human ratings.

Method

The proposed DialoFlow models the dynamic information flow in the whole dialogue history by addressing the semantic influence brought about by each utterance in sequence.

Before introducing the DialoFlow in detail, we first define some terms. Formally, let D={u1,u2,...,uN}\mathcal{D}=\{u_{1},u_{2},...,u_{N}\} denotes a whole dialogue. And for each utterance uk={uk1,uk2,...,ukT}u_{k}=\{u_{k}^{1},u_{k}^{2},...,u_{k}^{T}\} where uktu_{k}^{t} denotes the tt-th word in the kk-th utterance. We further denote u\textlessk={u1,u2,...,uk−1}u_{\textless k}=\{u_{1},u_{2},...,u_{k-1}\} as the dialogue history at the kk-th utterance. Besides, the dense representation of the dialogue history u<ku_{<k} at the kk-th utterance is represented as the context Ck\mathbf{C}_{k}. And the difference between the new context Ck+1\mathbf{C}_{k+1} at the (kk+1)-th utterance and the previous contexts Ck\mathbf{C}_{k} at the kk-th utterance can be defined as the semantic influence Ik\mathbf{I}_{k} of the kk-th utterance, which can be formulated as:

In our method, DialoFlow first encodes the dialogue history and predicts the future context Ck+1′\mathbf{C}_{k+1}^{\prime} according to all the previous history context C1,C2,...,Ck\mathbf{C}_{1},\mathbf{C}_{2},...,\mathbf{C}_{k}. Then at the response generation stage, the model acquires the predicted target semantic influence Ik′\mathbf{I}_{k}^{\prime}, and generate the target response uku_{k} auto-regressively considering both the predicted semantic influence and the historical sub-sentences. Specifically, as shown in Figure 2, DialoFlow models the context flow by designing a uni-directional Flow module upon the transformer, and we introduce three multi-task training objectives to supervise the context flow, semantic influence, and response generation, which are referred to as context flow modeling, semantic influence modeling, and response generation modeling, respectively.

2 Model Architecture

Figure 2 demonstrates the infrastructure of DialoFlow, which consists of the input embeddings, transformer blocks, a uni-directional Flow module, and a response generator.

Input Embedding. DialoFlow takes the sum of token embedding, segment embedding, and position embedding as the model input. In particular, we insert a special token “[C]” at the end of each utterance, which is used to capture the overall dense representation of the dialogue history. To enhance the modeling of different speakers, we utilize segment embedding containing two types: “[Speaker1]” and “[Speaker2]”.

Transformer Block. A transformer block consists of the following key components: layer normalization, multi-head attention, and feed-forward layers. We employ the pre-normalization used in GPT-2 Radford et al. (2019) instead of the post-normalization used in BERT Devlin et al. (2019), as (Shoeybi et al., 2019) show that the post-normalization leads to performance degradation when the model size increases while pre-normalization enables stable large-scale training. DialoFlow keeps the uni-directional dialogue encoding and enables training on the dialogue level rather than on the context-response setting. We can obtain the history context at the kk-th utterance encoded by the transformer blocks:

where Ck\mathbf{C}_{k} is the hidden states at the position of special token “[C]”. And the hidden states at the position of each token uktu_{k}^{t} in the input sequence are denoted as hkt\mathbf{h}_{k}^{t}.

Flow Module. To capture the dynamic information flow across the dialogue utterances, we design a Flow module to model the context changing scheme. The architecture of the Flow module is the same with one layer of transformer block. The Flow module takes all the previous context {C1,C2,...,Ck}\{\mathbf{C}_{1},\mathbf{C}_{2},...,\mathbf{C}_{k}\} as input and predicts the context at the (kk+1)-th utterance Ck+1′\mathbf{C}_{k+1}^{\prime}:

The predicted semantic influence brought about by the kk-th utterance can be computed as:

Response Generator. DialoFlow generates the utterance uku_{k} with the guidance of the predicted semantic influence Ik′\mathbf{I}_{k}^{\prime}. The response generator contains a feed-forward layer and a softmax layer to convert the hidden states to tokens. When generating the tt-th word, the response generator takes the predicted semantic influence Ik′\mathbf{I}_{k}^{\prime} and the hidden states hkt−1\mathbf{h}_{k}^{t-1} as input, and outputs the probability distribution of the tt-th word:

where ∣V∣|V| refers to the vocabulary size, W1W_{1} and b1b_{1} are learnable parameters.

3 Training Objectives

Different from traditional training approaches with context-response pair, DialoFlow is trained with the whole dialogue containing NN utterances. Correspondingly, we design three training tasks to optimize the model: 1) Context Flow Modeling, 2) Semantic Influence Modeling, and 3) Response Generation Modeling.

Context Flow Modeling. To capture the dynamic context flow, DialoFlow predicts the context at the kk-th utterance Ck′\mathbf{C}_{k}^{\prime} based on the previous context sequence {C1,...,Ck−1}\{\mathbf{C}_{1},...,\mathbf{C}_{k-1}\}. We minimize the L2 distance between the predicted context Ck′\mathbf{C}_{k}^{\prime} and the real context Ck\mathbf{C}_{k}:

Semantic Influence Modeling. To force the effectively modeling of semantic influence brought about by the nn-th utterance at the context Cn−1\mathbf{C}_{n-1}, we design a bag-of-words loss using the predicted semantic influence In′\mathbf{I}_{n}^{\prime}:

where fuktf_{u_{k}^{t}} denotes the estimated probability of the tt-th word uktu_{k}^{t} in the utterance uku_{k}. The function ff is used to predict the words in the utterance uku_{k} in a non-autoregressive way:

where ∣V∣|V| refers to the vocabulary size, W2W_{2} and b2b_{2} are learnable parameters.

Response Generation Modeling. The predicted semantic influence Ik′\mathbf{I}_{k}^{\prime} can also be regarded as a semantic expectation of the kk-th utterance. We incorporate the predicted semantic influence Ik′\mathbf{I}_{k}^{\prime} into the response generation stage to guide the generation. The response generation objective is as follows:

The overall training objective of DialoFlow can be computed as follows:

4 Flow Score

By optimizing with the aforementioned three training objectives, DialoFlow can capture the dynamic information flow across the dialogue history. As the DialoFlow is trained on human-human dialogues, the context flow scheme can be regarded as the general expectation of the dialogue development. Therefore, the closer gap between the semantic influence brought by the chatbot’s utterance and the expectation means the more human-likeness.

Based on the consideration, we propose an automatic reference-free metric Flow score for interactive dialogue evaluation based on DialoFlow. In the human-bot conversation, when the bot generates a new utterance uku_{k}, we measure the similarity between the predicted semantic influence Ik′\mathbf{I}_{k}^{\prime} and the real semantic influence Ik\mathbf{I}_{k} brought about by the utterance uku_{k}, which can be considered as the probability of the human-likeness of the utterance. To compute the similarity between the semantic influences, we measure both the cosine similarity and the length similarity:

Note that we introduce the length similarity to consider the influence of length difference on semantic similarity. For the overall quality of the chatbot in the dialogue, we design a metric, which can be regarded as the dialogue-level perplexity:

where MM denotes the turn numbers of the chatbot utterances and sk+12\frac{s_{k}+1}{2} is to scale the similarity value to $$. A lower Flow score corresponds to better dialogue quality.

Experiments

For model pre-training, we use the Reddit comments, which are collected by a third party and made publicly available on pushshift.io Baumgartner et al. (2020). We clean the data following the pipeline used in the DialoGPT.https://github.com/microsoft/DialoGPT

For response generation, we employ the multi-reference Reddit Test Dataset Zhang et al. (2020) which contains 6k examples with multiple references. We evaluate our pre-trained DialoFlow model on this dataset. The average length of the dialogue history in this dataset is 1.47. To further explore the dynamic information flow in the long dialogue history situation, we choose another popular open-domain dialogue dataset – DailyDialog Dataset Li et al. (2017), in which the average dialogue history length is about 4.66. DialoFlow is fine-tuned on the DailyDialog training set and evaluated on the DailyDialog multi-reference test set Gupta et al. (2019).

For interactive dialogue quality evaluation, we employ the collected data from the Interactive Evaluation of Dialog Track @ The Ninth Dialog System Technology Challenge (DSTC9) Gunasekara et al. (2021), which contains 2200 human-bot conversations from 11 chatbots. For each conversation, there are 3 human ratings on the overall quality (0-5). We calculate the correlation between the results of our proposed metric and the human ratings on the chatbot level. Human-human conversations are always regarded to be better than human-bot conversations. Therefore, we randomly sample 200 human-human dialogues from the BST Smith et al. (2020) dataset to see the metric’s performance on the real human-human conversations.

2 Experimental Setting

Pre-training Details. DialoFlow is pre-trained based on the pre-trained GPT-2 Radford et al. (2019), since Zhang et al. (2020) show that DialoGPT trained from the pre-trained GPT-2 is much better than from scratch. There are three different model sizes: DialoFlow-base, DialoFlow-medium, and DialoFlow-large, which are trained from the pre-trained GPT2-base, GPT2-medium, GPT2-large, respectively. We used AdamW optimizer Loshchilov and Hutter (2019) with 0.01 weight decay and linear learning rate scheduler with 12000 warm-up steps. The learning rate is 2e-4 for the base and medium version and 1e-4 for the large version. We use the batch size of 1024 for all model sizes. We trained the base and medium models for up to 4 epochs and trained the large model for 2 epochs. It costs about two months on 8 Nvidia V100 GPUs to train the large model.

Decoding Details. On the 6K Reddit multi-reference dataset, we use beam search (with beam width 10) on the DialoFlow-medium model and the DialoFlow-large model. We employ greedy search on the DialoFlow-base model, which keeps the same with (Zhang et al., 2020). On the DailyDialog dataset, we fine-tune the pre-trained DialoFlow and DialoGPT, select the checkpoint based on the validation loss, and then use beam search (with beam width 5) for decoding.

3 Baseline

For response generation, we compare our proposed DialoFlow with DialoGPT, a popular dialogue generation model pre-trained on the Reddit Comments. We choose the version trained from pre-trained OpenAI GPT-2 for comparison.

For interactive dialogue evaluation, we compare our metric with the following metrics: 1) FED score Mehri and Eskénazi (2020) is an automatic evaluation metric which uses DialoGPT-large, without any fine-tuning or supervision. FED takes the DialoGPT-large as the user and calculates the likelihood of follow-up utterances based on several pre-set usual human utterances. FED works under the pre-set common human utterances, which can reveal the dialogue quality. 2) Perplexity is used to measure the coherence of an utterance under the dialogue context. We employ DialoGPT-large to measure the perplexity for each utterance of the chatbot. We average the perplexity of all utterances in the whole dialogue as the baseline metric.

4 Evaluation Metrics

For dialogue response generation, we perform automatic evaluation using common reference-based metrics: BLEU Papineni et al. (2002), METEOR Lavie and Agarwal (2007), and NIST Lin and Och (2004). NIST is a variant of BLEU that weights nn-gram matches by their information gain, i.e., it indirectly penalizes uninformative nn-grams such as “I don’t know”, which is a more suitable metric than BLEU when dealing with multi-reference test sets. We also use Entropy Zhang et al. (2018) to evaluate the lexical diversity. We employ the evaluation scripts used by DialoGPT.

For interactive dialogue evaluation, we compute the Pearson and Spearman correlation between the automatic metrics and human ratings. We use the pre-trained DialoFlow-large to compute our proposed Flow score.

Results and Analysis

In this section, we show the performance of our pre-trained DialoFlow model on response generation as well as the performance of Flow score on interactive dialogue quality evaluation.

Table 1 lists the comparison of our pre-trained DialoFlow with the pre-trained DialoGPT on the Reddit multi-reference dataset. Generally, DialoFlow-large achieves the highest score on the NIST and METEOR, while DialoGPT-medium performs better on the BLEU. The performance of our DialoFlow increases with the model size, while the DialoGPT gets the best performance with the medium size rather than the large size. As NIST can effectively penalize common n-grams such as “I don’t know”, the results reveal that DialoGPT tends to generate general responses while our DialoFlow model can create more informative responses. The results also reflect that modeling the dynamic flow is helpful to boost the conversion quality and avoid converging to the general responses. For the lexical diversity, DialoFlow performs similarly with the DialoGPT on Entropy.

The average history length of the multi-reference Reddit dataset is only 1.45, which is a bit short. Thus, we conduct extensive experiments on the DailyDialog dataset (average history length = 4.66) to verify the performance gain on the long dialogue history. As shown in Table 1, DialoFlow shows significant improvements on all model sizes and on all metrics compared to the DialoGPT. The improvements on the DailyDialog dataset demonstrate that our DialoFlow model shows a great capacity to capture the dynamic information flow with a long history. Note that the performance improvement of the DailyDialog dataset is more remarkable than Reddit. In our opinion, conversations in Reddit are mainly the comments in forums, while in DailyDialog the dialogues are derived from daily life. Thus, in the DailyDialog dataset, the context flows are in the more similar schema, and the semantic influences are more predictable compared to the Reddit dataset.

Human Evaluation. We conduct human evaluation on 200 randomly sampled cases from the DailyDialog test dataset using crowd-sourcing. We compare DialoFlow and DialoGPT on the medium version. Each response pair is randomly presented to 3 judges, who rank them for relevance, informativeness, and human-likeness. The overall judge preferences are presented as a percentage of the total, as shown in Table 2. There is a strong preference for the responses generated by DialoFlow. The human evaluation demonstrates that modeling the dynamic information flow is effective for improving the quality of dialogue generation.

Analysis of dialogue history length. Figure 3 shows the performance of our DialoFlow and the DialoGPT on different history lengths. Overall, our DialoFlow achieves better performance on all history lengths. In particular, when history length equals 1, that is, the response is generated based on one history utterance, our DialoFlow also gains a prominent boosting. We attribute it to the guidance of predicted semantic inference.

Ablation Study. To explore the effect of the proposed training objectives, we conduct ablation studies on the medium version of DialoFlow, as shown in Table 1. With all three training objectives, DialoFlow model achives the best performance on NIST and METEOR. When we drop the Semantic Influence Modeling task, the performance slightly decreases. When we further drop the Context Flow Modeling task, which means the end-to-end training, the performance decreases again. The results reveal that the Context Influence Modeling task is effective for dialogue modeling and the Semantic Influence Modeling task can prompt the CIM task.

2 Dialogue Evaluation

Results. Table 4 shows the chatbot-level correlations of different automatic metrics with human ratings on the DSTC9 Interactive Conversation dataset. Our proposed Flow score achieves strong Spearman correlation of 0.90 (p\textless0.001p\textless 0.001) and strong Pearson correlation of 0.91 (p\textless0.001p\textless 0.001). FED only shows moderate correlations with a chatbot-level Spearman correlation of 0.56 (p\textless0.1p\textless 0.1). Perplexity score shows a very weak correlation. On the one hand, the results reveal that our proposed Flow score can effectively estimate the overall chatbot quality. On the other hand, high correlation also demonstrates that the DialoFlow model captures the general dynamic information flow in the natural human-human conversation.

Results Analysis. Table 3 shows the detailed human ratings, FED scores, perplexity, and our proposed Flow score for the 11 chatbots in the DSTC9 Interactive Dialogue Evaluation Track and the sampled human-human conversations. Good automatic metrics should perform well not only on human-bot conversations but also human-human conversations because the ultimate goal of the chatbot is to generate human-like responses. FED performs poorly on the human-human conversations compared to its performance on the other 11 chatbots. Our proposed Flow score takes the human-human conversations as the best one, and the Flow score gap between human-human conversations and the best chatbot is similar to the human rating gap.

Analysis about Flow score. The Flow score can be regarded as the perplexity on the utterance level. There are many different expressions for a specific semantic in natural conversations. Traditional word-level perplexity can estimate the coherence and fluency of the utterance but always performs unstably on variable expressions. The Flow score directly measures the semantic similarity and alleviates the problem with the traditional perplexity.

3 Case Study

Figure 4 shows the 2-D T-SNE visualization of the semantic context of a human-bot conversation encoded by our pre-trained DialoFlow model. The conversation can be split into four topics: greetings (1∼\sim4), talking about why bad day (5∼\sim13), explaining the terrible experience seeing the doctor (14∼\sim18), and discuss swimming (19∼\sim26). Correspondingly, in the visualization, the semantic context flow visualization changes a lot when the topic switches, revealing that DialoFlow can capture the dynamic information flow in the dialogue and effectively measure the semantic influence brought about by each utterance. Besides, it seems like that different speakers keep their own context flows.

Related Works

Multi-turn dialogue modeling. The modeling of multi-turn dialogue history mainly falls into two categories: 1) Flat concatenation. These works directly concatenate the dialogue history as the input sequence Zhang et al. (2020), which can not capture the information dynamics. 2) Hierarchical architectures. The hierarchical architecture is commonly used in the dialogue history understanding. Serban et al. (2016a) propose the hierarchical LSTM to generate responses. Li et al. (2019) introduce an incremental transformer to capture multi-turn dependencies. Shan et al. (2020); Gu et al. (2020) employ pre-trained BERT to encode individual utterances and design the utterance-level encoder to capture the turn-level structure. These methods suffer from the lack of context word-level information when encoding utterances. Different from these methods, our DialoFlow takes full advantage of both word-level information and utterance-level dynamic information. Besides, the proposed DialoFlow is pre-trained on the large-scale open-domain dialogue dataset.

Pre-trained models for dialogue generation. Recent advances in pre-trained language models have great success in dialogue response generation. DialoGPTZhang et al. (2020), Plato-2 Bao et al. (2020), MeenaAdiwardana et al. (2020), and BlenderSmith et al. (2020) achieve strong generation performances by training transformer-based language models on open-domain conversation corpus. In contrast, our proposed DialoFlow focuses on modeling the dynamic information flow in the pre-training process, and we design three training objectives to optimize the model.

Interactive Dialogue Evaluation. Evaluating the quality of interactive dialogue automatically is a challenging problem, as there is no gold reference for the utterances. Mehri and Eskénazi (2020) propose the FED score, an automatic dialogue evaluation metric using pre-trained DialoGPT-large, which works with pre-set common human comments, like “It is interesting to talk with you.”, revealing the dialogue quality. However, the FED score has limited performance on those dialogues without apparent comments. Our Flow score entirely depends on the pre-trained DialoFlow model with no need for human integration.

Conclusion and Future work

In this work, we proposed the DialoFlow to model the dynamic information flow across dialogue utterances by addressing the semantic influence brought about by each utterance. Specifically, we employed a uni-directional Flow module to model the context flow and designed three training objectives to optimize the DialoFlow model. Besides, upon the DialoFlow, we proposed the Flow score, an automatic reference-free evaluation metric for interactive dialogue evaluation, with the pre-trained DialoFlow. Experiments on response generation and dialogue evaluation all demonstrate that our method could effectively capture the dynamic information flow across utterances. For future work, we would like to apply the DialoFlow to the task-oriented dialogue and explore the application on the long text generation, such as the story generation.

Acknowledgement

We sincerely thank the anonymous reviewers for their thorough reviewing and valuable suggestions. This work is supported by National Key R&D Program of China (NO. 2018AAA0102502).

References