GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, Xiaodan Liang
Introduction
Coherence, what makes dialogue utterances unified rather than a random group of sentences, is an essential property to pursue an open-domain dialogue system aiming at conversing with humans. Although open-domain dialogue systems have achieved significant progress and performed much more human-like skills in recent years (Zhou et al., 2020; Adiwardana et al., 2020; Roller et al., 2020), automatically measuring dialogue coherence for state-of-the-art open-domain dialogue models is still an open and under-explored research problem attributing to the open-ended nature of dialogue (See et al., 2019).
Statistic-based automatic metrics, such as BLEU (Papineni et al., 2002), mostly rely on the degree of word overlap between a dialogue response and its corresponding gold response. However, due to the ignorance of the underlying semantic of a response, they are biased and correlate poorly with human judgements in terms of response coherence (Liu et al., 2016). To overcome this issue, some learning-based metrics were proposed to train a coherence scoring model by considering the utterance-level semantics, such as ADEM (Lowe et al., 2017), RUBER (Tao et al., 2018), and BERT-RUBER (Ghazarian et al., 2019). However, a coherent real-world dialogue should be not only coherent among utterances but also smooth at topic transition. As shown in Figure 1, the topics inside a coherent dialogue are close to each other in the commonsense graph, which embodies a smooth topic transition. Although the above metrics have demonstrated higher correlations with human judgements than statistic-based metrics, they only model dialogue coherence at utterance level without explicitly considering the fine-grained topic transition dynamics of dialogue flows.
To address the above problems, we propose a new automatic metric for open-domain dialogue systems, named as Graph-enhanced Representation for Automatic Dialogue Evaluation (GRADE), which explicitly models topic transition dynamics by reasoning over dialogue graphs and incorporates them into utterance-level contextualized representations. As a result, our method can capture more accurate semantic transition information, thus measuring dialogue coherence in a more human-like manner.
Specifically, our GRADE consists of two semantic extraction branches. One branch deploys BERT (Devlin et al., 2019) to learn the coarse-grained utterance-level contextualized representations, while another learns the fine-grained topic-level graph representations by constructing topic-level dialogue graphs and applying a graph neural network on the graphs to model the topic transition dynamics. As to the dialogue graph construction, we determine nodes and edges by utilizing the evidence from the commonsense knowledge graph, ConceptNet (Speer et al., 2017), including k-hop neighboring representations and hop-attention weights. GRADE is trained in an unsupervised manner with data automatically generated by a negative sampling strategy considering both lexical and semantic aspects rather than random sampling adopted by previous works (Tao et al., 2018; Ghazarian et al., 2019). Experimental results show that GRADE significantly outperforms other state-of-the-art metrics in terms of the Pearson and Spearman correlations with human judgements and can generalize to unseen chit-chat datasets well.
Our contributions are summarized as follows:
We propose GRADE, a novel automatic coherence metric for evaluating open-domain dialogue systems, which is the first attempt to introduce graph reasoning into dialogue evaluation.
We demonstrate the effectiveness of incorporating graph information into dialogue evaluation. Extensive experiments show that GRADE has significantly stronger correlations with human judgements than other state-of-the-art metrics.
We construct and release a new large-scale human evaluation benchmark with 11910 human annotations to the research community for encouraging future study on automatic metrics.
The code and data are available at https://github.com/li3cmz/GRADE.
Related Work
Automatic evaluation for open-domain dialogue systems is difficult since there are many appropriate responses for a dialogue context under the open-domain setting, known as the one-to-many problem (Zhao et al., 2017).
Initially, the statistic-based metrics in language generation tasks are adopted for dialogue evaluation, such as BLEU (Papineni et al., 2002), METEOR (Banerjee and Lavie, 2005) and ROUGE (Lin, 2004). These metrics use statistical rules to measure the surface similarity between generated responses and reference responses. For example, BLEU computes the geometric average of the n-gram precisions. However, they can not cope with the one-to-many problem and have weak correlations with human judgements (Liu et al., 2016).
In recent years, learning-based metrics have increasingly attracted interest from researchers. ADEM proposed by Lowe et al. (2017) achieves higher correlations with human judgements than the statistic-based metrics, which is trained with human-annotated data in a supervised manner. However, it is time-consuming and expensive to obtain large amounts of annotated data. To reduce the cost of obtaining annotated data, Tao et al. (2018) trained their metric RUBER with auto-constructed negative samples in an unsupervised manner.
With the advances of the pre-trained language model, BERT (Devlin et al., 2019) has been adopted for dialogue or NLG evaluation. Ghazarian et al. (2019) proposed BERT-RUBER, which outperforms RUBER significantly by incorporating BERT embeddings. BERTScore (Zhang et al., 2020) performs soft-overlap between candidate and reference sentences by using BERT embeddings directly without fine-tuning, and has been shown to correlate with human judgment robustly. Besides, Sellam et al. (2020) introduced BLEURT by further training regular pre-trained BERT with an elaborate pre-training scheme and fine-tuning on small amounts of rating data, which yields superior results.
Note that our model differs from the above learning-based metrics in two folds. First, our metric is trained with high-quality negative samples that are similar to the ground truths in both lexical and semantic aspects instead of randomly sampling. Second, different levels of representations are considered in our GRADE, especially the fine-grained topic-level graph representation.
GRADE Metric
In this paper, we focus on designing an evaluation metric that can automatically assess the coherence of responses produced by dialogue models. Formally, given a dialogue context and a response , where each is a token in the context and each is a token in the response, our goal is to learn a function that predicts the coherence score .
As illustrated in Figure 2, our GRADE predicts a coherence score between a context and a response in three steps: (1) producing the utterance-level contextualized representation (Section 3.1); (2) generating the topic-level graph representation (Section 3.2 and Section 3.3); (3) predicting the coherence score based on and (Section 3.4). The training details of our GRADE is elaborated in Section 3.5.
We use BERT (Devlin et al., 2019) to encode the context and the response . The pooled output feature of BERT is then taken as the utterance-level contextualized representation :
2 Dialogue Graph Construction
We construct a topic-level dialogue graph based on and , denoted as , where V is a set of topic nodes and E is a set of edges between topics. The details are described as follows.
where is the maximum number of hops taken into account and is set as 2, is the hop neighboring nodes of in the ConceptNet graph, and are the weight matrix and bias vector respectively.
Edges. Since our goal is to predict a coherence score of a response based on a context, we only consider the edges between the context nodes and the response nodes . In other words, the edges only exist between each context-topic node and each response-topic node . Moreover, we consider as a weighted undirected graph and assign a weight to each edge of by heuristically using the hop information in the ConceptNet commonsense graph, named as hop-attention weights. Specifically, let the weighted adjacency matrix of as , then the hop-attention weight of the edge between the nodes and (i.e., ) is determined by:
where indicates the shortest path between and over the ConceptNet graph. As a result, the distances between topic nodes are redefined and the nodes that are far away from each other will have low weight values. After determining the edges, we randomly deactivate a certain number of edges from at each training step to prevent over-smoothing, and normalize the adjacency matrix (Rong et al., 2020):
where is the augmented normalized adjacency matrix, is the corresponding degree matrix of and is the identity matrix.
3 Topic-level Graph Reasoning
We explicitly model the topic transition dynamics by reasoning over the constructed topic-level graph via two steps: aggregation and combination (Hamilton et al., 2017).
In the first step, we apply the graph attention network(GAT) (Veličković et al., 2018) to aggregate neighboring information of each node . The aggregated representation at the layer for the node is formulated as follows:
In the second step, the aggregated representation is combined with the node representation to get the updated node representation :
Finally, the topic-level graph representation is obtained by:
where is the node representation at the last layer, represents mean pooling and is a fully-connected layer with a ELU activation.
4 Coherence Scoring
To compute the coherence score , the contextualized representation and the graph representation are concatenated together and fed into a multi-layer perceptron(MLP) to transform the high-dimensional representation into a real number:
where , and are three different fully-connected layers whose activation functions are ELU, ELU and sigmoid, respectively.
5 Training
Training Objective. Inspired by Tao et al. (2018), we train our GRADE in an unsupervised manner. Given a dataset , where and are a ground-truth context-response pair and is a false response for the context selected by using negative sampling described in the next paragraph, then GRADE is trained to predict a higher score for each ground-truth response than its corresponding false response by minimizing the following margin ranking loss:
where N is the size of the dataset, is a margin value set as 0.1, and are the coherence scores of and respectively in the example.
Following Sato et al. (2020), we select the false response that is similar to the ground-truth response , instead of random sampling adopted in previous works (Tao et al., 2018; Ghazarian et al., 2019). Overall, we generate negative samples by two sampling methods: lexical sampling and embedding-based sampling. For lexical sampling, we use Lucenehttps://lucene.apache.org to retrieve utterances that are related to the ground-truth response from the training set, and select the middle one in the retrieved utterances as the false response . For embedding-based sampling, we first randomly sample 1000 utterances and take the utterances with the top-5 cosine similarity against the ground-truth response .All the utterances are encoded with BERT. The false response is then randomly selected from the top-5 utterances.
Experiments
Dialogue Models. We consider both retrieval-based and generation-based dialogue models to obtain diverse responses for metric evaluation so that the performance of the metrics can be assessed comprehensively. Specifically, we first deploy Transformer-Ranker and Transformer-Generator from the ParlAI platform (Miller et al., 2017), where the former is retrieval-based and the latter is generation-based. Besides, we also deploy two state-of-the-art dialogue models, BERT-Ranker (Urbanek et al., 2019) and DialoGPT (Zhang et al., 2019) that can output more human-like responses than Transformer-Ranker and Transformer-Generator.
Baseline Metrics. We compare our GRADE with seven dialogue metrics, consisting of three statistic-based metrics: BLEU (Papineni et al., 2002) ROUGE (Lin, 2004) and METEOR (Banerjee and Lavie, 2005), four learning-based metrics: ADEM (Lowe et al., 2017), BERT-RUBER (Ghazarian et al., 2019), BERTScore (Zhang et al., 2020) and BLEURT (Sellam et al., 2020). Note that, for comparison, we only present the BLEU-4 results for BLEU metric, and ROUGE-L for ROUGE, BERTScore-F1 for BERTScore.
Datasets. We use the DailyDialoghttp://yanran.li/dailydialog (Li et al., 2017) dataset which contains high-quality open-domain conversations about daily life including diverse topics, to learn our GRADE. In addition, another two chit-chat datasets, ConvAI2http://convai.io (Dinan et al., 2019) and EmpatheticDialogueshttps://github.com/facebookresearch/EmpatheticDialogues (Rashkin et al., 2019), are considered as unseen datasets to verify the transferability of the metrics. The details of the datasets are provided in Appendix A.
Implementation Details. We use for the utterance-level contextualized encoding. For the graph reasoning module, the GAT layer is set as 3 and the number of heads is 4, where both the input and output dimensions are 300. To train GRADE, we use Adam (Kingma and Ba, 2014) with , , and set batch size as 16, learning rate as 2e-5. Our GRADE is implemented with a natural language processing toolkit, Texar-Pytorch (Hu et al., 2019).
Human Judgements. We collected human judgements from Amazon Mechanical Turk (AMT). Each survey contained six questions, including five coherence questions and one attention check question. The submissions failed in the attention check are directly discarded. For each coherence question, workers were provided with a context-response pair and asked to assess the coherence between the context and the response on a scale of 1-5 (not coherent at all to very coherent). Each pair was assessed by 8 to 10 individual workers. In total, there are 1200 different pair and 11910 human annotations from 217 unique workers, as the final human judgements. As shown in Figure 3, the distributions of human judgements are balanced from score 1 to 5. Moreover, It also demonstrates that the dialogue models we selected are diverse in performance, which helps comprehensively assess the abilities of the metrics.
2 Experimental Results
DailyDialog Dataset. The test set results of the DailyDialog dataset are presented in Table 1. Overall, our GRADE obtains the highest correlations with human judgements in average. Although the Spearman value of GRADE on the Transformer-Ranker is lower than BLEURT which is trained on a very large-scale dataset, the averaged correlation result of GRADE is 1% higher than BLEURT. Besides, all the correlation results of GRADE are statistically significant with p-value <0.05, which is more reliable than the baselines.
Other Unseen Datasets. To verify the transferability of our GRADE, we further evaluate the human correlations of GRADE compared with other baselines on two unseen chit-chat datasets, ConvAI2 and EmpatheticDialogues. Results in Table 1 show that GRADE can easily adapt to other unseen datasets without any re-training and obtain more stable and higher correlations with human judgements than the baseline metrics. It is noteworthy that all Pearson and Spearman correlations of GRADE are statistically significant with p-value < 0.05, and most of them are with p-value < 0.01. Particularly, GRADE achieves a significant Pearson correlation of 0.606 and Spearman correlation of 0.617 for evaluating Transformer-Generator on the ConvAI2 dataset, bringing an improvement of 0.411 (Pearson) and 0.417 (Spearman) compared with BLEURT. Furthermore, Table 2 presents the correlation results of GRADE and other baselines for evaluating two state-of-the-art dialogue models, BERT-Ranker and DialoGPT. Our GRADE significantly outperforms the baseline metrics on human correlations, which shows that GRADE is better at evaluating the coherence of high-quality responses. Besides, Figure 4 illustrates the scatter plots against human judgements for DialoGPT on the ConvAI2 dataset. We can see that the scores predicted by GRADE are closer to the human scores than the baseline metrics, which intuitively shows the superiority of our GRADE.
3 Ablation Studies
We perform ablation studiesFor each ablation experiment, We run five times and take the averaged result since the results fluctuate over different runs (more details in Section 5). for the main components of GRADE to better analyze their relative contributions. The results are shown in Table 3.
Does the negative sampling strategy work? We first verify the effectiveness of our negative sampling strategy by replacing it with random sampling. As shown in Table 3, adopting the random sampling strategy hurts performance significantly with a 6.6% drop in average, which indicates the importance of our negative sampling strategy.
Does the graph work? To prove the contribution of our graph components, we perform three ablations respectively: 1) remove the entire graph branch of GRADE; 2) remove the k-hop neighboring representations used for initializing the node representations in the dialogue graph; 3) remove the hop-attention weights used for computing a weight for each edge in the dialogue graph. Consequently, the performance of GRADE decreased after removing the graph branch or one of the components in the graph branch.
How much graph information we need? Finally, we explore the number of k-hop neighboring representations needed for initializing the dialogue graph’s nodes in two aspects: the maximum number of hops (refer to the in Equation 3), and the number of neighboring nodes in the hop (denoted as , i.e., the number of nodes in in Equation 3). By comparing the results among the first row and the last three rows in Table 3, we confirm that incorporating both the hop and the hop neighboring nodes brings the best performance. Furthermore, we also observe that considering too much graph information may result in relatively poor performance, as shown in the last row. Therefore, the final version of GRADE adopts the 2-hop neighboring representations where .
4 Case Study
To more intuitively analyze the performance of our GRADE, three representative examples are shown in Figure 5. From the example in the first row, we can see that the score given by our metric is closer to the human score than the other two baseline metrics. However, in the second-row example, our metric performs poorly. The potential reason may be the lack of topics (i.e., keywords) in the model response, as illustrated in the graph that only contains context-topic nodes. As a result, the graph reasoning module in our GRADE fails to induce an appropriate graph representation, which harms the coherence scoring. Finally, the example in the last row shows a hard case that both our GRADE and the baseline metrics are failed to cope with. In this hard case, the topics of the model response are relevant to the dialogue context so that both our GRADE and BERT-RUBER, as learning-based metrics, deem that the response greatly matches the context. However, the truth is that the model response is more likely a response for the previous utterance U1 rather than U2, which is hard for metrics to recognize.
Conclusion and Discussion
In this paper, we proposed GRADE (Graph-enhanced Representations for Automatic Dialogue Evaluation), a novel metric for dialogue coherence evaluation of open-domain dialogue systems. Empirical results show that GRADE has stronger correlations with human judgements and can generalize to other unseen chit-chat datasets. Besides, we also release a new large-scale human evaluation benchmark to facilitate future research on automatic metrics.
A limitation of GRADE is the inconsistency between the training objective (relative ranking) and the expected behavior (absolute scoring). Specifically, the ranking loss we adopted only requires good responses to be ranked higher than bad responses, which is a relatively loose constraint compared with the absolute scoring that humans do. Therefore, GRADE may deviate from the human scoring criterion and fail to quantify the dialogue responses accurately, and that the human correlation results fluctuate over different runs. Overall, to develop a dialogue metric that can quantify in a more human-like manner, it is critical to reducing the gap between the training objective and the model behavior we truly care about.
Acknowledgments
We thank all anonymous reviewers for their constructive comments. This work was supported in part by National Key RD Program of China under Grant No. 2018AAA0100300, National Natural Science Foundation of China (NSFC) under Grant No.U19A2073 and No.61976233, Guangdong Province Basic and Applied Basic Research (Regional Joint Fund-Key) Grant No.2019B1515120039, Nature Science Foundation of Shenzhen Under Grant No. 2019191361, Zhijiang Lab’s Open Fund (No. 2020AA3AB14).
References
Appendix A Details of the Datasets
The detailed processing procedure of DailyDialog and the introduction of the other two unseen datasets are presented.
is a chit-chat dataset with strong annotations for topic, emotion and utterance act. It contains total 13,118 open-domain multi-turn dialogues. We use the initial split of DailyDialog where training/validation/test sets have 11,118/1,000/1,000 dialogues respectively. Next, we subdivide these dialogues into context-response pairs each of which is composed of a context with length = 2 and a ground-truth response . Therefore, the processed training/validation/test sets now have 59264/6015/5705 pairs respectively. Then, for each context-response pair, we obtain two false responses and based on the lexical sampling and embedding-based sampling methods respectively, and get two tuples , . In total, there are 118528/12030/11410 tuples as our final data for training GRADE.
ConvAI2
is a chit-chat dataset based on the PersonaChat dataset (Dinan et al., 2019) for a NIPS 2018 competition. The dataset was collected by asking workers to chat with each other naturally with a given persona. The conversations cover a broad range of topics and frequently change during the conversations since both the speakers want to say out their persona information.
EmpatheticDialogues
is a novel dataset of 25k conversations grounded in a wide range of emotions to facilitate training and evaluating dialogue systems. It has been verified that dialogue models trained on this dataset are perceived to be more empathetic by human evaluators.