Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation Evaluators

Sanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tur

Introduction

Due to recent advances in spoken language understanding and automatic speech recognition, conversational interfaces such as Alexa, Cortana, and Siri have become increasingly common. While these interfaces are task oriented, there is an increasing interest in building conversational systems that can engage in more social conversations. Building systems that can have a general conversation in an open domain setting is a challenging problem, but it is an important step towards more natural human-machine interactions.

Recently, there has been significant interest in building chatbots Sordoni et al. (2015); Wen et al. (2015) fueled by the availability of dialog data sets such as Ubuntu, Twitter, and Movie dialogs Lowe et al. (2015); Ritter et al. (2011); Danescu-Niculescu-Mizil and Lee (2011). However, as most chatbots are text-based, work on human-machine spoken dialog is relatively under-explored, partly due to lack of such dialog corpora. Spoken dialog poses additional challenges such as automatic speech recognition errors and divergence between spoken and written language.

Sequence-to-sequence (seq2seq) models Sutskever et al. (2014) and their extensions Luong et al. (2015); Sordoni et al. (2015); Li et al. (2015), which are used for neural machine translation (MT), have been widely adopted for dialog generation systems. In MT, given a source sentence, the correctness of the target sentence can be measured by semantic similarity to the source sentence. However, in open-domain conversations, a generic utterance such as “sounds good” could be a valid response to a large variety of statements. These seq2seq models are commonly trained on a maximum likelihood objective, which leads the models to place uniform importance on all user utterance and system response pairs. Thus, these models usually choose “safe” responses as they frequently appear in the dialog training data. This phenomenon is known as the generic response problem. These responses, while arguably correct, are bland and convey little information leading to short conversations and low user satisfaction.

Since response generation systems are trained by maximizing the average likelihood of the training data, they do not have a clear signal on how well the current conversation is going. We hypothesize that having a way to measure conversational success at every turn could be valuable information that can guide system response generation and help improving system quality. Such a measurement may also be useful for combining responses from various competing systems. To this end, we build a supervised conversational evaluator to assess two aspects of responses: engagement and coherence. The input to our evaluators are encoded conversations represented as fixed-length vectors as well as hand-crafted dialog and turn level features. The system outputs explicit scores on coherence and engagement of the system response.

We experiment with two ways to incorporate these explicit signals in response generation systems. First, we use the evaluator outputs as input to a reranking model, which are used to rescore the nn-best outputs obtained after beam search decoding. Second, we propose a technique to incorporate the evaluator loss directly into the conversational model as an additional discriminatory loss term. Using both human and automatic evaluations, we show that both of these methods significantly improve the system response quality. The combined model utilizing re-ranking and the composite loss outperforms models using either mechanism alone.

The contributions of this work are two-fold. First, we experiment with various hand-crafted features and conversational encoding schemes to build a conversational evaluation system that can provide explicit turn-level feedback to a response generation system on the highly subjective task. This system can be used independently to compare various response generation systems or as a signal to improve response generation. Second, we experiment with two complementary ways to incorporate explicit feedback to the response generation systems and show improvement in dialog quality using automatic metrics as well as human evaluation.

Related Works

There are two major themes in this work. The first is building evaluators that allow us to estimate human perceptions of coherence, topicality, and interestingness of responses in a conversational context. The second is the use of evaluators to guide the generation process. As a result, this work is related to two distinct bodies of work.

Automatic Evaluation of Conversations: Learning automatic evaluation of conversation quality has a long history Walker et al. (1997). However, we still do not have widely accepted solutions. Due to the similarity between conversational response generation and MT, automatic MT metrics such as BLEU Papineni et al. (2002) and METEOR Banerjee and Lavie (2005) are widely adopted for evaluating dialog generation. ROUGE Lin and Hovy (2003), which is also used for chatbot evaluation, is a popular metric for text summarization. These metrics primarily rely on token-level overlap over a corpus (also synonymy in the case of METEOR), and therefore are not well-suited for dialog generation since a valid conversational response may not have any token-level or even semantic-level overlap with the ground truths. While the shortcomings of these metrics are well known for MT Graham (2015); Espinosa et al. (2010), the problem is aggravated for dialog generation evaluation because of the much larger output space Liu et al. (2016); Novikova et al. (2017). However, due to the lack of clear alternatives, these metrics are still widely used for evaluating response generation Ritter et al. (2011); Lowe et al. (2017). To ensure comparability with other approaches, we report results on these metrics for our models.

To tackle the shortcomings of automatic metrics, there have been efforts to build models to score conversations. Lowe et al. (2017) train a model to predict the score of a system response given a dialog context. However, they work with tiny data sets (around 4000 sentences) in a non-spoken setting. Tao et al. (2017) address the expensive annotation process by adding in unsupervised data. However, their metric is not interpretable, and the results are also not shown on a spoken setting. Our work differs from the aforementioned works as the output of our system is interpretable at each dialog turn.

There has also been work on building evaluation systems that focus on specific aspects of dialog. Li et al. (2016c) use features for information flow, Yu et al. (2016) use features for turn-level appropriateness. However, these metrics are based on a narrow aspect of the conversation and fail to capture broad ranges of phenomena that lead to a good dialog.

Improving System Response Generation: Seq2Seq models have allowed researchers to train dialog models without relying on handcrafted dialog acts and slot values. Using maximum mutual information (MMI) Li et al. (2015) was one of the earlier attempts to make conversational responses more diverse Serban et al. (2016b, a). Shao et al. (2017) use a segment ranking beam search to produce more diverse responses. Our method extends the strategy employed by Shao et al. (2017) utilizing a trained model as the reranking function and is similar to Holtzman et al. (2018) but with different kind of trained model.

More recently, there have been works which aim to alleviate this problem by incorporating conversation-specific rewards in the learning process. Yao et al. (2016) use the IDF value of generated sentences as a reward signal. Xing et al. (2017) use topics as an additional input while decoding to produce more specific responses. Li et al. (2016b) add personal information to make system responses more user specific.Li et al. (2017) use distillation to train different models at different levels of specificity and use reinforcement learning to pick the appropriate system response. Zhou et al. (2017) and Zhang et al. (2018) introduce latent factors in the seq2seq models that control specificity in neural response generation. There has been recent work which combines responses from multiple sub-systems Serban et al. (2017); Papaioannou et al. (2017) and ranks them to output the final system response. Our method complements these approaches by introducing a novel learned-estimator model as the additional reward signal.

Data

The data used in this study was collected during the Alexa Prize Ram et al. (2017) competition and shared with the teams who were participating in the competition. Upon initiating the conversation, users were paired with a randomly selected chatbot built by the participants. At the end of the conversation, the users were prompted to rate the chatbot quality, from 1–5, with 5 being the highest.

We randomly sampled more than 15K conversations (approximately 160K turns) collected during the competition. These were annotated for coherence and engagement (See Section 3.1) and used to train the conversation evaluators. For training the response generators, we selected highly-rated user conversations, which resulted in around 370K conversations containing 4M user utterances and their corresponding system response. One notable statistic is that user utterances are typically very short (mean: 3.6 tokens) while the system responses generally are much longer (mean: 23.2 tokens).

Asking annotators to measure coherence and engagement directly is a time-consuming task. We observed that we could collect data much faster if we asked direct “yes” or “no” questions to our annotators. Hence, upon reviewing a user-chatbot interaction along with the entire conversation to the current turn, annotatorsThe data was collected through mechanical turk. Annotators were presented with the full context of the dialog up to the current turn. rated each chatbot response as “yes” or “no” on the following criteria:

The system response is comprehensible: The information provided by the chatbot made sense with respect to the user utterance and is syntactially correct.

The system response is on topic: The chatbot response was on the same topic as the user utterance or was relevant to the user utterance. For example, if a user asks about a baseball player on the LA Dodgers, then the chatbot mentions something about the baseball team.

The system response is interesting: The chatbot response contains information which is novel and relevant. For example, the chatbot would provide an answer about a baseball player and give some additional information to create a fleshed-out response.

I want to continue the conversation: Given the current state of the conversation and the system response, there is a natural way to continue the conversation. For example, this could be due to the system asking a question about the current conversation subject.

We use these questions as proxies for measuring coherence and engagement of responses. The answers to the first two questions (“comprehensible” and “on topic”) are used as a proxy for coherence. Similarly, the answer to the last two questions (“interesting” and “continue the conversation”) are used as a proxy for engagement.

Conversation Evaluators

We train conversational response evaluators to assess the state of a given conversation. Our models are trained on a combination of utterance and response pairs combined with context (past turn user utterances and system responses) along with other features, e.g., dialog acts and topics as described in Section 4.3. We experiment with different ways to encode the responses (Section 4.1) as well as with different feature combinations (Figure 1).

We pretrained models that produce sentence embeddings using the ParlAI chitchat data set Miller et al. (2017). We use the Quick-Thought (QT) loss Logeswaran and Lee (2018) to train the embeddings. Our word embeddings are initialized with FastText Bojanowski et al. (2016) to capture the sub-word features and then fine-tuned. We encode sentences into embeddings using the following methods:

The Transformer Network (1 layer, 600 dim) Vaswani et al. (2017)

Concatenated last states of a BiLSTM (1 layer, 600 dim)

The selected dimensions and network structures followed the original paper Vaswani et al. (2017). All models were trained with a batch size of 400 using Adam optimizer with learning rate of 5e-4.

To measure the sentence embedding quality, we evaluate our models on a few standard classification tasks. The models are used to get sentence representation, which are passed through feedforward networks that are trained for the following classification tasks: (i) Semantic Textual Similarity (STS) Marelli et al. (2014), (ii) Question Type Classification (TREC) Voorhees and Dang (2003), (iii) Subjectivity Classification (SUBJ) Pang and Lee (2004). Table 1 shows the different models’ performances on these tasks. Based on this, we choose the Transformer as our sentence encoder as it was overall the best performing while being fast.

2 Context

Given the contextual nature of the problem we extracted the sentence embeddings of user utterances and responses for the past 5 turns and used a 1 layer LSTM with 256 hidden units to encode conversational context. The last state of LSTM is used to obtain the encoded representation, which is then concatenated with other features (Section 4.3) in a fully-connected neural network.

3 Features

Apart from sentence embeddings and context, the following features are also used:

Dialog Act: Serban et al. (2017) show that dialog act (DA) features could be useful for response selection rankers. Following this, we use model Khatri et al. (2018)-predicted DAs Stolcke et al. (1998) of user utterances and system responses as an indicator feature.

Entity Grid: Cervone et al. (2018); Barzilay and Lapata (2008) show that entities and DA transitions across turns can be strong features for assessing dialog coherence. Starting from a grid representation of the turns of the conversation as a matrix (DAs ×\times entities), these features are designed to capture the patterns of topic and intent shift distribution of a dialog. We employ the same strategy for our models.

Named Entity (NE) Overlap: We use named entity overlap between user utterances and their corresponding system responses as a feature. Our named entities are obtained using SpaCyhttps://spacy.io/. Papaioannou et al. (2017) have also used similar NE features in their ranker.

Topic: We use a one-hot representation of a dialog turn topic predicted by a conversational topic model Guo et al. (2017) that classifies a given dialog turn into one of 26 pre-defined classes like Sports and Movies.

Response Similarity: Cosine similarity between user utterance embedding and system response embedding is used as a feature.

Length: We use the token-level length of the user utterance and the response as a feature.

The above features were selected from a large pool of features through significance testing on our development set. The effect of adding these features can be seen in Table 2. Some of the features such as Topic lack previous dialog context, which could be updated to include the context. We leave this extension for future work.

4 Models

Given the large number of features and their non-sequential nature, we train four binary classifiers using feedforward neural networks (FFNN). The input to these models is a dialog turn. Each output layer is a softmax function corresponding to a binary decision for each evaluation metric forming a four-dimensional vector. Each vector dimension corresponds to an evaluation metric (See Section 3.1). For example, one possible reference output would be , which corresponds to “not comprehensible,” “on topic,” “interesting,” and “I don’t want to continue.”

We experimented with training the evaluators jointly and separately and found that training them jointly led to better performance. We suspect this is due to the objectives of all evaluators being closely related. We concatenate the aforementioned features as an input to a 3-layer FFNN with 256 hidden units. Figure 1 depicts the architecture of the conversation evaluators.

Response Generation System

To incorporate the explicit turn level feedback provided by the conversation evaluators, we augment our baseline response generation system with the softmax scores provided by the conversation evaluators. Our baseline response generation system is described in Section 5.1. We then incorporate evaluators outputs using two techniques: reranking and fine-tuning.

We extended the approach of Yao et al. (2016) where the authors used Luong’s dot attention Luong et al. (2015). In our experiments, the decoder uses the same attention (Figure 2a). As we want to observe the full impact of conversational evaluators, we do not incorporate inverse document frequency (IDF) or conversation topics into our objective. Extending the objective to include these terms can be a good direction for future work.

To make the response generation system more robust, we added user utterances and system responses from the previous turn as context. The input to the response generation model is previous-turn user utterance, previous-turn system response, and current-turn user utterance concatenated sequentially. We insert a special transition token Serban et al. (2016c) between turns. We then use a single RNN to encode these sentences. Our word embeddings are randomly initialized and then fine-tuned during training. We used a 1-layer Gated Recurrent Neural network with 512 hidden units for both encoder and decoder to train the seq2seq model and MLE as our training objective.

2 Reranking (S2S_RR)

In this approach, we do not update the underlying encoder-decoder model. We maintain a beam to get 15-best candidates from the decoder. The top candidate out of the 15 candidates is equivalent to the output of the baseline model. Here, instead of selecting the top output, the final output response is chosen using a reranking model.

For our reranking model, we calculate BLEU scores for each of the 15 candidate responses against the ground truth response from the chatbot. We then sample two responses from the kk-best list and train a pairwise response reranker. The response with the higher BLEU is placed in the positive class (+1) and the one with lower BLEU is placed in the negative class (-1). We do this for all possible candidate combinations from the 15-best responses. We use the max-margin ranking loss to train the model. The model is a three-layered FFNN with 16 hidden units.

The input to the pairwise reranker is the softmax output of the 4 evaluators as shown in Figure 1. The input to the evaluators are described in Section 4. The output of the reranker is a scalar, which, if trained right, would give a higher value for responses with higher BLEU scores. Figure 2b depicts the architecture of this model.

3 Fine-tuning (S2S_FT)

In this approach, we use evaluators as a discriminatory loss to fine-tune the baseline encoder-decoder response generation system. We first train the baseline model and then, it is fine-tuned using the evaluator outputs in the hope of generating more coherent and engaging responses. One issue with MLE is that the learned models are not optimized for the final metric (e.g., BLEU). To combat this problem, we add a discriminatory loss in addition to the generative loss to the overall loss term as shown in Equation 1.

where zn=xn,yn−1,…,x0,y0z_{n}=x_{n},y_{n-1},\ldots,x_{0},y_{0} is the conversational context where nn is the context length. q∈R∣V∣×lenq\in R^{|V|\times len} of the first term corresponds to the softmax output generated by the response generation model. The term y^ni\hat{y}_{ni} refers to its corresponding decoder response at nthn_{th} conversation turn and ithi^{th} word generated. In the second term, the function EvalEval refers to the evaluator score produced for a user utterance, xnx_{n}, and decoder softmax output, qq.

The evaluator score is defined as the sum of softmax outputs of all 4 models.We keep the rest of the input (context and features) for the evaluator as is.

We weight the discriminator score by λ\lambda, which is a hyperparameter. We selected λ\lambda to be 10 using grid search to optimize for final BLEU on our development set. Figure 2c depicts the architecture of this approach. The decoder is fine-tuned to maximize evaluator scores along while minimizing the cross-entropy loss. The evaluator model is trained on the original annotated corpus and parameters are frozen.

4 Reranking + Fine-tuning (S2S_RR_FT)

We also combined fine-tuning with reranking, where we obtained the 15 candidates from the fine-tuned response generator and then we select the best response using the reranker, which is trained to maximize the BLEU score.

Experiments and Results

The conversation evaluators were trained using cross-entropy loss. We used a batch size of 128, dropout of 0.3 and Adam optimizer with a learning rate of 5e-5 for our conversational evaluators. Sentence embeddings for user utterances and system responses are obtained using the fast-text embeddings and Transformer network.

Table 2 shows the evaluator performance compared with a baseline with no handcrafted features. We present precision, recall, and f-score measures along with the accuracy. Furthermore, since the class distribution of the dataset is highly imbalanced, we also calculate Matthews correlation coefficient (MCC)Matthews (1975), which takes into account true and false positives and negatives. It is a balanced measure which can be used even if the classes sizes are very different. With the proposed features we observe significant improvement across all metrics.

We also performed a correlation study between the model predicted scores and human annotated scores (1 to 5) on 2000 utterances. The annotatorsSame setup as previously described were asked to answer a single question: “On a scale of 1–5, how coherent and engaging is this response given the previous conversation?” From Table 3, it can be observed that evaluator predicted scores has significant correlation (moderate to high) with the overall human evaluation score on this subjective task (0.2 – 0.4 Pearson correlation with turn-level ratings). Considering the substantial individual differences in evaluating open-domain conversations, we observe that our evaluators with moderate level of correlation can be used to provide turn-level feedback for a human-chatbot conversation.

2 Response Generation

We first trained the baseline model (S2S) on the conversational data set (4M utterance-response pairs from the competition. Section 3). The data were split into 80% training, 10% development, and 10% test sets. The baseline model was trained using Adam with learning rate of 1e-4 and batch size of 256 until the development loss converges. The vocabulary of 30K most frequent words were used. And the reranker was trained using the 20K number of beam outputs from the baseline model on the development set. Adam with learning rate of 1e-4 and batch size of 16 was used for the fine-tuning (S2S_FT).

Table 5 shows the performance comparison of different generation models (Section 5) on the Alexa Prize conversational data set. We observed that reranking nn-best responses using the evaluator-based reranker (S2S_RR) provides nearly 100% improvement in BLEU-4 scores.

Fine-tuning the generator by adding evaluator loss (S2S_FT) does improve the performance but the gains are smaller compared to reranking. We suspect that this is due to the reranker directly optimizing for BLEU. However, using a fine-tuned model and then reranking (S2S_RR_FT) complements each other and gives the best performance overall. Furthermore, we observe that even though the reranker is trained to maximize the BLEU scores, reranking shows significant gains in ROUGE scores as well. We also measured different systems performance using Distinct-2 Li et al. (2016a), which is the number of unique length-normalized bigrams in responses. The metric can be a surrogate for measuring diverse outputs. We see that our generators using reranking approaches improve on this metric as well. Table 4 also shows 2 sampled responses from different models.

To further analyze the impact of reranker trained to optimize on BLEU score, we trained a baseline response generation system on a Reddit data setWe use a publicly available data Baumgartner (2015)., which comprises of 9 million comments and corresponding response comments. All the hyperparameter setting followed the setting of training on the Alexa Prize conversational dataset.

We trained a new reranker for the Reddit data using the evaluator scores obtained from the models proposed in Section 4. We show in Table 6 that even though the evaluators are trained on a different data set, the reranker learns to select better responses nearly doubling the BLEU scores as well as improving on the Distinct-2 score. Thus, the evaluator generalizes in selecting more coherent and engaging responses in human-human interactions as well as human-computer interactions. As fine-tuning the evaluator is computationally expensive, we did not fine-tune it on the Reddit dataset.

The closest baseline that used BLEU scores for evaluation in open-domain setting is from Li et al. (2015) where they trained the models on Twitter data using Maximum Mutual Information (MMI) as the objective function. They obtained a BLEU score of 5.2 in their best setting on Twitter data (average length 23 chars), which is relatively less complex than Reddit (average length 75 chars).

3 Human Evaluation

As noted earlier, automatic evaluation metrics may not be the best way to measure chatbot response generation performance. Therefore, we performed human evaluation of our models. We asked annotators to provide ratings on the system responses from the models we evaluated, i.e., baseline model, S2S_RR, S2S_FT, and S2S_RR_FT. A rating was obtained on two metrics: coherence and engagement. Coherence measures how much the response is comprehensible and relevant to a user’s request and engagement shows interestingness of the response (Venkatesh et al. (2018)). We asked the annotators to provide the rating based on a scale of 1–5, with 5 being the best. We had four annotators rate 250 interactions. Table 7 shows the performance of the models on the proposed metrics. Our inter-annotator agreement is 0.42 on Cohen’s Kappa Coefficient, which implies moderate agreement. We believe this is because the task is relatively subjective and the conversations were performed in the challenging open-domain setting. The S2S_RR_FT model provides the best performance across all the metrics, followed by S2S_RR, followed by S2S_FT.

Conclusion

Human annotations for conversations show significant variance, but it is still possible to train models which can extract meaningful signal from the human assessment of the conversations. We show that these models can provide useful turn-level guidance to response generation models. We design a system using various features and context encoders to provide turn-level feedback in a conversational dialog. Our feedback is interpretable on 2 major axes of conversational quality: engagement and coherence. We also plan to provide similar evaluators to the university teams participating in the Alexa Prize competition. To show that such feedback is useful in building better conversational response systems, we propose 2 ways to incorporate this feedback, both of which help improve on the baselines. Combining both techniques results in the best performance. We view this work as complementary to other recent work in improving dialog systems such as Li et al. (2015) and Shao et al. (2017). While such open-domain systems are still in their infancy, we view the framework presented in this paper to be an important step towards building end-to-end coherent and engaging chatbots.

References