Multi-domain Dialogue State Tracking as Dynamic Knowledge Graph Enhanced Question Answering

Li Zhou, Kevin Small

Introduction

In a task-oriented dialogue system, the dialogue policy determines the next action to perform and next utterance to say based on the current dialogue state. A dialogue state defined by frame-and-slot semantics is a set of (key, value) pairs specified by the domain ontology (Jurafsky & Martin, 2019). A key is a (domain, slot) pair and a value is a slot value provided by the user. Figure 1 shows a dialogue and state in three domain contexts. Dialogue state tracking (DST) in multiple domains is a challenging problem. First of all, in production environments, the domain ontology is being continuously updated such that the model must generalize to new values, new slots, or even new domains during inference. Second, the number of slots and values in the training data are usually quite large. For example, the MultiWOZ 2.0/2.12.0/2.1 datasets (Budzianowski et al., 2018; Eric et al., 2019) have 3030 (domain, slot) pairs and more than 4,5004,500 values (Wu et al., 2019). As the model must understand slot and value paraphrases, it is infeasible to train each slot or value independently. Third, multi-turn inferences are often required as shown in the underlined areas of Figure 1.

Many single-domain DST algorithms have been proposed (Mrkšić et al., 2017; Ren et al., 2018; Zhong et al., 2018). For example, Zhong et al. (2018) learns a local model for each slot and a global model shared by all slots. However, single domain models are difficult to scale to multi-domain settings, leading to the development of multi-domain DST algorithms. For example, Nouri & Hosseini-Asl (2018) improves Zhong et al. (2018)’s work by removing local models and building a slot-conditioned global model to share parameters between domains and slots, thus computing a score for every (domain, slot, value) tuple. This approach remains problematic for settings with a large value set (e.g., user phone number). Wu et al. (2019) proposes an encoder-decoder architecture which takes dialogue contexts as source sentences and state annotations as target sentences, but does not explicitly use relationships between domains and slots. For example, if a user booked a restaurant and asks for a taxi, then the destination of the taxi is likely to be that restaurant, and if a user booked a 55 star hotel, then the user is likely looking for an expensive rather than a cheap restaurant. As we will show later, such relationships between domains and slots help improve model performance.

To tackle these challenges, we propose DSTQA (Dialogue State Tracking via Question Answering), a new multi-domain DST model inspired by recently developed reading comprehension and question answering models. Our model reads dialogue contexts to answer a series of questions that asks for the value of a (domain, slot) pair. Specifically, we construct two types of questions: 1) multiple choice questions for (domain, slot) pairs with a limited number of value options and 2) span prediction questions, of which the answers are spans in the contexts, designed for (domain, slot) pairs that have a large or infinite number of value options. Finally, we represent (domain, slot) pairs as a dynamically-evolving knowledge graph with respect to the dialogue context, and utilize this graph to drive improved model performance. Our contributions are as follows: (1) we propose to model multi-domain DST as a question answering problem such that tracking new domains, new slots and new values is simply constructing new questions, (2) we propose using a bidirectional attention (Seo et al., 2017) based model for multi-domain dialogue state tracking, and (3) we extend our algorithm with a dynamically-evolving knowledge graph to further exploit the structure between domains and slots.

Problem Formulation

In a multi-domain dialogue state tracking problem, there are MM domains D={d1,d2,...,dM}D=\{d_{1},d_{2},...,d_{M}\}. For example, in MultiWOZ 2.0/2.1 datasets, there are 7 domains: restaurant, hotel, train, attraction, taxi, hospital, and police. Each domain d∈Dd\in D has NdN^{d} slots Sd={s1d,s2d,...,sNdd}S^{d}=\{s^{d}_{1},s^{d}_{2},...,s^{d}_{N^{d}}\}, and each slot s∈Sds\in S^{d} has KsK^{s} possible values Vs={v1s,v2s,...,vKss}V^{s}=\{v^{s}_{1},v^{s}_{2},...,v^{s}_{K^{s}}\}. For example, the restaurant domain has a slot named price range, and the possible values are cheap, moderate, and expensive. Some slots do not have pre-defined values, that is, VsV^{s} is missing in the domain ontology. For example, the taxi domain has a slot named leave time, but it is a poor choice to enumerate all the possible leave times the user may request as the size of VsV^{s} will be very large. Meanwhile, the domain ontology can also change over time. Formally, we represent a dialogue XX as X={U1a,U1u,U2a,U2u,...,UTa,UTu}X=\{U^{a}_{1},U^{u}_{1},U^{a}_{2},U^{u}_{2},...,U^{a}_{T},U^{u}_{T}\}, where UtaU^{a}_{t} is the agent utterance in turn tt and UtuU^{u}_{t} is the user utterance in turn tt. Each turn tt is associated with a dialogue state yt\text{y}_{t}. A dialogue state yt\text{y}_{t} is a set of (domain, slot, value) tuples. Each tuple represents that, up to the current turn tt, a slot s∈Sds\in S^{d} of domain d∈Dd\in D, which takes the value v∈Vsv\in V^{s} has been provided by the user. Accordingly, yt\text{y}_{t}’s are targets that the model needs to predict.

Multi-domain Dialogue State Tracking via Question Answering (DSTQA)

We model multi-domain DST as a question answering problem and use machine reading methods to provide answers. To predict the dialogue state at turn tt, the model observes the context CtC_{t}, which is the concatenation of {U1a,U1u,...,Uta,Utu}\{U_{1}^{a},U_{1}^{u},...,U_{t}^{a},U_{t}^{u}\}. The context is read by the model to answer the questions defined as follows. First, for each domain d∈Dd\in D and each slot s∈Sds\in S^{d} where there exists a pre-defined value set VsV^{s}, we construct a question Qd,s={d,s,Vs,not mentioned,don’t care}Q_{d,s}=\{d,s,V^{s},\text{{\tt not mentioned}},\text{{\tt don't care}}\}. That is, a question is a set of words or phrases which includes a domain name, a slot name, a list of all possible values, and two special values not mentioned and don’t care. One example of the constructed question for restaurant domain and price range slot is Qd,s={\emrestaurant,\emprice range,cheap,moderate,expensive,not mentioned,don’t care}Q_{d,s}=\{\text{{\em restaurant}},\text{{\em price range}},\text{{\tt cheap}},\text{{\tt moderate}},\text{{\tt expensive}},\text{{\tt not mentioned}},\text{{\tt don't care}}\}. The constructed question represents the following natural language question:

“In the dialogue up to turn tt, did the user mention the ‘price range’ of the ‘restaurant’ he/she is looking for? If so, which of the following option is correct: A) cheap, B) moderate, C) expensive, D) don’t care.”

As we can see from the above example, instead of only using domains and slots to construct questions (corresponding to natural language questions what is the value of this slot?), we also add candidate values VsV^{s} into Qd,sQ_{d,s}, this is because values can be viewed as descriptions or complimentary information to domains and slots. For example, cheap, moderate and expensive explains what price range is. In this way, the constructed question Qd,sQ_{d,s} contains rich information about the domains and slots to predict, and easy to generalize to new values.

In the case that VsV^{s} is not available, the question is just the domain and slot names along with the special values, that is, Qd,s={d,s,not mentioned,don’t care}Q_{d,s}=\{d,s,\text{{\tt not mentioned}},\text{{\tt don't care}}\}. For example, the constructed question for train domain and leave time slot is Qd,s={\emtrain,\emleave time,not mentioned,don’t care}Q_{d,s}=\{\text{{\em train}},\text{{\em leave time}},\text{{\tt not mentioned}},\text{{\tt don't care}}\}, and represents the following natural language question:

“In the dialogue up to turn tt, did the user mention the ‘leave time’ of the ‘train’ he/she is looking for? If so, what is the ‘leave time’ the user preferred?”

The most important concept to note here is that the proposed DSTQA model can be easily extended to new domains, slots, and values. Tracking new domains and slots is simply constructing new queries, and tracking new values is simply extending the constructed question of an existing slot.

Although we formulate multi-domain dialogue state tracking as a question answering problem, we want to emphasize that there are some fundamental differences between these two settings. In a standard question answering problem, question understanding is a major challenge and the questions are highly dependent on the context where questions are often of many different forms (Rajpurkar et al., 2018). Meanwhile, in our formulation, the question forms are limited to two, every turn results in asking a restricted set of question types, and thus question understanding is straightforward. Conversely, our formulation has its own complicating characteristics including: (1) questions in consecutive turns tend to have the same answers, (2) an answer is either a span of the context or a value from a value set, and (3) the questions we constructed have some underlying connections defined by a dynamically-evolving knowledge graph (described in Section 4), which can help improve model performance. In any case, modeling multi-domain DST with this approach allows us to easily transfer knowledge to new domains, slots, and values simply by constructing new questions. Accordingly, many existing reading comprehension algorithms (Seo et al., 2017; Yu et al., 2018; Devlin et al., 2019; Clark & Gardner, 2018) can be directly applied here. In this paper, we propose a bidirectional attention flow (Seo et al., 2017) based model for multi-domain DST.

Figure 2 summarizes the DSTQA architecture, where notable subcomponents are detailed below.

We summarize context BcB^{c} into a single vector with respect to the domain and slot and then apply a bilinear function to calculate the score of each value. More specifically, We calculate the score of each value vv at turn tt by

Dynamic Knowledge Graph for Multi-domain dialogue State Tracking

In our problem formulation, at each turn, our proposed algorithm asks a set of questions, one for each (domain, slot) pair. In fact, the (domain, slot) pairs are not independent. For example, if a user requested a train for 33 people, then the number of people for hotel reservation may also be 33. If a user booked a restaurant, then the destination of the taxi is likely to be that restaurant. Specifically, we observe four types of relationships between (domain, slot) pairs in MultiWOZ 2.02.0/2.12.1 dataset:

(s,rv,s′)(s,r_{v},s^{\prime}): a slot s∈Sds\in S^{d} and another slot s′∈Sd′s^{\prime}\in S^{d^{\prime}} have the same set of possible values. That is, VsV^{s} equals to Vs′V^{s^{\prime}}. For example, in MultiWOZ 2.02.0/2.12.1 dataset, domain-slot pairs (restaurant, book day) and (hotel, book day) have this relationship.

(s,rs,s′)(s,r_{s},s^{\prime}): the value set of a slot s∈Sds\in S^{d} is a subset of the value set of s′∈Sd′s^{\prime}\in S^{d^{\prime}}. For example, in MultiWOZ 2.02.0/2.12.1 dataset, value sets of (restaurant, name), (hotel, name), (train, station) and (attraction, name) are subsets of the value set of (taxi, destination).

(s,rc,s′)(s,r_{c},s^{\prime}): the informed value v∈Vsv\in V^{s} of slot ss is correlated with the informed value v∈Vs′v\in V^{s^{\prime}} of slot s′s^{\prime} even though VsV^{s} and Vs′V^{s^{\prime}} do not overlap. For example, in MultiWOZ 2.02.0/2.12.1 dataset, the price range of a reserved restaurant is correlated with the star of the booked hotel. This relationship is not explicitly given in the ontology.

(s,ri,v)(s,r_{i},v): the user has informed value v∈Vsv\in V^{s} of slot s∈Sds\in S^{d}.

In this section, we propose using a dynamic knowledge graph to further improve model performance by exploiting this information. We represent (domain, slot) pairs and values as nodes in a graph linked by the relationship defined above, and then propagate information between them. The graph is dynamically evolving, since the fourth relationship above, rir_{i}, depends on the dialogue context.

The right-hand side of Figure 2 is an example of the graph we defined based on the ontology. There are two types of nodes {M,N}\{M,N\} in the graph. One is a (domain, slot) pair node representing a (domain, slot) pair in the ontology and another is a value node representing a value from a value set. For a domain d∈Dd\in D and a slot s∈Sds\in S^{d}, we denote the corresponding node by Md,sM_{d,s}, and for a value v∈Vsv\in V^{s}, we denote the corresponding node by NvN_{v}. There are also two types of edges. One type is the links between MM and NN. At each turn tt, if the answer to question Qd,sQ_{d,s} is v∈Vsv\in V^{s}, then NvN_{v} is added to the graph and linked to Md,sM_{d,s}. By default, Md,sM_{d,s} is linked to a special not mentioned node. The other type of edges is links between nodes in MM. Ideally we want to link nodes in MM based on the first three relationships described above. However, while rvr_{v} and rsr_{s} are known given the ontology, rcr_{c} is unknown and cannot be inferred just based on the ontology. As a result, we connect every node in MM (i.e. the (domain, slot) pair nodes) with each other, and let the model to learn their relationships with an attention mechanism, which will be described shortly.

2 Attention Over the Graph

We use an attention mechanism to calculate the importance of a node’s neighbors to that node, and then aggregate node embeddings based on attention scores. Veličković et al. (2018) describes a graph attention network, which performs self-attention over nodes. In contrast with their work, we use dialogue contexts to attend over nodes.

where γ=σ(Bc⊤⋅αb+zd,s)\gamma=\sigma({B^{c}}^{\top}\cdot\alpha^{b}+z_{d,s}) is the gate and controls how much graph information should flow to the context embedding given the dialogue context. Some utterances such as “book a taxi to Cambridge station” do not need information in the graph, while some utterances such as “book a taxi from the hotel to the restaurant” needs information from other domains. γ\gamma dynamically controls in what degree the graph embedding is used. and graph parameters are trained together with all other parameters.

Experiments

We evaluate our model on three publicly available datasets: (non-multi-domain) WOZ 2.02.0 (Mrkšić et al., 2017), MultiWOZ 2.02.0 (Budzianowski et al., 2018), and MultiWOZ 2.12.1 (Eric et al., 2019). Due to limited space, please refer to Appendix A.1 for results on (non-multi-domain) WOZ 2.02.0 dataset. MultiWOZ 2.02.0 dataset is collected from a Wizard of Oz style experiment and has 77 domains: restaurant, hotel, train, attraction, taxi, hospital, and police. Similar to Wu et al. (2019), we ignore the hospital and police domains because they only appear in training set. There are 3030 (domain, slot) pairs and a total of 1043810438 task-oriented dialogues. A dialogue may span across multiple domains. For example, during the conversation, a user may book a restaurant first, and then book a taxi to that restaurant. For both datasets, we use the train/test splits provided by the dataset. The domain ontology of the datasets is described in Appendix A.2. MultiWOZ 2.12.1 contains the same dialogues and ontology as MultiWOZ 2.02.0, but fixes some annotation errors in MultiWOZ 2.02.0.

Two common metrics to evaluate dialogue state tracking performance are Joint accruacy and Slot accuracy. Joint accuracy is the accuracy of dialogue states. A dialogue state is correctly predicted only if all the values of (domain, slot) pairs are correctly predicted. Slot accuracy is the accuracy of (domain, slot, value) tuples. A tuple is correctly predicted only if the value of the (domain, slot) pair is correctly predicted. In most literature, joint accuracy is considered as a more challenging and more important metric.

Existing dialogue state tracking datasets, such as MultiWOZ 2.02.0 and MultiWOZ 2.1, do not have annotated span labels but only have annotated value labels for slots. As a result, we preprocess MultiWOZ 2.02.0 and MultiWOZ 2.12.1 dataset to convert value labels to span labels: we take a value label in the annotation, and search for its last occurrence in the dialogue context, and use that occurrence as span start and end labels. There are 3030 slots in MultiWOZ 2.02.0/2.12.1 dataset, and 55 of them are time related slots such as restaurant book time and train arrive by, and the values are 24-hour clock time such as 08:15. We do span prediction for these 55 slots and do value prediction for the rest of slots because it is not practical to enumerate all time values. We can also do span prediction for other slots such as restaurant name and hotel name with the benefit of handling out-of-vocabulary values, but we leave these experiments as future work. WOZ 2.02.0 dataset only has one domain and 33 slots, and we do value prediction for all these slots without graph embeddings.

We implement our model using AllenNLP (Gardner et al., 2017) framework.https://github.com/alexa/dstqa For experiments with ELMo embeddings, we use a pre-trained ELMo modelhttps://allennlp.org/elmo in which the output size is DELMo=512D^{ELMo}=512. The dimension of character-level embeddings is DChar=100D^{Char}=100, making Dw=612D^{w}=612. ELMo embeddings are fixed during training. For experiments with GloVe embeddings, we use GloVe embeddings pre-trained on Common Crawl dataset.https://nlp.stanford.edu/projects/glove/ The dimension of GloVe embeddings is 300300, and the dimension of character-level embeddings is 100, such that Dw=400D^{w}=400. GloVe embeddings are trainable during training. The size of the role embedding is 128128. The dropout rate is set to 0.50.5. We use Adam as the optimizer and the learning rate is set to 0.0010.001. We also apply word dropout that randomly drop out words in dialogue context with probability 0.10.1.

When training DSTQA with the dynamic knowledge graph, in order to predict the dialogue state and calculate the loss at turn tt, we use the model with current parameters to predict the dialogue state up until turn t−1t-1, and dynamically construct a graph for turn tt. We have also tried to do teacher forcing which constructs the graph with ground truth labels (or sample ground truth labels with an annealed probability), but we observe a negative impact on joint accuracy. On the other hand, target network (Mnih et al., 2015) may be useful here and will be investigated in the future. More specifically, we can have a copy of the model that update periodically, and use this model copy to predict dialogue state up until turn t−1t-1 and construct the graph.

2 Results on MultiWoz 2.0 and MultiWOZ 2.1 dataset.

We first evaluate our model on MultiWOZ 2.0 dataset as shown in Table 1. We compare with five published baselines. TRADE (Wu et al., 2019) is the current published state-of-the-art model. It utilizes an encoder-decoder architecture that takes dialogue contexts as source sentences, and takes state annotations as target sentences. SUMBT (Lee et al., 2019) fine-tunes a pre-trained BERT model (Devlin et al., 2019) to learn slot and utterance representations. Neural Reading (Gao et al., 2019) learns a question embedding for each slot, and predicts the span of each slot value. GCE (Nouri & Hosseini-Asl, 2018) is a model improved over GLAD (Zhong et al., 2018) by using a slot-conditioned global module. Details about baselines are in Section 6.

For our model, we report results under two settings. In the DSTQA w/span setting, we do span prediction for the five time related slots as mentioned in Section 5.1. This is the most realistic setting as enumerating all possible time values is not practical in a production environment. In the DSTQA w/o span setting, we do value prediction for all slots, including the five time related slots. To do this, we collect all time values appeared in the training data to create a value list for time related slots as is done in baseline models. It works in these two datasets because there are only 173173 time values in the training data, and only 1414 out-of-vocabulary time values in the test data. Note that in all our baselines, values appeared in the training data are either added to the vocabulary or added to the domain ontology, so DSTQA w/o span is still a fair comparison with the baseline methods. Our model outperforms all models. DSTQA w/span has a 5.64%5.64\% relative improvement and a 2.74%2.74\% absolute improvement over TRADE. We also show the performance on each single domain in Appendix A.3. DSTQA w/o span has a 5.80%5.80\% relative improvement and a 2.82%2.82\% absolute improvement over TRADE. We can see that DSTQA w/o span performs better than DSTQA w/span, this is mainly because we introduce noises when constructing the span labels, meanwhile, span prediction cannot take the benefit of the bidirectional attention mechanism. However, DSTQA w/o span cannot handle out-of-vocabulary values, but can generalize to new values only by expanding the value sets, moreover, the performance of DSTQA w/o span may decrease when the size of value sets increases.

Table 2 shows the results on MultiWOZ 2.12.1 dataset. Compared with TRADE, DSTQA w/span has a 8.93%8.93\% relative improvement and a 4.07%4.07\% absolute improvement. DSTQA w/o span has a 12.21%12.21\% relative improvement and a 5.57%5.57\% absolute improvement. More baselines can be found at the leaderboard.http://dialogue.mi.eng.cam.ac.uk/index.php/corpus/ Our model outperforms all models on the leaderboard at the time of submission of this paper.

Ablation Study: Table 1 also shows the results of ablation study of DSTQA w/span on MultiWOZ 2.02.0 dataset. The first experiment completely removes the graph component, and the joint accuracy drops 0.47%0.47\%. The second experiment keeps the graph component but removes the gating mechanism, which is equivalent to setting γ\gamma in Equation (2) to 0.50.5, and the joint accuracy drops 0.98%0.98\%, demonstrating that the gating mechanism is important when injecting graph embeddings and simply adding the graph embeddings to context embeddings can negatively impact the performance. In the third experiment, we replace BiQDB_{i}^{QD} with the mean of query word embeddings and replace BjCDB_{j}^{CD} with the mean of context word embeddings. This is equivalent to setting the bi-directional attention scores uniformly. The joint accuracy significantly drops 1.62%1.62\%. The fourth experiment completely removes the bi-directional attention layer, and the joint accuracy drops 1.85%1.85\%. Both experiments show that bidirectional attention layer has a notably positive impact on model performance. The fifth experiment substitute ELMo embeddings with GloVe embeddings to demonstrate the benefit of using contextual word embeddings. We plan to try other state-of-the-art contextual word embeddings such as BERT (Devlin et al., 2019) in the future. We further show the model performance on different context lengths in Appendix A.4.

3 Generalization to New Domains

Table 3 shows the model performance on new domains. We take one domain in MultiWOZ 2.02.0 as the target domain, and the remaining 44 domains as source domains. Models are trained either from scratch using only 5%5\% or 10%10\% sampled data from the target domain, or first trained on the 44 source domains and then fine-tuned on the target domain with sampled data. In general, a model that achieves higher accuracy by fine-tuning is more desirable, as it indicates that the model can quickly adapt to new domains given limited data from the new domain. In this experiment, we compare DSTQA w/span with TRADE. As shown in Table 3, DSTQA consistently outperforms TRADE when fine-tuning on 5%5\% and 10%10\% new domain data. With 5%5\% new domain data, DSTQA fine-tuning has an average of 43.32%43.32\% relative improvement over DSTQA training from scratch, while TRADE fine-tuning only has an average of 19.99%19.99\% relative improvement over TRADE training from scratch. DSTQA w/ graph also demonstrates its benefit over DSTQA w/o graph, especially on the taxi domain. This is because the ‘taxi’ domain is usually mentioned at the latter part of the dialogue, and the destination and departure of the taxi are usually the restaurant, hotel, or attraction mentioned in the previous turns and are embedded in the graph.

4 Error Analysis

Figure 3 shows the different types of model prediction errors on MultiWOZ 2.12.1 dataset made by DSTQA w/span as analyzed by the authors. Appendix A.6 explains the meaning of each error type and also list examples for each error type. At first glance, annotation errors and annotation disagreements account for 56%56\% of total prediction errors, and are all due to noise in the dataset and thus unavoidable. Annotation errors are the most frequent errors and account for 28%28\% of total prediction errors. Annotation errors means that the model predictions are incorrect only because the corresponding ground truth labels in the dataset are wrong. Usually this happens when the annotators neglect the value informed by the user. Annotator disagreement on user confirmation accounts for 28%28\% (15%+13%15\%+13\%) of total errors. This type of errors comes from the disagreement between annotators when generating ground truth labels. All these errors are due to the noise in the dataset and unavoidable, which also explains why the task on MultiWOZ 2.12.1 dataset is challenging and the state-of-the-art joint accuracy is less than 50%50\%.

Values exactly matched but not recognized (10%10\%) and paraphrases not recognized (14%14\%) mean that the user mentions a value or a paraphrase of a value, but the model fails to recognize it. Multi-turn inferences failed (6%6\%) means that the model fails to refer to previous utterances when making prediction. User responses not understood (8%8\%) and implications not understood (3%3\%) mean that the model does not understand what the user says and fails to predict based on user responses. Finally, incorrect value references (2%2\%) means that there are multiple values of a slot in the context and the model refers to an incorrect one, and incorrect domain references (1%1\%) means that the predicted slot and value should belong to another domain. All these errors indicate insufficient understanding of agent and user utterances. A more powerful language model and a coreference resolution modules may help mitigate these problems. Please refer to Appendix A.6 for examples.

Related Works

Our work is most closely related to previous works in dialogue state tracking and question answering. Early models of dialogue state tracking (Thomson & Young, 2010; Wang & Lemon, 2013; Henderson et al., 2014) rely on handcrafted features to extract utterance semantics, and then use these features to predict dialogue states. Recently Mrkšić et al. (2017) propose to use convolutional neural network to learn utterance nn-gram representation, and achieve better performance than handcrafted features-based model. However, their model maintains a separate set of parameters for each slot and does not scale well. Models that handles scalable multi-domain DST have then been proposed (Ramadan et al., 2018; Rastogi et al., 2017). Zhong et al. (2018) and Nouri & Hosseini-Asl (2018) propose a global-local architecture. The global module is shared by all slots to transfer knowledge between them. Ren et al. (2018) propose to share all parameters between slots and fix the word embeddings during training, so that they can handle new slots and values during inference. However, These models do not scale when the sizes of value sets are large or infinite, because they have to evaluate every (domain, slot, tuple) during the training. Xu & Hu (2018) propose to use a pointer network with a Seq2Seq architecture to handle unseen slot values. Lee et al. (2019) encode slots and utterances with a pre-trained BERT model, and then use a slot utterance matching module, which is a multi-head attention layer, to compute the similarity between slot values and utterances. Rastogi et al. (2019) release a schema-guided DST dataset which contains natural language description of domains and slots. They also propose to use BERT to encode these natural language description as embeddings of domains and slots. Wu et al. (2019) propose to use an encoder-decoder architecture with a pointer network. The source sentences are dialogue contexts and the target sentences are annotated value labels. The model shares parameters across domains and does not require pre-defined domain ontology, so it can adapt to unseen domains, slots and values. Our work differs in that we formulate multi-domain DST as a question answering problem and use reading comprehension methods to provide answers. There have already been a few recent works focusing on using reading comprehension models for dialogue state tracking. For example, Perez & Liu (2017) formulate slot tracking as four different types of questions (Factoid, Yes/No, Indefinite knowledge, Counting and Lists/Sets), and use memory network to do reasoning and to predict answers. Gao et al. (2019) construct a question for each slot, which basically asks what is the value of slot i, then they predict the span of the value/answer in the dialogue history. Our model is different from these two models in question representation. We not only use domains and slots but also use lists of candidate values to construct questions. Values can be viewed as descriptions to domains and slots, so that the questions we formulate have richer information about domains and slots, and can better generalize to new domains, slots, and values. Moreover, our model can do both span and value prediction, depending on whether the corresponding value lists exists or not. Finally, our model uses a dynamically-involving knowledge graph to explicitly capture interactions between domains and slots.

In a reading comprehension (Rajpurkar et al., 2016) task, there is one or more context paragraphs and a set of questions. The task is to answer questions based on the context paragraphs. Usually, an answer is a text span in a context paragraph. Many reading comprehension models have been proposed (Seo et al., 2017; Yu et al., 2018; Devlin et al., 2019; Clark & Gardner, 2018; Chen et al., 2017). These models encode questions and contexts with multiple layers of attention-based blocks and predict answer spans based on the learned question and context embeddings. Some works also explore to further improve model performance by knowledge graph. For example Sun et al. (2018) propose to build a heterogeneous graph in which the nodes are knowledge base entities and context paragraphs, and nodes are linked by entity relationships and entity mentions in the contexts. Zhang et al. (2018) propose to use Open IE to extract relation triples from context paragraphs and build a contextual knowledge graph with respect to the question and context paragraphs. We would expect many of these technical innovations to apply given our QA-based formulation.

Conclusion

In this paper, we model multi-domain dialogue state tracking as question answering with a dynamically-evolving knowledge graph. Such formulation enables the model to generalize to new domains, slots and values by simply constructing new questions. Our model achieves state-of-the-art results on MultiWOZ 2.0 and MultiWOZ 2.1 dataset with a 5.80%5.80\% and a 12.21%12.21\% relative improvement, respectively. Also, our domain expansion experiments show that our model can better adapt to unseen domains, slots and values compared with the previous state-of-the-art model.

References

Appendix A Appendix

We also evaluate our algorithm on WOZ 2.02.0 dataset (Mrkšić et al., 2017)

WOZ 2.02.0 dataset has 12001200 restaurant domain task-oriented dialogues. There are three slots: ‘food’, ‘area’, ‘price range’, and a total of 9191 slot values. The dialogues are collected from a Wizard of Oz style experiment, in which the task is to find a restaurant that matches the slot values the user has specified. Each turn of a dialogue is annotated with a dialogue state, which indicates the slot values the user has informed. One example of the dialogue state is {‘food:Mexican’, ‘area’:‘east’, price range:‘moderate’}.

Table 4 shows the results on WOZ 2.02.0 dataset. We compare with four published baselines. SUMBT (Lee et al., 2019) is the current state-of-the-art model on WOZ 2.0 dataset. It fine-tunes a pre-trained BERT model (Devlin et al., 2019) to learn slot and utterance representations. StateNet PSI (Ren et al., 2018) maps contextualized slot embeddings and value embeddings into the same vector space, and calculate the Euclidean distance between these two. It also learns a joint model of all slots, enabling parameter sharing between slots. GLAD (Zhong et al., 2018) proposes to use a global module to share parameters between slots and a local module to learn slot-specific features. Neural Beflief Tracker (Mrkšić et al., 2017) applies CNN to learn n-gram utterance representations. Unlike prior works that transfer knowledge between slots by sharing parameters, our model implicitly transfers knowledge by formulating each slot as a question and learning to answer all the questions. Our model has a 1.24%1.24\% relative joint accuracy improvement over StateNet PSI. Although SUMBT achieves higher joint accuracy than DSTQA on WOZ 2.02.0 dataset, DSTQA achieves better performance than SUMBT on MultiWOZ 2.02.0 dataset, which is a more challenging dataset.

A.2 MultiWOZ 2.0/2.1 Ontology

The ontology of MultiWOZ 2.02.0 and MultiWOZ 2.12.1 datasets is shown in Table 5. There are 55 domains and 3030 slots in total. (two other domains ‘hospital’ and ‘police’ are ignored as they only exists in training set.)

A.3 Performance on Each Individual Domain

We show the performance of DSTQA w/span and TRADE on each single domain. We follow the same procedure as Wu et al. (2019) to construct training and test dataset for each domain: a dialogue is excluded from a domain’s training and test datasets if it does not mention any slots from that domain. During the training, slots from other domains are ignored. Table 6 shows the results. We can see that our model achieves better results on every domain, especially the hotel domain, which has a 11.24%11.24\% relative improvement. Hotel is the hardest domain as it has the most slots (10 slots) and has the lowest joint accuracy among all domains.

A.4 Joint Accuracy v.s. Context Length

We further show the model performance on different context lengths. Context lengths means the number of previous turns included in the dialogue context. Note that our baseline algorithms either use all previous turns as contexts to predict belief states or accumulate turn-level states of all previous turns to generate belief states. The results are shown in Figure 4. We can see that DSTQA with graph outperforms DSTQA without graph. This is especially true when the context length is short. This is because when the context length is short, graph carries information over multiple turns which can be used for multi-turn inference. This is especially useful when we want a shorter context length to reduce computational cost. In this experiment, the DSTQA model we use is DSTQA w/span.

A.5 Accuracy per Slot

The accuracy of each slot on MultiWOZ 2.02.0 and MultiWOZ 2.12.1 test set is shown in Figure 5 and Figure 6, respectively. Named related slots such as restaurant name, attraction name, hotel name has high error rate, because these slots have very large value set and high annotation errors.

A.6 Examples of Prediction Errors

This section describes prediciton errors made by DSTQA w/span. Incorrectly predicted (domain, slot, value) tuples are marked by underlines.