GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-Supervised Learning and Explicit Policy Injection
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, Jian Sun, Yongbin Li
Introduction
Task-oriented dialog (TOD) systems aim to help users accomplish certain tasks through conversations. Fundamental abilities of a TOD system include: (1) Dialog understanding: extracting structured semantics from user utterances; (2) Policy planning: determining a Dialog Act (DA) that leads to successful task completion; and (3) Dialog generation: producing appropriate responses (Figure 1). With the recent progress of Pre-trained Language Models (PLMs), remarkable performances improvements are achieved by casting TODs as generative language modeling tasks (Peng et al. 2020a; Lin et al. 2020), which benefit from the rich linguistic knowledge embedded in PLMs.
However, as reported in previous studies (Zhang et al. 2020b; Kulhánek et al. 2021), there are intrinsic differences between the distribution of human conversations and plain texts. Directly fine-tuning plain-text-trained PLMs on downstream dialog tasks hinders the model from effectively capturing conversational linguistic knowledge and thus leads to sub-optimal performances (Mehri et al. 2019; Zeng and Nie 2021; Wu and Xiong 2020). Current attempts to tackle this issue try to build Pre-trained Conversation Models (PCMs) by directly optimizing vanilla language model objectives on dialog corpora (Mehri, Eric, and Hakkani-Tur 2020; Zhang et al. 2020b; Henderson et al. 2019), which shows improved results on both dialog understanding (Wu et al. 2020) and generation (Peng et al. 2020b).
Despite these reported advances, few approaches are proposed to further enrich the pre-training process of PCMs with the knowledge of dialog policy. Specifically, existing methods either ignore explicit policy modeling or use latent variables without considering external dialog policy information (Bao et al. 2020), which hinders the possibility of learning controllable policy during pre-training. The optimization of dialog policy is usually formulated as a DA prediction task, which is crucial in TOD systems (Su et al. 2017; Liu et al. 2018). Therefore, we hypothesize that explicitly incorporating the DA annotations into the pre-training process can also facilitate learning better representations for policy optimization to improve the overall end-to-end performance.
A naive way to utilize these labels is to design a multi-task learning process (Sun et al. 2020) that directly combines vanilla unsupervised pre-training losses such as MLM (Devlin et al. 2018) with a supervised DA classification loss. However, this approach has several drawbacks when generalizing to large-scale pre-training paradigms: (1) The DA annotation schema is inconsistent among existing corpora, making it challenging to collect large-scale DA annotations; (2) A vast majority of available dialogs do not have DA labels. A naive joint training process without careful regularization would lead to highly over-fitting on those labeled samples, resulting in low performance; (3) All supervision signals from unlabeled data are self-supervised without any explicit inference over the DA space, so the linguistic knowledge PCMs can extract is only the general type, and the knowledge of dialog policy can not be effectively explored.
In this study, we propose a novel generative pre-trained model called GALAXY, aiming to inject the knowledge of dialog policy explicitly into pre-training at low cost while maintaining its strong ability on dialog understanding and generation. To begin with, we build a unified DA taxonomy for TOD and examine eight existing datasets to develop a new labeled dataset named UniDA with a total of 975K utterances. We also collect and process a large-scale unlabeled dialog corpus called UnDial with 35M utterances, whose scenarios ranging from online forums to customer services. Then, we propose a semi-supervised pre-training paradigm that applies consistency regularization (Verma et al. 2019) on all data. It minimizes the bi-directional KL-divergence between model predictions made on dropout-perturbed samples, which facilitates better representation learning from unlabeled dialog corpora. Since a large proportion of UnDial is from the Internet and not well-suited to our DA taxonomy, we add a learnable control gate on the KL loss of unlabeled data, so that only good samples are allowed for the consistent regularization, other samples are restricted back to normal self-supervised objectives. Experiments show that GALAXY substantially improves TOD systems and achieves new state-of-the-art results on In-Car, MultiWOZ2.0, and MultiWOZ2.1, pushing the end-to-end combined score to 107.45, 110.35, and 110.76, respectively. We also observe that GALAXY has a strong few-shot ability under various low-resource settings.
In summary, our main contributions are three-fold:
To the best of our knowledge, this is the first study to use semi-supervised pre-training to model explicit dialog policy for PCMs.
Experiments show our model has learned the knowledge of dialog policy, and achieves new state-of-the-art performance on several TOD benchmarks;
We collect a new labeled dataset UniDA as well as a large-scale unlabeled dialog corpus UnDial, hoping that can help bring forward the research in this area.
Related Work
are trained on large-scale textual corpora with Transformer (Devlin et al. 2018; Radford et al. 2019), which significantly improve dialog systems performance. Budzianowski and Vulić (2019) is the first work to validate the possibility of fine-tuning the information of all sub-tasks in a single paragraph of text on GPT-2. SimpleTOD (Hosseini-Asl et al. 2020) and SOLOIST (Peng et al. 2020a) further generalize this idea to an end-to-end setting where the semantic labels are generated instead of using ground truth values and also consider database results in the training process. Yang, Li, and Quan (2020) leverage the entire dialog session as the input sequence and demonstrate superior performance using self-generated responses during evaluation.
are variants of PLMs particularly adapted for conversational modeling. The main adaptation methods can be roughly divided into three types. The first is training PLMs on dialog corpora instead of plain texts with vanilla language model objectives. Recent work, such as DialoGPT (Zhang et al. 2020b), Meena (Adiwardana et al. 2020) and Blender (Roller et al. 2020) are trained on billions of open-domain dialogs, demonstrating powerful dialog generation performances. TOD-BERT (Wu et al. 2020) shows a great few-shot ability in various understanding tasks via pre-training BERT on extensive task-oriented dialog data. The second line is to design new dialog-oriented pre-training objectives (Bao et al. 2020; He et al. 2020, 2021; Xu and Zhao 2021; Su et al. 2021; Dai et al. 2021). Bao et al. (2020) use discrete latent variables to tackle the one-to-many mapping problem in open-domain dialog generation. Xu and Zhao (2021) propose to simulate the conversation features only using plain texts. The third is to integrate dialog annotations into the pre-training stage. Yu et al. (2020) use labels of dialog understanding as supervision to pre-train BERT. Peng et al. (2020b) use labeled conditional generation data to enhance dialog generation performance. Different from them, we are the first to utilize labels of dialog policy to improve PCMs.
learns from both unlabeled and labeled data. Approaches differ on what information to acquire from the structure of the unlabeled samples. Many initial results were based on generative models, such as variational autoencoders (Kingma and Welling 2019) and generative adversarial networks (Goodfellow et al. 2014). Pseudo-Labeling (Lee et al. 2013) is another widely used method, where unlabeled data is used as further training data after predicted by a model trained on labeled data. One line of recent research shows promising results by jointly training labeled data with supervised learning and unlabeled data with self-supervised learning (Sun et al. 2020). This lies in the paradigm of multi-task learning, where lower layers are often shared across all tasks while the top layers are task-specific. Consistency regularization (Verma et al. 2019) is also a prominent method in SSL, which improves classification performance by minimizing the discrepancy between predictions made on perturbed unlabeled data points. Recently, SimCSE (Gao, Yao, and Chen 2021) leverages dropout as the perturbed method and uses a contrastive objective as the regularization loss to learn sentence representations. Inspired by SimCSE, we adopt the same dropout method for perturbation, and use the bidirectional KL-divergence as in Liang et al. (2021) as our regularization loss, hoping to learn better representations that encodes the knowledge of dialog policy for downstream tasks. There are also some works (Jin et al. 2018; Zhang et al. 2020a; Liu et al. 2021) focusing on using latent variable models to alleviate the reliance on dialog labels via semi-supervised learning, but our work mainly targets the semi-supervised dialog pre-training.
Pre-training Dialog Datasets
In this section, we describe the new dialog datasets used for pre-training, including a labeled dialog dataset (UniDA) and a large-scale unlabeled dialog corpus (UnDial).
Dialog policyIn some datasets, the dialog act is defined as a combination of an act and its semantic contents. To unify different datasets, we neglect the contents and only use dialog acts as the annotations. We also focus on the text-in-text-out TOD systems in this paper, and leave the spoken DA in the future research. is tasked to predict dialog acts (DAs) given dialog context. Although DAs are general tags to describe speakers’ communicative behaviors (Bunt 2009), current DA annotations in task-oriented dialog are still limited and lack of unified taxonomy because each dataset is small and scattered. Recently, Paul, Goel, and Hakkani-Tür (2019) propose a universal task-oriented DA schema, but their dataset is still insufficient for pre-training purposes and the schema lacks some important features such as not_sure and dont_understand. To this end, we follow ISO (Bunt et al. 2010) and propose a more comprehensive unified DA taxonomy for task-oriented dialog, which consists of 20 frequently-used DAs. A complete description of the taxonomy is in Appendix A.1. Base on that, we align the annotations of eight existing benchmarks: MultiWOZ (Budzianowski et al. 2018), Frames (Asri et al. 2017), MSRe2e (Li et al. 2018), SGD (Rastogi et al. 2020), DSTC2 (Henderson, Thomson, and Williams 2014), SimJoint (Shah et al. 2018), STAR (Mosig, Mehri, and Kober 2020) and DailyDialog (Li et al. 2017). We add DailyDialog, an open-domain dialog dataset, to accommodate our dialog policy for more general types. Finally, a new dataset UniDA is obtained. Table 1 shows more detailed statistics.
2 Unlabeled Dataset: UnDial
Large clean dialogs are difficult to acquire. We build the unlabeled dialog corpora from various available sources, ranging from online forum chatting logs to customer service conversations. We select 14 existing dialog corpora and perform careful processing on all data. Then we acquire a large-scale unlabeled dialog dataset UnDial, which consists of 35M utterances. Table 2 shows the statistics of our final pre-training unlabeled data. For more details about the data statistics and the text processing method, please refer to Appendix A.2.
Method
In this section, we first introduce the model architecture. Then we describe each objective used in our pre-training and the proposed semi-supervised pre-training paradigm.
We choose UniLM (Dong et al. 2019) as our backbone model. It contains a bi-directional encoder for understanding and a uni-directional decoder for generation, which is naturally suitable for task-oriented dialog modeling. The encoder and the decoder are weight-shared. We adopt a similar scheme of input representation in Bao et al. (2020), where the input embeddings consist of four elements: tokens, roles, turns, and positions. Role embeddings are like segmentation embeddings in BERT and are used to differentiate which role the current token belongs to, either user or system. Turn embeddings are assigned to each token according to its turn number. Position embeddings are assigned to each token according to its relative position within its belonging sentence. More details can be found in Appendix B.1.
2 Pre-training Objectives
Four objectives are employed in our dialog pre-training process: response selection, response generation, DA prediction and consistency regularization. Figure 2 illustrates the procedure of pre-training.
Many work (Wu et al. 2020; Bao et al. 2020; Henderson et al. 2019) show that the response selection task can capture the coherency between dialog contexts and responses and thus benefit dialog understanding. We follow their implementation and model this task as a binary classification problem. Specifically, for a context response pair from the corpus, the positive example (with label ) is obtained by concatenating with its corresponding response , and the negative example (with label ) is constructed by concatenating with a response that is randomly selected from the corpus. A binary cross-entropy loss is defined as:
in which the classification probability is calculated by feeding the concatenated sequence of and into the bi-directional encoder and adding a binary classification head on the extracted representation of token [CLS] from the last transformer layer:
where is a fully-connected neural network with the output layer of size 1. is the sigmoid function acts on each dimension of the input vector.
The response generation task aims to predict the dialog response auto-regressively based on the dialog context . We adopt the standard negative log-likelihood loss for the generation task:
where is the -th word in , represents the words of previous steps.
For a context response pair sampled from UniDA, the DA prediction task aims to predict the DA label of the response based merely on the context . Note that, since there are some responses in UniDA are associated with multiple DAs, we model the DA prediction task as a multi-label classification problem. We denote , where is the total number of dialog acts. A multi-dimensional Bernoulli distribution is used for dialog acts: . Taking the dialog context as input, we add a multi-dimensional binary classifiers on to predict each act . The binary classification loss is:
where is a fully-connected neural network with the output layer of size . is the true label of .
For UnDial, the DA annotations are unavailable. In that case, we need to infer the DA labels based on the given dialog context . Instead of using in Eq. (5), we use a categorical distribution for dialog acts:
where is the softmax function, is the same feed-forward neural network in Eq. (5). So . Then we employ a dropout-based consistency regularization to learn better representations (Gao, Yao, and Chen 2021). Concretely, given the same dialog context , we feed to go through the forward pass of the model twice. Due to the randomness of the dropout mechanism in transformers, we can get two different sets of hidden features, and therefore, two different categorical distributions of dialog policy, denoted as and . Then the Kullback-Leibler (KL) divergence between these two output distributions is calculated as . We minimize the bidirectional KL divergence as in (Liang et al. 2021) between the two distributions to regularize the model predictions, which is defined as:
Figure 3 illustrate the procedure of computing .
3 Semi-supervised Pre-training Paradigm
For the unlabeled data UnDial, since some dialogs collected from the open-domain Internet are too noisy to be compatible with our DA taxonomy, we propose to use a gating mechanism to select a high-quality subset of UnDial for prediction. In practice, we compute a soft gating score based on the entropy of to control whether a data point is adopted for consistency regularization in the current iteration.
where is the Maximum Entropy of -dimensional probability distribution. is the current entropy of , i.e., . In practice, we use the perturbed distribution as the approximation of to calculate the gate score.
In the pre-training process, we mix and shuffle UniDA and UnDial, and randomly sample batches from the mixed corpus.
4 Fine-tuning and Inference
In the fine-tuning stage, we concentrate on task-oriented dialog tasks. For tasks that contained necessary semantic labels (e.g., belief states and dialog acts), we re-organize the response to contain those labels, and generate them together. Suppose the sequence of the labels is . Thus the new response is the concatenation of and and is generated in the downstream tasks. For tasks that do not have semantic labels, we generate the initial response . We also maintain the DA prediction task to alleviate the model discrepancy between pre-training and fine-tuning (Zeng and Nie 2021). Therefore, The fine-tuning loss is as follows:
where for tasks that provide DA annotations and for tasks that contain no DA annotations.
Experimental Settings
We evaluate the end-to-end dialog system performance of GALAXY on two well-studied task-oriented dialog benchmarks: Stanford In-Car Assistant (In-Car) (Eric and Manning 2017), MultiWOZ (Budzianowski et al. 2018). In-Car consists of dialogs between a user and an in-car assistant system covering three tasks: calendar scheduling, weather information retrieval, and point-of-interest navigation. Following the data processing in (Zhang et al. 2020a), we divide the dataset into training/validation/testing sets with 2425/302/304 dialogs respectively. MultiWOZ is a large-scale human-human dataset spanning seven domains, which is one of the most challenging datasets in task-oriented dialog due to its complex ontology and diverse language styles. We evaluate our model on MultiWOZ2.0 (the original version) and MultiWOZ2.1 (a revised version) since both are popular benchmarks with various competing models. Following the data processing in Yang, Li, and Quan (2020), we obtain 8438/1000/1000 dialogs for training/validation/testing respectively. We also adopt delexicalized responses for task-oriented generation, which allows the model to learn value-independent parameters (Zhang, Ou, and Yu 2020).
2 Evaluation Metrics
We use BLEU (Papineni et al. 2002) to measure the response generation quality. Metrics relate to task completion are used for separate datasets to facilitate comparison with prior works. For MultiWOZ, we report Inform, Success, as a combined score (Comb) is also computed via (Inform Success)BLEU as an overall quality measure as in Mehri, Srinivasan, and Eskenazi (2019). For In-Car, we use Match and SuccF1 following Lei et al. (2018), and calculate a similar combined score (Comb) via (Match SuccF1)BLEU.
Experimental Results
In our experiments, we focus on the setting of end-to-end dialog modeling (E2E), in which no ground-truth immediate labels are provided to the model. GALAXY is initialized with UniLM and then performs semi-supervised pre-training with UniDA and UnDial. Notably, we removed the validation and testing set of MultiWOZ from UniDA during pre-training for fairness. We compare GALAXY with all published work on respective datasets. We also compare different pre-trained conversation models (PCMs) and different semi-supervised pre-training methods to verify the efficacy of GALAXY. In addition, we conduct an extensive discussion and analysis to reveal the internal performance of GALAXY. More details about implementation can be found in Appendix B.2.
As shown in Table 3 and Table 4, GALAXY achieves new state-of-the-art combined scores on all datasets, improving In-Car by 2.5 points (from 104.95 to 107.45), MultiWOZ2.0 by 5.3 points (from 105.05 to 110.35), and MultiWOZ2.1 by 5.5 points (from 105.25 to 110.76). Note that in both tables, GALAXY is the only model that can obtain best Success while maintaining BLEU at a very high level, which means that GALAXY can take better dialog policy than other models to facilitate task completion, and therefore generate better responses. Our model can also achieve competitive results in Inform on par with other best baselines. We also report the results of GALAXY (w/o pre-train) without the pre-training procedure on more dialog corpora. From both tables, GALAXY also achieves comparable results with previous best models, indicating that our model architecture is competitive for dialog modeling. More E2E results given oracle belief states on MultiWOZ are shown in Appendix D.
2 Comparison with Other PCMs
We verify that GALAXY has a much better ability to fulfill task-oriented dialog tasks than other PCMs due to modeling dialog policy during pre-training. To alleviate the discrepancy brought from model structure, we use UniLM (Dong et al. 2019) and PLATO (Bao et al. 2020) as our baselines. We also train both models on our pre-training dialog datasets (UniDA and UnDial) with their original objectives and perform the same fine-tuning process on MultiWOZ2.0. We denote the new models as TOD-UniLM and TOD-PLATO, respectively. As shown in Table 5, the results of both models are worse than GALAXY due to the lack of using important information of dialog policy.
3 Comparison with Other Semi-supervised Pre-training Methods
4 Low Resource Evaluation
Many recent works (Peng et al. 2020b; Wu et al. 2020) have demonstrated that pre-trained models have a solid few-shot ability in the understanding and conditional generation tasks. We also evaluate GALAXY in the simulated low resource setting on MultiWOZ2.0, showing that it is more sample-efficiency than existing models. Specifically, we use 5%, 10%, 20%, and 50% of the training set data to train our models and baselines. To be fair, we discard the (1-X%) training data of MultiWOZ from UniDA in the pre-training process under each X% setting, eliminating the influence of using any external data. Compared baselines include: DAMD (Zhang, Ou, and Yu 2020), SOLOIST (Peng et al. 2020a), MinTL (Lin et al. 2020), PPTOD (Su et al. 2021) and UBAR (Yang, Li, and Quan 2020). Experimental results in Table 7 show that GALAXY significantly outperforms other models under all low-resource settings.
Analysis and Discussion
Figure 4 illustrates a case where GALAXY chooses correct dialog acts for the first two turns so that the whole conversation can steer towards successful task completion. On the contrary, UBAR takes a wrong DA notify-failure at the beginning turn and a redundant DA request at the second turn, which leads to a failure for the interaction.
Conclusion
In this paper, we propose GALAXY, a pre-trained conversation model that learns dialog policy explicitly in the pre-training process via semi-supervised learning. We introduce a dialog act prediction task for policy optimization and use a consistency regularization loss to learn better representations on unlabeled dialog corpora. A gating mechanism is also used to weigh suitable unlabeled samples. Experiments show that our model creates new SOTA results on several task-oriented dialog benchmarks and outperforms existing models by a large margin in various low-resource settings. We hope that GALAXY, and the newly collected labeled dataset UniDA and large-scale unlabeled corpus UnDial, can inspire researchers to explore the new paradigm to build pre-trained conversation models for task-oriented dialog.
Acknowledgement
This work was supported by Alibaba Group through Alibaba Research Intern Program. This work was partially supported by National Natural Science Foundation of China (No. 61906185), Youth Innovation Promotion Association of CAS China (No. 2020357), Shenzhen Science and Technology Innovation Program (Grant No. KQTD20190929172835662), Shenzhen Basic Research Foundation (No. JCYJ20200109113441941). We also thank Dr. Yichi Zhang for the cute cartoon in Figure 1.
References
Appendix
The hierarchical structure of our proposed unified DA taxonomy is illustrated in Figure 6. There are totally 20 labels.
This group consists of DAs about regular actions for social behaviors: hi, bye, thank_you, repeat, welcome, dont_understand.
hi means greeting responses, like ‘hello’, ‘how are you’.
bye means the responses for saying goodbye.
thank_you means the responses for appreciation.
repeat means asking the user to repeat what he/she said last turn again.
welcome denotes a paragraph of official texts to broadcast the information that the system can offer, like ‘welcome to Cambridge restaurant, we can help you to order food, you can find restaurants by talking about your favorite foods, area, price range.’
dont_understand means the system can not understand what the user says, which is normal when the user talk about something beyond the semantic scope that the system can process.
This group consists of DAs about providing suggestions or imperative orders.
propose means suggesting to do/offer/recommend something, in order to make the user consider the performance of a certain action, which the system believes is in the user’s interests. For example ‘How about we find a good place to have fun.’
direct means imperative responses that expresses an order, e.g., ‘you need to open the light before going to bad.’
This group consists of DAs that perform actions about asking.
request means asking the user about specific attributes, like ‘what area do you like?’
select means asking the user to choose a preferred choices from a set of candidates.
reqalts means asking the user for more information. e.g., ‘what else information do you want?’
This group consists of DAs that provides specific answers to the user.
affirm denotes the affirmative responses. e.g., ‘Yes, it is.’
not_sure means the system is not certain about the user’s confirmation.
negate denotes the negating responses. ‘Noe, it is not.’
inform denotes the normal answers to give the information required by the user. e.g., ‘The hotel is in the east area.’
offer means the system offer the current searching results from the database that match the user’s need. e.g., ‘There are 10 restaurants I’ve found for you.’
notify-success means the system notifies the user that his/her goal is finished successfully . e.g., ‘Sure, the XXX is a good one, I’ve booked it for you.’
notify-failure means the system notifies the user that his/her goal is not finished successfully . e.g., ‘Sorry, I can not book it for you now, because it is full’
This group consists of DAs that the system ask the user about something to confirm whether it is true or correct.
expl-confirm means to ask the user explicitly to check something. e.g. ‘Do you need to cheap restaurant ?’
impl-confirm means to check something implicitly, often in a statement that repeats what user says. e.g. ‘You want a cheap restaurant, OKay.’
A.2. Details for UnDial.
The Detailed statistics are given in Table 10. We totally aggregate 14 dialog corpora from the Internet. The processing methods includes: (1) Removing the instances where there is a URL in utterances. (2) Removing the instances containing word repetitions of at least three words (3) removing non-English sentences. (4) removing sentence containing special markers such as “[” or “]”, as this could be markup. (5) removing offensive language. (6) Replacing the non-unicode characters like emojis.
Appendix B
Figure 8 illustrates the input representations in the pre-training stage, we use special tokens [CLS], [BOS] and [EOS] to concatenate sentences in context and the response. Apart from token embeddings, we also have position emebddings, role embeddings and turn embeddings as in Bao et al. (2020).
For the fine-tuning stage, we need to consider the semantic labels, such as ‘belief states’ and ‘database results’, so we add more special tokens to concatenate them as in Yang, Li, and Quan (2020). Figure 9 shows the input sequence of GALAXY in downstream tasks: In-Car and MultiWOZ.
B.2. Implementation Details.
We introduce hyper-parameters used in pre-training and fine-tuning as follows. The number of transformer blocks in GALAXY is 12 and the hidden embedding dimension is 768. The total number of dialog acts is 20. In the pre-training stage, GALAXY is initialized with UniLM. The maximum sequence length of dialog context and response is set to 256 and 50, respectively. The batch size is set to 128 and AdamW optimizer is employed for optimization with an initial learning rate of 1e-5. The dropout rate is set to 0.3 for consistency regularization. For semi-supervised pre-training, at each iteration, we mix and shuffle the labeled dataset UniDA and unlabeled dataset UnDial, then randomly sample batches from the mixed corpus as the input of GALAXY. We use a random seed 11, and choose the model checkpoint at the 14th epoch as the final pre-trained model.
For the fine-tuning stage, the maximum sequence length of dialog context and response is set to 1024 and 100 due to longer responses including semantic labels. The grid search algorithm is applied on the validation set to automatically tune the hyper-parameters. We use AdamW optimizer with an initial learning rate of 1e-4. For MultiWOZ dataset, the batch size is set to 32 and the dropout rate is set to 0.1. For In-Car dataset, the batch size is set to 64 and the dropout rate is set to 0.35.
Appendix C
Figure 7 shows the framework for the VAE method. We leverage a hidden variable that has the same size as dialog act . For unlabeled data, the generative process of is (Figure 7 (a)):
Sample a latent variable z based on the dialog context and response for training: while only based on the dialog context for testing: .
Generate the response based on the dialog context and latent variable : .
For labeled data, the generative process of is (Figure 7 (b)):
Sample a latent variable z based on the dialog context , response and dialog act for training: while only based on the dialog context for testing: .
Predict the dialog act based on the dialog context and latent variable : .
Generate the response based on the dialog context , latent variable and dialog act : .
Appendix D
Table 11 shows the total end-to-end results given oracle belief states on MultiWOZ2.0 and MultiWOZ2.1.