Recent Advances in Deep Learning Based Dialogue Systems: A Systematic Survey
Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, Erik Cambria
Introduction
Dialogue systems (or chatbots) are playing a bigger role in the world. People may still have a stereotype that chatbots are those rigid agents in their phone calls to a bank. However, thanks to the revival of artificial intelligence, the modern chatbots can converse with rich topics ranging from your birthday party to a speech given by Biden, and, if you want, they can even book a place for your party or play the speech video. At present, dialogue systems are one of the hot topics in NLP and are highly demanded in industry and daily life. The market size of chatbot is projected to grow from 9.4 billion by 2024 at a compound annual growth rate (CAGR) of 29.7% Statistic source: https://markets.businessinsider.com and 80% of businesses are expected to be equipped with chatbot automation by the end of 2021 Statistic source: https://outgrow.co.
Traditional task-oriented dialogue systems are organized in a pipeline structure and consist of four functional modules: Natural Language Understanding, Dialogue State Tracking, Policy Learning, and Natural Language Generation, which will be discussed in detail in Section 3. Many state-of-the-art works design end-to-end task-oriented dialogue systems to achieve better optimization compared with pipeline methods. Open-domain dialogue systems are generally divided into three categories: generative systems, retrieval-based systems, and ensemble systems. Generative systems apply sequence-to-sequence models (see Section 2.2.5) to map the user message and dialogue history into a response sequence that may not appear in the training corpus. By contrast, retrieval-based systems try to select a pre-existing response from a certain response set. Ensemble systems combine generative methods and retrieval-based methods in two ways: retrieved responses can be compared with generated responses to choose the best among them; generative models can also be used to refine the retrieved responses (Zhu et al., 2018; Song et al., 2016; Qiu et al., 2017; Serban et al., 2017b). Generative systems can produce flexible and dialogue context-related responses while sometimes they lack coherence The quality of being logical and consistent not only between words/subwords but also between responses of different timesteps. and tend to make dull responses (Serban et al., 2016; Vinyals and Le, 2015; Sordoni et al., 2015b). Retrieval-based systems select responses from human response sets and thus are able to achieve better coherence in surface-level language. However, retrieval systems are restricted by the finiteness of the response sets and sometimes the responses retrieved show a weak correlation with the dialogue context (Zhu et al., 2018).
For dialogue systems, existing surveys (Arora et al., 2013; Wang and Yuan, 2016; Mallios and Bourbakis, 2016; Chen et al., 2017a; Gao et al., 2018) are either outdated or not comprehensive. Some definitions in these papers are no longer being used at present, and a lot of new works and topics are not covered. In addition, most of them lack a multi-angle analysis. Thus, in this survey, we comprehensively review high-quality works in recent years with a focus on deep learning-based approaches and provide insight into state-of-the-art research from both model angle and system angle. Moreover, this survey updates the definitions/names according to state-of-the-art research. E.g., we name "open-domain dialogue systems" instead of "chit-chat dialogue systems" because most of the articles (roughly 70% according to our survey) name them as the prior one. We also extensively cover the diverse hot topics in dialogue systems and extend some new topics that are popular in current research community (such as Domain Adaptation, Dialogue State Tracking Efficiency, End-to-end methods for task-oriented dialogue systems; Controllable Generation, Interactive Training, and Visual Dialogue for open-domain dialogue systems).
Traditional dialogue systems are mostly rule-based (Arora et al., 2013) and non-neural machine learning based systems. Rule-based systems are easy to implement and can respond naturally, which contributed to their popularity in earlier industry products. However, the dialogue flows of these systems are predetermined, which keeps the applications of the dialogue systems within certain scenarios. Non-neural machine learning based systems usually perform template filling to manage certain tasks. These systems are more flexible compared with rule-based systems because the dialogue flows are not predetermined. However, they cannot achieve high F1 scores (Powers, 2020) in template fillingTemplate filling is an efficient approach to extract and structure complex information from text to fill in a pre-defined template. They are mostly used in task-oriented dialogue systems. and are also restricted in application scenarios and response diversity because of the fixed templates. Most if not all state-of-the-art dialogue systems are deep learning-based systems (neural systems). The rapid growth of deep learning improves the performance of dialogue systems (Chen et al., 2017a). Deep learning can be viewed as representation learning with multilayer neural networks. Deep learning architectures are widely used in dialogue systems and their subtasks. Section 2 discusses various popular deep learning architectures.
Apart from dialogue systems, there are also many dialogue-related tasks in NLP, including but not limited to question answering, reading comprehension, dialogue disentanglement, visual dialogue, visual question answering, dialogue reasoning, conversational semantic parsing, dialogue relation extraction, dialogue sentiment analysis, hate speech detection, MISC detection, etc. In this survey, we also touch on some works tackling these dialogue-related tasks, since the design of dialogue systems can benefit from advances in these related areas.
We produced a diagram for this article to help readers familiarize the overall structure (Figure 1). In this survey, Section 1 briefly introduces dialogue systems and deep learning; Section 2 discusses the neural models popular in modern dialogue systems and the related work; Section 3 introduces the principles and related work of task-oriented dialogue systems and discusses the research challenges and hot topics; Section 4 briefly introduces the three kinds of systems and then focuses on hot topics in open-domain dialogue systems; Section 5 reviews the main evaluation methods for dialogue systems; Section 6 comprehensively summarizes the datasets commonly used for dialogue systems; finally, Section 7 concludes the paper and provides some insight on research trends.
Neural Models in Dialogue Systems
In this section, we introduce neural models that are popular in state-of-the-art dialogue systems and related subtasks. We also discuss the applications of these models or their variants in modern dialogue systems research to provide readers with a picture from the model’s perspective. This will help researchers acquaint these models and see how they are applied in state-of-the-art frameworks, which is rather helpful when designing a new dialogue system. The models discussed include: Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Vanilla Sequence-to-sequence Models, Hierarchical Recurrent Encoder-Decoder (HRED), Memory Networks, Attention Networks, Transformer, Pointer Net and CopyNet, Deep Reinforcement Learning models, Generative Adversarial Networks (GANs), Knowledge Graph Augmented Neural Networks. We start from some classical models (e.g., CNNs and RNNs), and readers who are familiar with their principles and corresponding applications in dialogue systems can choose to read selectively.
Deep neural networks have been considered as one of the most powerful models. ‘Deep’ refers to the fact that they are multilayer, which extracts features by stacking feed-forward layers. Feed-forward layers can be defined as: . Where the is an activation function; and are trainable parameters. The feed-forward layers are powerful due to the activation function, which makes the otherwise linear operation, non-linear. Whereas there exist some problems when using feed-forward layers. Firstly, the operations of feed-forward layers or multilayer neural networks are just template matching, where they do not consider the specific structure of data. Furthermore, the fully connected mechanism of traditional multilayer neural networks causes an explosion in the number of parameters and thus leads to generalization problems. LeCun et al. (1998) proposed LeNet-5, an early CNN. The invention of CNNs mitigates the above problems to some extent.
CNNs (Figure 2) usually consist of convolutional layers, pooling layers and feed-forward layers. Convolutional layers apply convolution kernels to perform the convolution operation:
Where and are respectively the indexes of rows and columns of the result matrix. denotes the input matrix and denotes the convolutional kernel. The pooling layers perform down-sampling on the result of convolutional layers to get a higher level of features and the feed-forward layers map them into a probability distribution to predict class scores.
A sliding window feature enables convolution layers to capture local features and the pooling layers can produce hierarchical features. These two mechanisms give CNNs the local perception and global perception ability, helping to capture some specific inner structures of data. The parameter sharing mechanism eases the parameter explosion problem and overfitting problem because the reduction of trainable parameters leads to less model complexity, improving the generalization ability.
Due to these good properties, CNNs have been widely applied in many works. Among them, the Computer Vision tasks benefit the most for that the Spatio-temporal data structures of images or videos are perfectly captured by CNNs. For more detailed mechanism illustrations and other variants of CNNs, readers can refer to these representative algorithm papers or surveys: (Krizhevsky et al., 2012; Zeiler and Fergus, 2014; Simonyan and Zisserman, 2014; Szegedy et al., 2015; He et al., 2016; Aloysius and Geetha, 2017; Rawat and Wang, 2017). In this survey, we focus on dialogue systems.
Recent years have seen a dramatic increase in applications of CNNs in NLP. Many tasks take words as basic units. However, phrases, sentences, or even paragraphs are also useful to semantic representations. As a result, CNNs are an ideal tool for the hierarchical modeling of language (Conneau et al., 2016).
CNNs are good textual feature extractors, but they may not be ideal sequential encoders. Some dialogue systems (Qiu et al., 2019; Bi et al., 2019; Ma et al., 2020a) directly used CNNs as the encoder of utterances or knowledge, but most of the state-of-the-art dialogue systems such as Feng et al. (2019); Wu et al. (2016); Tao et al. (2019); Wang et al. (2019b); Chauhan et al. (2019); Feldman and El-Yaniv (2019); Chen et al. (2019c); Lu et al. (2019b) and Coope et al. (2020) chose to use CNNs as a hierarchical feature extractor after encoding the text information, instead of directly applying them as encoders. This is due to the fixed input length and limited convolution span of CNNs. Generally, there are two main situations where CNNs are used to process encoded information in dialogue systems. The first situation is applying CNNs to extract features directly based on the feature vectors from the encoder (Wang et al., 2019b; Chauhan et al., 2019; Feldman and El-Yaniv, 2019; Chen et al., 2019c) and Coope et al. (2020). Within the works above, Feldman and El-Yaniv (2019) extracted features from character-level embeddings, illustrating the hierarchical extraction capability of CNNs. Another situation in which CNNs are used is extracting feature maps in response retrieval tasks. Some works built retrieval-based dialogue systems (Wu et al., 2016; Feng et al., 2019; Tao et al., 2019; Lu et al., 2019b). They used separate encoders to encode dialogue context and candidate responses and then used a CNN as an extractor of the similarity matrix calculated from the encoded dialogue context and candidate responses. Their experiments showed that this method can achieve good performance in response retrieval tasks.
The main reason why more recent works do not choose CNNs as dialogue encoders is that they fail to extract the information across temporal sequence steps continuously and flexibly (Krizhevsky et al., 2012). Some models introduced later do not process data points independently, which are desirable models for encoders.
2 Recurrent Neural Networks and Vanilla Sequence-to-sequence Models
NLP tasks including dialogue-related tasks try to process and analyze sequential language data points. Even though standard neural networks, as well as CNNs, are powerful learning models, they have two main limitations (Lipton et al., 2015). One is that they assume the data points are independent of each other. While it is reasonable if the data points are produced independently, essential information can be missed when processing interrelated data points (e.g., text, audio, video). Additionally, their inputs are usually of fixed length, which is a limitation when processing sequential data varying in length. Thus, a sequential model being able to represent the sequential information flow is desirable.
Markov models like Hidden Markov Models (HMMs) are traditional sequential models, but due to the time complexity of the inference algorithm (Viterbi, 1967) and because the size of transition matrix grows significantly with the increase of the discrete state space, in practice they are not applicable in dealing with problems involving large possible hidden states. The property that the hidden states of Markov models are only affected by the immediate hidden states further limits the power of this model.
RNN models are not proposed recently, but they greatly solve the above problems and some variants can amazingly achieve state-of-the-art performance in dialogue-related tasks as well as many other NLP tasks. The inductive bias of recurrent models is non-replaceable in many scenarios, and many up-to-date models incorporate the recurrence.
In 1982, Hopfield introduced an early family of RNNs to solve pattern recognition tasks (Hopfield, 1982). Jordan (1986) and Elman (1990) introduced two kinds of RNN architectures respectively. Generally, modern RNNs can be classified into Jordan-type RNNs and Elman-type RNNs.
The Jordan-type RNNs are shown in Figure 3(a). , , and are the inputs, hidden state, and output of time step , respectively. , and are weight matrixes. Each update of hidden state is decided by the current input and the output of last time step while each output is decided by current hidden state. Thus the hidden state and output of time step can be calculated as:
Where and are biases. and are activation functions.
The Elman-type RNNs are shown in Figure 3(b). The difference is that each hidden state is decided by the current input and the hidden state of last time step. Thus the hidden state and output of time step can be calculated as:
Simple RNNs can model long-term dependencies theoretically. But in practical training, long-range dependencies are difficult to learn (Bengio et al., 1994; Hochreiter et al., 2001). When backpropagating errors over many time steps, simple RNNs suffer from problems known as gradient vanishing and gradient explosion (Hochreiter and Schmidhuber, 1997). Some solutions were proposed to solve these problems (Williams and Zipser, 1989; Pascanu et al., 2013), which led to the inventions of some variants of traditional recurrent networks.
2.2 LSTM
Hochreiter and Schmidhuber (1997) introduced gate mechanisms in LSTM mainly to address the gradient vanishing problem. Input gate, forget gate and output gate were introduced to decide how much information from new inputs and past memories should be reserved. The model can be described by the following equations:
Where represents time step . , and are gates, denoting input gate, forget gate and output gate respectively. , , and are input, short-term memory, long-term memory and output respectively. is bias and is weight matrix. denotes element-wise multiplication.
The intuition of the term “Long Short-Term Memory" is that the proposed model applies both long-term and short-term memory vectors to encode the sequential data, and uses gate mechanisms to control the information flow. The performance of LSTM is impressive since that it achieved state-of-the-art results in many NLP tasks as a backbone model although this model was proposed in 1997.
2.3 GRU
Inspired by the gating mechanism, Cho et al. (2014b) proposed Gated Recurrent Unit (GRU), which can be modeled by the equations:
Where represents time step . and are gates, denoting update gate and reset gate respectively. , and are input, candidate activation vector and output respectively. is bias while and are weight matrixes. denotes element-wise multiplication.
LSTM and GRU, as two types of gating units, are very similar to each other (Chung et al., 2014). The most prominent common point between them is that from time step to time step , an additive component is introduced to update the state whereas simple RNNs always replace the activation. Both LSTM and GRU keep certain old components and mix them with new contents. This property enables the units to remember the information of history steps farther back and, more importantly, avoid gradient vanishing problems when backpropagating the error.
There also exist several differences between them. LSTM exposes its memory content under the control of the output gate, while the same content in GRU is in an uncontrolled manner. Additionally, different from LSTM, GRU does not independently gate the amount of new memory content being added. And if looking from experimental perspective, GRU has fewer parameters, which contributes to its faster convergence and better generalization ability. It has also been shown that GRU can achieve better performance in smaller datasets (Chung et al., 2014). However, Gruber and Jockisch (2020) showed that LSTM cells exhibited consistently better performance in a large-scale analysis of Neural Machine Translation.
2.4 Bidirectional Recurrent Neural Networks
In sequence learning, not only the past information is essential to the model inference, the future information should also be considered to achieve a better inference ability. Schuster and Paliwal (1997) proposed the bi-directional recurrent neural networks (BRNNs), which had two kinds of hidden layers: the first encoded information from past time steps while the second encoded information in a flipped direction. The model can be described using the equations:
Where and are the two hidden layers. Other variables are defined in the same way as in the case of LSTMs and GRUs.
2.5 Vanilla Sequence-to-sequence Models (Encoder-decoder Models)
Sutskever et al. (2014) first proposed the sequence-to-sequence model to solve the machine translation tasks. The sequence-to-sequence model aimed to map an input sequence to an output sequence by first using an encoder to map the input sequence into an intermediate vector and a decoder further generated the output based on the intermediate vector and history generated by the decoder. The equations below illustrate the encoder-decoder model:
Where is the time step, is the hidden vector and is the output vector. and are the sequential cells used by the encoder and decoder respectively. The last hidden state of the encoder is the intermediate vector, and this vector is usually used to initialize the first hidden state of the decoder. At encoding time, each hidden state is decided by the hidden state of the previous time step and the input at the current time step, while at decoding time, each hidden state is decided by the current hidden state and the output of the previous time step.
This model is powerful because it is not restricted to fixed-length inputs and outputs. Instead, the length of the source sequence and target sequence can differ. Based on this model, many more advanced sequence-to-sequence models have been developed, which will be discussed in this and subsequent sections.
RNNs play an essential role in neural dialogue systems for their strong ability to encode sequential text information. RNNs and their variants are found in many dialogue systems. Task-oriented systems apply RNNs as encoders of dialogue context, dialogue state, knowledge base entries, and domain tags (Moon et al., 2019; Chen et al., 2019b; Wu et al., 2019b, a). Open-domain systems apply RNNs as dialogue history encoders (Sankar et al., 2019; Du and Black, 2019; Ji et al., 2020; Chen et al., 2020b), among which retrieval-based systems model dialogue history and candidate responses together (Zhu et al., 2018; Tang et al., 2019; Feldman and El-Yaniv, 2019; Lu et al., 2019b). In knowledge-grounded systems, RNNs are encoders of outside knowledge sources (e.g., background, persona, topic, etc.) (Shuster et al., 2019; Majumder et al., 2020b; Chen et al., 2020b; Cho and May, 2020).
Furthermore, as the decoder of sequence-to-sequence models in dialogue systems (Huang et al., 2020c; Song et al., 2019; Liu et al., 2019; Lin et al., 2019), RNNs usually decode the hidden state of utterance sequences by greedy search or beam search (Aubert et al., 1994). These decoding mechanisms cause problems like generic responses, which will be discussed in later sections.
Some works (Liu et al., 2019; Mehri et al., 2019; Chen et al., 2019c; Ma et al., 2020a) combined RNNs as a part of dialogue representation models to train dialogue embeddings and further improved the performance of dialogue-related tasks. These embedding models were trained on dialogue tasks and present more dialogue features. They consistently outperformed state-of-the-art contextual representation models (e.g., BERT, ELMo, and GPT) in some dialogue tasks when these contextual representation models were not fine-tuned for the specific tasks.
3 Hierarchical Recurrent Encoder-Decoder (HRED)
Hierarchical Recurrent Encoder-Decoder (HRED) is a context-aware sequence-to-sequence model. It was first proposed by Sordoni et al. (2015a) to address the context-aware online query suggestion problem. It was designed to be aware of history queries and the proposed model can provide rare and high-quality results.
With the popularity of the sequence-to-sequence model, Serban et al. (2016) extended HRED to the dialogue domain and built an end-to-end context-aware dialogue system. HRED achieved noticeable improvements in dialogue and end-to-end question answering. This work attracted even more attention than the original paper for that dialogue systems are a perfect setting for the application of HRED. Traditional dialogue systems (Ritter et al., 2011) generated responses based on the single-turn messages, which sacrificed the information in the dialogue history. Sordoni et al. (2015b) combined dialogue history turns with a window size of 3 as the input of a sequence-to-sequence model for response generation, which is limited as well for that they encode the dialogue history only in token-level. The “turn-by-turn" characteristic of dialogue indicated that the turn-level information also matters. The HRED learned both token-level and turn-level representation, thus exhibiting promising dialogue context awareness.
Figure 4 represents the HRED in a dialogue setting. HRED models the token-level and turn-level sequences hierarchically with two levels of RNNs: a token-level RNN consisting of an encoder and a decoder, and a turn-level context RNN. The encoder RNN encodes the utterance of each turn token by token into a hidden state. This hidden state is then taken as the input of the context RNN at each turn-level time step. Thus the turn-level context RNN iteratively keeps track of the history utterances. The hidden state of context RNN at turn represents a summary of the utterances up to turn and is used to initialize the first hidden state of decoder RNN, which is similar to a standard decoder in sequence-to-sequence models (Sutskever et al., 2014). All of the three RNNs described above apply GRU cells as the recurrent unit, and the parameters of encoder and decoder are shared for each utterance.
Serban et al. (2017a) further proposed Latent Variable Hierarchical Recurrent Encoder-Decoder (VHRED) to model complex dependencies between sequences. Based on HRED, VHRED combined a latent variable into the decoder and turned the decoding process into a two-step generation process: sampling a latent variable at the first step and then generating the response conditionally. VHRED was trained with a variational lower bound on the log-likelihood and exhibited promising improvement in diversity, length, and quality of generated responses.
Many recent works in dialogue-related tasks apply HRED-based frameworks to capture hierarchical dialogue features. Zhang et al. (2019a) argued that standard HRED processed all contexts in dialogue history indiscriminately. Inspired by the architecture of Transformer (Vaswani et al., 2017), they proposed ReCoSa, a self-attention-based hierarchical model. It first applied LSTM to encode token-level information into context hidden vectors and then calculated the self-attention for both the context vectors and masked response vectors. At the decoding stage, the encoder-decoder attention was calculated to facilitate the decoding. Shen et al. (2019) proposed a hierarchical model consisting of 3 hierarchies: the discourse-level which captures the global knowledge, the pair-level which captured the topic information in utterance pairs, and the utterance level which captured the content information. Such a multi-hierarchy structure contributed to its higher quality responses in terms of diversity, coherence, and fluency. Chauhan et al. (2019) applied HRED and VGG-19 as a multimodal HRED (MHRED). The HRED encoded hierarchical dialogue context while VGG-19 extracted visual features for all images in the corresponding turn. With the addition of a position-aware attention mechanism, the model showed more diverse and accurate responses in a visually grounded setting. Mehri et al. (2019) learned dialogue context representations via four sub-tasks, three of which (next-utterance generation, masked-utterance retrieval, and inconsistency identification) made uses of HRED as the context encoder, and good performance was achieved. Cao et al. (2019) used HRED to encode the dialogue history between therapists and patients to categorize therapist and client MI behavioral codes and predict future codes. Qiu et al. (2020) applied an LSTM-based VHRED to address the two-agent and multi-agent dialogue structure induction problem in an unsupervised fashion. On top of that, they applied a Conditional Random Field model in two-agent dialogues and a non-projective dependency tree in multi-agent dialogues, both of them achieving better performance in dialogue structure modeling.
4 Memory Networks
Memory is a crucial component when addressing problems regarding past experiences or outside knowledge sources. The hippocampus of human brains and the hard disk of computers are the components that humans and computers depend on for reading and writing memories. Traditional models rarely have a memory component, thus lacking the ability of knowledge reusing and reasoning. RNNs iteratively pass history information across time steps, which, to some extent, can be viewed as a memory model. However, even for LSTM, which is a powerful variant of RNN equipped with a long-term and short-term memory, the memory module is too small and facts are not explicitly discriminated, thus not being able to compress specific knowledge facts and reuse them in tasks.
Weston et al. (2014) proposed memory networks, a model that is endowed with a memory component. As described in their work, a memory network has five modules: a memory module which stores the representations of memory facts; an ‘I’ module which maps the input memory facts into embedded representations; a ‘G’ module which decides the update of the memory module; an ‘O’ module which generates the output conditioned on the input representation and memory representation; an ‘R’ module which organizes the final response based on the output of ‘O’ module. This model needs a strong supervision signal for each module and thus is not practical to train in an end-to-end fashion.
Sukhbaatar et al. (2015) extended their prior work to an end-to-end memory network, which was commonly accepted as a standard memory network being easy to train and apply.
Figure 5 represents the proposed end-to-end memory networks. Its architecture consists of three stages: weight calculation, memory selection, and final prediction.
Weight calculation. The model first converts the input memory set into memory representations using a representation model . Then it maps the input query into its embedding space using another representation model , obtaining an embedding vector . The final weights are calculated as follows:
Where is the weight corresponding to each input memory conditioned on the query.
Memory selection. Before generating the final prediction, a selected memory vector is generated by first encoding the input memory into an embedded vector using another representation model , then calculating the weighted sum over the using the weights calculated in the previous stage:
Where o represents the selected memory vector. This vector cannot be found in memory representations. The soft memory selection facilitates differentiability in gradient computing, which makes the whole model end-to-end trainable.
Final prediction. The final prediction is obtained by mapping the sum vector of the selected memory and the embedded query into a probability vector :
Many dialogue-related works incorporate memory networks into their framework, especially for tasks involving an external knowledge base like task-oriented dialogue systems, knowledge-grounded dialogue systems, and QA.
Chen et al. (2019c) argued that state-of-the-art task-oriented dialogue systems tended to combine dialogue history and knowledge base entries in a single memory module, which influenced the response quality. They proposed a task-oriented system that consists of three memory modules: two long-term memory modules storing the dialogue history and the knowledge base respectively; a working memory module that memorizes two distributions and controls the final word prediction. He et al. (2020a) trained a task-oriented dialogue system with a “Two-teacher-one-student" framework to improve the knowledge retrieval and response quality of their memory networks. They first trained two teacher networks using reinforcement learning with complementary goal-specific reward functions respectively. Then with a GAN framework, they trained two discriminators to teach the student memory network to generate responses similar to those of the teachers, transferring the expert knowledge from the two teachers to the student. The advantage is that this training framework needs only weak supervision and the student network can benefit from the complementary targets of teacher networks. Kim et al. (2019) solved the dialogue state tracking in task-oriented dialogue systems with a memory network that memorized the dialogue states. Different from other works, they did not update all dialogue states in the memory module from scratch. Instead, their model first predicted which states needed to be updated and then overwrote the target states. By selectively overwriting the memory module, they improved the efficiency of the dialogue state tracking task. Dai et al. (2020) applied the MemN2N (Sukhbaatar et al., 2015) as task-oriented utterance encoder, memorizing the existing responses and dialogue history. Then they used model-agnostic meta-learning (MAML) (Finn et al., 2017) to train the framework to retrieve correct responses in a few-shot fashion.
Tian et al. (2019) proposed a knowledge-grounded chit-chat system. A memory network was used to store query-response pairs and at the response generation stage, the generator produced the response conditioned on both the input query and memory pairs. It extracted key-value information from the query-response pairs in memory and combined them into token prediction. Xu et al. (2019) proposed to use meta-words to generate responses in open-domain systems in a controllable way. Meta-words are phrases describing response attributes. Using a goal-tracking memory network, they memorized the meta-words and generated responses based on the user message while incorporating meta-words at the same time. Gan et al. (2019) performed multi-step reasoning conditioned on a dialogue history memory module and a visual memory module. Two memory modules recurrently refined the representation to perform the next reasoning process. Experimental results illustrated the benefits of combining image and dialogue clues to improve the performance of visual dialogue systems. Han et al. (2019) trained a reinforcement learning agent to decide which memory vector can be replaced when the memory module is full to improve the accuracy and efficiency of the document-grounded question-answering task. They solved the scalability problem of memory networks by learning the query-specific value corresponding to each memory. Gao et al. (2020c) solved the same problem in a conversational machine reading task. They proposed an Explicit Memory Tracker (EMT) to decide whether the provided information in memory is enough for final prediction. Furthermore, a coarse-to-fine strategy was applied for the agent to make clarification questions to request additional information and refine the reasoning.
5 Attention and Transformer
As introduced in Section 2.2, traditional sequence-to-sequence models decode the token conditioning on the current hidden state and output vector of last time step, which is formulated as:
Where g is a sequential model which maps the input vectors into a probability vector.
However, such a decoding scheme is limited when the input sentence is long. RNNs are not able to encode all information into a fixed-length hidden vector. Cho et al. (2014a) proved via experiments that a sequence-to-sequence model performed worse when the input sequence got longer. Also, for the limited-expression ability of a fixed-length hidden vector, the performance of the decoding scheme in Equation (24) largely depends on the first few steps of decoding, and if the decoder fails to have a good start, the whole sequence would be negatively affected.
Bahdanau et al. (2014) proposed the attention mechanism in the machine translation task. They described the method as “jointly align and translate", which illustrated the sequence-to-sequence translation model as an encoder-decoder model with attention. At the decoding stage, each decoding state would consider which parts of the encoded source sentence are correlated, instead of depending only on the immediate prior output token. The output probability distribution can be described as:
Where denotes the time step; is the output token, is the decoder hidden state and is the weighted source sentence:
Where is the normalized weight score:
is the similarity score between and encoder hidden state , where the score is predicted by the similarity model :
Figure 6 illustrates the attention model, where t and T denote time steps of decoder and encoder respectively.
Memory networks are similar to attention networks in the way they operate, except for the choice of the similarity model. In memory networks, the encoded memory can be viewed as the encoded source sentence in attention. However, the memory model proposed by Sukhbaatar et al. (2015) chose cosine distance as the similarity model while the attention proposed by Bahdanau et al. (2014) used a feed-forward network which is trainable together with the whole sequence-to-sequence model.
5.2 Transformer
Before transformers, most works combined attention with recurrent units, except for few works such as Parikh et al. (2016) and Gehring et al. (2017). Recurrent models condition each hidden state on the previous hidden state and the current input and are flexible in sequence length. However, due to their sequential nature, recurrent models cannot be trained in parallel, which severely undermines their potential. Vaswani et al. (2017) proposed Transformer, which entirely utilized attention mechanisms without any recurrent units and deployed more parallelization to speed up training. It applied self-attention and encoder-decoder attention to achieve local and global dependencies respectively.
Figure 7 represents the transformer. The following details its key mechanisms.
The Transformer consists of an encoder and a decoder. The encoder maps the input sequence into continuous hidden states . The decoder further generates the output sequence based on the hidden states of the encoder. The probability model of the Transformer is in the same form as that of the vanilla sequence-to-sequence model introduced in Section 2.2.5. Vaswani et al. (2017) stacked 6 identical encoder layers and 6 identical decoder layers. An encoder layer consists of a multi-head attention component and a simple feed-forward network, both of which apply residual structure. The structure of a decoder layer is almost the same as that of an encoder layer, except for an additional encoder-decoder attention layer, which computes the attention between decoder hidden states of the current time step and the encoder output vectors. The input of the decoder is partially masked to make sure that each prediction is based on the previous tokens, avoiding predicting with the presence of future information. Both inputs of encoder and decoder use a positional encoding mechanism.
For an input sentence , each token corresponds to three vectors: query, key, and value. The self-attention computes the attention weight for every token against all other tokens in by multiplying the query of with the keys of all the remaining tokens one-by-one. For parallel computing, the query, key ,and value vectors of all tokens are combined into three matrices: Query (Q), Key (K) ,and Value (V). The self-attention of an input sentence is computed by the following formula:
Where is the dimension of queries or keys.
To jointly consider the information from different subspaces of embedding, query, key, and value vectors are mapped into vectors of identical shapes by using different linear transformations, where denotes the number of heads. Attention is computed on each of these vectors in parallel, and the results are concatenated and further projected. The multi-head attention can be described as:
Where and denotes the linear transformations.
The proposed transformer architecture has no recurrent units, which means that the order information of sequence is dismissed. The positional encoding is added with input embeddings to provide positional information. The paper chooses cosine functions for positional encoding:
Where denotes the position of the target token and denotes the dimension, which means that each dimension of the positional matrix uses a different wavelength for encoding.
Recently, many transformer-based pretrain models have been developed. Unlike Embeddings from Language Model (ELMo) proposed by Peters et al. (2018), which is an LSTM-based contextual embedding model, transformer-based pretrain models are more powerful. Two most popular models are GPT-2 https://openai.com/blog/better-language-models/ and BERT (Devlin et al., 2018). GPT-2 and BERT both consist of 12 transformer blocks and BERT is further improved by making the training bi-directional. They are powerful due to their capability of adapting to new tasks after pretraining. This property helped achieve significant improvements in many NLP tasks. There also evolve many Transformer variants (Zaheer et al., 2020; Dai et al., 2019; Guo et al., 2019), which are designed to reduce the model parameters/computational complexity, or improve performance of the original Transformer in diverse scenarios. Lin et al. (2021) and Tay et al. (2020) systematically summarize the state-of-the-art Transformer variants for academics that are interested.
Attention is a mechanism to catch the importance of different parts in the target sequence. Zhu et al. (2018) applied a two-level attention to generate words. Given the user message and candidate responses selected by a retrieval system, the generator first computes word-level attention weights, then uses sentence-level attention to rescale the weights. This two-level attention helps the generator catch different importance given the encoded context. Liu et al. (2019) used an attention-based recurrent architecture to generate responses. They designed a multi-level encoder-decoder of which the multi-level encoder tries to map raw words, low-level clusters, and high-level clusters into hierarchical embedded representations while the multi-level decoder leveraged the hierarchical representations using attention and then generated responses. At each decoding stage, the model calculated two attention weights for the output of the higher-level decoder and the hidden state of the current level’s encoder. Chen et al. (2019b) computed multi-head self-attention for the outputs of a dialogue act predictor. Unlike the transformer, which concatenates the outputs of different heads, they passed the outputs directly to the next multi-head layer. The stacked multi-head layers then generated the responses with dialogue acts as the input.
Transformers are powerful sequence-to-sequence models and meanwhile, their encoders also serve as good dialogue representation models. Henderson et al. (2019b) built a transformer-based response retrieval model for task-oriented dialogue systems. A two-channel transformer encoder was designed for encoding user messages and responses, both of which were initially presented as unigrams and bigrams. A simple cosine distance was then applied to calculate the semantic similarity between the user message and the candidate response. Li et al. (2019d) built multiple incremental transformer encoders to encode multi-turn conversations and their related document knowledge. The encoded utterance and related document of the previous turn were treated as a part of the input of the next turn’s transformer encoder. The pretrained model was adaptable to multiple domains with only a small amount of data from the target domain. Bao et al. (2019b) used stacked transformers for dialogue generation pretraining. Besides the response generation task, they also pretrained the model together with a latent act prediction task. A latent variable was applied to solve the “one-to-many" problem in response generation. The multi-task training scheme improved the performance of the proposed transformer pretraining model.
Large transformer-based pretrain models are adaptable to many tasks and are thus popular in recent works. Golovanov et al. (2019) used GPT as a sequence-to-sequence model to directly generate utterances and compared the performances under single- and multi-input settings. Majumder et al. (2020b) first used a probability model to retrieve related news corpus and then combined the news corpus and dialogue context as input of a GPT-2 generator for response generation. They proposed that by using discourse pattern recognition and interrogative type prediction as two subtasks for multi-task learning, the dialogue modeling could be further improved. Wu et al. (2019c) used BERT as an encoder of context and candidate responses in their goal-based response retrieval system while Zhong et al. (2020) built Co-BERT, a BERT-based response selection model, to retrieve empathetic responses given persona-based training corpus. Zhao et al. (2020b) built a knowledge-grounded dialogue system in a synthesized fashion. They used both BERT and GPT-2 to perform knowledge selection and response generation jointly, where BERT was for knowledge selection and GPT-2 generated responses based on dialogue context and the selected knowledge.
6 Pointer Net and CopyNet
In some NLP tasks like dialogue systems and question-answering, the agents sometimes need to directly quote from the user message. Pointer Net (Oriol et al., 2015) (Figure 8) solved the problem of directly copying tokens from the input sentence.
Traditional sequence-to-sequence models (Sutskever et al., 2014; Graves et al., 2014) with an encoder-decoder structure map a source sentence to a target sentence. Generally, these models first map source sentence into hidden state vectors with an encoder, and then predict the output sequence based on the hidden states. The sequence prediction is accomplished step-by-step, each step predicting one token using greedy search or beam search. The overall sequence-to-sequence model can be described by the following probability model:
Where constitutes a training pair, = denotes the input sequence and = denotes the ground target sequence. is a decoder model.
The sequence-to-sequence models have the vanilla backbones and attention-based backbones. Vanilla models predict the target sequence based only on the last hidden state of the encoder and pass it across different decoder time steps. Such a mechanism restricts the information received by the decoder at each decoding stage. Attention-based models consider all hidden states of the encoder at each decoding step and calculate their importance when utilizing them. To compare the mechanism of Pointer Net and Attention, we present the equations explained in Section 2.2 here again. The decoder predicts the token conditioned partially on the weighted sum of encoder hidden states :
Where is the normalized weight score:
is the similarity score between and encoder hidden state , where the score is predicted by the similarity model :
At each decoding step, both vanilla and attention-based sequence-to-sequence models predict a distribution over a fixed dictionary , where denotes the tokens and denotes the total count of different tokens in the training corpus. However, when copying words from the input sentence, we do not need such a large dictionary. Instead, equals to the number of tokens in the input sequence (including repeated ones) and is not fixed since it changes according to the length of the input sequence. Pointer Net made a simple change to the attention-based sequence-to-sequence models: instead of predicting the token distribution based on the weighted sum of encoder hidden states , it directly used the normalized weights as predicted distribution:
Where is a set of probability numbers which represents the probability distribution over the tokens of the input sequence. Obviously, the token prediction problem is now transformed into position prediction problem, where the model only needs to predict a position in the input sequence. This mechanism is like a pointer that points to its target, hence the name “Pointer Net".
6.2 CopyNet
In real-world applications, simply copying from the source message is not enough. Instead, in tasks like dialogue systems and QA, agents also require the ability to generate words that are not in the source sentence. CopyNet (Gu et al., 2016) (Figure 9) was proposed to incorporate the copy mechanism into traditional sequence-to-sequence models. The model decides at each decoding stage whether to copy from the source or generate a new token not in the source.
The encoder of CopyNet is the same as that of a traditional sequence-to-sequence model, whereas the decoder has some differences compared with a traditional attention-based decoder. When predicting the token at time step , it combines the probabilistic models of generate-mode and copy-mode:
Where is the time step. is the decoder hidden state and is the predicted token. and represent weighted sum of encoder hidden states and encoder hidden states respectively. and are generate-mode and copy-mode respectively.
Besides, though it still uses and weighted attention vector to update the decoder hidden state, is uniquely encoded with both its embedding and its location-specific hidden state; also, CopyNet combines attentive read and selective read to capture information from the encoder hidden states, where the selective read is the same method used in Pointer Net. Different from the Neural Turing Machines (Graves et al., 2014; Kurach et al., 2015), the CopyNet has a location-based mechanism that enables the model to be aware of some specific details in training data in a more subtle way.
Copy mechanism is suitable for dialogues involving terminologies or external knowledge sources, and it is popular in knowledge-grounded or task-oriented dialogue systems.
For knowledge-grounded systems, external documents or dialogues are sources to copy from. Lin et al. (2020a) combined a recurrent knowledge interactive decoder with a knowledge-aware pointer network to achieve both knowledge-grounded generation and knowledge copy. In the proposed model, they first calculated the attention distribution over external knowledge, then used two pointers referring to dialogue context and knowledge source respectively to copy out-of-vocabulary (OOV) words. Wu et al. (2020b) applied a multi-class classifier to flexibly fuse three distributions: generated words, generated knowledge entities, and copied query words. They used Context-Knowledge Fusion and Flexible Mode Fusion to perform the knowledge retrieval, response generation, and copying jointly, making the generated responses precise, coherent, and knowledge-infused. Ji et al. (2020) proposed a Cross Copy Network to copy from internal utterance (dialogue history) and external utterance (similar cases) respectively. They first used pretrained language models for similar case retrieval, then combined the probability distribution of two pointers to make a prediction. They only experimented with court debate and customer service content generation tasks, where similar cases were easy to obtain.
Many dialogue state tracking tasks generate slots and slot values using a copy component (Wu et al., 2019a; Ouyang et al., 2020; Gangadharaiah and Narayanaswamy, 2020; Chen et al., 2020a; Zhang et al., 2020; Li et al., 2020d). Among them Wu et al. (2019a), Ouyang et al. (2020) and Chen et al. (2020a) solved the problem of multi-domain dialogue state tracking. Wu et al. (2019a) proposed TRAnsferable Dialogue statE generator (TRADE), a copy-based dialogue state generator. The generator decoded the slot value multiple times for each possible (domain, slot) pair, then a slot gate was applied to decide which pair belonged to the dialogue. The output distribution was a copy of the slot values belonging to the selected (domain, slot) pairs from vocabulary and dialogue history. Chen et al. (2020a) used a different copy strategy from TRADE. Instead of using the whole dialogue history as the copy source, they copied state values from user utterances and system messages respectively, which took the slot-level context as input. Ouyang et al. (2020) proposed slot connection mechanism to efficiently utilize existing states from other domains. Attention weights were calculated to measure the connection between the target slot and related slot-value tuples in other domains. Three distributions over token generation, dialogue context copying, and past state copying were finally gated and fused to predict the next token. Gangadharaiah and Narayanaswamy (2020) combined a pointer network with a template-based tree decoder to fill the templates recursively and hierarchically. Copy mechanisms also alleviated the problem of expensive data annotation in end-to-end task-oriented dialogue systems. Copy-augmented dialogue generation models were proven to perform significantly better than strong baselines with limited domain-specific or multi-domain data (Zhang et al., 2020; Li et al., 2020d; Gao et al., 2020a).
Pointer networks and CopyNet are also used to solve other dialogue-related tasks. Yu and Joty (2020) applied a pointer net for online conversation disentanglement. The pointer module pointed to the ancestor message to which the current message replies and a classifier predicted whether two messages belonged to the same thread. In dialogue parsing tasks, the pointer net is used as the backbone parsing model to construct discourse trees (Aghajanyan et al., 2020; Lin et al., 2019). Tay et al. (2019) used a pointer-generator framework to perform machine reading comprehension over a long span, where the copy mechanism reduced the demand of including target answers in context.
7 Deep Reinforcement Learning Models and Generative Adversarial Networks
In recent years, two exciting approaches exhibit the potential of artificial intelligence. The first one is deep reinforcement learning, which outperforms humans in many complex problems such as large-scale games, conversations, and car-driving. Another technique is GAN, showing amazing capability in generation tasks. The data samples generated by GAN models like articles, paintings, and even videos, are sometimes indistinguishable from human creations.
AlphaGo (Silver et al., 2016) stimulated the research interests again in reinforcement learning in recent years (Graves et al., 2016; Mnih et al., 2016; Wang et al., 2016; Tamar et al., 2016; Jaderberg et al., 2016; Mirowski et al., 2016). Reinforcement learning is a branch of machine learning aiming to train agents to perform appropriate actions while interacting with a certain environment. It is one of the three fundamental machine learning branches, with supervised learning and unsupervised learning being the other two. It can also be seen as an intermediate between supervised learning and unsupervised learning because it only needs weak signals for training (Wang et al., 2016).
Figure 10 illustrates the reinforcement learning framework, consisting of an agent and an environment. The framework is a Markov Decision Process (MDP) (Puterman, 2014), which can be described by a five-tuple M = . denotes an infinite set of environment states; denotes a set of actions that agent chooses from conditioned on a given environment state ; is the transition probability matrix in MDP, denoting the probability of an environment state transfer after agent takes an action; is an average reward the agent receives from the environment after taking an action under state ; is a discount factor. The flow of this framework is a loop of the following two steps: the agent first makes an observation on the current environment state and chooses an action based on its policy; then according to the transition probability matrix , the environment’s state transfers to , and simultaneously provides a reward .
Reinforcement learning is applicable to solve many challenges in dialogue systems because of the agent-environment nature of a dialogue system. A two-party dialogue system consists of an agent, which is an intelligent chatbot, and an environment, which is usually a user or a user simulator. Here we mainly discuss deep reinforcement learning.
Deep reinforcement learning means applying deep neural networks to model the value function or policy of the reinforcement learning framework. “Deep model" is in contrast to the “shallow model". The shallow model normally refers to traditional machine learning models like Decision Trees or KNN. Feature engineering, which is usually based on shallow models, is time and labor consuming, and also over-specified and incomplete. Different from that, deep neural models are easy to design and have a strong fitting capability, which contributes to many breakthroughs in recent research. Deep representation learning gets rid of human labor and exploits hierarchical features in data automatically, which strengthens the semantic expressiveness and domain correlations significantly.
We discuss two typical reinforcement models: Deep Q-Networks (Mnih et al., 2015) and REINFORCE (Williams, 1992; Sutton et al., 1999). They belong to Q-learning and policy gradient respectively, which are two families of reinforcement learning.
A Deep Q-Network is a value-based RL model. It determines the best policy according to the Q-function:
Where is an optimal Q-function and is the corresponding optimal policy. In Deep Q-Networks, the Q function is modeled using a deep neural network, such as CNNs, RNNs, etc.
As in Gao et al. (2018), the parameters of the Q model are updated using the rule:
Where the is an observed trajectory. denotes step-size and the parameter update is calculated using temporal difference (Sutton, 1988). However, this update mechanism suffers from unstableness and demands a large number of training samples. There are two typical tricks for a more efficient and stable parameter update.
The first method is experience replay (Lin, 1992; Mnih et al., 2015). Instead of using one training sample at a time to update the parameters, it uses a buffer to store training samples, and iteratively retrieves training samples from the buffer pool to perform parameter updates. It avoids encountering training samples that change too fast in distribution during training time, which increases the learning stability; further, it uses each training sample multiple times, which improves the efficiency.
The second is two-network implementation (Mnih et al., 2015). This method uses two networks in Q-function optimization, one being the Q-network, another being a target network. The target network is used to calculate the temporal difference, and its parameters are frozen while training, aligning with periodically. The parameters are then updated with the following rule:
Since does not change in a period of time, the target network calculates the temporal difference in a stable manner, which facilitates the convergence of training.
7.2 REINFORCE
REINFORCE is a policy-based RL algorithm that has no value network. It optimizes the policy directly. The policy is parameterized by a policy network, whose output is a distribution over continuous or discrete actions. A long-term reward is computed for evaluation of the policy network by collecting trajectory samples of length :
denotes a long-term reward and the goal is to optimize the policy network in order to maximize . Here stochastic gradient ascentStochastic gradient ascent simply uses the negated objective function of stochastic gradient descent. is used as an optimizer:
Where is computed by:
Both models have their advantages: Deep Q-Networks are more sample efficient while REINFORCE is more stable (Li, 2017). REINFORCE is more popular in recent works. Modern research involves larger action spaces, which means that value-based RL models like Deep Q-Networks are not suitable for problem-solving. Value-based methods “select an action to maximize the value", which means that their action sets should be discrete and moderate in scale; while policy gradient methods such as REINFORCE are different, they predict the action via policy networks directly, which sets no restriction on the action space. As a result, policy gradient methods are more suitable for tasks involving a larger action space.
Considering the respective benefits brought by the Q-learning and policy gradient, some work has been done combining the value- and policy-based methods. Actor-critic algorithm (Konda and Tsitsiklis, 2000; Sutton et al., 1999) was proposed to alleviate the severe variance problem when calculating the gradient in policy gradient methods. It estimates a value function for term in Equation (45) and incorporates it in policy optimization. Equation (45) is then transformed into the formula below:
Where stands for the value function estimated.
7.3 GANs
It is easy to link the actor-critic model with another framework - GANs (Goodfellow et al., 2014; Zhang et al., 2018c; Feng et al., 2020a) because of their similar inner structure and logic (Pfau and Vinyals, 2016). Actually, there are quite a few recent works in dialogue systems that train GANs with reinforcement learning framework (Zhu et al., 2018; Wu et al., 2019b; He et al., 2020a; Zhu et al., 2020; Qin et al., 2020).
Figure 11 represents the GAN consisting of a generator and a discriminator where the training process can be viewed as a competition between them: the generator tries to generate data distributions to fool the discriminator while the discriminator attempts to distinguish between real data (real) and generated data (fake). During training, the generator takes noise as input and generates data distribution while the discriminator takes real and fake data as input and the binary annotation as the label. The whole GAN model is trained end-to-end as a connection of generator and discriminator to minimize the following cross-entropy losses:
Where and denote a bilevel loss, where and being discriminator and generator respectively. is the noise input of the generator and is the input of the discriminator.
GAN can be viewed as a special actor-critic (Pfau and Vinyals, 2016). In the learning architecture of GAN, the generator acts as the actor and the discriminator acts as the critic or environment which gives the real/fake feedback as a reward. However, the actions taken by the actor cannot change the states of the environment, which means that the learning architecture of GAN is a stateless Markov decision process. Also, the actor has no access to the state of the environment and generates data distribution simply conditioned on Gaussian noise, which means that the generator in the GAN framework is a blind actor/agent. In a nutshell, GAN is a special actor-critic where the actor is blind and the whole process is a stateless MDP.
The interactive nature of dialogue systems motivates the wide application of reinforcement learning and GAN models in its research.
One common application of reinforcement learning in dialogue systems is the reinforced dialogue management in task-oriented systems. Dialogue state tracking and policy learning are two typical modules of a dialogue manager. Huang et al. (2020c) and Li et al. (2020d) trained the dialogue state tracker with reinforcement learning. Both of them combined a reward manager into their tracker to enhance tracking accuracy. For the policy learning module, reinforcement learning seems to be the best choice since almost all recent related works learned policy with reinforcement learning (Zhang et al., 2019c; Wang et al., 2020d; Zhu et al., 2020; Wang et al., 2020a; Takanobu et al., 2020; Huang et al., 2020b; Xu et al., 2020a). The increasing preference of reinforcement learning in policy learning tasks attributes to the characteristic of them: in policy learning tasks, the model predicts a dialogue action (action) based on the states from the DST module (state), which perfectly accords with the function of the agent in the reinforcement learning framework.
Due to the huge action space needed to generate language directly, many open-domain dialogue systems trained with reinforcement learning framework do not generate responses but instead select responses. Retrieval-based systems have a limited action set and are suitable to be trained in a reinforcement learning scheme. Some works achieved promising performance in retrieval-based dialogue tasks (Bouchacourt and Baroni, 2019; Li et al., 2016a; Zhao and Eskenazi, 2016). However, retrieval systems fail to generalize in all user messages and may give unrelated responses (Qiu et al., 2017), which makes generation-based dialogue systems preferable. Still considering the action space problem, some works build their systems combining retrieval and generative methods (Zhu et al., 2018; Serban et al., 2017b). Zhu et al. (2018) chose to first retrieve a set of n-best response candidates and then generated responses based on the retrieved results and user message. Comparatively, Serban et al. (2017b) first generated and retrieved candidate responses with different dialogue models and then trained a scoring model with online reinforcement learning to select responses from both generated and retrieved responses. Since training a generative dialogue agent using reinforcement learning from scratch is particularly difficult, first pretraining the agent with supervised learning to warm-start is a good choice. Wu et al. (2019b), He et al. (2020a), Williams and Zweig (2016) and Yao et al. (2016) applied this pretrain-and-finetune strategy on dialogue learning and achieved outstanding performance, which proved that the reinforcement learning can improve the response quality of data-driven chatbots. Similarly, pretrain-and-finetune was also applicable to domain transfer problems. Some works pretrained the model in a source domain and expanded the domain area with reinforcement training (Mo et al., 2018; Li et al., 2016d).
Some systems use reinforcement learning to select from outside information like persona, document, knowledge graph, etc., and generate responses accordingly. Majumder et al. (2020a) and Jaques et al. (2020) performed persona selection and persona-based response generation simultaneously and trained their agents with a reinforcement framework. Bao et al. (2019a) and Zhao et al. (2020b) built document-grounded systems. Similarly, they used reinforcement learning to accomplish document selection and knowledge-grounded response generation. There were also some works combining knowledge graphs into the dialogue systems and treated them as outside knowledge source (Moon et al., 2019; Xu et al., 2020a). In a reinforced training framework, the agent chooses an edge based on the current node and state for each step and then combines the knowledge into the response generation process.
Dialogue-related tasks like dialogue relation extraction (Li et al., 2019c), question answering (Hua et al., 2020) and machine reading comprehension (Guo et al., 2020) benefit from reinforcement learning as well because of their interactive nature and the scarcity of annotated data.
The application of GAN in dialogue systems is divided into two streams. The first sees the GAN framework applied to enhance response generation (Li et al., 2017a; Zhu et al., 2018; Wu et al., 2019b; He et al., 2020a; Zhu et al., 2020; Qin et al., 2020). The discriminator distinguishes generated responses from human responses, which incentivizes the agent, which is also the generator in GAN, to generate higher-quality responses. Another stream uses GAN as an evaluation tool of dialogue systems (Kannan and Vinyals, 2017; Bruni and Fernandez, 2017). After training the generator and discriminator as a whole framework, the discriminator is used separately as a scorer to evaluate the performance of a dialogue agent and was shown to achieve a higher correlation with human evaluation compared with traditional reference-based metrics like BLEU, METEOR, ROUGE-L, etc. We discuss the evaluation of dialogue systems as a challenge in Section 5.
8 Knowledge Graph Augmented Neural Networks
Supervised training with annotated data tries to learn the knowledge distribution of a dataset. However, a dataset is comparatively sparse and thus learning a reliable knowledge distribution needs a huge amount of annotated data (Annervaz et al., 2018).
Knowledge Graph (KG) is attracting more and more research interests in recent years. KG is a structured knowledge source consisting of entities and their relationships (Ji et al., 2022). In other words, KG is the knowledge facts presented in graph format.
Figure 12 shows an example of a KG consisting of entities and their relationships. A KG is stored in triples under the Resource Description Framework (RDF). For example, Albert Einstein, University of Zurich, and their relationship can be expressed as .
Knowledge graph augmented neural networks first represent the entities and their relations in a lower dimension space, then use a neural model to retrieve relevant facts (Ji et al., 2022). Knowledge graph representation learning can be generally divided into two categories: structure-based representations and semantically-enriched representations. Structure-based representations use multi-dimensional vectors to represent entities and relations. Models such as TransE (Bordes et al., 2013), TransR (Lin et al., 2015), TransH (Wang et al., 2014), TransD (Ji et al., 2015), TransG (Xiao et al., 2015), TransM (Fan et al., 2014), HolE (Nickel et al., 2016) and ProjE (Shi and Weninger, 2017) belong to this category. The semantically-enriched representation models like NTN (Socher et al., 2013), SSP (Xiao et al., 2017) and DKRL (Xie et al., 2016) combine semantic information into the representation of entities and relations. The neural retrieval models also have two main directions: distance-based matching model and semantic matching model. Distance-based matching models (Bordes et al., 2013) consider the distance between projected entities while semantic matching models (Bordes et al., 2014) calculate the semantic similarity of entities and relations to retrieve facts.
Knowledge-grounded dialogue systems benefit greatly from the structured knowledge format of KG, where facts are widely intercorrelated. Reasoning over a KG is an ideal approach for combining commonsense knowledge into response generation, resulting in accurate and informative responses (Young et al., 2018). Jung et al. (2020) proposed AttnIO, a bi-directional graph exploration model for knowledge retrieval in knowledge-grounded dialogue systems. Attention weights were calculated at each traversing step, and thus the model could choose a broader range of knowledge paths instead of choosing only one node at a time. In such a scheme, the model could predict adequate paths even when only having the destination node as the label. Zhang et al. (2019b) built ConceptFlow, a dialogue agent that guided to more meaningful future conversations. It traversed in a commonsense knowledge graph to explore concept-level conversation flows. Finally, it used a gate to decide to generate among vocabulary words, central concept words, and outer concept words. Majumder et al. (2020a) proposed to generate persona-based responses by first using COMET (Bosselut et al., 2019) to expand a persona sentence in context along 9 relation types and then applied a pretrained model to generate responses based on dialogue history and the persona variable. Yang et al. (2020) used knowledge graph as an external knowledge source in task-oriented dialogue systems to incorporate domain-specified knowledge in the response. First, the dialogue history was parsed as a dependency tree and encoded into a fixed-length vector. Then they applied multi-hop reasoning over the graph using the attention mechanism. The decoder finally predicted tokens either by copying from graph entities or generating vocabulary words. Moon et al. (2019) proposed DialKG Walker for the conversational reasoning task. They computed a zero-shot relevance score between predicted KG embedding and ground KG embedding to facilitate cross-domain predictions. Furthermore, they applied an attention-based graph walker to generate graph paths based on the relevance scores. Huang et al. (2020a) evaluated the dialogue systems by combining the utterance-level contextualized representation and topic-level graph representation. They first constructed the dialogue graph based on encoded (context, response) pairs and then reasoned over the graph to get a topic-level graph representation. The final score was calculated by passing the concatenated vector of contextualized representation and graph representation to a feed-forward network.
Task-oriented Dialogue Systems
This section introduces task-oriented dialogue systems including modular and end-to-end systems. Task-oriented systems solve specific problems in a certain domain such as movie ticket booking, restaurant table reserving, etc. We focus on deep learning-based systems due to the outstanding performance. For readers who want to learn more about traditional rule-based and statistical models, there are several surveys to refer to (Theune, 2003; Lemon and Pietquin, 2007; Mallios and Bourbakis, 2016; Chen et al., 2017a; Santhanam and Shaikh, 2019).
This section is organized as follows. We first discuss modular and end-to-end systems respectively by introducing the principles and reviewing recent works. After that, we comprehensively discuss related challenges and hot topics for task-oriented dialogue systems in recent research to provide some important research directions.
A task-oriented dialogue system requires stricter response constraints because it aims to accurately handle the user message. Therefore, modular methods were proposed to generate responses in a more controllable way. The architecture of a modular-based system is depicted in Figure 13. It consists of four modules:
Natural Language Understanding (NLU). This module converts the raw user message into semantic slots, together with classifications of domain and user intention. However, some recent modular systems omit this module and use the raw user message as the input of the next module, as shown in Figure 13. Such a design aims to reduce the propagation of errors between modules and alleviate the impact of the original error (Kim et al., 2018).
Dialogue State Tracking (DST). This module iteratively calibrates the dialogue states based on the current input and dialogue history. The dialogue state includes related user actions and slot-value pairs.
Dialogue Policy Learning. Based on the calibrated dialogue states from the DST module, this module decides the next action of a dialogue agent.
Natural Language Generation (NLG). This module converts the selected dialogue actions into surface-level natural language, which is usually the ultimate form of response.
Among them, Dialogue State Tracking and Dialogue Policy Learning constitute the Dialogue Manager (DM), the central controller of a task-oriented dialogue system. Usually, a task-oriented system also interacts with an external Knowledge Base (KB) to retrieve essential knowledge about the target task. For example, in a movie ticket booking task, after understanding the requirement of the user message, the agent interacts with the movie knowledge base to search for movies with specific constraints such as movie name, time, cinema, etc.
It has been proven that the NLU module impacts the whole system significantly in the term of response quality (Li et al., 2017b). The NLU module converts the natural language message produced by the user into semantic slots and performs classification. Table 2 shows an example of the output format of the NLU module. The NLU module manages three tasks: domain classification, intent detection, and slot filling. Domain classification and intent detection are classification problems, which use classifiers to predict a mapping from the input language sequence to a predefined label set. In the given example, the predicted domain is “movie" and the intent is “find_movie". Slot filling is a tagging problem, which can be viewed as a sequence-to-sequence task. It maps a raw user message into a sequence of slot names. In the example, the NLU module reads the user message “Recommend a movie at Golden Village tonight." and outputs the corresponding tag sequence. It recognizes “Golden Village" as the place to go, which is tagged as “B_desti" and “I_desti" for the two words respectively. Similarly, the token “tonight" is converted into “B_time". ‘B’ represents the beginning of a chunk, and ‘I’ indicates that this tag is inside a target chunk. For those unrelated tokens, an ‘O’ is used indicating that this token is outside of any chunk of interest. This tagging method is called Inside-Outside-Beginning (IOB) tagging (Ramshaw and Marcus, 1999), which is a common method in Named-Entity Recognition (NER) tasks.
Domain classification and intent detection belong to the same category of tasks. Deep learning methods are proposed to solve the classification problems of dialogue domain and intent. Deng et al. (2012) and Tur et al. (2012) were the first who successfully improved the recognition accuracy of dialogue intent. They built deep convex networks to combine the predictions of a prior network and the current utterances as an integrated input of a current network. A deep learning framework was also used to classify the dialogue domain and intent in a semi-supervised fashion (Yann et al., 2014). To solve the difficulty of training a deep neural network for domain and intent prediction, Restricted Boltzmann Machine (RBM) and Deep Belief Networks (DBNs) were applied to initialize the parameters of deep neural networks (Sarikaya et al., 2014). To make use of the strengths of RNNs in sequence processing, some works used RNNs as utterance encoders and made predictions for intent and domain categories (Ravuri and Stolcke, 2015, 2016). Hashemi et al. (2016) used a CNN to extract hierarchical text features for intent detection and illustrated the sequence classification capabilities of CNNs. Lee and Dernoncourt (2016) proposed a model for intent classification of short utterances. Short utterances are hard for intent detection because of the lack of information in a single dialogue turn. This paper used RNN and CNN architectures to incorporate the dialogue history, thus obtaining the context information as an additional input besides the current turn’s message. The model achieved promising performances on three intent classification datasets. More recently, Wu et al. (2020a) pretrained Task-Oriented Dialogue BERT (TOD-BERT) and significantly improved the accuracy in the intent detection sub-task. The proposed model also exhibited a strong capability of few-shot learning and could effectively alleviate the data insufficiency issue in a specific domain.
The slot filling problem is also called semantic tagging, a sequence classification problem. It is more challenging for that the model needs to predict multiple objects at a time. Deep Belief Nets (DBNs) exhibit promising capabilities in the learning of deep architectures and have been applied in many tasks including semantic tagging. Sarikaya et al. (2011) used a DBN-initialized neural network to complete slot filling in the call-routing task. Deoras and Sarikaya (2013) built a DBN-based sequence tagger. In addition to the NER input features used in traditional taggers, they also combined part of speech (POS) and syntactic features as a part of the input. The recurrent architectures benefited the sequence tagging task in that they could keep track of the information along past timesteps to make the most of the sequential information. Yao et al. (2013) first argued that instead of simply predicting words, RNN Language Models (RNN-LMs) could be applied in sequence tagging. On the output side of RNN-LMs, tag labels were predicted instead of normal vocabularies. Mesnil et al. (2013) and Mesnil et al. (2014) further investigated the impact of different recurrent architectures in the slot filling task and found that all RNNs outperformed the Conditional Random Field (CRF) baseline. As a powerful recurrent model, LSTM showed promising tagging accuracy on the ATIS dataset owing to the memory control of its gate mechanism (Yao et al., 2014). Gangadharaiah and Narayanaswamy (2020) argued that the shallow output representations of traditional semantic tagging lacked the ability to represent the structured dialogue information. To improve, they treated the slot filling task as a template-based tree decoding task by iteratively generating and filling in the templates. Different from traditional sequence tagging methods, Coope et al. (2020) tackled the slot filling task by treating it as a turn-based span extraction task. They applied the conversational pretrained model ConveRT and utilized the rich semantic information embedded in the pretrained vectors to solve the problem of in-domain data insufficiency. The inputs of ConveRT are the requested slots and the utterance, while the output is a span of interest as the slot value.
Some works choose to combine domain classification, intent detection, and slot filling into a multitask learning framework to jointly optimize the shared latent space. Hakkani-Tür et al. (2016) applied a bi-directional RNN-LSTM architecture to jointly perform three tasks. Liu and Lane (2016) augmented the traditional RNN encoder-decoder model with an attention mechanism to manage intent detection and slot filling. The slot filling applied explicit alignment. Chen et al. (2016) proposed an end-to-end memory network and used a memory module to store user intent and slot values in history utterances. Attention was further applied to iteratively select relevant intent and slot values at the decoding stage. Multi-task learning of three NLU subtasks contributed to the domain scaling and facilitated the zero-shot or few-shot training when transferring to a new domain (Bapna et al., 2017; Lee and Jha, 2019). Zhang et al. (2018a) captured the hierarchical structure of dialogue semantics in NLU multi-task learning by applying a capsule-based neural network. With a dynamic routing-by-agreement strategy, the proposed architecture raised the accuracy of both intent detection and slot filling on the SNIPS-NLU and ATIS dataset.
More recently, some novel ideas appear in NLU research, which provides new possibilities for further improvements. Traditional NLU modules rely on the text converted from the audio message of the user using the Automatic Speech Recognition (ASR) module. However, Singla et al. (2020) jumped over the ASR module and directly used audio signals as the input of NLU. They found that by reducing the module numbers of a pipeline system, the predictions were more robust since fewer errors were broadcasted. Su et al. (2019b) argued that Natural Language Understanding (NLU) and Natural Language Generation (NLG) were reversed processes. Thus, their dual relationship could be exploited by training with a dual-supervised learning framework. The experiments exhibited improvement in both tasks.
2 Dialogue State Tracking
Dialogue State Tracking (DST) is the first module of a dialogue manager. It tracks the user’s goal and related details every turn based on the whole dialogue history to provide the information based on which the Policy Learning module (next module) decides the agent action to make.
The NLU and DST modules are closely related. Both NLU and DST perform slot filling for the dialogue. However, they actually play different roles. The NLU module tries to make classifications for the current user message such as the intent and domain category as well as the slot each message token belongs to. For example, given a user message “Recommend a movie at Golden Village tonight.", the NLU module will convert the raw message into “", where the slots are usually filled by tagging each word of the user message as described in Section 3.1. However, the DST module does not classify or tag the user message. Instead, it tries to find a slot value for each slot name in a pre-existing slot list based on the whole dialogue history. For example, there is a pre-existing slot list “", where the underscore behind the colon is a placeholder denoting that this place can be filled with a value. Every turn, the DST module will look up the whole dialogue history up to the current turn and decide which content can be filled in a specific slot in the slot list. If the user message “Recommend a movie at Golden Village tonight." is the only message in a dialogue, then the slot list can be filled as “", where the slots unspecified by the user up to current turn can be filled with “". To conclude, the NLU module tries to tag the user message while the DST module tries to find values from the user message to fill in a pre-existing form. Some dialogue systems took the output of the NLU module as the input of DST module (Williams et al., 2013; Henderson et al., 2014a, b), while others directly used raw user messages to track the state (Kim et al., 2019; Wang et al., 2020e; Hu et al., 2020).
Dialogue State Tracking Challenges (DSTCs), a series of popular challenges in DST, provides benchmark datasets, standard evaluation frameworks, and test-beds for research (Williams et al., 2013; Henderson et al., 2014a, b; Kim et al., 2016, 2017). The DSTCs cover many domains such as restaurants, tourism, etc.
A dialogue state contains all essential information to be conveyed in the response (Henderson, 2015). As defined in DSTC2 (Henderson et al., 2014a), the dialogue state of a given dialogue turn consists of informable slots Sinf and requestable slots Sreq. Informable slots are attributes specified by users to constrain the search of the database while requestable slots are attributes whose values are queried by the user. For example, the serial number of a movie ticket is usually a requestable slot because users seldom assign a specific serial number when booking a ticket. Specifically, the dialogue state has three components:
Goal constraint corresponding with informable slots. The constraints can be specific values mentioned by the user in the dialogue or a special value. Special values include Dontcare indicating the user’s indifference about the slot and None indicating that the user has not specified the value in the conversation yet.
Requested slots. It can be a list of slot names queried by the user seeking answers from the agent.
Search method of current turn. It consists of values indicating the interaction categories. By constraints denotes that the user tries to specify constraint information in his requirement; by alternatives denotes that the user requires an alternative entity; finished indicates that the user intends to end the conversation.
However, considering the numerous challenges such as tracking efficiency, tracking accuracy, domain adaptability, and end-to-end training, many alternative representations have been proposed recently, which will be discussed later.
Figure 14 is an example of the DST process for 4 dialogue turns in a restaurant table booking task. The first column includes the raw dialogue utterances, with denoting the system message and denoting the user message. The second column includes the N-best output lists of the NLU module and their corresponding confidence scores. The third column includes the labels of a turn, indicating the ground truth slot-value pairs. The fourth column includes the example DST outputs and their corresponding confidence scores. The fifth column indicates the correctness of the tracker output.
Earlier works use hand-craft rules or statistical methods to solve DST tasks. While widely used in industry dialogue systems, rule-based DST methods (Goddeau et al., 1996) have many restrictions such as limited generalization, high error rate, low domain adaptability, etc (Williams, 2014). Statistical methods (Lee, 2013; Lee and Eskenazi, 2013; Ren et al., 2013; Williams, 2013, 2014) also suffer from noisy conditions and ambiguity (Young et al., 2010).
Recently, many neural trackers have emerged. Neural trackers have multiple advantages over rule-based and statistical trackers. In general, they are categorized into two streams. The first stream has predefined slot names and values, and each turn the DST module tries to find the most appropriate slot-value pairs based on the dialogue history; the second stream does not have a fixed slot value list, so the DST module tries to find the values directly from the dialogue context or generate values based on the dialogue context. Obviously, the latter one is more flexible and in fact, more and more works are solving DST in the second way. We discuss the works of both categories here.
The first stream can be viewed as a multi-class or multi-hop classification task. For multi-class classification DST, the tracker predicts the correct class from multiple values but this method suffers from high complexity when the value set grows large. On the other hand, for the multi-hop classification tasks, the tracker reads only one slot-value pair at a time and performs binary prediction. Working in this fashion reduces the model complexity but raises the system reaction time since for each slot there will be multiple tracking processes. Henderson et al. (2013) was the first who used a deep learning model in the DST tasks. They integrated many feature functions (e.g., SLU score, Rank score, Affirm score, etc.) as the input of a neural network, then predict the probability of each slot-value pair. Mrkšić et al. (2015) applied an RNN as a neural tracker to gain awareness on dialogue context. Mrkšić et al. (2016) proposed a multi-hop neural tracker which took the system output and user utterances as the first two inputs (to model the dialogue context), and the candidate slot-value pairs as the third input. The tracker finally made a binary prediction on the current slot-value pair based on the dialogue history.
The second stream attracts more attention because it not only reduces the model and time complexity of DST tasks but also facilitates end-to-end training of task-oriented dialogue systems. Moreover, it is also flexible when the target domain changes. Lei et al. (2018) proposed belief span, a text span of the dialogue context corresponding to a specific slot. They built a two-stage CopyNet to copy and store slot values from the dialogue history. The slots were stored to prepare for neural response generation. The belief span facilitated the end-to-end training of dialogue systems and increased the tracking accuracy in out-of-vocabulary cases. Based on this, Lin et al. (2020c) proposed the minimal belief span and argued that it was not scalable to generate belief states from scratch when the system interacted with APIs from diverse domains. The proposed MinTL framework operated insertion (INS), deletion (DEL) and substitution (SUB) on the dialogue state of last turn based on the context and the minimal belief span. Wu et al. (2019a) proposed the TRADE model. The model also applied the copy mechanism and used a soft-gated pointer-generator to generate the slot value based on the domain-slot pair and encoded dialogue context. Quan and Xiong (2020) argued that simply concatenating the dialogue context was not preferable. Alternatively, they used [sys] and [usr] to discriminate the system and user messages. This simple long context modeling method achieved a 7.03% improvement compared with the baseline. Cheng et al. (2020) proposed Tree Encoder-Decoder (TED) architecture which utilized a hierarchical tree structure to represent the dialogue states and system acts. The TED generated tree-structured dialogue states of the current turn based on the dialogue history, dialogue action, and dialogue state of the last turn. This approach led to a 20% improvement on the state-of-the-art DST baselines which represented dialogue states and user goals in a flat space. Chen et al. (2020a) built an interactive encoder to exploit the dependencies within a turn and between turns. Furthermore, they used the attention mechanism to construct the slot-level context for user and system respectively, which were embedding vectors based on which the generator copied values from the dialogue context. Shan et al. (2020) applied BERT to perform multi-task learning and generated the dialogue state. They first encoded word-level and turn-level contexts. Then they retrieved the relevant information for each slot from the context by applying both word-level and turn-level attention. Furthermore, the slot values were predicted based on the retrieved information. Similarly, Wang et al. (2020e) used BERT for slot value prediction. They performed Slot Attention (SA) to retrieve related spans and Value Normalization (VN) to convert the spans into final values. Huang et al. (2020c) proposed Meta-Reinforced MultiDomain State Generator (MERET), which was a dialogue state generator further finetuned with policy gradient reinforcement learning.
3 Policy Learning
The Policy learning module is the other module of a dialogue manager. This module controls which action will be taken by the system based on the output dialogue states from the DST module. Assuming that we have the dialogue state of the current turn and the action set , the task of this module is to learn a mapping function : . This module is comparatively simpler than other modules in the term of task definition but actually, the task itself is challenging (Peng et al., 2017). For example, in the tasks of movie ticket and restaurant table booking, if the user books a two-hour movie slot and intends to go for dinner after that, then the agent should be aware that the time gap between movie slot and restaurant slot has to be more than two hours since the commuting time from the cinema to the restaurant should be considered.
Supervised learning and reinforcement learning are mainstream training methods for dialogue policy learning (Chen et al., 2017a). Policies learned in a supervised fashion exhibit great decision-making ability (Su et al., 2016; Dhingra et al., 2016; Williams et al., 2017; Liu and Lane, 2017). In some specific tasks, the supervised policy model can complete tasks precisely, but the training process totally depends on the quality of training data. Moreover, the annotated datasets require intensive human labor, and the decision ability is restricted by the specific task and domain, showing weak transferring capability. With the prevalence of reinforcement learning methods, more and more task-oriented dialogue systems use reinforcement learning to learn the policy. The dialogue policy learning fits the reinforcement learning setting since the agent of reinforcement learning learns a policy to map environment states to actions as well.
Usually, the environment of reinforce policy learning is a user or a simulated user in which setting the training is called online learning. However, it is data- and time-consuming to learn a policy from scratch in the online learning scenario, so the warm-start method is needed to speed up the training process. Henderson et al. (2008) used expert data to restrict the initial action space exploration. Chen et al. (2017b) applied teacher-student learning framework to transfer the teacher expert knowledge to the target network in order to warm-start the system.
Almost all recent dialogue policy learning works are based on reinforcement learning methods. Online learning is an ideal approach to get training samples iteratively for a reinforcement learning agent, but human labor is very limited. Zhang et al. (2019c) proposed Budget-Conscious Scheduling (BCS) to better utilize limited user interactions, where the user interaction is seen as the budget. The BCS used a probability scheduler to allocate the budget during training. Also, a controller decided whether to use real user interactions or simulated ones. Furthermore, a goal-based sampling model was applied to simulate the experiences for policy learning. Such a budget-controlling mechanism achieved ideal performance in the practical training process. Considering the difficulty of getting real online user interactions and the huge amount of annotated data required for training user simulators, Takanobu et al. (2020) proposed Multi-Agent Dialog Policy Learning, where they have two agents interacting with each other, performing both user and agent, learning the policy simultaneously. Furthermore, they incorporated a role-specific reward to facilitate role-based response generation. A High task completion rate was observed in experiments. Wang et al. (2020d) introduced Monte Carlo Tree Search with Double-q Dueling network (MCTS-DDU), where a decision-time planning was proposed instead of background planning. They used the Monte Carlo simulation to perform a tree search of the dialogue states. Gordon-Hall et al. (2020) trained expert demonstrators in a weakly supervised fashion to perform Deep Q-learning from Demonstrations (DQfD). Furthermore, Reinforced Fine-tune Learning was proposed to facilitate domain transfer. In reinforce dialogue policy learning, the agent usually receives feedback at the end of the dialogue, which is not efficient for learning. Huang et al. (2020b) proposed an innovative reward learning method that constrains the dialogue progress according to the expert demonstration. The expert demonstration could either be annotated or not, so the approach was not labor intensive. Wang et al. (2020b) proposed to co-generate the dialogue actions and responses to maintain the inherent semantic structures of dialogue. Similarly, Le et al. (2020b) proposed a unified framework to simultaneously perform dialogue state tracking, dialogue policy learning, and response generation. Experiments showed that unified frameworks have a better performance both in their sub-tasks and in their domain adaptability. Xu et al. (2020a) used a knowledge graph to provide prior knowledge of the action set and solved policy learning task in a graph-grounded fashion. By combining a knowledge graph, a long-term reward was obtained to provide the policy agent with a long-term vision while choosing actions. Also, the candidate actions were of higher quality due to prior knowledge. The policy learning was further performed in a more controllable way.
4 Natural Language Generation
Natural Language Generation (NLG) is the last module of a task-oriented dialogue system pipeline. It manages to convert the dialogue actions generated from the dialogue manager into a final natural language representation. E.g., Assuming “Inform (name = Wonder Woman; genre = Action; desti = Golden Village)" to be the dialogue action from policy learning module, then the NLG module converts it into language representations such as “There is an action movie named Wonder Woman at Golden Village."
Traditional NLG modules are pipeline systems. Defined by Siddharthan (2001), the standard pipeline of NLG consists of four components, as shown in Figure 15.
The core modules of this pipeline are Content Determination, Sentence Planning, and Surface Realization, as proposed by Reiter (1994). Cahill et al. (1999) further improved the NLG pipeline by adding three more components: lexicalization, referring expression generation, and aggregation. However, this model has a drawback that the input of the system is ambiguous.
Deep learning methods were further applied to enhance the NLG performance and the pipeline is collapsed into a single module. End-to-end natural language generation has achieved promising improvements and is the most popular way to perform NLG in recent work. Wen et al. (2015a) argued that language generation should be fully data-driven and not depend on any expert rules. They proposed a statistical language model based on RNNs to learn response generation with semantic constraints and grammar trees. Additionally, they used a CNN reranker to further select better responses. Similarly, an LSTM model was used by Wen et al. (2015b) to learn sentence planning and surface realization simultaneously. Tran and Nguyen (2017) further improved the generation quality on multiple domains using GRU. The proposed generator consistently generated high-quality responses on multiple domains. To improve the domain adaptability of recurrent models, Wen et al. (2016b) proposed to first train the recurrent language model on data synthesized from out-of-domain datasets, then finetune on a comparatively smaller in-domain dataset. This training strategy was proved effective in human evaluation. Context-awareness is important in dialogue response generation because only depending on the dialogue action of the current turn may cause illogical responses. Zhou et al. (2016) built an attention-based Context-Aware LSTM (CA-LSTM) combining target user questions, all semantic values, and dialogue actions as input to generate context-aware responses in QA. Likewise, Dušek and Jurčíček (2016a) concatenated the preceding user utterance with the dialogue action vector and fed it into an LSTM model. Dušek and Jurčíček (2016b) put a syntax constraint upon their neural response generator. A two-stage sequence generation process was proposed. First, a syntax dependency tree was generated to have a structured representation of the dialogue utterance. The generator in the second stage integrated sentence planning and surface realization and produced natural language representations.
More recent works have focused on the reliability and quality of generated responses. A tree-structured semantic representation was proposed by Balakrishnan et al. (2019) to achieve better content planning and surface realization performance. They further designed a novel beam search algorithm to improve the semantic correctness of the generated response. To avoid mistakes such as slot value missing or redundancy in generated responses, Li et al. (2020d) proposed Iterative Rectification Network (IRN), a framework trained with supervised learning and finetuned with reinforcement learning. It iteratively rectified generated tokens by incorporating slot inconsistency penalty into its reward. Golovanov et al. (2019) applied large-scale pretrained models for NLG tasks. After comparing single-input and multi-input methods, they concluded that different types of input context will cause different inductive biases in generated responses and further proposed to utilize this characteristic to better adapt a pretrained model to a new task. Baheti et al. (2020) solved NLG reliability problem in conversational QA. Though with different pipeline structures, they used similar methods to increase the fluency and semantic correctness of the generated response. They proposed Syntactic Transformations (STs) to generate candidate responses and used a BERT to rank their qualities. These generated responses can be viewed as an augmentation of the original dataset to be further used in NLG model learning. Oraby et al. (2019) proposed a method to create datasets with rich style markups from easily available user reviews. They further trained multiple NLG models based on generated data to perform joint control of semantic correctness and language style. Similarly, Elder et al. (2020) put forward a data augmentation approach which put a restriction on response generation. Though this restriction caused dull and less diverse responses, they argued that in task-oriented systems, reliability was more important than diversity.
5 End-to-end Methods
The modules discussed above can achieve good performance in their respective tasks, with the help of recent relevant advances. However, there exist two significant drawbacks in modular systems (Zhao and Eskenazi, 2016): (1) Modules in many pipeline systems are sometimes not differentiable, which means that errors from the end are not able to be propagated back to each module. In real dialogue systems training, usually the only signal is the user response, while other supervised signals like dialogue states and dialogue actions are scarce. (2) Though the modules jointly contribute to the success of a dialogue system, the improvement of one module may not necessarily raise the response accuracy or quality of the whole system. This causes additional training of other modules, which is labor intensive and time-consuming. Additionally, due to the handcrafted features in pipeline task-oriented systems such as dialogue states, it is usually hard to transfer modular systems to another domain, since the predefined ontologies require modification.
There exist two main methods for the end-to-end training of task-oriented dialogue systems. One is to make each module of a pipeline system differentiable, then the whole pipeline can be viewed as a large differentiable system and the parameters can be optimized by back-propagation in an end-to-end fashion (Le et al., 2020b). Another way is to use only one end-to-end module to perform both knowledge base retrieval and response generation, which is usually a multi-task learning neural model.
The increasing applications of neural models have made it possible for modules to be differentiable. While many modules are easily differentiable, there remains one task that makes differentiation challenging: the knowledge base query. Many task-oriented dialogue systems require an external knowledge source to retrieve related knowledge facts required by the user. For example, in the restaurant table booking task, the knowledge fact can be an available slot of one specific restaurant. Traditional methods use a symbolic query to match entries based on their attributes. The system performs semantic parsing on the user message to represent a symbolic query according to the user goal (Li et al., 2017b; Williams and Zweig, 2016; Wen et al., 2016c). However, this retrieval process is not differentiable, which prevents the whole framework from being end-to-end trainable. With the application of key-value memory networks (Miller et al., 2016), Eric and Manning (2017) used the key-value retrieval mechanism to retrieve relevant facts. The proposed architecture was augmented with the attention mechanism to compute the relevance between utterance representations of dialogue and key representations of the knowledge base. Dhingra et al. (2016) presented a soft retrieval mechanism that uses a “soft" posterior distribution over the knowledge base to replace the symbolic queries. They further combined this soft retrieval mechanism into a reinforcement learning framework to achieve complete end-to-end training based on user feedback. Williams et al. (2017) proposed Hybrid Code Networks (HCNs), which encoded domain-specific knowledge into software and system action templates, achieving the differentiability of the knowledge retrieval module. They did not explicitly model the dialogue states but instead learned the latent representation and optimized the HCN using supervised learning and reinforcement learning jointly. Ham et al. (2020) used GPT-2 to form a neural pipeline and perform domain prediction, dialogue state tracking, policy learning, knowledge retrieval, and response generation in a pipeline fashion. The system could easily interact with external systems because it outputs explicit intermediate results from each module and thus being interpretable. Likewise, Hosseini-Asl et al. (2020) built a neural pipeline with GPT-2 and explicitly generated results for each neural module as well.
More recent works tend not to build their end-to-end systems in a pipeline fashion. Instead, they use complex neural models to implicitly represent the key functions and integrate the modules into one. Research in task-oriented end-to-end neural models focuses either on training methods or model architecture, which are the keys to response correctness and quality. Wang et al. (2019a) proposed an incremental learning framework to train their end-to-end task-oriented system. The main idea is to build an uncertainty estimation module to evaluate the confidence of appropriate responses generated. If the confidence score was higher than a threshold, then the response would be accepted, while a human response would be introduced if the confidence score was low. The agent could also learn from human responses using online learning. Dai et al. (2020) used model-agnostic meta-learning (MAML) to improve the adaptability and reliability jointly with only a handful of training samples in a real-life online service task. Similarly, Qian and Yu (2019) also trained the end-to-end neural model using MAML to facilitate the domain adaptation, which enables the model to first train on rich-resource tasks and then on new tasks with limited data. Lin et al. (2020c) proposed Minimalist Transfer Learning (MinTL) to plug-and-play large-scale pretrained models for domain transfer in dialogue task completion. To maintain the sequential correctness of generated responses, Wu et al. (2019b) trained an inconsistent order detection module in an unsupervised fashion. This module detected whether an utterance pair is ordered or not to guide the task-completing agent towards generating more coherent responses. He et al. (2020a) proposed a “Two-Teacher One-Student" training framework. At the first stage, the two teacher models were trained in a reinforcement learning framework, with the objective of retrieving knowledge facts and generating human-like responses respectively. Then at the second stage, the student network was forced to mimic the output of the teacher networks. Thus, the expert knowledge of the two teacher networks was transferred to the student network. Balakrishnan et al. (2019) introduced a constrained decoding method to improve the semantic correctness of the responses generated by the proposed end-to-end system. Many end-to-end task-oriented systems used a memory module to store relevant knowledge facts and dialogue history. Chen et al. (2019c) argued that a single memory module was not enough for precise retrieval. They used two long-term memory modules to store the knowledge tuples and dialogue history respectively, and then a working memory was applied to control the token generation. Zhang et al. (2020) proposed LAtent BElief State (LABES) model, which treated the dialogue states as discrete latent variables to reduce the reliance on turn-level DST labels. To solve the data insufficiency problem in some tasks, Gao et al. (2020a) augmented the response generation model with a paraphrase model in their end-to-end system. The paraphrase model was jointly trained with the whole framework and it aimed to augment the training samples. Yang et al. (2020) leveraged the graph structure information of both a knowledge graph and the dialogue context-dependency tree. They proposed a recurrent cell architecture to learn representations on the graph and performed multi-hop reasoning to exploit the entity links in the knowledge graph. With the augmentation of graph information, consistent improvement was achieved on two task-oriented datasets.
6 Research Challenges and Hot Topics
In this section, we review recent works in task-oriented dialogue systems and point out the frequently studied topics to provide some important research directions. This section can be seen as an augmentation of the literature review in previous sections discussing techniques developed for each module, and focuses more on some specific problems to be solved in the current research community.
The Natural Language Understanding task converts the user message into a predefined format of semantic slots. A popular way to perform NLU is by finetuning large-scale pretrained language models. Wu and Xiong (2020) compared many pretrained language models including BERT-based and GPT-based systems in three subtasks of task-oriented dialogue systems - domain identification, intent detection, and slot tagging. This empirical paper is aimed to provide insights and guidelines in pretrained model selection and application for related research. Wu et al. (2020a) pretrained TOD-BERT and outperformed strong baselines in the intent detection task. The model proposed also had a strong few-shot learning ability to alleviate the data insufficiency problem. Coope et al. (2020) proposed Span-ConveRT, which was a pretrained model designed for slot filling task. It viewed the slot filling task as a turn-based span extraction problem and also performed well in the few-shot learning scenario.
6.2 Domain Transfer for NLU
Another challenge or hot topic in NLU research is the domain transfer problem, which is also the key issue of task-oriented dialogue systems. Hakkani-Tür et al. (2016) built an RNN-LSTM architecture for multitask learning of domain classification, intent detection, and slot-filling problem. Training samples from multiple domains were combined in a single model where respective domain data reinforces each other. Bapna et al. (2017) used a multi-task learning framework to leverage slot name encoding and slot description encoding, thus implicitly aligning the slot-filling model across domains. Likewise, Lee and Jha (2019) also applied slot description to exploit the similar semantic concepts between slots of different domains, which solved the sub-optimal concept alignment and long training time problems encountered in past works involving multi-domain slot-filling.
6.3 Domain Transfer for DST
Domain adaptability is also a significant topic for dialogue state trackers. The domain transfer in DST is challenging due to three main reasons (Ren et al., 2018): (1) Slot values in ontologies are different when the domain changes, which accounts for the incompatibility of models. (2) When the domain changes, the slot number will also change, causing different numbers of model parameters. (3) Hand-crafted lexicons make it difficult for generalization over domains. Mrkšić et al. (2015) used delexicalized n-gram features to solve the domain incompatibility problem by replacing all specified slot names and values with generic symbols. Lin et al. (2020c) introduced Levenshtein belief spans (Lev), which were short context spans relating to the user message. Different from previous methods which generated dialogue state from scratch, they performed substitution (SUB), deletion (DEL), and insertion (INS) based on past states to alleviate the dependency on annotated in-domain training samples. Huang et al. (2020c) applied model-agnostic meta-learning (MAML) to first learn on several source domains and then adapt on the target domain, while Campagna et al. (2020) improved the zero-shot transfer learning by synthesizing in-domain data using an abstract conversation model and the domain ontology. Ouyang et al. (2020) modeled explicit slot connections to exploit the existing slots appearing in other domains. Thus, the tracker could copy slot values from the connected slots directly, alleviating the burden of reasoning and learning. Wang et al. (2020e) proposed Value Normalization (VN) to convert supporting dialogue spans into state values and could achieve high accuracy with only 30% available ontology.
6.4 Tracking Efficiency for DST
Tracking efficiency is another hot topic in dialogue state tracking challenges. Usually, there are multiple states within a dialogue, so how to compute the slot values without any redundant steps becomes very significant when attempting to reduce the reaction time of a system. Kim et al. (2019) argued that predicting the dialogue state from scratch at every turn was not efficient. They proposed to first predict the operations to be taken on each of the slots (i.e., Carryover, Delete, Dontcare, Update), and then perform respective operations as predicted. Ouyang et al. (2020) used a slot connection mechanism to directly copy slot values from the source slot, which reduced the expense of reasoning. Hu et al. (2020) and Wang et al. (2020e) proposed slot attention to calculate the relations between the slot and dialogue context, thus only focusing on the relevant slots at each turn.
6.5 Training Environment for PL
The environment of the Policy Learning framework has been a long-existing problem. Li et al. (2017b) built a user simulator to model the user feedback as the reward signal of an environment. They modeled a stack-like user agenda to iteratively change the user goal and thus shifting the dialogue states. While using a user simulator for environment modeling seems to be promising for that it involves less human interaction, Zhang et al. (2019c) argued that training a user simulator required a large amount of annotated data. Takanobu et al. (2020) proposed Multi-Agent Dialog Policy Learning, where they have two agents interact with each other, performing both user and agent, learning policy simultaneously. Furthermore, they incorporated a role-specific reward to facilitate role-based response generation and here both agents also acted as the environment of the other one.
6.6 Response Consistency for NLG
Response consistency in NLG is a challenging problem since it cannot be solved by simply augmenting the training samples. Instead, additional corrections or regulations should be designed. Wen et al. (2015b) proposed the Semantically Controlled LSTM (SC-LSTM) which used a semantic planning gate to control the retention or abandonment of dialogue actions thus ensuring the response consistency. Likewise, Tran and Nguyen (2017) also applied a gating mechanism to jointly perform sentence planning and surface realization where dialogue action features were gated before entering GRU cells. Li et al. (2020d) proposed Iterative Rectification Network (IRN), which combined a slot inconsistency reward into the reinforcement learning framework. Thus, the model iteratively checked the correctness of slots and corresponding values.
6.7 End-to-end Task-oriented Dialogue Systems
End-to-end systems are usually fully data-driven, which contributes to their robust and natural responses. However, because of the finiteness of annotated training samples, a hot research topic is figuring out how to increase the response quality of end-to-end task-oriented dialogue systems with limited data. Using rule-based methods to constrain response generation is a way to improve response quality. Balakrishnan et al. (2019) used linearized tree-structured representation as input to obtain control over discourse-level and sentence-level semantic concepts. Kale and Rastogi (2020) used templates to improve the semantic correctness of generated responses. They broke down the response generation into a two-stage process: first generating semantically correct but possibly incoherent responses based on the slots, with the constraint of templates; then in the second stage, pretrained language models were applied to re-organize the generated utterances into coherent ones. Training the network with reinforcement learning was another strategy to alleviate the reliance on annotated data. He et al. (2020a) trained two teacher networks using a reinforcement learning framework with the objectives of knowledge retrieval and response generation respectively. Then the student network learns to produce responses by mimicking the output of teacher networks. Training the network in a supervised way, Dai et al. (2020) alternatively tried to optimize the learning strategy to improve the learning efficiency of models given limited data. They combined the meta-learning algorithm with human-machine interaction and achieved significant improvement compared with strong baselines not trained with the meta-learning algorithms. A more direct way to solve the data finiteness problem in supervised learning was augmenting the dataset (Elder et al., 2020), which also improved the response quality to some extent. Additionally, pretraining large-scale models on common corpus and then applying them in a domain that lacks annotated data is a popular approach in recent years (Henderson et al., 2019b; Mehri et al., 2019; Bao et al., 2019b).
6.8 Retrieval Methods for Task-oriented Dialogue Systems
Retrieval-based methods are rare in task-oriented systems for the insufficiency of candidate entries to cover all possible responses which usually involve specific knowledge from external knowledge-base. However, Henderson et al. (2019b) argued that in some situations not relating with specific knowledge facts, retrieval-based methods were more precise and effective. They first pretrained the response selection model on general domain corpora and then finetuned on small target domain data. Experiments on six datasets from different domains proved the effectiveness of the pretrained response selection model. Lu et al. (2019b) constructed Spatio-temporal context features to facilitate response selection, and achieved significant improvements on the Ubuntu IRC dataset.
Open-Domain Dialogue Systems
This section discusses open-domain dialogue systems, which are also called chit-chat dialogue systems or non-task-oriented dialogue systems. Almost all state-of-the-art open-domain dialogue systems are based on neural methods. We organize this section by first briefly introducing the concepts of different branches of open-domain dialogue systems, and then we focus on different research challenges and hot topics. We view these challenges and hot topics as different research directions in open-domain dialogue systems.
Instead of managing to complete tasks, open-domain dialogue systems aim to perform chit-chat with users without the task and domain restriction (Ritter et al., 2011) and are usually fully data-driven. Open-domain dialogue systems are generally divided into three categories: generative systems, retrieval-based systems, and ensemble systems. Generative systems apply sequence-to-sequence models to map the user message and dialogue history into a response sequence that may not appear in the training corpus. By contrast, retrieval-based systems try to find a pre-existing response from a certain response set. Ensemble systems combine generative methods and retrieval-based methods in two ways: retrieved responses can be compared with generated responses to choose the best among them; generative models can also be used to refine the retrieved responses (Zhu et al., 2018; Song et al., 2016; Qiu et al., 2017; Serban et al., 2017b). Generative systems can produce flexible and dialogue context-related responses while sometimes they lack coherence and tend to make dull responses. Retrieval-based systems select responses from human response sets and thus are able to achieve better coherence in surface-level language. However, retrieval systems are restricted by the finiteness of the response sets and sometimes the responses retrieved show a weak correlation with the dialogue context (Zhu et al., 2018).
In the next few subsections, we discuss some research challenges and hot topics in open-domain dialogue systems. We aim to to help researchers quickly grasp the current research trends via a systematic discussion on articles solving certain problems.
Dialogue context consists of user and system messages and is an important source of information for dialogue agents to generate responses because dialogue context decides the conversation topic and user goal (Serban et al., 2017a). A context-aware dialogue agent responds not only depending on the current message but also based on the conversation history. The earlier deep learning-based systems added up all word representations in dialogue history or used a fixed-size window to focus on the recent context (Sordoni et al., 2015b; Li et al., 2015). Serban et al. (2016) proposed Hierarchical Recurrent Encoder-Decoder (HRED), which was ground-breaking in building context-awareness dialogue systems. They built a word-level encoder to encode utterances and a turn-level encoder to further summarize and deliver the topic information over past turns. Xing et al. (2018) augmented the hierarchical neural networks with the attention mechanism to help the model focus on more meaningful parts of dialogue history.
Both generative and retrieval-based systems rely heavily on dialogue context modeling. Shen et al. (2019) proposed Conversational Semantic Relationship RNN (CSRR) to model the dialogue context in three levels: utterance-level, pair-level, and discourse-level, capturing content information, user-system topic, and global topic respectively. Zhang et al. (2019a) argued that the hierarchical encoder-decoder does not lay enough emphasis on certain parts when the decoder interacted with dialogue contexts. Also, they claimed that attention-based HRED models also suffered from position bias and relevance assumption insufficiency problems. Therefore, they proposed ReCoSa, whose architecture was inspired by the transformer. The model first used a word-level LSTM to encode dialogue contexts, and then self-attention was applied to update the utterance representations. In the final stage, an encoder-decoder attention was computed to facilitate the response generation process. Additionally, Mehri et al. (2019) examined several applications of large-scale pretrained models in dialogue context learning, providing guidance for large-scale network selection in context modeling.
Some works propose structured attention to improve context-awareness. Qiu et al. (2020) learned structured dialogue context by combining structured attention with a Variational Recurrent Neural Network (VRNN). Comparatively, Ferracane et al. (2019) examined the RST discourse tree model proposed by Liu and Lapata (2018) and observed little or even no discourse structures in the learned latent tree. Thus, they argued that structured attention did not benefit dialogue modeling and sometimes might even harm the performance.
Interestingly, Feng et al. (2020b) not only utilized dialogue history, but also future conversations. Considering that in real inference situations dialogue agents cannot be explicitly aware of future information, they first trained a scenario-based model jointly on past and future context and then used an imitation framework to transfer the scenario knowledge to a target network.
Better context modeling improves the response selection performance in retrieval-based dialogue systems (Jia et al., 2020). Tao et al. (2019) proposed Interaction-over-Interaction network (IoI), which consisted of multiple interaction blocks to perform deeper interactions between dialogue context and candidate responses. Jia et al. (2020) organized the dialogue history into conversation threads by performing classifications on their dependency relations. They further used a pretrained Transformer model to encode the threads and candidate responses to compute the matching score. Lin et al. (2020b) argued that response-retrieval datasets should not only be annotated with relevant or irrelevant responses. Instead, a greyscale metric should be used to measure the relevance degree of a response given the dialogue context, thus increasing the context-awareness ability of retrieval models.
Dialogue rewriting problem aims to convert several messages into a single message conveying the same information and dialogue context awareness is very crucial to this task (Xu et al., 2020b). Su et al. (2019a) modeled multi-turn dialogues via dialogue rewriting and benefited from the conciseness of rewritten utterances.
2 Response Coherence
Coherence is one of the qualities that a good generator seeks (Stent et al., 2005). Coherence means maintaining logic and consistency in a dialogue, which is essential in an interaction process for that a response with weak consistency in logic and grammar is hard to understand. Coherence is a hot topic in generative systems but not in retrieval-based systems because candidate responses in retrieval methods are usually human responses, which are naturally coherent.
Refining the order or granularity of sentence functions is a popular strategy for improving the language coherence. Wu et al. (2019b) improved the response coherence via the task of inconsistent order detection. The dialogue systems learned response generation and order detection jointly, which was self-supervised multi-task learning. Xu et al. (2019) presented the concept of meta-words. Meta-words were diverse attributes describing the response. Learning dialogue based on meta-words helped promote response generation in a more controllable way. Liu et al. (2019) used three granularities of encoders to encode raw words, low-level clusters, and high-level clusters. The architecture was called Vocabulary Pyramid Network (VPN), which performed a multi-pass encoding and decoding process on hierarchical vocabularies to generate coherent responses. Shen et al. (2019) also built a three-level hierarchical dialogue model to capture richer features and improved the response quality. Ji et al. (2020) built Cross Copy Networks (CCN), which used a copy mechanism to copy from similar dialogues based on the current dialogue context. Thus, the system benefited from the pre-existing coherent responses, which alleviated the need of performing the reasoning process from scratch.
Many work employ strategies to achieve response coherence on a higher level, which improves the overall quality of the generated responses. Li et al. (2019b) improved the logical consistency of generated utterances by incorporating an unlikelihood loss to control the distribution mismatches. Bao et al. (2019a) proposed a Generation-Evaluation framework that evaluated the qualities, including coherence, of the generated response. The feedback was further seen as a reward signal in the reinforcement learning framework and guided to a better dialogue strategy via policy gradient, thus improving the response quality. Gao et al. (2020b) raised response quality by ranking generated responses based on user feedbacks like upvotes, downvotes, and comments on social networks. Zhu et al. (2018) built a retrieval-enhanced generation model, which enhanced the generated responses in two ways. First, a discriminator was trained with the help of a retrieval system, and then the generator was trained in a GAN framework under the supervision signal of a discriminator. Second, retrieved responses were also used as a part of the generator input to provide a coherent example for the generator. Xu et al. (2020a) achieved a global coherent dialogue by constructing a knowledge graph from corpora. They further performed graph walks to decide “what to say" and “how to say", thus improving the dialogue flow coherence. Mesgar et al. (2019) proposed an assessment approach for dialogue coherence evaluation by combining the dialogue act prediction in a multi-task learning framework and learned rich dialogue representations.
There also evolve some data-wise methods for better response coherence. Bi et al. (2019) proposed to annotate sentence functions in existing conversation datasets to improve the sentence logic and coherence of generated responses. Akama et al. (2020) focused on data effectiveness as well. They filtered out low-quality utterance pairs by scoring the relatedness and connectivity, which was proved to be effective in improving the response coherence. Akama et al. (2020) presented a method for evaluating dataset utterance pairs’ quality in terms of connectedness and relatedness. The proposed scoring technique is based on research findings that have been widely disseminated in the conversation and linguistics communities. Lison and Bibauw (2017) included a weighting model in their neural architecture. The weighting model, which is based on conversation data, assigns a numerical weight to each training sample that reflects its intrinsic quality for dialogue modeling and achieved good result in experiments.
3 Response Diversity
The Bland and generic response is a long-existing problem in generative dialogue systems. Because of the high frequency of generic responses like “I don’t know" in training samples and the beam search decoding scheme of neural sequence-to-sequence models, generative dialogue systems tend to respond with universally acceptable but meaningless utterances (Serban et al., 2016; Vinyals and Le, 2015; Sordoni et al., 2015b). For example, to respond to the user message “I really want to have a meal", the agent tends to choose simple responses like “It’s OK" instead of responding with more complicated sentences like recommendations and suggestions.
Earlier works solve this challenge by modifying the decoding objective or adding a reranking process. Li et al. (2015) replaced the traditional likelihood objective with mutual information. The optimization of mutual information objective aims to achieve a Maximum Mutual Information (MMI). Specifically, the task is to find a best response based on the dialogue context , in order to maximize their mutual information:
The objective causes the model to choose responses with high probability even if the response is unconditionally frequent in the dataset, thus causing it to ignore the content of . Maximizing the mutual information as Equation (49) solves this issue by achieving a trade-off between safety and relativity.
With a similar intuition as described above, increasing response diversity by modifying the decoding scheme at inference time has been explored in earlier works. Vijayakumar et al. (2016) combined a dissimilarity term into the beam search objective and proposed Diverse Beam Search (DBS) to promote diversity. Similarly, Shao et al. (2017) proposed a stochastic beam search algorithm by performing stochastic sampling when choosing top-B responses. In the beam search algorithm, siblings sharing the same parent nodes tended to guide to similar sequences. Inspired by this, Li et al. (2016c) penalized siblings sharing the same parent nodes using an additional term in the beam search objective. This encouraged the algorithm to search more diverse paths by expanding from different parent nodes. Some works further added a reranking stage to select more diverse responses in the generated N-best list (Li et al., 2015; Sordoni et al., 2015b; Shao et al., 2017).
A user message can be mapped into multiple acceptable responses, which is also known as the one-to-many mapping problem. Qiu et al. (2019) considered the one-to-many mapping problem in open-domain dialogue systems and proposed a two-stage generation model to increase response diversity - the first stage extracting common features of multiple ground truth responses and the second stage extracting the distinctive ones. Ko et al. (2020) solved the one-to-many mapping problem via a classification task to learn latent semantic representations. So that given one example response, different ones could be generated by exploring the semantically close vectors in the latent space.
Different training strategies have been proposed to increase response diversity. Bao et al. (2019a) used human instinct or pre-defined objective as a reward signal in a reinforcement learning setting to prompt the agent to avoid generating dull responses. Still, in a reinforcement learning framework, Zhu et al. (2020) performed counterfactual reasoning to explore the potential response space. Given a pre-existing response, the model inferred another policy, which represented another possible response, thus increasing the response diversity. He and Glass (2019) used a negative training method to minimize the generation of bland responses. They first collected negative samples and then gave negative training signals based on these samples to fine-tune the model, impeding the model to generate bland responses. To achieve a better performance, Du and Black (2019) synthesized different dialogue models designed for response diversity based on boosting training. The ensemble model significantly outperformed each of its base models.
Utilizing external knowledge sources is another way to improve the diversity of generated responses because it can enrich the content. Wu et al. (2020b) built a common-sense dialogue generation model which seeks highly related knowledge facts based on the dialogue history. Likewise, Su et al. (2020) incorporated external knowledge sources to diversify the response generation, but the difference was that they utilized non-conversational texts like news articles as relevant knowledge facts, which were obviously easier to obtain. Tian et al. (2019) used a memory module to abstract and store useful information in the training corpus for generating diverse responses.
Another approach to diversify the response generation is to make modifications to the training corpus. Csáky et al. (2019) solved the challenge by filtering out the generic responses in the dataset using an entropy-based algorithm, which was simple but effective. Augmented with human feedback data, Gao et al. (2020b) proposed that the generated responses could be reranked via a response ranking framework trained on the human feedback data and responses with higher quality including diversity were selected. Stasaski et al. (2020) proposed to change the data collection pipeline by iteratively computing the diversity of responses from different human participants in dataset construction and selected those participants who tend to generate informative and diverse responses.
4 Speaker Consistency and Personality-based Response
In open-domain dialogue systems, one big issue is that the responses are entirely learned from training data. The inconsistent response may be received when asking the system about some personal facts (e.g., age, hobbies). If the dataset contains multiple utterance pairs about the query of age, then the response generated tends to be shifting, which is unacceptable because personal facts are usually not random. Thus, for a data-driven chatbot, it is necessary to be aware of its role and respond based on a fixed persona.
Explicitly modeling the persona is the main strategy in recent works. Liu et al. (2020b) proposed a persona-based dialogue generator consisting of a Receiver and a Transmitter. The receiver was responsible for modeling the interlocutor’s persona through several turns’ chat while Transmitter generated utterances based on the persona of agent and interlocutor, together with conversation content. The proposed model supported conversations between two persona-based chatbots by modeling each other’s persona. Without training with additional Natural Language Inference labels, Kim et al. (2020) built an imaginary listener following a normal generator, which reasoned over the tokens generated by the generator and predicted a posterior distribution over the personas in a certain space. After that, a self-conscious speaker generated tokens aligned with the predicted persona. Likewise, Boyd et al. (2020) used an augmented GPT-2 to reason over the past conversations and model the target actor’s persona, conditioning on which persona consistency was achieved.
Responding with personas needs to condition on some persona descriptions. For example, to build a generous agent, descriptions like “I am a generous person" are needed as a part of the model input. However, these descriptions require hand-crafted feature design, which is labor intensive. Madotto et al. (2019) proposed to use Model-Agnostic Meta-Learning (MAML) to adapt to new personas with only a few training samples and needed no persona description. Majumder et al. (2020a) relied on external knowledge sources to expand current persona descriptions so that richer persona descriptions were obtained, and the model could associate current descriptions with some commonsense facts.
Song et al. (2020a) argued that traditional persona-based systems were one-stage systems and the responses they generated still contain many persona inconsistent words. To tackle this issue, they proposed a three-stage architecture to ensure persona consistency. A generate-delete-rewrite mechanism was implemented to remove the unacceptable words generated in prototype responses and rewrite them.
5 Empathetic Response
Empathy means being able to sense other people’s feelings (Ma et al., 2020b). An empathetic dialogue system can sense the user’s emotional changes and produce appropriate responses with a certain sentiment. This is an essential topic in chit-chat systems because it directly affects the user’s feeling and to some extent decides the response quality. Industry systems such as Microsoft’s Cortana, Facebook M, Google Assistant, and Amazon’s Alexa are all equipped with empathy modules (Wang et al., 2020g).
There are two ways to generate utterances with emotion: one is to use explicit sentiment words as a part of input; another is to implicitly combine neural words (Song et al., 2019). Song et al. (2019) proposed a unified framework that uses a lexicon-based attention to explicitly plugin emotional words and a sequence-level emotion classifier to classify the output sequence, implicitly guiding the generator to generate emotional responses through backpropagation. Zhong et al. (2020) used CoBERT for persona-based empathetic response selection and further investigated the impact of persona on empathetic responses. Smith et al. (2020) blended the skills of being knowledgeable, empathetic, and role-aware in one open-domain conversation model and overcame the bias issue when blending these skills.
Since the available datasets for empathetic conversations are scarce, Rashkin et al. (2018) provided a new benchmark and dataset for empathetic dialogue systems. Oraby et al. (2019) constructed a dialogue dataset with rich emotional markups from user reviews and further proposed a novel way to generate similar datasets with rich markups.
6 Controllable Generation
Controllable dialogue generation is an important line of work in open-domain dialogue systems since solely learning from data sample distributions causes many uncertain responses. Some of the dialogue systems are grounded on some external knowledge such as knowledge graph and documents. However, grounding alone without explicit control and semantic targeting may induce output that is accurate but vague.
We may get some inspirations from the prior work on language generation and machine translation since similarly to dialogue systems they are generation-based or seq-to-seq problems. Some related work aimed to enforce user-specified constraints, most notably using lexical constraints (Hokamp and Liu, 2017; Hu et al., 2019; Miao et al., 2019). These methods exclusively use constraints at inference time. Constraints can be included into the latent space during training, resulting in better predictions. Other studies (See et al., 2019; Keskar et al., 2019; Tang et al., 2019) have looked at non-lexical constraints, but they haven’t looked into how they can help with grounding external knowledge. These publications also assume that the system can always be given (gold) constraints, which limits the ability to demonstrate larger benefits of the approaches.
Controllable text generation has also been used to extract high-level style information from contextual information in text style transfer (Hu et al., 2017) and other tasks (Ficler and Goldberg, 2017; Dong et al., 2017; Gao et al., 2019), allowing the former to be independently modified. Zhao et al. (2018) learns an interpretable representation for dialogue systems using discrete latent actions. While existing studies employ “style" descriptors (e.g., positive/negative, formal/informal) as control signals, Wu et al. (2020c) use specific lexical constraints to regulate creation, allowing for finer semantic control. Content planned generation (Wiseman et al., 2017; Hua and Wang, 2019) focuses response generation on a small number of essential words or table entries. This line of work, on the other hand, does not require consideration of the discourse context, which is critical for response generation.
7 Conversation Topic
Daily chats of people usually involve a topic or goal. Actually, a topic or goal is the key to keep each participant engaged in conversations and thus being essential to a chatbot. In real applications, a good topic model helps to retrieve related knowledge and guide the conversation instead of passively responding to the user’s message (Xing et al., 2017). For example, if the user mentions “I like sunny days", a topic-aware system may reason over relevant external knowledge and produce responses like “I know there is a nice park near the seaside, have you ever been there before?". Thus, the agent pushes the conversation to a more engaging stage and enriches the dialogue content.
Almost all topic-aware dialogue agents need to model explicit topics, which can be entities from external knowledge-base, or topic embeddings that have some semantic meaning. Wu et al. (2019c) tried to change the traditional passive response fashion and radically pursue active guidance of conversation. The dialogue agent consists of a leader and a follower, where the leader reasons over a knowledge graph and decides the conversation topic. Likewise, a common-sense knowledge graph was used by Liu et al. (2020c) to lead the conversation topic and make recommendations. Tang et al. (2019) built a topic-aware retrieval-based chatbot. It aimed to guide the conversation topic to the target one step by step. It used a keyword predictor to predict turn-level keywords and selected the discourse-level keyword based on that. The discourse-level keyword was further fed into the retrieval model to retrieve responses regarding a certain topic. Chen and Yang (2020) built a multi-view sequence-to-sequence model to learn dialogue topics by first extracting dialogue structures of unstructured chit-chat dialogues, then generating topic summaries using BART decoder.
In some applications of certain scenarios the conversation topic is essential, and these are where the topic-aware dialogue agents can be applied to. Zhang and Danescu-Niculescu-Mizil (2020) studied the topic-aware chatbot in counseling conversations. In counseling conversations, the agent led the dialogue topic by deciding between empathetically addressing a situation within the current range and moving on to a new target resolution. Cao et al. (2019) studied chatbots in the psychotherapy treatment area and built a topic prediction model to forecast the behavior codes for upcoming conversations, thus guiding the dialogue.
8 Knowledge-Grounded System
External knowledge such as common-sense knowledge is a significant source of information when organizing an utterance. Humans associate current conversation context with their experiences and memories and produce meaningful related responses, such capability results in the gap between human and machine chit-chat systems. As discussed, the earlier chit-chat systems are simply variants of machine translation systems, which can be viewed as sequence-to-sequence language models. However, dialogue generation is much more complicated than machine translation because of the higher freedom and vaguer constraints. Thus, chit-chat systems cannot simply consist of a sequence-to-sequence mapping since appropriate and informative responses are always related to some external common-sense knowledge. Instead, there must be a module incorporating world knowledge.
Many researchers devoted their research efforts to building knowledge-grounded dialogue systems. A representative model is memory networks introduced in Section 2.4. Knowledge grounded systems use Memory Networks to store external knowledge and the generator retrieves relevant knowledge facts from it at the generation stage (Ghazvininejad et al., 2018; Vougiouklis et al., 2016; Yin et al., 2015). Tian et al. (2019) built a memory-augmented conversation model. The proposed model abstracted from the training samples and stored useful ones in the memory module. Zhao et al. (2020b) built a knowledge-grounded dialogue generation system based on GPT-2. They combined a knowledge selection module into the language model and learned knowledge selection and response generation simultaneously. Lin et al. (2020a) proposed Knowledge-Interaction and knowledge Copy (KIC). They performed recurrent knowledge interactions during the decoding phase to compute an attention distribution over the memory. Then they performed knowledge copy using a knowledge-aware pointer network to copy knowledge words according to the attention distribution computed.
Documents contain large amount of knowledge facts, but they have a drawback that they are usually too long to retrieve useful information from (Li et al., 2019d). Li et al. (2019d) built a multi-turn document-grounded system. They used an incremental transformer to encode multi-turns’ dialogue context and respective documents retrieved. In the generation phase, they designed a two-stage generation scheme. The first stage took dialogue context as input and generated coherent responses; the second stage utilized both the utterance from the first stage and the document retrieved for the current turn for response generation. In this case, selecting knowledge based on both dialogue context and generated response was called posterior knowledge selection, while selecting knowledge with only dialogue context was called prior knowledge selection, which only utilized prior information. Wang et al. (2020c) built a document quotation model in online conversations and investigated the consistency between quoted sentences and latent dialogue topics.
Knowledge graph is another source of external information, which is becoming more and more popular in knowledge-grounded systems because of their structured nature. Jung et al. (2020) proposed a dialogue-conditioned graph traversal model for knowledge-grounded dialogue systems. The proposed model leveraged attention flows of two directions and fully made use of the structured information of knowledge graph to flexibly decide the expanding range of nodes and edges. Likewise, Zhang et al. (2019b) applied graph attention to traverse the concept space, which was a common-sense knowledge graph. The graph attention helped to move to more meaningful nodes conditioning on dialogue context. Xu et al. (2020a) applied knowledge graphs as an external source to control a coarse-level utterance generation. Thus, the conversation was supported by common-sense knowledge, and the agent guided the dialogue topic in a more reasonable way. Moon et al. (2019) built a retrieval system retrieving responses based on the graph reasoning task. They used a graph walker to traverse the graph conditioning on symbolic transitions of the dialogue context. Huang et al. (2020a) proposed Graph-enhanced Representations for Automatic Dialogue Evaluation (GRADE), a novel evaluation metric for open-domain dialogue systems. This metric considered both contextualized representations and topic-level graph representations. The main idea was to use an external knowledge graph to model the conversation logic flow as a part of the evaluation criteria.
Knowledge-grounded datasets containing context-knowledge-response triples are scarce and hard to obtain. Cho and May (2020) collected a large dataset consisting of more than 26000 turns of improvised dialogues which were further grounded with a larger movie corpus as external knowledge. Also tackling the data insufficiency problem, Li et al. (2020b) proposed a method that did not require context-knowledge-response triples for training and was thus data-efficient. They viewed knowledge as a latent variable to bridge the context and response. The variational approach learned the parameters of the generator from both a knowledge corpus and a dialogue corpus which were independent of each other.
9 Interactive Training
Interactive training, also called human-in-loop training, is a unique training method for dialogue systems. Annotated data is fixed and limited, not being able to cover all dialogue settings. Also, it takes a long time to train a good system. But in some industrial products, the dialogue systems need not be perfect when accomplishing their tasks. Thus, interactive training is desirable because the dialogue systems can improve themselves via interactions with users anywhere and anytime, which is a more flexible and cheap way to finetune the parameters.
Training schemes with the above intuition have been developed in recent years. Li et al. (2016a) introduced a reinforcement learning-based online learning framework. The agent interacted with a human dialogue partner and the partner provided feedback as a reward signal. Asghar et al. (2016) first trained the agent with two-stage supervised learning, and then used an interaction-based reinforcement learning to finetune. Every time the user chose the best one from K responses generated by the pretrained model and then responded to this selected response. Instead of learning through being passively graded, Li et al. (2016b) proposed a model that actively asked questions to seek improvement. Active learning was applicable to both offline and online learning settings. Hancock et al. (2019) argued that most conversation samples an agent saw happened after it was pretrained and deployed. Thus, they proposed a framework to train the agent from the real conversations it participated in. The agent evaluated the satisfaction score of the user from the user’s response of each turn and explicitly requested the user feedback when it thought that a mistake has been made. The user feedback was further used for learning. Bouchacourt and Baroni (2019) placed the interactive learning in a cooperative game and tried to learn a long-term implicit strategy via Reinforce algorithm. Some of these work has been adopted by industry products and is a very promising direction for study.
10 Visual Dialogue
More and more researchers cast their eyes to a broader space and are not only restricted to NLP. The combination of CV and NLP giving rise to tasks like visual question answering attracted lots of interest. The VQA task is to answer a question based on the content of a picture or video. Recently, this has evolved into a more challenging task: visual dialogue, which conditions a dialogue on the visual information and dialogue history. The dialogue consists of a series of queries, and the query form is usually more informal, which is why it is more complicated than VQA.
Visual dialogue can be seen as a multi-step reasoning process over a series of questions (Gan et al., 2019). Gan et al. (2019) learned semantic representation of the question based on dialogue history and a given image, and recurrently updated the representation. Shuster et al. (2019) proposed a set of image-based tasks and provided strong baselines. Wang et al. (2020f) employed R-CNN as an image encoder and fused the visual and dialogue modality with a VD-BERT. The proposed architecture achieved sufficient interactions between multi-turn dialogue and images. The proposed architecture is shown as an example model for Visual Dialogue tasks in Figure 16.
Compared with image-grounded dialogue systems, video-grounded systems are more interesting but also more challenging. There are two main challenges of video dialogue, as claimed by Le and Hoi (2020). One is that both spatial and temporal features exist in the video, which increases the difficulty of feature extraction. Another is that video dialogue features span across multiple conversation turns and thus are more complicated. A GPT-2 model was applied by Le and Hoi (2020), being able to fuse multi-modality information over different levels. Likewise, Le et al. (2019) built a multi-modal transformer network to incorporate information from different modalities and further applied a query-aware attention to extract context-related features from non-text modalities. Le et al. (2020a) proposed a Bi-directional Spatio-Temporal Learning (BiST) leveraging temporal-to-spatial and spatial-to-temporal reasoning process and could adapt to the dynamically evolving semantics in the video.
Some researchers hold different opinions on the effectiveness of dialogue history in visual dialogue. Takmaz et al. (2020) proposed that many expressions were already mentioned in previous turns and they built a visual dialogue model grounded on both image and conversation history. They further proved that better performance was achieved when grounding the model on dialogue context. However, Agarwal et al. (2020) argued that though with dialogue history the visual dialogue model could achieve better results, in fact only a small proportion of cases benefited from the history. Furthermore, they proved that existing evaluation metrics for visual dialogue promoted generic responses.
The visual dialogue task benefits a lot from the pretraining-based learning. The popularity of NLP pretraining sparked interest in multi-modal pretraining. VideoBERT (Sun et al., 2019b) is widely recognized as the pioneering work in the field of multimodal pretraining. It’s a model that’s been pre-trained on video frame features and text. CBT (Sun et al., 2019a), which is similarly pretrained on video-text pairs, is a contemporary work of VideoBERT. For video representation learning, Miech et al. (2020) used unlabeled narrated films. More researchers have focused their attention on visual-linguistic pretraining, inspired by the early work in multi-modal pretraining. For this objective, there are primarily two types of model designs. The single-stream model (Alberti et al., 2019; Chen et al., 2019d; Gan et al., 2020; Li et al., 2020a, 2019a, c; Su et al., 2019c; Zhou et al., 2020b) is one example. (Li et al., 2020a) used a BERT model to process the concatenation of objects and words and pre-trained it with three standard tasks. Similar methods were proposed by Chen et al. (2019d) and Qi et al. (2020), but with more pretraining tasks and larger datasets. With an adversarial training technique, Gan et al. (2020) further enhanced the model. Su et al. (2019c) employed the same architecture, but incorporated single-modal data and pre-trained the object detector. Instead of using recognized objects, Huang et al. (2020d) sought to enter pixels directly. The object labels were used by Li et al. (2020c) to improve cross-modal alignment. Zhou et al. (2020b) suggested a single-stream model that learns both caption generation and VQA tasks at the same time. The two-stream model (Lu et al., 2019a, 2020; Tan and Bansal, 2019; Yu et al., 2020) is another type of model architecture. Tan and Bansal (2019) suggested a two-stream model with co-attention and solely used in-domain data to train the model. Lu et al. (2019a) introduced a similar architecture with a more complex co-attention model, which they pretrained with out-of-domain data, and Lu et al. (2020) improved VilBERT with multi-task learning. Yu et al. (2020) recently added the scene graph to the model, which improved performance. Aside from these studies, Singh et al. (2020) looked at the impact of pretraining dataset selection on downstream task performance.
The annotation of visual dialogue is laborious and thus the datasets are scarce. Recently, some researchers have tried to tackle the data insufficiency problem. Shuster et al. (2020) collected a dataset (IMAGE-CHAT, shown in Figure 17) of image-grounded human-human conversations in which speakers are asked to perform role-playing based on an emotional mood or style offered, since the usage of such characteristics is also a significant factor in engagingness. Kamezawa et al. (2020) constructed a visual-grounded dialogue dataset. Interestingly, it additionally annotated the eye-gaze locations of the interlocutor in the image to provide information on what the interlocutor was paying attention to. Cogswell et al. (2020) proposed a method to utilize the VQA data when adapting to a new task, minimizing the requirement of dialogue data which is expensive to annotate.
Evaluation Approaches
Evaluation is an essential part of research in dialogue systems. It is not only a way to assess the performance of agents, but it can also be a part of the learning framework which provides signals to facilitate the learning (Bao et al., 2019a). This section discusses the evaluation methods in task-oriented and open-domain dialogue systems.
Task-oriented systems aim to accomplish tasks and thus have more direct metrics evaluating their performance such as task completion rate and task completion cost. Some evaluation methods also involve metrics like BLEU to compare system responses with human responses, which will be discussed later. In addition, human-based evaluation and user simulators are able to provide real conversation samples.
Task Completion Rate is the rate of successful events in all task completion attempts. It measures the task completion ability of a dialogue system. For example, in movie ticket booking tasks, the Task Completion Rate is the fraction of dialogues that meet all requirements specified by the user, such as movie time, cinema location, movie genre, etc. The task completion rate was applied in many task-oriented dialogue systems (Walker et al., 1997; Williams, 2007; Peng et al., 2017). Additionally, some works (Singh et al., 2002; Yih et al., 2015) used partial success rate.
Task Completion Cost is the resources required when completing a task. Time efficiency is a significant metric belonging to Task Completion Cost. In dialogue-related tasks, the number of conversation turns is usually used to measure the time efficiency and dialogue with fewer turns is preferred when accomplishing the same task.
Human-based Evaluation provides user dialogues and user satisfaction scores for system evaluation. There are two main streams of human-based evaluation. One is to recruit human labor via crowdsourcing platforms to test-use a dialogue system. The crowdsource workers converse with the dialogue systems about predefined tasks and then metrics like Task Completion Rate and Task Completion Cost can be calculated. Another is computing the evaluation metrics in real user interactions, which means that evaluation is done after the system is deployed in real use.
User Simulator provides simulated user dialogues based on pre-defined rules or models. Since recruiting human labor is expensive and real user interactions are not available until a mature system is deployed, user simulators are able to provide task-oriented dialogues at a lower cost. There are two kinds of user simulators. One is agenda-based simulators (Schatzmann and Young, 2009; Li et al., 2016e; Ultes et al., 2017), which only feed dialogue systems with the pre-defined user goal as a user message, without surface realization. Another is model-based simulators (Chandramohan et al., 2011; Asri et al., 2016), which generate user utterances using language models given constraint information.
2 Evaluation Methods for Open-domain Dialogue Systems
Evaluation of open-domain dialogue systems has long been a challenging problem. Unlike task-oriented systems, there is no clear metric like task completion rate or task completion cost. Both human and automatic evaluation methods are developed for ODD during these years. Human evaluation has been adopted by many works (Ritter et al., 2011; Shang et al., 2015; Sordoni et al., 2015b) to converse with and rate dialogue agents. However, human evaluation is not an ideal approach for that human labor is expensive and the evaluation results are highly subjective, varying from person to person. Researchers tend to hire crowd source workers (Ritter et al., 2011; Shang et al., 2015; Sordoni et al., 2015b) or random people (Moon et al., 2019; Jung et al., 2020) to conduct human evaluation, both of which have two main drawbacks: 1. The evaluator group is highly random, and there exists huge gap between people with different knowledge levels or from different domains. 2. Though individual bias could be weakened by increasing the number of evaluators, the evaluator group cannot be very large because of the limited budgets (in articles mentioned above the sizes of human evaluator groups are usually 5-20). Thus, automatic and objective metrics are desirable. In general, there are two categories of automatic metrics in recent research: word-overlap metrics and neural metrics.
Word-overlap Metrics are widely used in Machine Translation and Summarization tasks, which calculate the similarity between the generated sequence and the ground truth sequence. Representative metrics like BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004) are n-gram matching metrics. METEOR (Banerjee and Lavie, 2005) was further proposed with an improvement based on BLEU. It identified the paraphrases and synonyms between the generated sequence and the ground truth. Galley et al. (2015) extended the BLEU by exploiting numerical ratings of responses. Liu et al. (2016) argued that word-overlap metrics were not correlated well with human evaluation. These metrics are effective in Machine Translation because each source sentence has a ground truth to compare with, whereas in dialogues there may be many possible responses corresponding with one user message, and thus an acceptable response may receive a low score if simply computing word-overlap metrics.
Neural Metrics are metrics computed by neural models. Neural methods improve the evaluation effectiveness in terms of adaptability compared with word-overlap metrics, but they require an additional training process. Su et al. (2015) used an RNN and a CNN model to extract turn-level features in a sequence and give the score. Tao et al. (2018) proposed Ruber, which was an automatic metric combining referenced and unreferenced components. The referenced one computed the similarity between generated response representations and ground truth representations, while the unreferenced one learned a scoring model to rate the query-response pairs. Lowe et al. (2017) learned representations of dialogue utterances using an RNN and then computed the dot-product between generated response and ground truth response as an evaluation score. Kannan and Vinyals (2017) and Bruni and Fernandez (2017) used the discriminator of a GAN framework to distinguish the generated responses from human responses. If a generated response achieved a high confidence score, this was indicative of a human-like response, thus desirable.
Evaluation of open-domain dialogue systems is a hot topic at present and many researchers cast their eyes on this task recently. Some papers introduce two or more custom evaluation metrics for better evaluation, such as response diversity, response consistency, naturalness, knowledgeability, understandability, etc., to study "what to evaluate". Bao et al. (2019a) evaluated the generated responses by designing two metrics. One was the informativeness metric calculating information utilization over turns. Another was the coherence metric, which was predicted by GRUs, given the response, context, and background as input. Likewise, Akama et al. (2020) designed scoring functions to compute connectivity of utterance pairs and content relatedness as two evaluation metrics and used another fusion function to combine the metrics. Pang et al. (2020) combined four metrics in their automatic evaluation framework: the context coherence metric based on GPT-2; phrase fluency metric based on GPT-2; diversity metric based on n-grams; logical self-consistency metric based on textual-entailment-inference. Mehri and Eskenazi (2020) proposed a reference-free evaluation metric. They annotated responses considering the following qualities: Understandable (0-1), Maintains Context (1-3), Natural (1-3), Uses Knowledge (0-1), Interesting (1-3), Overall Quality (1-5). Furthermore, a transformer was trained on these annotated dialogues to compute the score of quality.
Apart from "what to evaluate", there are also a multitude of papers studying "how to evaluate", which focus more on refining the evaluation process. Liang et al. (2020) proposed a three-stage framework to denoise the self-rating process. They first performed dialogue flow anomaly detection via self-supervised representation learning, and then the model was fine-tuned with smoothed self-reported user ratings. Finally, they performed a denoising procedure by calculating the Shapley value and removed the samples with negative values. Zhao et al. (2020a) trained RoBERTa as a response scorer to achieve reference-free and semi-supervised evaluation. Sato et al. (2020) constructed a test set by first generating several responses based on one user message and then human evaluation was performed to annotate each response with a score, where the response with the highest score was taken as a true response and the remainder taken as false responses. Dialogue systems were further evaluated by comparing the response selection accuracy on the test set, where a cross-entropy loss was calculated between the generated response and candidate responses to perform the selection operation. Likewise, Sinha et al. (2020) trained a BERT-based model to discriminate between true and false responses, where false responses were automatically generated. The model was further used to predict the evaluation score of a response based on dialogue context. Huang et al. (2020a) argued that responses should not be simply evaluated based on their surface-level features, and instead the topic-level features were more essential. They incorporated a common-sense graph in their evaluation framework to obtain topic-level graph representations. The topic-level graph representation and utterance-level representation were jointly considered to evaluate the coherence of responses generated by open-domain dialogue systems.
Ranking is also an approach that evaluates dialogue systems effectively. Gao et al. (2020b) leveraged large-scale human feedback data such as upvotes, downvotes, and replies to learn a GPT-2-based response ranker. Thus, responses were evaluated by their rankings given by the ranker. Deriu et al. (2020) also evaluated the dialogue systems by ranking. They proposed a low-cost human-involved evaluation framework, in which different conversational agents conversed with each other and the human’s responsibility was to annotate whether the generated utterance was human-like or not. The systems were evaluated by comparing the number of turns their responses were judged as human-like responses.
Datasets
The dataset is one of the most essential components in dialogue systems study. Nowadays the datasets are not enough no matter for task-oriented or open-domain dialogue systems, especially for those tasks requiring additional annotations (Novikova et al., 2017). For task-oriented dialogue systems, data can be collected via two main methods. One is to recruit human labor via crowdsourcing platforms to produce dialogues in a given task. Another is to collect dialogues in real task completions like film ticket booking. For open-domain dialogue systems, apart from dialogues collected in real interactions, social media is also a significant source of data. Some social media companies such as Twitter and Reddit provide API access to a small proportion of posts, but these services are restricted by many legal terms which affect the reproducibility of research. As a result, many recent works in dialogue systems collect their own datasets for train and test.
In this section, we review and categorize these datasets and make a comprehensive summary. To our best knowledge, Table LABEL:Datasets_for_Task-oriented_dialogue_systems and LABEL:Datasets_for_Open-domain_dialogue_systems cover almost all available datasets used in recent task-oriented or open-domain dialogue systems.
2 Datasets for Open-domain Dialogue Systems
Conclusions and Trends
More and more researchers are investigating conversational tasks. One factor contributing to the popularity of conversational tasks is the increasing demand for chatbots in industry and daily life. Industry agents like Apple’s Siri, Microsoft’s Cortana, Facebook M, Google Assistant, and Amazon’s Alexa have brought huge convenience to people’s lives. Another reason is that a considerable amount of natural language data is in the form of dialogues, which contributes to the efforts in dialogue research.
In this paper we discuss dialogue systems from two perspectives: model and system type. Dialogue systems are a complicated but promising task because it involves the whole process of communication between agent and human. The works of recent years show an overwhelming preference towards neural methods, no matter in task-oriented or open-domain dialogue systems. Neural methods outperform traditional rule-based methods, statistical methods and machine learning methods for that neural models have the stronger fitting ability and require less hand-crafted feature engineering.
We systematically summarized and categorized the latest works in dialogue systems, and also in other dialogue-related tasks. We hope these discussions and insights provide a comprehensive picture of the state-of-the-art in this area and pave the way for further research. Finally, we discuss some possible research trends arising from the works reviewed:
The world is multimodal and humans observe it via multiple senses such as vision, hearing, smell, taste, and touch. In a conversational interaction, humans tend to make responses not only based on text, but also on what they see and hear. Thus, some researchers argue that chatbots should also have such abilities to blend information from different modalities. There are some recent works trying to build multimodal dialogue systems (Le et al., 2019; Chauhan et al., 2019; Saha et al., 2020; Singla et al., 2020; Young et al., 2020), but these systems are still far from mature.
Dialogue systems are categorized into task-oriented and open-domain systems. Such a research boundary has existed for a long time because task-oriented dialogue systems involve dialogue states, which constrain the decoding process. However, works in end-to-end task-oriented dialogue systems and knowledge-grounded open-domain systems provide a possibility of blending these two categories into a single framework, or even a single model. Such blended dialogue systems perform as assistants and chatbots simultaneously.
In Section 6 we reviewed many datasets for dialogue systems training. However, data is still far from enough to train a perfect dialogue system. Many learning techniques are designed to alleviate this problem, such as reinforcement learning, meta-learning, transfer learning, and active learning. But many works ignore a significant source of information, which is the dialogue corpus on the Internet. There is a large volume of conversational corpus on the Internet but people have no access to the raw corpus because much of it is in a messy condition. In the future, dialogue agents should be able to explore useful corpus on the Internet in real-time for training. This can be achieved by standardizing online corpus access and their related legal terms. Moreover, real-time conversational corpus exploration can be an independent task that deserves study.
User modeling is a hot topic in both dialogue generation (Gür et al., 2018; Serras et al., 2019) and dialogue systems evaluation (Kannan and Vinyals, 2017). Basically, the user modeling module tries to simulate the real decisions and actions of a human user. It makes decisions based on the dialogue state or dialogue history. In dialogue generation tasks, modeling the user helps the agent converse more coherently, based on the background information or even speaking habits. Besides that, a mature user simulator can provide an interactive training environment, which reduces the reliance on annotated training samples when training a dialogue system. In dialogue systems evaluation tasks, a user simulator provides user messages to test a dialogue agent. More recent user simulators also give feedback concerning the responses generated by the dialogue agent. However, user modeling is a challenging task since no matter explicit user simulation or implicit user modeling is actually the same in difficulty as response generation. Since response generation systems are not perfect yet, user modeling can still be a topic worthy of study.
Most of our daily conversations are chitchats without any purpose. However, there are quite a few scenarios when we purposely guide the conversation content to achieve a specific goal. Current open-domain dialogue systems tend to model the conversation without a long-term goal, which does not exhibit enough intelligence. There are some recent works that apply reinforcement policy learning to model a long-term reward which encourages the agent to converse with a long-term goal, such as the work of Xu et al. (2020a). This topic will lead to strong artificial intelligence, which is useful in some real-life applications such as negotiation or story-telling chatbots.
Acknowledgements
This research/project is supported by A*STAR under its Industry Alignment Fund (LOA Award I1901E0046).