The Adapter-Bot: All-In-One Controllable Conversational Model
Andrea Madotto, Zhaojiang Lin, Yejin Bang, Pascale Fung
Introduction
Large pre-trained language models Peters et al. (2018); Radford et al. (2019); Brown et al. (2020) have greatly improved the state-of-the-art in many down-stream tasks. Similarly, transformer-based conversational models trained on large unlabeled human-to-human conversation (i.e. Reddit comments) Zhang et al. (2019); Adiwardana et al. (2020); Roller et al. (2020) have shown excellent performance in modelling human responses. These models are capable of generating coherent and fluent responses.
Despite their capabilities, existing large conversational models are unable to on-demand control generated responses. For instance, once conversational language models Zhang et al. (2019); Adiwardana et al. (2020); Roller et al. (2020) are fine-tuned on multiple conversational datasets, there is no mechanism (e.g., control codes or latent variables) for controlling which response ‘skill’ to use. Furthermore, these conversational models are unable to add dialogue ‘skills’ continuously without retraining all the model parameters. In advanced smart speakers such as Alexa,https://www.amazon.com/alexa-skills/b?ie=UTF8&node=13727921011, multiple skills can be added easily since the overall infrastructure is hard-coded, but in deep learning models, adding new conversational skills, without catastrophically forgetting all the previous ones, is challenging French (1999); Kirkpatrick et al. (2017). Finally, pre-trained conversational models Zhang et al. (2019); Adiwardana et al. (2020) are restricted to open-domain chats, and they often do not use knowledge to ground their responses. Ideally, a conversational model has to generalize over knowledge of multiple types, e.g., text, tables, images, and graphs.
To overcome these challenges, we propose the Adapter-Bot, a dialogue model that uses a fixed backbone conversational model such as DialGPT Zhang et al. (2019) and triggers on-demand dialogue skills (e.g., emphatic response, weather information etc.) via different Adapters Houlsby et al. (2019). Each adapter can be trained independently, thus allowing a continual integration of skills without retraining the entire conversational model. Moreover, we propose two interaction modes, automatic and manual. In the first, we train a Dialogue Manager (BERT Devlin et al. (2019)) to select which adapter to use give a certain dialogue history. In the second, we let the user decide what style or skill to use for the response, to show high-level control over the chat-bot.
We implemented 22 dialogue skills by leveraging multiple datasets to train each of the adapters (Section 7.1). The datasets cover multiple dialogue skills, such as movie/book/music/sport recommendation, knowledge-grounded responses, weather forecast goal-oriented domains, personalized and empathetic responses, and 12 different response styles (e.g., positive/negative/questions etc.). We compare each of the dialogue skills with the current state-of-the-art models based on automatic metrics. Finally, we present our interactive system which enable to chat with the Adapter-Bot.
Related Work
Generating human-like responses involves overcoming a variety of challenges such as personalization Li et al. (2016b); Zhang et al. (2018a); Dinan et al. (2019); Wolf et al. (2019), emotions Li et al. (2017); Rashkin et al. (2019); Fan et al. (2018, 2020a, 2020b); Zhou et al. (2018), diversity Li et al. (2016a, c); Ghandeharioun et al. (2019); Serban et al. (2017); Gao et al. (2018); Lin et al. (2020c) and so on. In terms of controlled dialogue generation, See et al. (2019) studied conditional generative models Kikuchi et al. (2016) and weighted decoding Ghazvininejad et al. (2017a) in controlling models trained on Persona-Chat. More recently, large pre-trained conversational models have been able to achieve a high humanness score Adiwardana et al. (2020); Roller et al. (2020); Zhang et al. (2019). Instead of training a large model from scratch, in this paper, we propose a lightweight method to include the many dialogue skills in existing conversational models.
Knowledge-Grounded Dialogues
Generating responses grounded on knowledge (e.g., graphs, documents, tables, images etc.) has been explored in different dialogue scenarios, such as document-grounded conversations Dinan et al. (2018); Gopalakrishnan et al. (2019); Ghazvininejad et al. (2018); Moghe et al. (2018); Wu et al. (2020), knowledge bases (e.g., tables) for end-to-end task-oriented dialogue systems Bordes and Weston (2017); Eric et al. (2017a); Eric and Manning (2017); Reddy et al. (2019); Madotto et al. (2018); Wu et al. (2019a); Neelakantan et al. (2019); Qin et al. (2019, 2020); Raghu et al. (2019); Haihong et al. (2019), dialogue-based recommendation-systems Moon et al. (2019); Zhou et al. (2020); Wu et al. (2019b), and image-grounded conversation Shuster et al. (2020a); Mostafazadeh et al. (2017). In this paper, we propose a general model that is able to process any kind of text-based knowledge. While Shuster et al. (2020b) has also recently explored a multi-task model that includes images as input, we leave the exploration of image-based response generation for future work.
Mixture-of-Experts
The idea of having specialized parameters, or so-called experts, has been a widely studied topic in the last two decades (Jacobs et al., 1991; Jordan and Jacobs, 1994). More recently, the Mixture-of-Experts (Shazeer et al., 2017; Kaiser et al., 2017) model, which adds a large number of experts between two LSTMs was proposed. In dialogue systems, the idea of selecting different experts for different dialogue skills has been explored for task-oriented dialogue systems Madotto et al. (2020b), emotional response generation Lin et al. (2019), and to select different retrieval dialogue models Smith et al. (2020). In this paper, we propose to encode the specialized parameters directly onto a large pre-trained model using adapter layers Houlsby et al. (2019) and to use a dialogue manager, directly trained on the dialogue history, to select the experts.
Methodology
The Adapter-Bot is made of a frozen backbone conversational model, a set of trainable independent residual adapters Houlsby et al. (2019), and a trainable dialogue manager. As shown in Figure 1, the Adapter-Bot processes the dialogue history with the required meta-knowledge, and it generates a response using the appropriate adapter layer selected by the dialogue manager (BERT in the figure). The meta-knowledge is retrieved or accessed via API calls.
Let us define a dialogue as the alternation of utterances between two speakers, where an utterance is a sequence of words. Then, we denote the meta-knowledge as the set of external information used at each turn. This set can be either empty, when no external knowledge is required, or made of different knowledge data types: text, tables, and graphs.
We define the backbone conversational model as a function that gets the dialogue history and the external knowledge , and generates the system response word by word. Formally,
In our proposed methodology, we use a pre-trained conversational model such as DialoGPT Zhang et al. (2019) as backbone, but in general, any transformer-based architecture can be used (e.g. BlenderBot, Meena).
where and are trainable parameters in of dimension and respectively, and LN denotes layer normalization Ba et al. (2016). The bottleneck dimension is tunable and it allows adjustment of the capacity of the adapter according the to complexity of the dialogue skills.
2 The Adapter-Bot
We model the Adapter-Bot as an extension of Equation 1 where a further input, namely, the adapter skill-ids , is added to . We add a set of residual adapters, parameterized as , to the back-bone model, each of which is indexed by its own skill-id (i.e., index refers to adapter with parameters ). The skill-id can be either selected by the user or by the dialogue manager (Section 4). Formally, the Adapter-Bot models the system response as
We define a set of dialogue datasets , where each dataset is made of dialogues with their corresponding meta-knowledge aligned. Then, we optimize the adapter parameters in to minimize the negative log-likelihood over the dataset of dialogues and its knowledge .
Dialogue Manager
The dialogue manager is trained to select the right dialogue skill by predicting the index of the residual adapter. More formally, given the dialogue history the dialogue manger predicts an index in . The dialogue adapter can be any classifier, and it is trained using the same set of dialogue datasets , but instead of using the response as supervision, we use the adapter index of the corresponding dialogue. For example, to select adapter , we train the dialogue manager to predict the index from the dialogue histories in .
Knowledge Retrieval
We apply different strategies to retrieve knowledge from different sources. To fetch the relevant information from Wikipedia, we use the TF-IDF retriever implemented by Chen et al. (2017), which computes the dot product of the TF-IDF weighted vector between the last user utterance and the Wikipedia articles. Then, the first paragraph of the highest score article is used as meta-information.
To retrieve information from a knowledge graph, we first extract the entities from user utterances and match them with the entity node. Then we return the first-order neighbours. We store the extracted sub-graph as set of triples in the form (entity1, relation, entity2).
To query the online API (e.g., weather API), we use heuristic rules to extract the slot values (e.g., location), or GPS location if available.
Response Re-ranking
We consider 11 different response style/topics (e.g., positive, negative, scientific etc.). As shown by Dathathri et al. (2019) and Adiwardana et al. (2020), sampling multiple responses and re-ranking them based on a certain criterion (e.g., discriminator loss, word overlapping, etc.) is very helpful for generating more diverse and fluent answers. Therefore, we use additional classifier for each style/topic More information about the dataset used to train the classifiers can be found in Madotto et al. (2020a) and use it to score each of the generated responses. The response with the higher score is returned to final user.
Experimental Setup
We train each of the adapters using different dialogue datasets and settings. In this section, we summarize each of the datasets, which skill they represent, and how they have been pre-processed. Table 2 summarizes the main statistics of each dataset.
(ED) Rashkin et al. (2019) is a benchmark dataset for emotional responses consisting of 25K one-to-one open-domain conversations grounded in emotional situations. In this paper, we use the pre-processing provided by Lin et al. (2020d), in which multiple custom sentences have been added to make the resulting bot more empathetic and fluent.
Persona Chat
(PER) Zhang et al. (2018a); Dinan et al. (2019) is a multi-turn conversational dataset, where two speakers are paired and a persona description (4–5 sentences) is randomly assigned to each of them. The two speakers get to know each other, and they use the persona description to ground the conversation.
Plug-and-Play Style
(PP) Madotto et al. (2020a) is a single-turn synthetically generated dataset for style/topic-controlled generation. This dataset has been created using the Plug-and-Play Language-Model (PPLM) Dathathri et al. (2019) primed with 1K turns of open-ended human-to-human conversation Adiwardana et al. (2020). Madotto et al. (2020a) generated six response style/topics, positive, negative, question, sport, business/finance, and science. In this paper, we further generate traces for five emotional styles angry, fearful, surprised, joyous and sad Saravia et al. (2018). To summarize, we consider 11 styles/topics (POS, NEG, QUE, SPO, FIN, SCI, ANG, FEA, SUR, JOY and SAD) each of which has 1K training samples.
Wizard-of-Wikipedia
(WoW) Dinan et al. (2018) is an open-domain dialogue dataset where two speakers are paired and a conversational topic is assigned. One of the speakers has access to a Wikipedia page and has to ground his or her responses on a Wikipedia sentence.
Stanford-Multi-Domain
(SMD) Eric et al. (2017c) is a multi-domain multi-turn task-oriented dialogue dataset. This dataset covers three domains, weather (WEA), navigation (POI) and calendar (SCH). Differently from modularized task-oriented systems, this dataset provides a small knowledge base for each dialogue, and thus it can be used to train end-to-end dialogue systems Eric et al. (2017c).
OpenDialKG
Moon et al. (2019) is a human-to-human- collected dataset consisting of four domains: music, sport, book, and movie. The dataset provides a large knowledge graph with 100K entities and 1.1M relations, extracted from freebase, and an annotated entity path that connects the user and the system utterance.
2 Settings
We use DialoGPT Zhang et al. (2019) medium size (345M) as the back-bone model and BERT Devlin et al. (2019) base as the pre-trained classifier for the dialogue manager. We consider two settings of our model: 1) manual-mode, where the user can choose the desired skills or response styles by giving the corresponding adapter skill-id; 2) auto-mode, where the dialogue manager will predict the adapter skill-id by conditioning on the dialogue history. For each adapter, we use bottleneck size , resulting in 9.83M (2.8%) additional parameters. We use batch size 16, learning rate , and early stop according to the performance in the validation sets.
Our model is evaluated using automatic metrics such as Perplexity (Ppl), BLEU, Avg-BLEU, n-gram F1 (F1), entity-F1, distinct n-grams (Dist 1/2/3) in each dataset. The metric used by each task is listed in the results table caption.
3 Results
Table 3, 4, 5, 6, 7 and 8 compare the Adapter-Bot with state-of-the-art models. Overall the Adapter-Bot achieves comparable performance to strong baselines by fine-tuning a fraction of the parameters.
In open-chat tasks, such as ED, WoW, and Per, the Adapter-Bot has a comparable perplexity (1%-3% higher) to large conversational models such as BST Roller et al. (2020) and DDD Shuster et al. (2020b), and lower perplexity than many other existing models. In task-oriented tasks, such as WEA, POI, SCH, in which the meta-knowledge plays a key role, the Adapter-Bot performs slightly worse than task-specific models (e.g., GLMP Wu et al. (2019a) and DFF Qin et al. (2020)) in terms of F1 score, but comparably well in terms of fluency (BLEU). In DialKG, there are no existing baselines for the response generation task; thus we report a fine-tuned GPT-2 baseline, and the Adapter-Bot performs as well as the baseline under the same settings. Finally, for the style/topic generation, we follow the evaluation in Madotto et al. (2020a) in which adapter-based controlled response generation performs better than re-ranking only DialGPT, Weight Decoding Ghazvininejad et al. (2017b) and PPLM Dathathri et al. (2019).
For the dialogue manager, we experimented with single-turn and multiple-turn dialogue history. The BERT-based dialogue manager achieves 95.35% test-set accuracy on the multi-turn and 92.92% test-set accuracy on the single-turn history.
Interactive Systems
To make the system easily accessible, we establish a web-based demo, based on BotUI, for chatting with the Adapter-Bot. The demo supports both manual-mode and auto-mode. In addition to the above-mentioned dialogue skills, we also add multiple features:
Emotion Face Recognition We deploy a javascript-based face emotion recognition model (face-api.js), for monitoring the involuntary reaction of the user while interacting with the Adapter-Bot. This model runs directly on the user browser; thus it does not require sending any images to the server. This is important for two reasons: privacy and real-time performance.
Text Emotion Recognition We deploy a text-based emoji-classifier for text using deepMoji Felbo et al. (2017) (torchMoji). This is used to make the chat-bot more empathetic by showing the corresponding emoji turn by turn in the interface.
Toxic Classifier We deploy a toxic classifier to detect possible offensive responses from the model. The classifier, BERT-base, is trained using the Toxic Comment Classification Dataset and deployed using IBM docker-container.
Covid-19 To show the flexibility of our model, we implement two further skills: Covid-19 QA Su et al. (2020) and Covid-19 fact-checker Lee et al. (2020). The first is accessed with an API-call that, given a question about Covid-19, returns the answer based on a large repository of scientific articles. The second is deploy with an adapter, but instead of training it to generate a response, it is used to score the falseness of a given claim. These two skills can be triggered in manual-mode only, and they show how the same backbone model can also be used to deploy non-dialogue skills.
Localization We deploy a localization feature which is used to obtain weather information without specifying a location. Our system uses the free weather API from rapidAPI, which provides real-time data from more than 40,000 weather stations.
Conclusion
In this paper, we presented the Adapter-Bot, a dialogue model that is built with a fixed pre-trained conversational model and multiple trainable light-weight adapters. The model allows high-level control of different dialogue skills and continuous skills integration. We preliminarily showed 6 goal-oriented skills, 12 response styles, and personalized and emphatic responses. A web-based demo is established to make the system easily accessible.