Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tur
Introduction
Building conversational bots that can interact with humans in natural language (also known as conversational AI) has been of interest to researchers since the early days of computing, as exemplified by text-based systems such as ELIZA . Work on conversational AI generally belongs in one of the following two categories: task-oriented and open-domain. Task-oriented bots aim to help humans accomplish a specific task through multi-turn interactions, whereas open-domain bots aim to serve as social conversation partners with whom humans can have natural and engaging conversations. In addition to mastering traditional language skills like comprehension, open-domain bots (also known as socialbots) need to perfect several conversational skills that come naturally to humans: recalling from world knowledge, reasoning in conjunction with conversational history and constructing valid responses. Socialbots also need to be able to have adequate topical breadth and depth and perform smooth topic transitions.
A critical limiting factor for research into learning these conversational skills is the scarcity of datasets of knowledge-grounded conversations and associated knowledge sources. We introduce Topical-Chat, a dataset of 11K human-human conversations about knowledge spanning 8 broad topics. Figure 1 depicts a conversation snippet from Topical-Chat, with the full conversation available in Appendix B. The dataset was collected by partnering up Amazon Mechanical Turk workers, providing them topical reading sets and asking partners to have naturally coherent and engaging conversations grounded in their provided reading sets. Partners do not have explicitly defined roles they need to serve during a conversation and the reading sets provided to them could be symmetric or asymmetric to varying degrees, which accurately reflects real-world conversations where the world knowledge that both partners gained prior to a conversation may or may not be symmetric. Partners are also asked to annotate each turn of their conversation on several dimensions, such as reading set utilization and sentiment.
In order to create benchmarks for future research with Topical-Chat, we trained several encoder-decoder conversational models on Topical-Chat, each of which aims to generate a response grounded in a reading set and conditioned on conversational history. We specifically leverage the Transformer architecture similar to . We demonstrate the ability of our models to have engaging conversations grounded in knowledge through automated and human evaluation.
Related Work
Recent interest in knowledge-grounded conversations has led to the release of multiple datasets. A dataset of 4K conversations was released , where Wikipedia articles about 30 movies served as the knowledge base. The collection was performed with portions of the articles shown to conversation partners in a scheduled way. A similar dataset of conversations about movies was also released , where the knowledge base comprises Wikipedia articles, reviews and comments mined from the web about 1K movies. The collection involved self-dialogues, where one crowdworker generates utterances for both sides. More recently, the Wizard of Wikipedia (WoW) dataset was released, where the focus, similar to ours, is on collecting open-domain knowledge-grounded conversations. A key difference is their knowledge base comprises Wikipedia articles, whereas we relied on multiple data sources, specifically Washington Post articles and Reddit fun facts in addition to Wikipedia articles about entities, to enable lively interactions.
Sequence-to-sequence generative modeling approaches have become popular for response generation, where the goal is to generate a response given the previous turn in a conversation . However, responses generated by these sequence-to-sequence models are not always coherent or contextually appropriate and are noted to be often generic and lacking interesting content . Such approaches don’t explicitly ground responses on relevant knowledge. This has led to work on approaches that include world knowledge into conversational response generation. End-to-end memory networks have been used to condition the generated responses on knowledge, where attention over the knowledge relevant to the conversation context is estimated, and multiple knowledge representations are included as input during response decoding. Other work retrieves relevant knowledge graphs given the conversation context and encodes the graphs with a static graph attention mechanism. The decoder attentively reads the retrieved knowledge graphs and the knowledge triples within each graph. More recently, a Transformer Memory Network was used to encode knowledge sentences and conversation context and decode a response.
Topical-Chat
Workers on Amazon Mechanical Turk (also known as Turkers) are partnered up and provided highly topical reading sets, and each pair of workers is asked to have a naturally coherent and engaging conversation grounded in their provided reading sets. In our setting, the reading sets provided to conversation partners could be symmetric or have varying degrees of asymmetry, where a pair of reading sets is called symmetric if they contain the exact same information and asymmetric otherwise. This serves as a generalization of the Wizard-Apprentice setting . Unlike most (knowledge-grounded or otherwise) conversation settings , the partners do not have explicitly defined roles they need to serve during their conversation. We leverage information asymmetry to implicitly cause both partners to serve dual roles of a teacher and a participant during their conversation. This setting more accurately reflects real-world conversations, where the world knowledge that both partners have gained prior to a conversation may or may not be symmetric. This makes the Topical-Chat datasetDataset available at www.github.com/alexa/Topical-Chatversatile, realistic and enables the modeling of both partners.
To construct reading sets, we created a knowledge base composed of three primitives: entities, facts and articles.
Entity Selection: We first selected 300 popular entities spanning 8 topics from a prior human-bot conversational dataset collected during a large-scale open-domain socialbot competition between academic research groups . We specifically selected the entities from all user utterances in this prior dataset, since user utterances inform us what users are interested in talking to socialbots about. To maintain topic diversity, we considered the frequency distribution of the 8 topics across all user utterances to allocate an entity budget for each topic (with all budgets adding up to 300). We then picked the top- most frequent entities for each topic . The topics and their respective budgets are provided in Table 1.
Fact Selection: We fetched the Wikipedia lead sections of the 300 entities and crowdsourced 8-10 fun facts for each entity using Reddit . For each entity, we maintained two versions of the fetched Wikipedia lead sections. The first is a shortened version that consists of the first paragraph of the lead section and optionally the second paragraph if the first paragraph contains less than 50 words. The second is a summarized version created by extractively summarizing the entire lead section using TextRank into 150 words or less.
Article Selection: We fetched Washington Post articles from 2018 that each referenced 3 or more of the 300 entities and contained 600-1000 words. We removed articles with profane language and then considered the topic-entity budgets to finalize 3088 articles, ensuring adequate coverage for all topics.
2 Reading Sets Creation
Using the created knowledge base, we construct a pair of reading sets real-time to provide to partners in a conversation. The foundation of a pair of reading sets is an article. For each conversation to be collected, we randomly selected an article from our knowledge base that has not already been used at most 4 times to collect an acceptable conversation. We then apply a random configuration from a pre-defined list of configurations to that article. Configurations are defined to impose varying degrees of information symmetry or asymmetry between partners, leading to the collection of a wide variety of conversations.
Config A: Both Turkers get a Washington Post article and shortened Wikipedia lead sections about the top 3 entities by frequency of occurrence in the article. However, they each get a different set of fun facts about these entities. This enables asymmetry in entity-level fun facts.
Config B: Both Turkers get a Washington Post article and 4-5 fun facts about the top 3 entities by frequency of occurrence in the article. However, one Turker gets shortened Wikipedia lead sections and the other gets summarized Wikipedia lead sections about these entities. This enables asymmetry in entity-level Wikipedia descriptions.
2.2 Symmetric Configurations
Config C: Both Turkers get shortened Wikipedia lead sections and 4-5 fun facts corresponding to the top 3 entities by frequency of occurrence in a Washington Post article. However, the Washington Post article itself is not shown to either Turker. Config D: Both Turkers get a Washington Post article, shortened Wikipedia lead sections and 4-5 fun facts corresponding to the top 3 entities by frequency of occurrence in the article.
3 Conversation Collection
Qualified workers on Mechanical Turk who take up our Human Intelligence Tasks (also known as HITs) are partnered up and provided topical reading sets to read and consequently chat about. The reading sets are also displayed on the Turkers’ screens, near the chat window, during the conversation for reference. All information about an entity E1 (shortened/summarized Wikipedia lead sections and fun facts) are displayed as a group titled Factual Section 1. The Washington Post article about entities E1, E2 and E3 is chunked into 4 similar-sized sections, which are displayed with the titles Article Section 1-4. Turkers qualify for our HITs if their past approved HITs and approval rates are at least 1000 and 99% respectively, ensuring our conversations involve experienced Turkers. We used a customized version of the ParlAI framework to collect conversations. We allow partner Turkers to submit their conversation only if they have conversed for at least 20 turns. At each turn during a conversation, while they are waiting for their partner to respond, we ask each partner to: annotate the sentiment of their message on an 8-point scale (Angry, Disgusted, Fearful, Sad, Happy, Surprised, Curious to Dive Deeper, Neutral), specify the knowledge source used to generate their message (Factual Section 1-3, Article Section 1-4 and/or Personal Knowledge) and rate the quality of their partner’s previous message on a 5-point scale (Poor, Not Good, Passable, Good and Excellent). At the end of a conversation, we ask both partners to rate the quality of the conversation on the same 5-point scale.
We relied on a mix of manual reviewing and automated checks to ensure the conversations we were collecting were acceptable. The automated checks involved computing and verifying our quality metrics were above tuned thresholds. Turkers who had conversations of exceptionally high quality were awarded bonuses. Some statistics about our dataset are shown in Table 2. We created two versions of the validation and test set: frequent and rare. The frequent set contains entities frequently seen in the training set, while the rare set contains entities that were infrequently or never seen in the training set. The presence of multiple entities per conversation by design made it harder to perform a perfect entity-level split of our dataset unlike in , where this is much easier to accomplish since each conversation is associated with a single entity (referred to as a topic in their paper). Appendix A describes our approach.
Models
Let us denote a partial conversation , where for , is the turn in the conversation. Our conversation history is denoted as , which is a flattened sequence of all tokens in . , the ground-truth response at turn , is our target sequence to be predicted for all models. Denote the reading set corresponding to the Turker associated with turn as , which we tokenize into a series of knowledge candidate sentences [], . Denote as a truncate parameter for a knowledge sentence , which retains at most tokens from the start in . Denote as a truncate parameter for a conversation history , which retains at most tokens from the end in .
We train a Transformer with () pairs. During inference, it decodes a response given a conversation history .
2 Transformer with Knowledge
and a selected sentence from [] are encoded with a shared Transformer, concatenated and passed to the Transformer decoder. Knowledge selection in the absence of ground-truth response is an open problem. We currently utilize in the argmax oracle to select , as follows:
and are TF-IDF vectors for and . The TF-IDF vectorizer is learned by sentence-tokenizing all reading sets in Topical-Chat and treating each sentence as a document.
Experiments
All models were trained using ParlAI . Our Transformer contains two layers with two attention heads and a feed-forward hidden layer size of 300 with dropout 0.2. We randomly initialized 300-dimensional word embeddings, which are learned during training. We do not learn positional embeddings and encode position using one-hot vectors. We use a batch size of 32, stochastic gradient descent for optimization with a gradient clip of 0.1 and learning rate scheduler decay 0.5 with patience 3. We stop training when perplexity on the validation frequent set does not decrease for 10 epochs. We use beam search with a beam size of 5 for decoding.
We also experimented with pre-training the Transformer on BookCorpus using a language modeling objective of maximizing the log-likelihood of the next token given a context window of tokens . We use byte-pair encoding (BPE) when pre-training (vocabulary size 37758). When not pre-training, we do not use BPE (vocabulary size 49957).
Results
We use the following acronyms for models for the sake of brevity: TF = Transformer, w/ p.t. = with pre-training, w/ k. = with knowledge. We used a large = 128 when using knowledge, effectively making the parameter irrelevant in our setting since most knowledge sentences have fewer than 128 tokens. In order to decide on an appropriate , we tried training a Transformer that uses knowledge with varying and evaluated them on automated metrics described below (Table 5). We observe that = 32 works best. We believe this reflects our knowledge model’s inability to attend to important tokens in the dialog context when a large is used. Consequently, we used = 32 in Tables 3 and 4.
For automated evaluation, we consider metrics such as perplexity (PPL), unigram F1 of model prediction with ground-truth response and -gram diversity (Div.) . In Table 3, we observe that all our models have high unigram and bigram diversity, demonstrating that the models learn to decode responses that are lexically informative and diverse. We also observe an improvement in unigram F1 and increase in PPL when knowledge is used.
We performed human evaluation of our models by first creating 150 evaluation snippets, each comprising {, , []}, = , where [] is a set of responses ( from trained models and one ground-truth response ) given a partial conversation and selected sentence . The partial conversation corresponding to each snippet came from a distinct conversation in the Topical-Chat test frequent set. For each in each snippet, we asked two humans to separately annotate (possible values in parentheses) whether is comprehensible (0/1), on-topic (0/1) and interesting (0/1). We also asked them to annotate how effectively is utilized in (0-3) and if they would have liked to continue the conversation after (0/1). We computed Cohen’s kappa for binary and Fleiss’ kappa for nominal-scale annotations as measures of reliability of agreement and observed poor agreement for interesting (0.29) and continue conversation (0.27). Consequently, we aggregate and report mean annotation scores for parameters with high agreement in Table 4. We use the following acronyms for the sake of brevity: comprehensible = comp., on-topic = o.t., leverage knowledge = l.k. We observe that all models are rated to mostly produce comprehensible responses and the models that ingest knowledge are rated to produce responses that leverage them, albeit only somewhat effectively.
Conclusion
We introduce Topical-Chat, an open-domain knowledge-grounded conversation dataset without explicit roles for conversation partners and containing depth and breadth of topical coverage with transitions in conversations. We train simple Transformer-based models for response generation and evaluate them using automated metrics for benchmarking. We also provide evidence of qualitative value through human evaluation of these models. We hope that the release of Topical-Chat fosters data-driven research in open-domain knowledge-grounded conversational AI.
References
Appendix A Valid / Test Set Creation Strategy
We adopt a greedy strategy for creating the validation and test splits from the collected data. We first create a list of all entity triplets for all collected conversations. We define entity frequency as the number of triplets containing an entity and consequently assign a score for each triplet as the sum of its constituent entity frequencies. We sort the list of all triplets in increasing order of their scores, extract the top 10% of the triplets and create two equal or off-by-one sized partitions (validation rare) and (test rare) of the corresponding conversations. Next, we randomly select 80% of the triplets and create partition (train) of the corresponding conversations. Finally, we create two equal or off-by-one sized partitions (validation frequent) and (test frequent) of the remaining conversations.