Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading

Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, Jianfeng Gao

Introduction

While end-to-end neural conversation models (Shang et al., 2015; Sordoni et al., 2015; Vinyals and Le, 2015; Serban et al., 2016; Li et al., 2016a; Gao et al., 2019a, etc.) are effective in learning how to be fluent, their responses are often vacuous and uninformative. A primary challenge thus lies in modeling what to say to make the conversation contentful. Several recent approaches have attempted to address this difficulty by conditioning the language decoder on external information sources, such as knowledge bases Agarwal et al. (2018); Liu et al. (2018a), review posts Ghazvininejad et al. (2018); Moghe et al. (2018), and even images Das et al. (2017); Mostafazadeh et al. (2017). However, empirical results suggest that conditioning the decoder on rich and complex contexts, while helpful, does not on its own provide sufficient inductive bias for these systems to learn how to achieve deep and accurate integration between external knowledge and response generation.

We posit that this ongoing challenge demands a more effective mechanism to support on-demand knowledge integration. We draw inspiration from how humans converse about a topic, where people often search and acquire external information as needed to continue a meaningful and informative conversation. Figure 1 illustrates an example human discussion, where information scattered in separate paragraphs must be consolidated to compose grounded and appropriate responses. Thus, the challenge is to connect the dots across different pieces of information in much the same way that machine reading comprehension (MRC) systems tie together multiple text segments to provide a unified and factual answer (Seo et al., 2017, etc.).

We introduce a new framework of end-to-end conversation models that jointly learn response generation together with on-demand machine reading. We formulate the reading comprehension task as document-grounded response generation: given a long document that supplements the conversation topic, along with the conversation history, we aim to produce a response that is both conversationally appropriate and informed by the content of the document. The key idea is to project conventional QA-based reading comprehension onto conversation response generation by equating the conversation prompt with the question, the conversation response with the answer, and external knowledge with the context. The MRC framing allows for integration of long external documents that present notably richer and more complex information than relatively small collections of short, independent review posts such as those that have been used in prior work (Ghazvininejad et al., 2018; Moghe et al., 2018).

We also introduce a large dataset to facilitate research on knowledge-grounded conversation (2.8M turns, 7.4M sentences of grounding) that is at least one order of magnitude larger than existing datasets (Dinan et al., 2019; Moghe et al., 2018). This dataset consists of real-world conversations extracted from Reddit, linked to web documents discussed in the conversations. Empirical results on our new dataset demonstrate that our full model improves over previous grounded response generation systems and various ungrounded baselines, suggesting that deep knowledge integration is an important research direction.Code for reproducing our models and data is made publicly available at https://github.com/qkaren/converse_reading_cmr.

Task

We propose to use factoid- and entity-rich web documents, e.g., news stories and Wikipedia pages, as external knowledge sources for an open-ended conversational system to ground in.

Formally, we are given a conversation history of turns X=(x1,…,xM)X=(\bm{x}_{1},\dots,\bm{x}_{M}) and a web document D=(s1,…,sN)D=(\bm{s}_{1},\dots,\bm{s}_{N}) as the knowledge source, where si\bm{s}_{i} is the iith sentence in the document. With the pair (X,D)(X,D), the system needs to generate a natural language response y\bm{y} that is both conversationally appropriate and reflective of the contents of the web document.

Approach

Our approach integrates conversation generation with on-demand MRC. Specifically, we use an MRC model to effectively encode the conversation history by treating it as a question in a typical QA task (e.g., SQuAD (Rajpurkar et al., 2016)), and encode the web document as the context. We then replace the output component of the MRC model (which is usually an answer classification module) with an attentional sequence generator that generates a free-form response. We refer to our approach as CMR (Conversation with on-demand Machine Reading). In general, any off-the-shelf MRC model could be applied here for knowledge comprehension. We use Stochastic Answer Networks (SAN)https://github.com/kevinduh/san_mrc (Liu et al., 2018b), a performant machine reading model that until very recently held state-of-the-art performance on the SQuAD benchmark. We also employ a simple but effective data weighting scheme to further encourage response grounding.

We adapt the SAN model to encode both the input document and conversation history and forward the digested information to a response generator. Figure 2 depicts the overall MRC architecture. Different blocks capture different concepts of representations in both the input conversation history and web document. The leftmost blocks represent the lexicon encoding that extracts information from XX and DD at the token level. Each token is first transformed into its corresponding word embedding vector, and then fed into a position-wise feed-forward network (FFN) Vaswani et al. (2017) to obtain the final token-level representation. Separate FFNs are used for the conversation history and the web document.

The next block is for contextual encoding. The aforementioned token vectors are concatenated with pre-trained 600-dimensional CoVe vectors McCann et al. (2017), and then fed to a BiLSTM that is shared for both conversation history and web document. The step-wise outputs of the BiLSTM carry the information of the tokens as well as their left and right context.

2 Response Generation

Having read and processed both the conversation history and the extra knowledge in the document, the model then produces a free-form response y=(y1,…,yT){\bf y}=(y_{1},\dots,y_{T}) instead of generating a span or performing answer classification as in MRC tasks.

We use an attentional recurrent neural network decoder (Luong et al., 2015) to generate response tokens while attending to the memory. At the beginning, the initial hidden state h0\bm{h}_{0} is the weighted sum of the representation of the history XX. For each decoding step tt with a hidden state ht\bm{h}_{t}, we generate a token yty_{t} based on the distribution:

where τ>0\tau>0 is the softmax temperature. The hidden state ht\bm{h}_{t} is defined as follows:

Here, [⋅+ ⁣ ⁣+⋅][\cdot+\!\!+\cdot] indicates a concatenation of two vectors; fattentionf_{attention} is a dot-product attention Vaswani et al. (2017); and zt\bm{z}_{t} is a state generated by GRU(et−1,ht−1)\text{GRU}(\bm{e}_{t-1},\bm{h}_{t-1}) with et−1\bm{e}_{t-1} being the embedding of the word yt−1y_{t-1} generated at the previous (t−1t-1) step. In practice, we use top-kk sample decoding to draw yty_{t} from the above distribution p(yt)p(y_{t}). Section 5 provides more details about the experimental configuration.

3 Data Weighting Scheme

Dataset

To create a grounded conversational dataset, we extract conversation threads from Reddit, a popular and large-scale online platform for news and discussion. In 2015 alone, Reddit hosted more than 73M conversations.https://redditblog.com/2015/12/31/reddit-in-2015/ On Reddit, user submissions are categorized by topics or “subreddits”, and a submission typically consists of a submission title associated with a URL pointing to a news or background article, which initiates a discussion about the contents of the article. This article provides framing for the conversation, and this can naturally be seen as a form of grounding. Another factor that makes Reddit conversations particularly well-suited for our conversation-as-MRC setting is that a significant proportion of these URLs contain named anchors (i.e., ‘#’ in the URL) that point to the relevant passages in the document. This is conceptually quite similar to MRC data (Rajpurkar et al., 2016) where typically only short passages within a larger document are relevant in answering the question.

We reduce spamming and offensive language by manually curating a list of 178 relatively “safe” subreddits and 226 web domains from which the web pages are extracted. To convert the web page of each conversation into a text document, we extracted the text of the page using an html-to-text converter,https://www.crummy.com/software/BeautifulSoup while retaining important tags such as , <h1> to <h6>, and <p>. This means the entire text of the original web page is preserved, but these main tags retain some high-level structure of the article. For web URLs with named anchors, we preserve that information by indicating the anchor text in the document with tags <anchor> and </anchor>. As the whole documents in the dataset tend to be lengthy, anchors offer important hints to the model about which parts of the documents should likely be focused on in order to produce a good response. We considered it sensible to keep them as they are also available to the human reader.</p><p class="leading-[1.75]">After filtering short or redacted turns, or which quote earlier turns, we obtained 2.8M conversation instances respectively divided into train, validation, and test (Table 1). We used different date ranges for these different sets: years 2011-2016 for train, Jan-Mar 2017 for validation, and the rest of 2017 for test. For the test set, we select conversational turns for which 6 or more responses were available, in order to create a multi-reference test set. Given other filtering criteria such as turn length, this yields a 6-reference test set of size 2208. For each instance, we set aside one of the 6 human responses to assess human performance on this task, and the remaining 5 responses serve as ground truths for evaluating different systems.While this is already large for a grounded dataset, we could have easily created a much bigger one given how abundant Reddit data is. We focused instead on filtering out spamming and offensive language, in order to strike a good balance between data quality and size. Table 1 provides statistics for our dataset, and Figure 1 presents an example from our dataset that also demonstrates the need to combine conversation history and background information from the document to produce an informative response.</p><p class="leading-[1.75]">To enable reproducibility of our experiments, we crawled web pages using Common Crawl (http://commoncrawl.org), a service that crawls web pages and makes its historical crawls available to the public. We also release the code (URL redacted for anonymity) to recreate our dataset from both a popular Reddit dumphttp://files.pushshift.io/reddit/ and Common Crawl, and the latter service ensures that anyone reproducing our data extraction experiments would retrieve exactly the same web pages. We made a preliminary version of this dataset available for a shared task Galley et al. (2019) at Dialog System Technology Challenges (DSTC) Yoshino et al. (2019). Back-and-forth with participants helped us iteratively refine the dataset. The code to recreate this dataset is included.We do not report on shared task systems here, as these systems do not represent our work and some of these systems have no corresponding publications. Along with the data described here, we provided a standard Seq2Seq baseline to the shared task, which we improved for the purpose of this paper (improved Bleu, Nist and Meteor). Our new Seq2Seq baseline is described in Section 5.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Experiments</h2><p class="leading-[1.75]">We evaluate our systems and several competitive baselines: Seq2Seq (Sutskever et al., 2014) We use a standard LSTM Seq2Seq model that only exploit the conversation history for response generation, without any grounding. This is a competitive baseline initialized using pretrained embeddings. MemNet: We use a Memory Network designed for grounded response generation (Ghazvininejad et al., 2018). An end-to-end memory network (Sukhbaatar et al., 2015) encodes conversation history and sentences in the web documents. Responses are generated with a sequence decoder. CMR-f : To directly measure the effect of incorporating web documents, we compare to a baseline which omits the document reading component of the full model (Figure 2). As with the Seq2Seq approach, the resulting model generates responses solely based on conversation history. CMR: To measure the effect of our data weighting scheme, we compare to a system that has identical architecture to the full model, but is trained without associating weights to training instances. CMR+w: As described in section 3, the full model reads and comprehends both the conversation history and document using an MRC component, and sequentially generates the response. The model is trained with the data weighting scheme to encourage grounded responses. Human: To get a better sense of the systems’ performance relative to an upper bound, we also evaluate human-written responses using different metrics. As described in Section 4, for each test instance, we set aside one of the 6 human references for evaluation, so the ‘human’ is evaluated against the other 5 references for automatic evaluation. To make these results comparable, all the systems are also automatically evaluated against the same 5 references.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Experiment Details</h2><p class="leading-[1.75]">For all the systems, we set word embedding dimension to 300 and used the pretrained GloVehttps://nlp.stanford.edu/projects/glove/ for initialization. We set hidden dimensions to 512 and dropout rate to 0.4. GRU cells are used for Seq2Seq and MemNet (we also tested LSTM cells and obtained similar results). We used the Adam optimizer for model training, with an initial learning rate of 0.0005. Batch size was set to 32. During training, all responses were truncated to have a maximum length of 30, and maximum query length and document length were set to 30, 500, respectively. we used regular teacher-forcing decoding during training. For inference, we found that top-<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> random sample decoding Fan et al. (2018) provides the best results for all the systems. That is, at each decoding step, a token was drawn from the <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> most likely candidates according to the distribution over the vocabulary. Similar to recent work (Fan et al., 2018; Edunov et al., 2018), we set <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi><mo>=</mo><mn>20</mn></mrow><annotation encoding="application/x-tex">k=20</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">20</span></span></span></span> (other common <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> values like 10 gave similar results). We selected key hyperparameter configurations on the validation set.</p><p class="leading-[1.75]">Table 2 shows automatic metrics for quantitative evaluation over three qualities of generated texts. We measure the overall relevance of the generated responses given the conversational history by using standard Machine Translation (MT) metrics, comparing generated outputs to ground-truth responses. These metrics include Bleu-4 (Papineni et al., 2002), Meteor (Lavie and Agarwal, 2007). and Nist (Doddington, 2002). The latter metric is a variant of Bleu that weights <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-gram matches by their information gain by effectively penalizing uninformative <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-grams (such as “I don’t know”), which makes it a relevant metric for evaluating systems aiming diverse and informative responses. MT metrics may not be particularly adequate for our task Liu et al. (2016), given its focus on the informativeness of responses, and for that reason we also use two other types of metrics to measure the level of grounding and diversity.</p><p class="leading-[1.75]">As a diversity metric, we count all <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-grams in the system output for the test set, and measure: (1) Entropy-<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span> as the entropy of the <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-gram count distribution, a metric proposed in Zhang et al. (2018b); (2) Distinct-<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span> as the ratio between the number of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-gram types and the total number of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>-grams, a metric introduced in Li et al. (2016a).</p><p class="leading-[1.75]">For the grounding metrics, we first compute ‘#match,’ the number of non-stopword tokens in the response that are present in the document but not present in the context of the conversation. Excluding words from the conversation history means that, in order to produce a word of the document, the response generation system is very likely to be effectively influenced by that document. We then compute both precision as ‘#match’ divided by the total number of non-stop tokens in the response, and recall as ‘#match’ divided by the total number of non-stop tokens in the document. We also compute the respective F1 score to combine both. Looking only at exact unigram matches between the document and response is a major simplifying assumption, but the combination of the three metrics offers a plausible proxy for how greatly the response is grounded in the document. It seems further reasonable to assume that these can serve as a surrogate for less quantifiable forms of grounding such as paraphrase – e.g., US <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover><mo stretchy="true" minsize="3.0em">→</mo><mpadded width="+0.6em" lspace="0.3em"><mrow></mrow></mpadded></mover></mrow><annotation encoding="application/x-tex">\xrightarrow{}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.633em;vertical-align:-0.011em;"></span><span class="mrel x-arrow"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.622em;"><span style="top:-3.144em;"><span class="pstrut" style="height:2.522em;"></span><span class="sizing reset-size6 size3 mtight x-arrow-pad"><span class="mord mtight"></span></span></span><span class="svg-align" style="top:-2.511em;"><span class="pstrut" style="height:2.522em;"></span><span class="hide-tail" style="height:0.522em;min-width:1.469em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.522em" viewBox="0 0 400000 522" preserveAspectRatio="xMaxYMin slice"><path d="M0 241v40h399891c-47.3 35.3-84 78-110 128 -16.7 32-27.7 63.7-33 95 0 1.3-.2 2.7-.5 4-.3 1.3-.5 2.3-.5 3 0 7.3 6.7 11 20 11 8 0 13.2-.8 15.5-2.5 2.3-1.7 4.2-5.5 5.5-11.5 2-13.3 5.7-27 11-41 14.7-44.7 39-84.5 73-119.5s73.7-60.2 119-75.5c6-2 9-5.7 9-11s-3-9-9-11c-45.3-15.3-85 -40.5-119-75.5s-58.3-74.8-73-119.5c-4.7-14-8.3-27.3-11-40-1.3-6.7-3.2-10.8-5.5 -12.5-2.3-1.7-7.5-2.5-15.5-2.5-14 0-21 3.7-21 11 0 2 2 10.3 6 25 20.7 83.3 67 151.7 139 205zm0 0v40h399900v-40z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.011em;"><span></span></span></span></span></span></span></span></span> American – when the statistics are aggregated on a large test dataset.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">2 Automatic Evaluation</h2><p class="leading-[1.75]">Table 2 shows automatic evaluation results for the different systems. In terms of appropriateness, the different variants of our models outperform the Seq2Seq and MemNet baselines, but differences are relatively small and, in case of one of the metrics (Nist), the best system does not use grounding. Our goal, we would note, is not to specifically improve response appropriateness, as many responses that completely ignore the document (e.g., I don’t know) might be perfectly appropriate. Our systems fare much better in terms of Grounding and Diversity: our best system (CMR+w) achieves an F1 score that is more than three times (0.38% vs. 0.12%) higher than the most competitive non-MRC system (MemNet).</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">3 Human Evaluation</h2><p class="leading-[1.75]">We sampled 1000 conversations from the test set. Filters were applied to remove conversations containing ethnic slurs or other offensive content that might confound judgments. Outputs from systems to be compared were presented pairwise to judges from a crowdsourcing service. Four judges were asked to compare each pair of outputs on Relevance (the extent to which the content was related to and appropriate to the conversation) and Informativeness (the extent to which the output was interesting and informative). Judges were asked to agree or disagree with a statement that one of the pair was better than the other on the above two parameters, using a 5-point Likert scale.The choices presented to the judges were Strongly Agree, Agree, Neutral, Disagree, and Strongly Disagree. Pairs of system outputs were randomly presented to the judges in random order in the context of short snippets of the background text. These results are presented in summary form in Table 3, which shows the overall preferences for the two systems expressed as a percentage of all judgments made. Overall inter-rater agreement measured by Fliess’ Kappa was 0.32 (“fair"). Nevertheless, the differences between the paired model outputs are statistically significant (computed using 10,000 bootstrap replications).</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">4 Qualitative Study</h2><p class="leading-[1.75]">Table 4 illustrates how our best model (CMR+w) tends to produce more contentful and informative responses compared to the other systems. In the first example, our system refers to a particular episode mentioned in the article, and also uses terminology that is more consistent with the article (e.g., series). In the second example, humorous song seems to positively influence the response, which is helpful as the input doesn’t mention singing at all. In the third example, the CMR+w model clearly grounds its response to the article as it states the fact (Steve Jobs: CEO of Apple) retrieved from the article. The outputs by the other two baseline models are instead not relevant in the context.</p><p class="leading-[1.75]">Figure 3 displays the attention map of the generated response and (part of) the document from our full model. The model successfully attends to the key words (e.g., 36th, episode) of the document. Note that the attention map is unlike what is typical in machine translation, where target words tend to attend to different portions of the input text. In our task, where alignments are much less one-to-one compared to machine translation, it is common for the generator to retain focus on the key information in the external document to produce semantically relevant responses.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Related Work</h2><p class="leading-[1.75]">Traditional dialogue systems (see Jurafsky and Martin (2009) for an historical perspective) are typically grounded, enabling these systems to be reflective of the user’s environment. The lack of grounding has been a stumbling block for the earliest end-to-end dialogue systems, as various researchers have noted that their outputs tend to be bland Li et al. (2016a); Gao et al. (2019b), inconsistent Zhang et al. (2018a); Li et al. (2016b); Zhang et al. (2019), and lacking in factual content Ghazvininejad et al. (2018); Agarwal et al. (2018). Recently there has been growing interest in exploring different forms of grounding, including images, knowledge bases, and plain texts Das et al. (2017); Mostafazadeh et al. (2017); Agarwal et al. (2018); Yang et al. (2019). A recent survey is included in Gao et al. (2019a).</p><p class="leading-[1.75]">Prior work, e.g, (Ghazvininejad et al., 2018; Zhang et al., 2018a; Huang et al., 2019), uses grounding in the form of independent snippets of text: Foursquare tips and background information about a given speaker. Our notion of grounding is different, as our inputs are much richer, encompassing the full text of a web page and its underlying structure. Our setting also differs significantly from relatively recent work (Dinan et al., 2019; Moghe et al., 2018) exploiting crowdsourced conversations with detailed grounding labels: we use Reddit because of its very large scale and better characterization of real-world conversations. We also require the system to learn grounding directly from conversation and document pairs, instead of relying on additional grounding labels. Moghe et al. (2018) explored directly using a span-prediction QA model for conversation. Our framework differs in that we combine MRC models with a sequence generator to produce free-form responses.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Machine Reading Comprehension:</h2><p class="leading-[1.75]">MRC models such as SQuAD-like models, aim to extract answer spans (starting and ending indices) from a given document for a given question Seo et al. (2017); Liu et al. (2018b); Yu et al. (2018). These models differ in how they fuse information between questions and documents. We chose SAN Liu et al. (2018b) because of its representative architecture and competitive performance on existing MRC tasks. We note that other off-the-shelf MRC models, such as BERT (Devlin et al., 2018), can also be plugged in. We leave the study of different MRC architectures for future work. Questions are treated as entirely independent in these “single-turn” MRC models, so recent work (e.g., CoQA Reddy et al. (2019) and QuAC Choi et al. (2018)) focuses on multi-turn MRC, modeling sequences of questions and answers in a conversation. While multi-turn MRC aims to answer complex questions, that body of work is restricted to factual questions, whereas our work—like much of the prior work in end-to-end dialogue—models free-form dialogue, which also encompasses chitchat and non-factual responses.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Conclusions</h2><p class="leading-[1.75]">We have demonstrated that the machine reading comprehension approach offers a promising step to generating, on the fly, contentful conversation exchanges that are grounded in extended text corpora. The functional combination of MRC and neural attention mechanisms offers visible gains over several strong baselines. We have also formally introduced a large dataset that opens up interesting challenges for future research.</p><p class="leading-[1.75]">The CMR (Conversation with on-demand machine reading) model presented here will help connect the many dots across multiple data sources. One obvious future line of investigation will be to explore the effect of other off-the-shelf machine reading models such as BERT Devlin et al. (2018) within the CMR framework.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">Acknowledgements</h2><p class="leading-[1.75]">We are grateful to the anonymous reviewers, as well as to Vighnesh Shiv, Yizhe Zhang, Chris Quirk, Shrimai Prabhumoye, and Ziyu Yao for helpful comments and suggestions on this work. This research was supported in part by NSF (IIS- 1524371), DARPA CwC through ARO (W911NF- 15-1-0543), and Samsung AI Research.</p><h2 class="font-display text-xl font-medium text-[#fcfdff] mt-8 mb-2" style="letter-spacing:-0.3px">References</h2></div></article><aside class="w-72 shrink-0 hidden lg:block"><div class="sticky top-20 space-y-8"><div class="space-y-3"><p class="text-[10px] font-medium tracking-widest text-[#888e90] uppercase">Annotations (<!-- -->0<!-- -->)</p><div class="rounded-xl p-4 space-y-1" style="background:#0a0a0c;border:1px solid rgba(255,255,255,0.06)"><p class="text-[12px] text-[#888e90]">No annotations yet.</p><p class="text-[11px] text-[#464a4d]">Select any passage to add the first one.</p></div></div></div></aside></div><footer class="px-4 sm:px-8 py-10" style="border-top:1px solid rgba(255,255,255,0.06)"><div class="max-w-5xl mx-auto flex flex-col sm:flex-row items-start sm:items-center justify-between gap-6"><div class="flex items-center gap-6"><a class="font-display text-sm font-medium text-[#fcfdff] hover:opacity-80 transition-opacity" href="/">paper7</a><a class="text-[12px] text-[#464a4d] hover:text-[#888e90] transition-colors" href="/">papers</a><a class="text-[12px] text-[#464a4d] hover:text-[#888e90] transition-colors" href="/feed">following</a><a class="text-[12px] text-[#464a4d] hover:text-[#888e90] transition-colors" href="/opensource">open source</a></div><div class="flex items-center gap-4"><a href="https://github.com/p7dotorg" target="_blank" rel="noopener noreferrer" class="text-[12px] text-[#464a4d] hover:text-[#888e90] transition-colors">GitHub</a></div></div></footer></div><!--$--><!--/$--><script src="/_next/static/chunks/0xyj7g8po3_e2.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C" id="_R_" async=""></script><script>(self.__next_f=self.__next_f||[]).push([0])</script><script>self.__next_f.push([1,"1:\"$Sreact.fragment\"\n3:I[39756,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"default\"]\n4:I[37457,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"default\"]\n6:I[97367,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"OutletBoundary\"]\n7:\"$Sreact.suspense\"\na:I[97367,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"ViewportBoundary\"]\nc:I[97367,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"MetadataBoundary\"]\ne:I[68027,[\"/_next/static/chunks/3y9fbee6ilrjr.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/1kve9pjlan_ev.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"default\",1]\n:HL[\"/_next/static/chunks/01p-yq1wd4g7c.css?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"style\"]\n:HL[\"/_next/static/media/2a65768255d6b625-s.p.3u4lli0-axodc.woff2?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"font\",{\"crossOrigin\":\"\",\"type\":\"font/woff2\"}]\n:HL[\"/_next/static/media/70e3db2de7f94926-s.p.39pl-v7c3qrze.woff2?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"font\",{\"crossOrigin\":\"\",\"type\":\"font/woff2\"}]\n:HL[\"/_next/static/media/caa3a2e1cccd8315-s.p.0wgildi0cnwt9.woff2?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"font\",{\"crossOrigin\":\"\",\"type\":\"font/woff2\"}]\n:HL[\"/_next/static/chunks/1hzpfqbo46xnf.css?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"style\"]\n"])</script><script>self.__next_f.push([1,"0:{\"P\":null,\"c\":[\"\",\"1906.02738\"],\"q\":\"\",\"i\":false,\"f\":[[[\"\",{\"children\":[[\"arxivId\",\"1906.02738\",\"d\",null],{\"children\":[\"__PAGE__\",{}]}]},\"$undefined\",\"$undefined\",16],[[\"$\",\"$1\",\"c\",{\"children\":[[[\"$\",\"link\",\"0\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/01p-yq1wd4g7c.css?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"$undefined\"}],[\"$\",\"script\",\"script-0\",{\"src\":\"/_next/static/chunks/3y9fbee6ilrjr.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"async\":true,\"nonce\":\"$undefined\"}],[\"$\",\"script\",\"script-1\",{\"src\":\"/_next/static/chunks/1kve9pjlan_ev.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"async\":true,\"nonce\":\"$undefined\"}]],\"$L2\"]}],{\"children\":[[\"$\",\"$1\",\"c\",{\"children\":[null,[\"$\",\"$L3\",null,{\"parallelRouterKey\":\"children\",\"error\":\"$undefined\",\"errorStyles\":\"$undefined\",\"errorScripts\":\"$undefined\",\"template\":[\"$\",\"$L4\",null,{}],\"templateStyles\":\"$undefined\",\"templateScripts\":\"$undefined\",\"notFound\":\"$undefined\",\"forbidden\":\"$undefined\",\"unauthorized\":\"$undefined\"}]]}],{\"children\":[[\"$\",\"$1\",\"c\",{\"children\":[\"$L5\",[[\"$\",\"link\",\"0\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/1hzpfqbo46xnf.css?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"$undefined\"}],[\"$\",\"script\",\"script-0\",{\"src\":\"/_next/static/chunks/18w4ci49njz81.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"async\":true,\"nonce\":\"$undefined\"}],[\"$\",\"script\",\"script-1\",{\"src\":\"/_next/static/chunks/3evshbwpjzg9r.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"async\":true,\"nonce\":\"$undefined\"}]],[\"$\",\"$L6\",null,{\"children\":[\"$\",\"$7\",null,{\"name\":\"Next.MetadataOutlet\",\"children\":\"$@8\"}]}]]}],{},null,false,null]},null,false,\"$@9\"]},null,false,null],[\"$\",\"$1\",\"h\",{\"children\":[null,[\"$\",\"$La\",null,{\"children\":\"$Lb\"}],[\"$\",\"div\",null,{\"hidden\":true,\"children\":[\"$\",\"$Lc\",null,{\"children\":[\"$\",\"$7\",null,{\"name\":\"Next.Metadata\",\"children\":\"$Ld\"}]}]}],[\"$\",\"meta\",null,{\"name\":\"next-size-adjust\",\"content\":\"\"}]]}],false]],\"m\":\"$undefined\",\"G\":[\"$e\",[[\"$\",\"link\",\"0\",{\"rel\":\"stylesheet\",\"href\":\"/_next/static/chunks/01p-yq1wd4g7c.css?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"precedence\":\"next\",\"crossOrigin\":\"$undefined\",\"nonce\":\"$undefined\"}]]],\"S\":false,\"h\":null,\"s\":\"$undefined\",\"l\":\"$undefined\",\"p\":\"$undefined\",\"d\":\"$undefined\"}\n"])</script><script>self.__next_f.push([1,"f:[]\n9:\"$Wf\"\n"])</script><script>self.__next_f.push([1,"2:[\"$\",\"html\",null,{\"lang\":\"en\",\"className\":\"geist_9c6cb61b-module__8NX9hq__variable playfair_display_d1a90cca-module__u1YQ1W__variable h-full antialiased\",\"children\":[\"$\",\"body\",null,{\"className\":\"min-h-full flex flex-col bg-black text-[#fcfdff]\",\"children\":[\"$\",\"$L3\",null,{\"parallelRouterKey\":\"children\",\"error\":\"$undefined\",\"errorStyles\":\"$undefined\",\"errorScripts\":\"$undefined\",\"template\":[\"$\",\"$L4\",null,{}],\"templateStyles\":\"$undefined\",\"templateScripts\":\"$undefined\",\"notFound\":[[[\"$\",\"title\",null,{\"children\":\"404: This page could not be found.\"}],[\"$\",\"div\",null,{\"style\":{\"fontFamily\":\"system-ui,\\\"Segoe UI\\\",Roboto,Helvetica,Arial,sans-serif,\\\"Apple Color Emoji\\\",\\\"Segoe UI Emoji\\\"\",\"height\":\"100vh\",\"textAlign\":\"center\",\"display\":\"flex\",\"flexDirection\":\"column\",\"alignItems\":\"center\",\"justifyContent\":\"center\"},\"children\":[\"$\",\"div\",null,{\"children\":[[\"$\",\"style\",null,{\"dangerouslySetInnerHTML\":{\"__html\":\"body{color:#000;background:#fff;margin:0}.next-error-h1{border-right:1px solid rgba(0,0,0,.3)}@media (prefers-color-scheme:dark){body{color:#fff;background:#000}.next-error-h1{border-right:1px solid rgba(255,255,255,.3)}}\"}}],[\"$\",\"h1\",null,{\"className\":\"next-error-h1\",\"style\":{\"display\":\"inline-block\",\"margin\":\"0 20px 0 0\",\"padding\":\"0 23px 0 0\",\"fontSize\":24,\"fontWeight\":500,\"verticalAlign\":\"top\",\"lineHeight\":\"49px\"},\"children\":404}],[\"$\",\"div\",null,{\"style\":{\"display\":\"inline-block\"},\"children\":[\"$\",\"h2\",null,{\"style\":{\"fontSize\":14,\"fontWeight\":400,\"lineHeight\":\"49px\",\"margin\":0},\"children\":\"This page could not be found.\"}]}]]}]}]],[]],\"forbidden\":\"$undefined\",\"unauthorized\":\"$undefined\"}]}]}]\n"])</script><script>self.__next_f.push([1,"b:[[\"$\",\"meta\",\"0\",{\"charSet\":\"utf-8\"}],[\"$\",\"meta\",\"1\",{\"name\":\"viewport\",\"content\":\"width=device-width, initial-scale=1\"}]]\n"])</script><script>self.__next_f.push([1,"10:I[27201,[\"/_next/static/chunks/05-c3ty_6dwfk.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/14mrh2-p_w84d.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"IconMark\"]\n8:null\nd:[[\"$\",\"title\",\"0\",{\"children\":\"Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading — p7\"}],[\"$\",\"meta\",\"1\",{\"name\":\"description\",\"content\":\"Highlight passages in research papers, annotate what matters, discuss with other readers.\"}],[\"$\",\"meta\",\"2\",{\"property\":\"og:title\",\"content\":\"paper7 — Annotate academic papers\"}],[\"$\",\"meta\",\"3\",{\"property\":\"og:description\",\"content\":\"Highlight passages in research papers, annotate what matters, discuss with other readers.\"}],[\"$\",\"meta\",\"4\",{\"property\":\"og:site_name\",\"content\":\"paper7\"}],[\"$\",\"meta\",\"5\",{\"name\":\"twitter:card\",\"content\":\"summary\"}],[\"$\",\"meta\",\"6\",{\"name\":\"twitter:title\",\"content\":\"paper7 — Annotate academic papers\"}],[\"$\",\"meta\",\"7\",{\"name\":\"twitter:description\",\"content\":\"Highlight passages in research papers, annotate what matters, discuss with other readers.\"}],[\"$\",\"link\",\"8\",{\"rel\":\"icon\",\"href\":\"/icon?4955f7db4679c841\",\"alt\":\"$undefined\",\"type\":\"image/png\",\"sizes\":\"32x32\"}],[\"$\",\"$L10\",\"9\",{}]]\n"])</script><script>self.__next_f.push([1,"11:I[27997,[\"/_next/static/chunks/3y9fbee6ilrjr.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/1kve9pjlan_ev.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/18w4ci49njz81.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\",\"/_next/static/chunks/3evshbwpjzg9r.js?dpl=dpl_FDPxwzq5JWNpfYJqSkE537ffVE7C\"],\"default\"]\n12:T441,Although neural conversation models are effective in learning how to produce fluent responses, their primary challenge lies in knowing what to say to make the conversation contentful and non-vacuous. We present a new end-to-end approach to contentful neural conversation that jointly models response generation and on-demand machine reading. The key idea is to provide the conversation model with relevant long-form text on the fly as a source of external knowledge. The model performs QA-style reading comprehension on this text in response to each conversational turn, thereby allowing for more focused integration of external knowledge than has been possible in prior approaches. To support further research on knowledge-grounded conversation, we introduce a new large-scale conversation dataset grounded in external web pages (2.8M turns, 7.4M sentences of grounding). Both human evaluation and automated metrics show that our approach results in more contentful responses compared to a variety of previous methods, improving both the informativeness and diversity of generated output.13:T5fe3,"])</script><script>self.__next_f.push([1,"## Introduction\n\nWhile end-to-end neural conversation models (Shang et al., 2015; Sordoni et al., 2015; Vinyals and Le, 2015; Serban et al., 2016; Li et al., 2016a; Gao et al., 2019a, etc.) are effective in learning how to be fluent, their responses are often vacuous and uninformative. A primary challenge thus lies in modeling what to say to make the conversation contentful. Several recent approaches have attempted to address this difficulty by conditioning the language decoder on external information sources, such as knowledge bases Agarwal et al. (2018); Liu et al. (2018a), review posts Ghazvininejad et al. (2018); Moghe et al. (2018), and even images Das et al. (2017); Mostafazadeh et al. (2017). However, empirical results suggest that conditioning the decoder on rich and complex contexts, while helpful, does not on its own provide sufficient inductive bias for these systems to learn how to achieve deep and accurate integration between external knowledge and response generation.\n\nWe posit that this ongoing challenge demands a more effective mechanism to support on-demand knowledge integration. We draw inspiration from how humans converse about a topic, where people often search and acquire external information as needed to continue a meaningful and informative conversation. Figure 1 illustrates an example human discussion, where information scattered in separate paragraphs must be consolidated to compose grounded and appropriate responses. Thus, the challenge is to connect the dots across different pieces of information in much the same way that machine reading comprehension (MRC) systems tie together multiple text segments to provide a unified and factual answer (Seo et al., 2017, etc.).\n\nWe introduce a new framework of end-to-end conversation models that jointly learn response generation together with on-demand machine reading. We formulate the reading comprehension task as document-grounded response generation: given a long document that supplements the conversation topic, along with the conversation history, we aim to produce a response that is both conversationally appropriate and informed by the content of the document. The key idea is to project conventional QA-based reading comprehension onto conversation response generation by equating the conversation prompt with the question, the conversation response with the answer, and external knowledge with the context. The MRC framing allows for integration of long external documents that present notably richer and more complex information than relatively small collections of short, independent review posts such as those that have been used in prior work (Ghazvininejad et al., 2018; Moghe et al., 2018).\n\nWe also introduce a large dataset to facilitate research on knowledge-grounded conversation (2.8M turns, 7.4M sentences of grounding) that is at least one order of magnitude larger than existing datasets (Dinan et al., 2019; Moghe et al., 2018). This dataset consists of real-world conversations extracted from Reddit, linked to web documents discussed in the conversations. Empirical results on our new dataset demonstrate that our full model improves over previous grounded response generation systems and various ungrounded baselines, suggesting that deep knowledge integration is an important research direction.Code for reproducing our models and data is made publicly available at https://github.com/qkaren/converse_reading_cmr.\n\n## Task\n\nWe propose to use factoid- and entity-rich web documents, e.g., news stories and Wikipedia pages, as external knowledge sources for an open-ended conversational system to ground in.\n\nFormally, we are given a conversation history of turns $X=(\\bm{x}_{1},\\dots,\\bm{x}_{M})$ and a web document $D=(\\bm{s}_{1},\\dots,\\bm{s}_{N})$ as the knowledge source, where $\\bm{s}_{i}$ is the $i$th sentence in the document. With the pair $(X,D)$, the system needs to generate a natural language response $\\bm{y}$ that is both conversationally appropriate and reflective of the contents of the web document.\n\n## Approach\n\nOur approach integrates conversation generation with on-demand MRC. Specifically, we use an MRC model to effectively encode the conversation history by treating it as a question in a typical QA task (e.g., SQuAD (Rajpurkar et al., 2016)), and encode the web document as the context. We then replace the output component of the MRC model (which is usually an answer classification module) with an attentional sequence generator that generates a free-form response. We refer to our approach as CMR (Conversation with on-demand Machine Reading). In general, any off-the-shelf MRC model could be applied here for knowledge comprehension. We use Stochastic Answer Networks (SAN)https://github.com/kevinduh/san_mrc (Liu et al., 2018b), a performant machine reading model that until very recently held state-of-the-art performance on the SQuAD benchmark. We also employ a simple but effective data weighting scheme to further encourage response grounding.\n\nWe adapt the SAN model to encode both the input document and conversation history and forward the digested information to a response generator. Figure 2 depicts the overall MRC architecture. Different blocks capture different concepts of representations in both the input conversation history and web document. The leftmost blocks represent the lexicon encoding that extracts information from $X$ and $D$ at the token level. Each token is first transformed into its corresponding word embedding vector, and then fed into a position-wise feed-forward network (FFN) Vaswani et al. (2017) to obtain the final token-level representation. Separate FFNs are used for the conversation history and the web document.\n\nThe next block is for contextual encoding. The aforementioned token vectors are concatenated with pre-trained 600-dimensional CoVe vectors McCann et al. (2017), and then fed to a BiLSTM that is shared for both conversation history and web document. The step-wise outputs of the BiLSTM carry the information of the tokens as well as their left and right context.\n\n## 2 Response Generation\n\nHaving read and processed both the conversation history and the extra knowledge in the document, the model then produces a free-form response ${\\bf y}=(y_{1},\\dots,y_{T})$ instead of generating a span or performing answer classification as in MRC tasks.\n\nWe use an attentional recurrent neural network decoder (Luong et al., 2015) to generate response tokens while attending to the memory. At the beginning, the initial hidden state $\\bm{h}_{0}$ is the weighted sum of the representation of the history $X$. For each decoding step $t$ with a hidden state $\\bm{h}_{t}$, we generate a token $y_{t}$ based on the distribution:\n\nwhere $\\tau\u003e0$ is the softmax temperature. The hidden state $\\bm{h}_{t}$ is defined as follows:\n\nHere, $[\\cdot+\\!\\!+\\cdot]$ indicates a concatenation of two vectors; $f_{attention}$ is a dot-product attention Vaswani et al. (2017); and $\\bm{z}_{t}$ is a state generated by $\\text{GRU}(\\bm{e}_{t-1},\\bm{h}_{t-1})$ with $\\bm{e}_{t-1}$ being the embedding of the word $y_{t-1}$ generated at the previous ($t-1$) step. In practice, we use top-$k$ sample decoding to draw $y_{t}$ from the above distribution $p(y_{t})$. Section 5 provides more details about the experimental configuration.\n\n## 3 Data Weighting Scheme\n\n## Dataset\n\nTo create a grounded conversational dataset, we extract conversation threads from Reddit, a popular and large-scale online platform for news and discussion. In 2015 alone, Reddit hosted more than 73M conversations.https://redditblog.com/2015/12/31/reddit-in-2015/ On Reddit, user submissions are categorized by topics or “subreddits”, and a submission typically consists of a submission title associated with a URL pointing to a news or background article, which initiates a discussion about the contents of the article. This article provides framing for the conversation, and this can naturally be seen as a form of grounding. Another factor that makes Reddit conversations particularly well-suited for our conversation-as-MRC setting is that a significant proportion of these URLs contain named anchors (i.e., ‘#’ in the URL) that point to the relevant passages in the document. This is conceptually quite similar to MRC data (Rajpurkar et al., 2016) where typically only short passages within a larger document are relevant in answering the question.\n\nWe reduce spamming and offensive language by manually curating a list of 178 relatively “safe” subreddits and 226 web domains from which the web pages are extracted. To convert the web page of each conversation into a text document, we extracted the text of the page using an html-to-text converter,https://www.crummy.com/software/BeautifulSoup while retaining important tags such as \u003ctitle\u003e, \u003ch1\u003e to \u003ch6\u003e, and \u003cp\u003e. This means the entire text of the original web page is preserved, but these main tags retain some high-level structure of the article. For web URLs with named anchors, we preserve that information by indicating the anchor text in the document with tags \u003canchor\u003e and \u003c/anchor\u003e. As the whole documents in the dataset tend to be lengthy, anchors offer important hints to the model about which parts of the documents should likely be focused on in order to produce a good response. We considered it sensible to keep them as they are also available to the human reader.\n\nAfter filtering short or redacted turns, or which quote earlier turns, we obtained 2.8M conversation instances respectively divided into train, validation, and test (Table 1). We used different date ranges for these different sets: years 2011-2016 for train, Jan-Mar 2017 for validation, and the rest of 2017 for test. For the test set, we select conversational turns for which 6 or more responses were available, in order to create a multi-reference test set. Given other filtering criteria such as turn length, this yields a 6-reference test set of size 2208. For each instance, we set aside one of the 6 human responses to assess human performance on this task, and the remaining 5 responses serve as ground truths for evaluating different systems.While this is already large for a grounded dataset, we could have easily created a much bigger one given how abundant Reddit data is. We focused instead on filtering out spamming and offensive language, in order to strike a good balance between data quality and size. Table 1 provides statistics for our dataset, and Figure 1 presents an example from our dataset that also demonstrates the need to combine conversation history and background information from the document to produce an informative response.\n\nTo enable reproducibility of our experiments, we crawled web pages using Common Crawl (http://commoncrawl.org), a service that crawls web pages and makes its historical crawls available to the public. We also release the code (URL redacted for anonymity) to recreate our dataset from both a popular Reddit dumphttp://files.pushshift.io/reddit/ and Common Crawl, and the latter service ensures that anyone reproducing our data extraction experiments would retrieve exactly the same web pages. We made a preliminary version of this dataset available for a shared task Galley et al. (2019) at Dialog System Technology Challenges (DSTC) Yoshino et al. (2019). Back-and-forth with participants helped us iteratively refine the dataset. The code to recreate this dataset is included.We do not report on shared task systems here, as these systems do not represent our work and some of these systems have no corresponding publications. Along with the data described here, we provided a standard Seq2Seq baseline to the shared task, which we improved for the purpose of this paper (improved Bleu, Nist and Meteor). Our new Seq2Seq baseline is described in Section 5.\n\n## Experiments\n\nWe evaluate our systems and several competitive baselines: Seq2Seq (Sutskever et al., 2014) We use a standard LSTM Seq2Seq model that only exploit the conversation history for response generation, without any grounding. This is a competitive baseline initialized using pretrained embeddings. MemNet: We use a Memory Network designed for grounded response generation (Ghazvininejad et al., 2018). An end-to-end memory network (Sukhbaatar et al., 2015) encodes conversation history and sentences in the web documents. Responses are generated with a sequence decoder. CMR-f : To directly measure the effect of incorporating web documents, we compare to a baseline which omits the document reading component of the full model (Figure 2). As with the Seq2Seq approach, the resulting model generates responses solely based on conversation history. CMR: To measure the effect of our data weighting scheme, we compare to a system that has identical architecture to the full model, but is trained without associating weights to training instances. CMR+w: As described in section 3, the full model reads and comprehends both the conversation history and document using an MRC component, and sequentially generates the response. The model is trained with the data weighting scheme to encourage grounded responses. Human: To get a better sense of the systems’ performance relative to an upper bound, we also evaluate human-written responses using different metrics. As described in Section 4, for each test instance, we set aside one of the 6 human references for evaluation, so the ‘human’ is evaluated against the other 5 references for automatic evaluation. To make these results comparable, all the systems are also automatically evaluated against the same 5 references.\n\n## Experiment Details\n\nFor all the systems, we set word embedding dimension to 300 and used the pretrained GloVehttps://nlp.stanford.edu/projects/glove/ for initialization. We set hidden dimensions to 512 and dropout rate to 0.4. GRU cells are used for Seq2Seq and MemNet (we also tested LSTM cells and obtained similar results). We used the Adam optimizer for model training, with an initial learning rate of 0.0005. Batch size was set to 32. During training, all responses were truncated to have a maximum length of 30, and maximum query length and document length were set to 30, 500, respectively. we used regular teacher-forcing decoding during training. For inference, we found that top-$k$ random sample decoding Fan et al. (2018) provides the best results for all the systems. That is, at each decoding step, a token was drawn from the $k$ most likely candidates according to the distribution over the vocabulary. Similar to recent work (Fan et al., 2018; Edunov et al., 2018), we set $k=20$ (other common $k$ values like 10 gave similar results). We selected key hyperparameter configurations on the validation set.\n\nTable 2 shows automatic metrics for quantitative evaluation over three qualities of generated texts. We measure the overall relevance of the generated responses given the conversational history by using standard Machine Translation (MT) metrics, comparing generated outputs to ground-truth responses. These metrics include Bleu-4 (Papineni et al., 2002), Meteor (Lavie and Agarwal, 2007). and Nist (Doddington, 2002). The latter metric is a variant of Bleu that weights $n$-gram matches by their information gain by effectively penalizing uninformative $n$-grams (such as “I don’t know”), which makes it a relevant metric for evaluating systems aiming diverse and informative responses. MT metrics may not be particularly adequate for our task Liu et al. (2016), given its focus on the informativeness of responses, and for that reason we also use two other types of metrics to measure the level of grounding and diversity.\n\nAs a diversity metric, we count all $n$-grams in the system output for the test set, and measure: (1) Entropy-$n$ as the entropy of the $n$-gram count distribution, a metric proposed in Zhang et al. (2018b); (2) Distinct-$n$ as the ratio between the number of $n$-gram types and the total number of $n$-grams, a metric introduced in Li et al. (2016a).\n\nFor the grounding metrics, we first compute ‘#match,’ the number of non-stopword tokens in the response that are present in the document but not present in the context of the conversation. Excluding words from the conversation history means that, in order to produce a word of the document, the response generation system is very likely to be effectively influenced by that document. We then compute both precision as ‘#match’ divided by the total number of non-stop tokens in the response, and recall as ‘#match’ divided by the total number of non-stop tokens in the document. We also compute the respective F1 score to combine both. Looking only at exact unigram matches between the document and response is a major simplifying assumption, but the combination of the three metrics offers a plausible proxy for how greatly the response is grounded in the document. It seems further reasonable to assume that these can serve as a surrogate for less quantifiable forms of grounding such as paraphrase – e.g., US $\\xrightarrow{}$ American – when the statistics are aggregated on a large test dataset.\n\n## 2 Automatic Evaluation\n\nTable 2 shows automatic evaluation results for the different systems. In terms of appropriateness, the different variants of our models outperform the Seq2Seq and MemNet baselines, but differences are relatively small and, in case of one of the metrics (Nist), the best system does not use grounding. Our goal, we would note, is not to specifically improve response appropriateness, as many responses that completely ignore the document (e.g., I don’t know) might be perfectly appropriate. Our systems fare much better in terms of Grounding and Diversity: our best system (CMR+w) achieves an F1 score that is more than three times (0.38% vs. 0.12%) higher than the most competitive non-MRC system (MemNet).\n\n## 3 Human Evaluation\n\nWe sampled 1000 conversations from the test set. Filters were applied to remove conversations containing ethnic slurs or other offensive content that might confound judgments. Outputs from systems to be compared were presented pairwise to judges from a crowdsourcing service. Four judges were asked to compare each pair of outputs on Relevance (the extent to which the content was related to and appropriate to the conversation) and Informativeness (the extent to which the output was interesting and informative). Judges were asked to agree or disagree with a statement that one of the pair was better than the other on the above two parameters, using a 5-point Likert scale.The choices presented to the judges were Strongly Agree, Agree, Neutral, Disagree, and Strongly Disagree. Pairs of system outputs were randomly presented to the judges in random order in the context of short snippets of the background text. These results are presented in summary form in Table 3, which shows the overall preferences for the two systems expressed as a percentage of all judgments made. Overall inter-rater agreement measured by Fliess’ Kappa was 0.32 (“fair\"). Nevertheless, the differences between the paired model outputs are statistically significant (computed using 10,000 bootstrap replications).\n\n## 4 Qualitative Study\n\nTable 4 illustrates how our best model (CMR+w) tends to produce more contentful and informative responses compared to the other systems. In the first example, our system refers to a particular episode mentioned in the article, and also uses terminology that is more consistent with the article (e.g., series). In the second example, humorous song seems to positively influence the response, which is helpful as the input doesn’t mention singing at all. In the third example, the CMR+w model clearly grounds its response to the article as it states the fact (Steve Jobs: CEO of Apple) retrieved from the article. The outputs by the other two baseline models are instead not relevant in the context.\n\nFigure 3 displays the attention map of the generated response and (part of) the document from our full model. The model successfully attends to the key words (e.g., 36th, episode) of the document. Note that the attention map is unlike what is typical in machine translation, where target words tend to attend to different portions of the input text. In our task, where alignments are much less one-to-one compared to machine translation, it is common for the generator to retain focus on the key information in the external document to produce semantically relevant responses.\n\n## Related Work\n\nTraditional dialogue systems (see Jurafsky and Martin (2009) for an historical perspective) are typically grounded, enabling these systems to be reflective of the user’s environment. The lack of grounding has been a stumbling block for the earliest end-to-end dialogue systems, as various researchers have noted that their outputs tend to be bland Li et al. (2016a); Gao et al. (2019b), inconsistent Zhang et al. (2018a); Li et al. (2016b); Zhang et al. (2019), and lacking in factual content Ghazvininejad et al. (2018); Agarwal et al. (2018). Recently there has been growing interest in exploring different forms of grounding, including images, knowledge bases, and plain texts Das et al. (2017); Mostafazadeh et al. (2017); Agarwal et al. (2018); Yang et al. (2019). A recent survey is included in Gao et al. (2019a).\n\nPrior work, e.g, (Ghazvininejad et al., 2018; Zhang et al., 2018a; Huang et al., 2019), uses grounding in the form of independent snippets of text: Foursquare tips and background information about a given speaker. Our notion of grounding is different, as our inputs are much richer, encompassing the full text of a web page and its underlying structure. Our setting also differs significantly from relatively recent work (Dinan et al., 2019; Moghe et al., 2018) exploiting crowdsourced conversations with detailed grounding labels: we use Reddit because of its very large scale and better characterization of real-world conversations. We also require the system to learn grounding directly from conversation and document pairs, instead of relying on additional grounding labels. Moghe et al. (2018) explored directly using a span-prediction QA model for conversation. Our framework differs in that we combine MRC models with a sequence generator to produce free-form responses.\n\n## Machine Reading Comprehension:\n\nMRC models such as SQuAD-like models, aim to extract answer spans (starting and ending indices) from a given document for a given question Seo et al. (2017); Liu et al. (2018b); Yu et al. (2018). These models differ in how they fuse information between questions and documents. We chose SAN Liu et al. (2018b) because of its representative architecture and competitive performance on existing MRC tasks. We note that other off-the-shelf MRC models, such as BERT (Devlin et al., 2018), can also be plugged in. We leave the study of different MRC architectures for future work. Questions are treated as entirely independent in these “single-turn” MRC models, so recent work (e.g., CoQA Reddy et al. (2019) and QuAC Choi et al. (2018)) focuses on multi-turn MRC, modeling sequences of questions and answers in a conversation. While multi-turn MRC aims to answer complex questions, that body of work is restricted to factual questions, whereas our work—like much of the prior work in end-to-end dialogue—models free-form dialogue, which also encompasses chitchat and non-factual responses.\n\n## Conclusions\n\nWe have demonstrated that the machine reading comprehension approach offers a promising step to generating, on the fly, contentful conversation exchanges that are grounded in extended text corpora. The functional combination of MRC and neural attention mechanisms offers visible gains over several strong baselines. We have also formally introduced a large dataset that opens up interesting challenges for future research.\n\nThe CMR (Conversation with on-demand machine reading) model presented here will help connect the many dots across multiple data sources. One obvious future line of investigation will be to explore the effect of other off-the-shelf machine reading models such as BERT Devlin et al. (2018) within the CMR framework.\n\n## Acknowledgements\n\nWe are grateful to the anonymous reviewers, as well as to Vighnesh Shiv, Yizhe Zhang, Chris Quirk, Shrimai Prabhumoye, and Ziyu Yao for helpful comments and suggestions on this work. This research was supported in part by NSF (IIS- 1524371), DARPA CwC through ARO (W911NF- 15-1-0543), and Samsung AI Research.\n\n## References"])</script><script>self.__next_f.push([1,"5:[\"$\",\"$L11\",null,{\"paper\":{\"arxivId\":\"1906.02738\",\"title\":\"Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading\",\"authors\":\"Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, Jianfeng Gao\",\"abstract\":\"$12\",\"markdown\":\"$13\",\"version\":\"v1\",\"categories\":\"cs.CL,cs.AI,cs.LG\",\"fetchedAt\":\"$D2026-08-05T23:19:57.691Z\"},\"initialAnnotations\":[],\"related\":[]}]\n"])</script></body></html>