Multi-style Generative Reading Comprehension

Kyosuke Nishida, Itsumi Saito, Kosuke Nishida, Kazutoshi Shinoda, Atsushi Otsuka, Hisako Asano, Junji Tomita

Introduction

Question answering has been a long-standing research problem. Recently, reading comprehension (RC), a challenge to answer a question given textual evidence provided in a document set, has received much attention. Current mainstream studies have treated RC as a process of extracting an answer span from one passage (Rajpurkar et al., 2016, 2018) or multiple passages (Joshi et al., 2017; Yang et al., 2018), which is usually done by predicting the start and end positions of the answer (Yu et al., 2018; Devlin et al., 2018).

The demand for answering questions in natural language is increasing rapidly, and this has led to the development of smart devices such as Alexa. In comparison with answer span extraction, however, the natural language generation (NLG) capability for RC has been less studied. While datasets such as MS MARCO (Bajaj et al., 2018) and NarrativeQA (Kociský et al., 2018) have been proposed for providing abstractive answers, the state-of-the-art methods for these datasets are based on answer span extraction (Wu et al., 2018; Hu et al., 2018). Generative models suffer from a dearth of training data to cover open-domain questions.

Moreover, to satisfy various information needs, intelligent agents should be capable of answering one question in multiple styles, such as well-formed sentences, which make sense even without the context of the question and passages, and concise phrases. These capabilities complement each other, but previous studies cannot use and control different styles within a model.

In this study, we propose Masque, a generative model for multi-passage RC. It achieves state-of-the-art performance on the Q&A task and the Q&A + NLG task of MS MARCO 2.1 and the summary task of NarrativeQA. The main contributions of this study are as follows.

We introduce the pointer-generator mechanism (See et al., 2017) for generating an abstractive answer from the question and multiple passages, which covers various answer styles. We extend the mechanism to a Transformer (Vaswani et al., 2017) based one that allows words to be generated from a vocabulary and to be copied from the question and passages.

We introduce multi-style learning that enables our model to control answer styles and improves RC for all styles involved. We also extend the pointer-generator to a conditional decoder by introducing an artificial token corresponding to each style, as in (Johnson et al., 2017). For each decoding step, it controls the mixture weights over three distributions with the given style (Figure 1).

Problem Formulation

Given a question with JJ words xq={x1q,…,xJq}x^{q}=\{x^{q}_{1},\ldots,x^{q}_{J}\}, a set of KK passages, where the kk-th passage is composed of LL words xpk={x1pk,…,xLpk}x^{p_{k}}=\{x^{p_{k}}_{1},\ldots,x^{p_{k}}_{L}\}, and an answer style label ss, an RC model outputs an answer y={y1,…,yT}y=\{y_{1},\ldots,y_{T}\} conditioned on the style.

In short, given a 3-tuple (xq,{xpk},s)(x^{q},\{x^{p_{k}}\},s), the system predicts P(y)P(y). The training data is a set of 6-tuples: (xq,{xpk},s,y,a,{rpk})(x^{q},\{x^{p_{k}}\},s,y,a,\{r^{p_{k}}\}), where aa and {rpk}\{r^{p_{k}}\} are optional. Here, aa is 11 if the question is answerable with the provided passages and otherwise, and rpkr^{p_{k}} is 11 if the kk-th passage is required to formulate the answer and otherwise.

Proposed Model

We propose a Multi-style Abstractive Summarization model for QUEstion answering, called Masque. Masque directly models the conditional probability p(y∣xq,{xpk},s)p(y|x^{q},\{x^{p_{k}}\},s). As shown in Figure 2, it consists of the following modules.

The question-passages reader (§3.1) models interactions between the question and passages.

The passage ranker (§3.2) finds passages relevant to the question.

The answer possibility classifier (§3.3) identifies answerable questions.

The answer sentence decoder (§3.4) outputs an answer sentence conditioned on the target style.

Our model is based on multi-source abstractive summarization: the answer that it generates can be viewed as a summary from the question and passages. The model also learns multi-style answers together. With these two characteristics, we aim to acquire the style-independent NLG ability and transfer it to the target style. In addition, to improve natural language understanding in the reader module, our model considers RC, passage ranking, and answer possibility classification together as multi-task learning.

The reader module is shared among multiple answer styles and the three task-specific modules.

1.2 Shared Encoder Layer

1.3 Dual Attention Layer

This layer uses a dual attention mechanism to fuse information from the question to the passages as well as from the passages to the question.

Here, Aˉpk=EqApk\bar{A}^{p_{k}}=E^{q}A^{p_{k}}, Bˉpk=EpkBpk\bar{B}^{p_{k}}=E^{p_{k}}B^{p_{k}}, Aˉˉpk=BˉpkApk\bar{\bar{A}}^{p_{k}}=\bar{B}^{p_{k}}A^{p_{k}}, Bˉˉpk=AˉpkBpk\bar{\bar{B}}^{p_{k}}=\bar{A}^{p_{k}}B^{p_{k}}, Bˉ=max⁡k(Bˉpk)\bar{B}=\max_{k}(\bar{B}^{p_{k}}), and Bˉˉ=max⁡k(Bˉˉpk)\bar{\bar{B}}=\max_{k}(\bar{\bar{B}}^{p_{k}}).

1.4 Modeling Encoder Layer

2 Passage Ranker

3 Answer Possibility Classifier

4 Answer Sentence Decoder

Given the outputs provided by the reader module, the decoder generates a sequence of answer words one element at a time. It is auto-regressive (Graves, 2013), consuming the previously generated words as additional input at each decoding step.

Let yy represent one-hot vectors of the words in the answer. This layer has the same components as the word embedding layer of the reader module, except that it uses a unidirectional ELMo to ensure that the predictions for position tt depend only on the known outputs at positions previous to tt.

To be able to use multiple answer styles within a single system, our model introduces an artificial token corresponding to the style at the beginning of the answer (y1y_{1}), as done in (Johnson et al., 2017; Takeno et al., 2017). At test time, the user can specify the first token to control the style. This modification does not require any changes to the model architecture. Note that introducing the token at the decoder prevents the reader module from depending on the answer style.

4.2 Attentional Decoder Layer

This layer uses a stack of Transformer decoder blocks on top of the embeddings provided by the word embedding layer. The input is immediately mapped to a dd-dimensional vector by a linear transformation, and the output is a sequence of dd-dimensional vectors: {s1,…,sT}\{s_{1},\ldots,s_{T}\}.

Here, the $$ operator denotes vector concatenation across the columns. This attention for the concatenated passages produces attention weights that are comparable between passages.

4.3 Multi-source Pointer-Generator

Our extended mechanism allows both words to be generated from a vocabulary and words to be copied from both the question and multiple passages (Figure 3). We expect that the capability of copying words will be shared among answer styles.

A recent Transformer-based pointer-generator randomly chooses one of the attention-heads to form a copy distribution; that approach gave no significant improvements in text summarization (Gehrmann et al., 2018).

where k(l)k(l) means the passage index corresponding to the ll-th word in the concatenated passages.

The final distribution of yty_{t} is defined as a mixture of the three distributions:

4.4 Combined Attention

In order not to attend words in irrelevant passages, our model introduces a combined attention. While the original technique combined word and sentence level attentions (Hsu et al., 2018), our model combines the word and passage level attentions. The word attention, Eq. 1, is re-defined as

5 Loss Function

We define the training loss as the sum of losses via

Experiments on MS MARCO 2.1

We evaluated our model on MS MARCO 2.1 Bajaj et al. (2018). It is the sole dataset providing abstractive answers with multiple styles and serves as a great test bed for building open-domain QA agents with the NLG capability that can be used in smart devices. The details of our setup and output examples are in the supplementary material.

MS MARCO 2.1 provides two tasks for generative open-domain QA: the Q&A task and the Q&A + Natural Language Generation (NLG) task. Both tasks consist of questions submitted to Bing by real users, and each question refers to ten passages. The dataset also includes annotations on the relevant passages, which were selected by humans to form the final answers, and on whether there was no answer in the passages.

We associated the two tasks with two answer styles. The NLG task requires a well-formed answer that is an abstractive summary of the question and passages, averaging 16.6 words. The Q&A task also requires an abstractive answer but prefers it to be more concise than in the NLG task, averaging 13.1 words, and many of the answers do not contain the context of the question. For the question “tablespoon in cup”, a reference answer in the Q&A task is “16,” while that in the NLG task is “There are 16 tablespoons in a cup.”

In addition to the ALL dataset, we prepared two subsets for ablation tests as listed in Table 1. The ANS set consisted of answerable questions, and the NLG set consisted of the answerable questions and well-formed answers, so that NLG ⊂\subset ANS ⊂\subset ALL. We note that multi-style learning enables our model to learn from different answer styles of data (i.e., the ANS set), and multi-task learning with the answer possibility classifier enables our model to learn from both answerable and unanswerable data (i.e., the ALL set).

We trained our model with mini-batches consisting of multi-style answers that were randomly sampled. We used a greedy decoding algorithm and did not use any beam search or random sampling, because they did not provide any improvements.

ROUGE-L and BLEU-1 were used to evaluate the models’ RC performance, where ROUGE-L is the main metric on the official leaderboard. We used the reported scores of extractive (Seo et al., 2017; Yan et al., 2019; Wu et al., 2018), generative (Tan et al., 2018), and unpublished RC models at the submission time.

In addition, to evaluate the individual contributions of our modules, we used MAP and MRR for the ranker and F1F_{1} for the classifier, where the positive class was the answerable questions.

2 Results

Table 2 shows the performance of our model and competing models on the leaderboard. Our ensemble model of six training runs, where each model was trained with the two answer styles, achieved state-of-the-art performance on both tasks in terms of ROUGE-L. In particular, for the NLG task, our single model outperformed competing models in terms of both ROUGE-L and BLEU-1.

Table 3 lists the results of an ablation test for our single model (controlled with the NLG style) on the NLG dev. setWe confirmed with the organizer that the dev. results were much better than the test results, but there was no problem.. Our model trained with both styles outperformed the model trained with the single NLG style. Multi-style learning enabled our model to improve its NLG performance by also using non-sentence answers.

Table 3 shows that our model also outperformed the model that used RNNs and self-attentions instead of Transformer blocks as in MCAN (McCann et al., 2018). Our deep decoder captured the multi-hop interaction among the question, the passages, and the answer better than a single-layer LSTM decoder could.

Furthermore, Table 3 shows that our model (jointly trained with the passage ranker and answer possibility classifier) outperformed the model that did not use the ranker and classifier. Joint learning thus had a regularization effect on the question-passages reader.

We also confirmed that the gold passage ranker, which can perfectly predict the relevance of passages, significantly improved the RC performance. Passage ranking will be a key to developing a system that can outperform humans.

Table 4 lists the passage ranking performance on the ANS dev. setThis evaluation requires our ranker to re-rank 10 passages. It is not the same as the Passage Re-ranking task.. The ranker shares the question-passages reader with the answer decoder, and this sharing contributed to improvements over the ranker trained without the answer decoder. Also, our ranker outperformed the initial ranking provided by Bing by a significant margin.

Figure 4 shows the precision-recall curve for answer possibility classification on the ALL dev. set. Our model identified the answerable questions well. The maximum F1F_{1} score was 0.7893, where the threshold of answer possibility was 0.44110.4411. This is the first report on answer possibility classification with MS MARCO 2.1.

Figure 5 shows the lengths of the answers generated by our model broken down by the answer style and query type. The generated answers were relatively shorter than the reference answers, especially for the Q&A task, but well controlled with the target style for every query type. The short answers degraded our model’s BLEU scores in the Q&A task (Table 2) because of BLEU’s brevity penalty Papineni et al. (2002).

Experiments on NarrativeQA

Next, we evaluated our model on NarrativeQA Kociský et al. (2018). It requires understanding the underlying narrative rather than relying on shallow pattern matching. Our detailed setup and output examples are in the supplementary material.

We only describe the settings specific to this experiment.

Following previous studies, we used the summary setting for the comparisons with the reported baselines, where each question refers to one summary (averaging 659 words), and there is no unanswerable questions. Our model therefore did not use the passage ranker and answer possibility classifier.

The NarrativeQA dataset does not explicitly provide multiple answer styles. In order to evaluate the effectiveness of multi-style learning, we used the NLG subset of MS MARCO as additional training data. We associated the NarrativeQA and NLG datasets with two answer styles. The answer style of NarrativeQA (NQA) is different from that of MS MARCO (NLG) in that the answers are short (averaging 4.73 words) and contained frequently pronouns. For instance, for the question “Who is Mark Hunter?”, a reference is “He is a high school student in Phoenix.”

BLEU-1 and 4, METEOR, and ROUGE-L were used in accordance with the evaluation in the dataset paper (Kociský et al., 2018). We used the reports of top-performing extractive (Seo et al., 2017; Tay et al., 2018; Hu et al., 2018) and generative (Bauer et al., 2018; Indurthi et al., 2018) models.

2 Results

Table 5 shows that our single model, trained with two styles and controlled with the NQA style, pushed forward the state-of-the-art by a significant margin. The evaluation scores of the model controlled with the NLG style were low because the two styles are different. Also, our model without multi-style learning (trained with only the NQA style) outperformed the baselines in terms of ROUGE-L. This indicates that our model architecture itself is powerful for natural language understanding in RC.

Related Work and Discussion

Recent breakthroughs in transfer learning demonstrate that pre-trained language models perform well on RC with minimal modifications (Peters et al., 2018; Devlin et al., 2018; Radford et al., 2018, 2019). In addition, our model also uses ELMo (Peters et al., 2018) for contextualized embeddings.

Multi-task learning is a transfer mechanism to improve generalization performance (Caruana, 1997), and it is generally applied by sharing the hidden layers between all tasks, while keeping task-specific layers. Wang et al. (2018) and Nishida et al. (2018) reported that the sharing of the hidden layers between the multi-passage RC and passage ranking tasks was effective. Our results also showed the effectiveness of the sharing of the question-passages reader module among the RC, passage ranking, and answer possibility classification tasks.

In multi-task learning without task-specific layers, Devlin et al. (2018) and Chen et al. (2017) improved RC performance by learning multiple datasets from the same extractive RC setting. McCann et al. (2018) and Yogatama et al. (2019) investigated multi-task and curriculum learning on many different NLP tasks; their results were below task-specific RC models. Our multi-style learning does not use style-specific layers; instead uses a style-conditional decoder.

S-Net (Tan et al., 2018) used an extraction-then-synthesis mechanism for multi-passage RC. The models proposed by McCann et al. (2018), Bauer et al. (2018), and Indurthi et al. (2018) used an RNN-based pointer-generator mechanism for single-passage RC. Although these mechanisms can alleviate the lack of training data, large amounts of data are still required. Our multi-style learning will be a key technique enabling learning from many RC datasets with different styles.

In addition to MS MARCO and NarrativeQA, there are other datasets that provide abstractive answers. DuReader (He et al., 2018), a Chinese multi-document RC dataset, provides longer documents and answers than those of MS MARCO. DuoRC (Saha et al., 2018) and CoQA (Reddy et al., 2018) contain abstractive answers; most of the answers are short phrases.

Many studies have been carried out in the framework of style transfer, which is the task of rephrasing a text so that it contains specific styles such as sentiment. Recent studies have used artificial tokens (Sennrich et al., 2016; Johnson et al., 2017), variational auto-encoders (Hu et al., 2017), or adversarial training (Fu et al., 2018; Tsvetkov et al., 2018) to separate the content and style on the encoder side. On the decoder side, conditional language modeling has been used to generate output sentences with the target style. In addition, output length control with conditional language modeling has been well studied (Kikuchi et al., 2016; Takeno et al., 2017; Fan et al., 2018). Our style-controllable RC relies on conditional language modeling in the decoder.

The simplest approach is to concatenate the passages and find the answer from the concatenation, as in (Wang et al., 2017). Earlier pipelined models found a small number of relevant passages with a TF-IDF based ranker and passed them to a neural reader (Chen et al., 2017; Clark and Gardner, 2018), while more recent models have used a neural re-ranker to more accurately select the relevant passages (Wang et al., 2018; Nishida et al., 2018). Also, non-pipelined models (including ours) consider all the provided passages and find the answer by comparing scores between passages (Tan et al., 2018; Wu et al., 2018). The most recent models make a proper trade-off between efficiency and accuracy (Yan et al., 2019; Min et al., 2018).

The previous work of (Levy et al., 2017; Clark and Gardner, 2018) outputted a no-answer score depending on the probability of all answer spans. Hu et al. (2019) proposed an answer verifier to compare an answer with the question. Sun et al. (2018) jointly learned an RC model and an answer verifier. Our model introduces a classifier on top of the question-passages reader, which is not dependent on the generated answer.

Current state-of-the-art models use the pointer-generator mechanism (See et al., 2017). In particular, content selection approaches, which decide what to summarize, have recently been used with abstractive models. Most methods select content at the sentence level (Hsu et al., 2018; Chen and Bansal, 2018) or the word level (Pasunuru and Bansal, 2018; Li et al., 2018; Gehrmann et al., 2018). Our model incorporates content selection at the passage level in the combined attention.

Query-based summarization has rarely been studied because of a lack of datasets. Nema et al. (2017) proposed an attentional encoder-decoder model; however, Saha et al. (2018) reported that it performed worse than BiDAF on DuoRC. Hasselqvist et al. (2017) proposed a pointer-generator based model; however, it does not consider copying words from the question.

Conclusion

This study sheds light on multi-style generative RC. Our proposed model, Masque, is based on multi-source abstractive summarization and learns multi-style answers together. It achieved state-of-the-art performance on the Q&A task and the Q&A + NLG task of MS MARCO 2.1 and the summary task of NarrativeQA. The key to its success is transferring the style-independent NLG capability to the target style by use of the question-passages reader and the conditional pointer-generator decoder. In particular, the capability of copying words from the question and passages can be shared among the styles, while the capability of controlling the mixture weights for the generative and copy distributions can be acquired for each style. Our future work will involve exploring the potential of our multi-style learning towards natural language understanding.

References

Appendix A Supplementary Material

We used a modified version of the L2 regularization proposed in (Loshchilov and Hutter, 2017), with w=0.01w=0.01 on all non-bias. We additionally used a dropout (Srivastava et al., 2014) rate of 0.3 for all highway networks and residual and scaled dot-product attention operations in the multi-head attention mechanism. We also used one-sided label smoothing (Szegedy et al., 2016) for the passage relevance and answer possibility labels. We smoothed only the positive labels to 0.9.

The ensemble model consisted of six training runs with identical architectures and hyperparameters but with different weight initializations. The final answer was decided with a weighted majority, where we used the ROUGE-L score for the dev. set as the weight of each model.

We used the official evaluation script. The answers were normalized by making words lowercase.

A.2 Experimental Setup for NarrativeQA

Our best model was jointly trained with the NarrativeQA and MS MARCO NLG datasets for a total of seven epochs with a batch size of 64, where each batch consisted of multi-style answers that were randomly sampled. For efficient multi-style learning, each summary in the NarrativeQA dataset was divided into ten passages (size of 130 words) with sentence-level overlaps such that each sentence in the summary was entirely contained in a passage. Each passage from MS MARCO was also truncated to 130 words. The rest of the configuration was the same as in the MS MARCO experiments.

An official evaluation script is not provided, so we used the evaluation script created by Bauer et al. (2018)https://github.com/yicheng-w/CommonSenseMultiHopQA/. The answers were normalized by making words lowercase and removing punctuation marks.

A.3 Output Examples Generated by Masque

Tables A.3 and A.3 list the generated examples for questions from MS MARCO 2.1 and NarrativeQA, respectively. We can see from the examples that our model could control answer styles appropriately for various question and reasoning types. We did find some important errors: style errors, yes/no classification errors, copy errors with respect to numerical values, grammatical errors, and multi-hop reasoning errors.