SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering
Chenguang Zhu, Michael Zeng, Xuedong Huang
Introduction
Traditional machine reading comprehension (MRC) tasks share the single-turn setting of answering a single question related to a passage. There is usually no connection between different questions and answers to the same passage. However, the most natural way humans seek answers is via conversation, which carries over context through the dialogue flow.
To incorporate conversation into reading comprehension, recently there are several public datasets that evaluate QA model’s efficacy in conversational setting, such as CoQA (Reddy et al., 2018), QuAC (Choi et al., 2018) and QBLink (Elgohary et al., 2018). In these datasets, to generate correct responses, models are required to fully understand the given passage as well as the context of previous questions and answers. Thus, traditional neural MRC models are not suitable to be directly applied to this scenario. Existing approaches to conversational QA tasks include BiDAF++ (Yatskar, 2018), FlowQA (Huang et al., 2018), DrQA+PGNet (Reddy et al., 2018), which all try to find the optimal answer span given the passage and dialogue history.
In this paper, we propose SDNet, a contextual attention-based deep neural network for the task of conversational question answering. Our network stems from machine reading comprehension models, but has several unique characteristics to tackle contextual understanding during conversation. Firstly, we apply both inter-attention and self-attention on passage and question to obtain a more effective understanding of the passage and dialogue history. Secondly, SDNet leverages the latest breakthrough in NLP: BERT contextual embedding (Devlin et al., 2018). Different from the canonical way of appending a thin layer after BERT structure according to (Devlin et al., 2018), we innovatively employed a weighted sum of BERT layer outputs, with locked BERT parameters. Thirdly, we prepend previous rounds of questions and answers to the current question to incorporate contextual information. Empirical results show that each of these components has substantial gains in prediction accuracy.
We evaluated SDNet on CoQA dataset, which improves the previous state-of-the-art model’s result by 1.6% (from 75.0% to 76.6%) overall score. The ensemble model further increase the score to . Moreover, SDNet is the first model ever to pass on CoQA’s in-domain dataset.
Approach
In this section, we propose the neural model, SDNet, for the conversational question answering task, which is formulated as follows. Given a passage , and history question and answer utterances , the task is to generate response given the latest question . The response is dependent on both the passage and history utterances.
To incorporate conversation history into response generation, we employ the idea from DrQA+PGNet (Reddy et al., 2018) to prepend the latest rounds of utterances to the current question . The problem is then converted into a machine reading comprehension task. In other words, the reformulate question is . To differentiate between question and answering, we add symbol before each question and before each answer in the experiment.
Encoding layer encodes each token in passage and question into a fixed-length vector, which includes both word embeddings and contextualized embeddings. For contextualized embedding, we utilize the latest result from BERT (Devlin et al., 2018). Different from previous work, we fix the parameters in BERT model and use the linear combination of embeddings from different layers in BERT.
Integration layer uses multi-layer recurrent neural networks (RNN) to capture contextual information within passage and question. To characterize the relationship between passage and question, we conduct word-level attention from question to passage both before and after the RNNs. We employ the idea of history-of-word from FusionNet (Huang et al., 2017) to reduce the dimension of output hidden vectors. Furthermore, we conduct self-attention to extract relationship between words at different positions of context and question.
Output layer computes the final answer span. It uses attention to condense the question into a fixed-length vector, which is then used in a bilinear projection to obtain the probability that the answer should start and end at each position.
An illustration of our model SDNet is in Figure 1.
2 Encoding layer
We use 300-dim GloVe (Pennington et al., 2014) embedding and contextualized embedding for each word in context and question. We employ BERT (Devlin et al., 2018) as contextualized embedding. Instead of adding a scoring layer to BERT structure as proposed in (Devlin et al., 2018), we use the transformer output from BERT as contextualized embedding in our encoding layer. BERT generates layers of hidden states for all BPE tokens (Sennrich et al., 2015) in a sentence/passage and we employ a weighted sum of these hidden states to obtain contextualized embedding. Furthermore, we lock BERT’s internal weights, setting their gradients to zero. In ablation studies, we will show that this weighted sum and weight-locking mechanism can significantly boost the model’s performance.
In detail, suppose a word is tokenized to BPE tokens , and BERT generates hidden states for each BPE token, . The contextual embedding for word is then a per-layer weighted sum of average BERT embedding, with weights .
3 Integration layer
To simplify notation, we define the attention function above as , meaning we compute the attention score based on two sets of vectors and , and use that to linearly combine vector set . So the word-level attention above can be simplified as .
For each context word in , we also include a feature vector including 12-dim POS embedding, 8-dim NER embedding, a 3-dim exact matching vector indicating whether each context word appears in the question, and a normalized term frequency, following the approach in DrQA (Chen et al., 2017).
RNN. In this component, we use two separate bidirectional RNNs (BiLSTMs (Hochreiter & Schmidhuber, 1997)) to form the contextualized understanding for and .
where and is the number of RNN layers. We use variational dropout (Kingma et al., 2015) for input vector to each layer of RNN, i.e. the dropout mask is shared over different timesteps.
Question Understanding. For each question word in , we employ one more layer of RNN to generate a higher level of understanding of the question.
Self-Attention on Question. As the question has integrated previous utterances, the model needs to directly relate previously mentioned concept with the current question. This is helpful for concept carry-over and coreference resolution. We thus employ self-attention on question. The formula is the same as word-level attention, except that we are attending a question to itself: . The final question representation is thus .
Multilevel Inter-Attention. After multiple layers of RNN extract different levels of understanding of each word, we conduct multilevel attention from question to context based on all layers of generated representations.
However, the aggregated dimensions can be very large, which is computationally inefficient. We thus leverage the history-of-word idea from FusionNet (Huang et al., 2017): we use all previous levels to compute attentions scores, but only linearly combine RNN outputs.
In detail, we conduct times of multilevel attention from each RNN layer output of question to context.
where history-of-word vectors are defined as
An additional RNN layer is applied to obtain the contextualized representation for each word in .
Self Attention on Context. Similar to questions, we conduct self attention on context to establish direct correlations between all pairs of words in . Again, we use the history of word concept to reduce the output dimension by linearly combining .
The self-attention is followed by an additional layer of RNN to generate the final representation of context:
4 Output layer
Generating Answer Span. This component is to generate two scores for each context word corresponding to the probability that the answer starts and ends at this word, respectively.
Firstly, we condense the question representation into one vector: , where and is a parametrized vector.
Secondly, we compute the probability that the answer span should start at the -th word:
where is a parametrized matrix. We further fuse the start-position probability into the computation of end-position probability via a GRU, . Thus, the probability that the answer span should end at the -th word is:
where is another parametrized matrix.
For CoQA dataset, the answer could be affirmation “yes”, negation “no” or no answer “unknown”. We separately generate three probabilities corresponding to these three scenarios, , respectively. For instance, to generate the probability that the answer is “yes”, , we use:
where and are parametrized matrix and vector, respectively.
Training. For training, we use all questions/answers for one passage as a batch. The goal is to maximize the probability of the ground-truth answer, including span start/end position, affirmation, negation and no-answer situations. Equivalently, we minimize the negative log-likelihood function :
where and are the ground-truth span start and end position for the -th question. indicate whether the -th ground-truth answer is a passage span, “yes”, “no” and “unknown”, respectively. More implementation details are in Appendix.
Prediction. During inference, we pick the largest span/yes/no/unknown probability. The span is constrained to have a maximum length of 15.
Experiments
We evaluated our model on CoQA (Reddy et al., 2018), a large-scale conversational question answering dataset. In CoQA, many questions require understanding of both the passage and previous questions and answers, which poses challenge to conventional machine reading models. Table 1 summarizes the domain distribution in CoQA. As shown, CoQA contains passages from multiple domains, and the average number of question answering turns is more than 15 per passage. Many questions require contextual understanding to generate the correct answer.
For each in-domain dataset, 100 passages are in the development set, and 100 passages are in the test set. The rest in-domain dataset are in the training set. The test set also includes all of the out-of-domain passages.
Baseline models and metrics. We compare SDNet with the following baseline models: PGNet (Seq2Seq with copy mechanism) (See et al., 2017), DrQA (Chen et al., 2017), DrQA+PGNet (Reddy et al., 2018), BiDAF++ (Yatskar, 2018) and FlowQA (Huang et al., 2018). Aligned with the official leaderboard, we use as the evaluation metric, which is the harmonic mean of precision and recall at word level between the predicted answer and ground truth.According to official evaluation of CoQA, when there are more than one ground-truth answers, the final score is the average of max against all-but-one ground-truth answers.
Results. Table 2 report the performance of SDNet and baseline models.Result was taken from official CoQA leaderboard on Nov. 30, 2018. As shown, SDNet achieves significantly better results than baseline models. In detail, the single SDNet model improves overall by 1.6%, compared with previous state-of-art model on CoQA, FlowQA. Ensemble SDNet model further improves overall score by 2.7%, and it’s the first model to achieve over 80% score on in-domain datasets (80.7%).
Figure 2 shows the score on development set over epochs. As seen, SDNet overpasses all but one baseline models after the second epoch, and achieves state-of-the-art results only after 8 epochs.
Ablation Studies. We conduct ablation studies on SDNet model and display the results in Table 3. The results show that removing BERT can reduce the score on development set by . Our proposed weight sum of per-layer output from BERT is crucial, which can boost the performance by , compared with using only last layer’s output. This shows that the output from each layer in BERT is useful in downstream tasks. This technique can also be applied to other NLP tasks. Using BERT-base instead of BERT-large pretrained model hurts the score by , which manifests the superiority of BERT-large model. Variational dropout and self attention can each improve the performance by 0.24% and 0.75%, respectively.
Contextual history. In SDNet, we utilize conversation history via prepending the current question with previous rounds of questions and ground-truth answers. We experimented the effect of and show the result in Table 4. Excluding dialogue history () can reduce the score by as much as , showing the importance of contextual information in conversational QA task. The performance of our model peaks when , which was used in the final SDNet model.
Conclusions
In this paper, we propose a novel contextual attention-based deep neural network, SDNet, to tackle conversational question answering task. By leveraging inter-attention and self-attention on passage and conversation history, the model is able to comprehend dialogue flow and fuse it with the digestion of passage content. Furthermore, we incorporate the latest breakthrough in NLP, BERT, and leverage it in an innovative way. SDNet achieves superior results over previous approaches. On the public dataset CoQA, SDNet outperforms previous state-of-the-art model by 1.6% in overall metric.
Our future work is to apply this model to open-domain multiturn QA problem with large corpus or knowledge base, where the target passage may not be directly available. This will be an even more realistic setting to human question answering.
References
Appendix A Implementation Details
We use spaCy for tokenization. As BERT use BPE as the tokenizer, we did BPE tokenization for each token generated by spaCy. In case a token in spaCy corresponds to multiple BPE sub-tokens, we average the BERT embeddings of these BPE sub-tokens as the embedding for the token. We fix the BERT weights and use the BERT-Large-Uncased model.
During training, we use a dropout rate of 0.4 for BERT layer outputs and 0.3 for other layers. We use variational dropout (Kingma et al., 2015), which shares the dropout mask over timesteps in RNN. We batch the data according to passages, so all questions and answers from the same passage make one batch.
We use Adamax (Kingma & Ba, 2014) as the optimizer, with a learning rate of and . We train the model using 30 epochs, with each epoch going over the data once. We clip the gradient at length .
The word-level attention has a hidden size of 300. The flow module has a hidden size of 300. The question self attention has a hidden size of 300. The RNN for both question and context has layers and each layer has a hidden size of 125. The multilevel attention from question to context has a hidden size of 250. The context self attention has a hidden size of 250. The final layer of RNN for context has a hidden size of 125.