Let's Ask Again: Refine Network for Automatic Question Generation

Preksha Nema, Akash Kumar Mohankumar, Mitesh M. Khapra, Balaji Vasan Srinivasan, Balaraman Ravindran

Introduction

Over the past few years, there has been a growing interest in Automatic Question Generation (AQG) from text - the task of generating a question from a passage and optionally an answer. AQG is used in curating Question Answering datasets, enhancing user experience in conversational AI systems Shum et al. (2018) and for creating educational materials Heilman and Smith (2010). For the above applications, it is essential that the questions are (i) grammatically correct (ii) answerable from the passage and (iii) specific to the answer. Existing approaches focus on encoding the passage, the answer and the relationship between them using complex functions and then generate the question in one single pass. However, by carefully analysing the generated questions, we observe that these approaches tend to miss one or more of the important aspects of the question. For instance, in Table 1, the question generated by the single-pass baseline model for the first passage is grammatically correct but is not specific to the answer. In the second example, the generated question is both syntactically incorrect and incomplete.

The above examples indicate that there is clear scope of improving the general quality of the questions. Additionally, the quality can be specifically improved in terms of aspects like: fluency (Example 2) and answerability (Example 1). One way to approach this is by re-visiting the passage and answer with the aim to refine the initial draft by generating a better question in the second pass and then improving it with respect to a certain aspect. We can draw a comparison between this process and how humans tend to write a rough initial draft first and then refine it over multiple passes, where the later revisions focus on improving the draft aiming at certain aspects like fluency or completeness. With this motivation, we propose Refine Network (RefNet), which examines the initially generated question and performs a second pass to generate a revised question. Furthermore, we propose Reward-RefNet which uses explicit reward signals to achieve refinement focused on specific properties of the question such as fluency and answerability.

Our RefNet is a seq2seq based model that comprises of two decoders: Preliminary and Refinement Decoder. The Refinement Decoder takes the initial draft of the question generated by the Preliminary decoder as an input along with passage and answer, and generates the refined question by attending onto both the passage and the initial draft using a Dual Attention Network. The proposed dual attention aids RefNet to generate the final question by revisiting the appropriate parts of the input passage and initial draft. From Table 1, we can infer that our RefNet model is able to generate better questions in the second pass by fixing the errors in the initial draft. Our Reward-RefNet model uses REINFORCE with a baseline algorithm to explicitly reward the Refinement Decoder for generating a better question as compared to the Preliminary Decoder based on certain desired parameters like fluency and answerability. This leads to more answerable (see Reward-RefNet example for passage 11 in Table 1) and fluent (see Reward-RefNet example for passage 22 in Table 1) questions as compared to vanilla RefNet model.

Our experiments show that the proposed RefNet model outperforms existing state-of-the-art models on the SQuAD dataset by 12.312.3% and 3.73.7% (on BLEU) given the relevant sentence and passage respectively. We also achieve state-of-the-art results on HOTPOT-QA and DROP datasets with an improvement of 7.577.57% and 15.2515.25% respectively over the single-decoder baseline (on BLEU). Our human evaluations further validate these results. We further analyze and explain the impact of including the Refinement Decoder by examining the interaction between both the decoders. Interestingly, we observe that the inclusion of the Refinement Decoder boosts the quality of the questions generated by the initial decoder also. Lastly, our human evaluation of the questions generated by Reward-RefNet corroborate empirical results, i.e., it improves the question w.r.t. to fluency and answerability as compared to RefNet questions.

Refine Networks (RefNet) Model

We then use explicit rewards to enforce refinement on a desired metric, such as, fluency or answerability through our Reward-RefNet model. In the following sub-sections, we describe the passage encoder, preliminary and refinement decoders and our reward mechanism.

We use a 3 layered encoder consisting of: (i) Embedding, (ii) Contextual and (iii) Passage-Answer Fusion layers as described below. To capture interaction between passage and answer, we ensure that the passage and answer representations are fused together at every layer.

In this layer, we compute a dd-dimensional embedding for every word in the passage and the answer. This embedding is obtained by concatenating the word’s Glove embedding Pennington et al. (2014) with its character based embedding as discussed in Seo et al. (2016). Additionally, for passage words, we also compute a positional embedding based on the relative position of the word w.r.t. the answer span as described in Zhao et al. (2018). For every passage word, this positional embedding is also concatenated to the word and character based embeddings. We discuss the impact of character embeddings and answer tagging in Appendix A. In the subsequent sections, we will refer to embedding of the ii-th passage word wipw^{p}_{i} as e(wip)\mathbf{e(w^{p}_{i})} and the jj-th answer word wjaw^{a}_{j} as e(wja)\mathbf{e(w^{a}_{j})}.

In this layer, we compute a contextualized representation for every word in the passage by passing the word embeddings (as computed above) through a bidirectional-LSTM Hochreiter and Schmidhuber (1997):

where htp→\overrightarrow{\mathbf{h}^{p}_{t}} is the hidden state of the forward LSTM at time tt. We then concatenate the forward and backward hidden states as htp=[htp→;htp←]\mathbf{h}^{p}_{t}=[\overrightarrow{\mathbf{h}^{p}_{t}};\overleftarrow{\mathbf{h}^{p}_{t}}].

The answer could correspond to a span in the passage. Let j+1j+1 and j+nj+n be the start and end indices of the answer span in the passage respectively. We can thus refer to {hj+1p,…,hj+np}\{\mathbf{h}^{p}_{j+1},\dots,\mathbf{h}^{p}_{j+n}\} as the representation of the answer words in the context of the passage. We then obtain contextualized representations for the nn answer words by passing them through LSTM as follows:

The final state ha=[hna→;hna←]\mathbf{h}^{a}=[\overrightarrow{\mathbf{h}^{a}_{n}};\overleftarrow{\mathbf{h}^{a}_{n}}] of this Bi-LSTM is used as the answer representation in the subsequent stages. When the answer is not present in the passage, only e(wta)\mathbf{e(w^{a}_{t})} is passed to the LSTM.

In this layer, we refine the representations of the passage words based on the answer representation as follows:

2 Preliminary and Refinement Decoders

As discussed earlier, RefNet has two decoders, viz., Preliminary Decoder and Refinement Decoder, as described below:

This decoder generates an initial draft of the question, one word at a time, using an LSTM as follows:

Refinement Decoder: Once the preliminary decoder generates the entire question, the refinement decoder uses it to generate an updated version of the question using a Dual Attention Network. It first computes an attention weighted sum of the embeddings of the words generated by the first decoder as:

where βit\beta^{t}_{i} are parameterized and normalized attention weights computed by attention network A3\mathbf{A}_{3}. Since the initial draft could be erroneous or incomplete, we obtain additional information from the passage instead of only relying on the output of the first decoder. We do so by computing a context vector ct\mathbf{c}_{t} as

where γit\gamma^{t}_{i} are parameterized and normalized attention weights computed by attention network A3\mathbf{A}_{3}. The hidden state of the refinement decoder at time tt is computed as follows:

3 Reward-RefNet

4 Copy Module

Along with the above-mentioned three modules, we adopt the pointer-network and coverage mechanism from See et al. (2017). We use it to (i) handle Out-of-Vocabulary words and (ii) avoid repeating phrases in the generated questions.

Experimental Details

In this section, we discuss (i) the datasets for which we tested our proposed model, (ii) implementation details and (iii) evaluation metrics used to compare our model with the baseline and existing works.

SQuAD Rajpurkar et al. (2016): It contains 100100K (question, answer) pairs obtained from 536536 Wikipedia articles, where the answers are a span in the passage. For SQuAD, AQG has been tried from both sentences and passages. In the former case, only the sentence which contains the answer span is used as input, whereas in the latter case the entire passage is used. We use the same train-validation-test splits as used in Zhao et al. (2018).

Hotpot QA Yang et al. (2018) : Hotpot-QA is a multi-document and multi-hop QA dataset. Along with the triplet (P, A, Q), the authors also provide supporting facts that potentially lead to the answer. The answers here are either yes/no or answer span in P. We concatenate these supporting facts to form the passage. We use 1010% of the training data for validation and use the original dev set as test set.

DROP Dua et al. (2019): The DROP dataset is a reading comprehension benchmark which requires discrete reasoning over passage. It contains 9696K questions which require discrete operations such as addition, counting, or sorting to obtain the answer. We use 1010% of the original training data for validation and use the original dev set as test set.

2 Implementation Details

We use 300300 dimensional pre-trained Glove word embeddings, which are fixed during training. For character-level embeddings, we initially use a 2020 dimensional embedding for the characters which is then projected to 100100 dimensions. For answer-tagging, we use embedding size of 33. The hidden size for all the LSTMs is fixed to 512512. We use 2-layer, 1-layer and 2-layer stacked BiLSTM for the passage encoder, answer encoder and the decoders (both) respectively. We take the top 30,00030,000 frequent words as the vocabulary. We use Adam optimizer with a learning rate of 0.00040.0004 and train our models for 1010 epochs using cross entropy loss. For the Reward-RefNet model, we fine-tune the pre-trained model with the loss function mentioned in Section 2.3 for 33 epochs. The best model is chosen based on the BLEU Papineni et al. (2002) score on the validation split. For all the results we use beam search decoding with a beam size of 55.

3 Evaluation

We evaluate our models, based on nn-gram similarity metrics BLEU Papineni et al. (2002), ROUGE-L Lin (2004), and METEOR Lavie and Denkowski (2009) using the package released in Sharma et al. (2017)https://github.com/Maluuba/nlg-eval. We also quantify the answerability of our models using QBLEU-4https://github.com/PrekshaNema25/Answerability-MetricNema and Khapra (2018).

Results and Discussions

In this section, we present the results and analysis of our proposed model RefNet. Throughout this section, we refer to our models as follows:

is our single decoder model containing the encoder and the Preliminary Decoder described earlier. Note that the performance of this model is comparable to our implementation of the model proposed in Zhao et al. (2018).

includes the encoder, the Preliminary Decoder and the Refinement Decoder. We will (i) compare RefNet’s performance with EAD and existing models across all the mentioned datasets (ii) report human evaluations to compare RefNet and EAD (iii) analyze Refinement and Preliminary Decoders iv) present the performance of Reward RefNet with two different reward signal (fluency and answerability).

In Table 2, we compare the performance of RefNet with existing single decoder architectures across different datasets. On BLEU-4 metric, RefNet beats the existing state-of-the-art model by 12.3012.30%, 9.749.74%, 17.4817.48%, and 3.713.71% respectively on SQuAD (sentence), HOTPOT-QA, DROP and SQuAD (passage) dataset. Also it outperforms EAD by 7.837.83%, 7.577.57%, 15.2515.25% and 3.853.85% respectively on SQuAD (sentence), HOTPOT-QA, DROP and SQuAD (passage). In general, RefNet is consistently better than existing models across all nn-gram scores (BLEU, ROUGE-L and METEOR). Along with nn-gram scores, we also observe improvements on Q-BLEU4 as well, which as described earlier, gives a measure of both answerability and fluency.

2 Human Evaluations

We conducted human evaluations to analyze the quality of the questions produced by EAD and RefNet. We randomly sampled 500500 questions generated from the SQuAD (sentence level) dataset and asked the annotators to compare the quality of the generated questions. The annotators were shown a pair of questions, one generated by EAD and one by RefNet from the same sentence, and were asked to decide which one was better in terms of Fluency, Completeness, and Answerability. They were allowed to skip the question pairs where they could not make a clear choice. Three annotators rated each question and the final label was calculated based on majority voting. We observed that the RefNet model outperforms the EAD model across all three metrics. Over 68.668.6%, 66.766.7% and 64.264.2% of the generated questions from RefNet were respectively more fluent, complete and answerable when compared to the EAD model. However, there are some cases where EAD does better than RefNet. For example, in Table 3, we show that while trying to generate a more elaborate question, RefNet introduces an additional phrase “in the united” which is not required. Due to such instances, annotators preferred the EAD model in around 3030% of the instances.

3 Analysis of Refinement Decoder and Preliminary Decoder

The two decoders impact each other through two paths: (i) indirect path, where they share the encoder and the output projection to the vocabulary VV, (ii) direct path, via the dual attention network, where the initial draft of the question is attended by the Refinement Decoder. When RefNet has only indirect path, we can infer from row 1 of Table 4 that the performance of Preliminary Decoder improves when compared to the EAD model (16.8416.84 v/s 17.5917.59 BLEU). This suggests that generating two variants of the question improves the performance of the first decoder pass as well. This is perhaps due to the additional feedback that the shared encoder and output layer get from the Refinement Decoder. When we add the direct path (attention network) between the two decoders, the performance of the Refinement Decoder improves as compared to the Preliminary Decoder as shown in rows 3 and 4 of the Table 4 Comparison on Answerability: We also evaluate both the initial and refined draft using QBLEU4. As discussed earlier, Q-Metric measures Answerability using four components, viz., Named Entities, Important Words, Function Words, and Question Type. We observe that the increase in Q-Metric for refined questions is because the RefNet model can correct/add the relevant Named Entities in the question. In particular, we observe that the Named Entity component score in Q-Metric increases from 32.4232.42 for the first draft to 37.8137.81 for the refined draft.

Qualitative Analysis: Figure 2 shows that the RefNet model indeed generates more elaborate questions when compared to the Preliminary Decoder. As shown in Table 5, the quality of the refined question is better than the initial draft of the questions. Here RefNet adds the phrase “multi-tape Turing Machine,” (row 2) which removes any ambiguity in the question.

4 Analysis of Reward-RefNet

In this section, we analyze the impact of employing different reward signals in Reward-RefNet. As discussed earlier in section 2.3, we use fluency and answerability scores as reward signals. As shown in Table 6, when BLEU-4 (fluency) is used as a reward signal, there is improvement in BLEU-4 scores of Reward-RefNet as compared to RefNet model. We validated these results through human evaluations across 200200 samples. Annotators prefer the Reward-RefNet model in 6767% of the cases for fluency. Similarly when we use Answerability score as a reward signal, answerability improves for the model and annotators prefer the Reward-RefNet in 7070% of the cases for answerability. The performance of Reward-RefNet on fluency and answerability is similar for other datasets (see Appendix B).

Case Study: Originality of the Questions We observe that current state-of-the-art models perform very well in terms of BLEU/QBLEU scores when the actual question has significant overlap with the passage. For example, consider a passage from the SQuAD dataset in Table 7, where except the question word who, the model sequentially copies everything from the passage and achieves a QBLEU score of 92.492.4. However, the model performs poorly in situations where the true question is novel and does not contain a large sequence of words from the passage itself. In order to quantify this, we first sort the true questions based on its BLEU-2 overlap with the passage in ascending order. We then select the first NN true questions and compute the QBLEU score with the generated questions. The results are shown in red in Figure 3. Towards the left, where there are true questions with low overlap with the passage, the performance is poor, but it gradually improves as the overlap increases.

The task of generating questions with high originality (where the model phrases the question in its own words) is a challenging aspect of AQG since it requires complete understating of the semantics and syntax of the language. In order to improve questions generated on originality, we explicitly reward our model for having low nn-gram score with the passage as compared to the initial draft. As a result we observe that with Reward-RefNet(Originality), there is an improvement in the performance where the overlap with the passage was less (as shown in blue in Figure 3).As shown in Table 8, although both questions are answerable given the passage, the question generated from Reward-RefNet(Originality) is better.

Related Work

Early works on Question Generation were essentially rule based systems Heilman and Smith (2010); Mostow and Chen (2009); Lindberg et al. (2013); Labutov et al. (2015). Current models for AQG are based on the encode-attend-decode paradigm and they either generate questions from the passage alone Du and Cardie (2017); Du et al. (2017); Yao et al. (2018) or from the passage and a given answer (in which case the generated question must result in the given answer). Over the past couple of years, several variants of the encode-attend-decode model have been proposed. For example, Zhou et al. (2018) proposed a sequential copying mechanism to explicitly select a sub-span from the passage. Similarly, Zhao et al. (2018) mainly focuses on efficiently incorporating paragraph level content by using Gated Self Attention and Maxout pointer networks. Some works Yuan et al. (2017) even use Question Answering as a metric to evaluate the generated questions. There has also been some work on generating questions from images Jain et al. (2017); Li et al. (2017) and from knowledge bases Serban et al. (2016); Reddy et al. (2017). The idea of multi pass decoding which is central to our work has been used by Xia et al. (2017) for machine translation and text summarization albeit with a different objective. Some works have also augmented seq2seq models Rennie et al. (2017); Paulus et al. (2018); Song et al. (2017) with external reward signals using REINFORCE with baseline algorithm Williams (1992). The typical rewards used in these works are BLEU and ROUGE scores. Our REINFORCE loss is different from the previous ones as it uses the first decoder’s reward as the baseline instead of reward of the greedy policy.

Conclusion and Future Work

In this work, we proposed Refine Networks (RefNet) for Question Generation to focus on refining and improving the initial version of the generated question. Our proposed RefNet model consisting of a Preliminary Decoder and a Refinement Decoder with Dual Attention Network outperforms the existing state-of-the-art models on the SQuAD, HOTPOT-QA and DROP datasets. Along with automated evaluations, we also conducted human evaluations to validate our findings. We further showed that using Reward-RefNet improves the initial draft on specific aspects like fluency, answerability and originality. As a future work, we would like to extend RefNet to have the ability to decide whether a refinement is needed on the generated initial draft.

Acknowledgements

We thank Amazon Web Services for providing free GPU compute and Google for supporting Preksha Nema’s contribution in this work through Google Ph.D. Fellowship programme. We would like to acknowledge Department of Computer Science and Engineering, IIT Madras and Robert Bosch Center for Data Sciences and Artificial Intelligence, IIT Madras (RBC-DSAI) for providing us sufficient resources. We would also like to thank Patanjali SLPSK, Sahana Ramnath, Rahul Ramesh, Anirban Laha, Nikita Moghe and the anonymous reviewers for their valuable and constructive suggestions.

References

Appendix A Impact of Various Embeddings

We perform an ablation study to identify the impact of various word embeddings used in RefNet. When character embedding is not used in RefNet, the performance on SQuAD sentence-level drops from 18.1618.16 to 17.9717.97 BLEU-4 score. Meanwhile, when positional embeddings are dropped the performance decreases to 17.8717.87 BLEU-4 score.

Appendix B Reward-RefNet on Various Datasets

Table 9 shows the comparison between RefNet and Reward-RefNet on BLEU-4 score and answerability score when the respective scores are used as rewards in Reward-RefNet. We can infer from Table 9 that there is improvement in fluency and answerability across all the datasets.

Appendix C Visualization of Attention Weights

We plot the aggregated attention given to the passage and initial draft of the generation question across the various time-steps of the decoder in Figure 4. Although, both the questions are specific to the answer, A2\mathbf{A_{2}} pays some attention to the context surrounding the answer, which leads to a complete question. Also, note that in A3\mathbf{A_{3}}, while attending on to initial draft “oncogenic” word is not paid attention to and thus the final draft revises over the initial draft by correcting it to generate a better question.