Neural Question Generation from Text: A Preliminary Study

Qingyu Zhou, Nan Yang, Furu Wei, Chuanqi Tan, Hangbo Bao, Ming Zhou

Introduction

Automatic question generation from natural language text aims to generate questions taking text as input, which has the potential value of education purpose (Heilman, 2011). As the reverse task of question answering, question generation also has the potential for providing a large scale corpus of question-answer pairs.

Previous works for question generation mainly use rigid heuristic rules to transform a sentence into related questions (Heilman, 2011; Chali and Hasan, 2015). However, these methods heavily rely on human-designed transformation and generation rules, which cannot be easily adopted to other domains. Instead of generating questions from texts, Serban et al. (2016) proposed a neural network method to generate factoid questions from structured data.

In this work we conduct a preliminary study on question generation from text with neural networks, which is denoted as the Neural Question Generation (NQG) framework, to generate natural language questions from text without pre-defined rules. The Neural Question Generation framework extends the sequence-to-sequence models by enriching the encoder with answer and lexical features to generate answer focused questions. Concretely, the encoder reads not only the input sentence, but also the answer position indicator and lexical features. The answer position feature denotes the answer span in the input sentence, which is essential to generate answer relevant questions. The lexical features include part-of-speech (POS) and named entity (NER) tags to help produce better sentence encoding. Lastly, the decoder with attention mechanism Bahdanau et al. (2015) generates an answer specific question of the sentence.

Large-scale manually annotated passage and question pairs play a crucial role in developing question generation systems. We propose to adapt the recently released Stanford Question Answering Dataset (SQuAD) (Rajpurkar et al., 2016) as the training and development datasets for the question generation task. In SQuAD, the answers are labeled as subsequences in the given sentences by crowed sourcing, and it contains more than 100K questions which makes it feasible to train our neural network models. We conduct the experiments on SQuAD, and the experiment results show the neural network models can produce fluent and diverse questions from text.

Approach

In this section, we introduce the NQG framework, which consists of a feature-rich encoder and an attention-based decoder. Figure 1 provides an overview of our NQG framework.

In the NQG framework, we use Gated Recurrent Unit (GRU) Cho et al. (2014) to build the encoder. To capture more context information, we use bidirectional GRU (BiGRU) to read the inputs in both forward and backward orders. Inspired by Chen and Manning (2014); Nallapati et al. (2016), the BiGRU encoder not only reads the sentence words, but also handcrafted features, to produce a sequence of word-and-feature vectors. We concatenate the word vector, lexical feature embedding vectors and answer position indicator embedding vector as the input of BiGRU encoder. Concretely, the BiGRU encoder reads the concatenated sentence word vector, lexical features, and answer position feature, x=(x1,x2,…,xn)x=(x_{1},x_{2},\dots,x_{n}), to produce two sequences of hidden vectors, i.e., the forward sequence (h⃗1,h⃗2,…,h⃗n)(\vec{h}_{1},\vec{h}_{2},\dots,\vec{h}_{n}) and the backward sequence (\reflectbox{\vec{\reflectbox{hh}}}_{1},\reflectbox{\vec{\reflectbox{hh}}}_{2},\dots,\reflectbox{\vec{\reflectbox{hh}}}_{n}). Lastly, the output sequence of the encoder is the concatenation of the two sequences, i.e., h_{i}=[\vec{h}_{i};\reflectbox{\vec{\reflectbox{hh}}}_{i}].

To generate a question with respect to a specific answer in a sentence, we propose using answer position feature to locate the target answer. In this work, the BIO tagging scheme is used to label the position of a target answer. In this scheme, tag B denotes the start of an answer, tag I continues the answer and tag O marks words that do not form part of an answer. The BIO tags of answer position are embedded to real-valued vectors throu and fed to the feature-rich encoder. With the BIO tagging feature, the answer position is encoded to the hidden vectors and used to generate answer focused questions.

Lexical Features

Besides the sentence words, we also feed other lexical features to the encoder. To encode more linguistic information, we select word case, POS and NER tags as the lexical features. As an intermediate layer of full parsing, POS tag feature is important in many NLP tasks, such as information extraction and dependency parsing (Manning et al., 1999). Considering that SQuAD is constructed using Wikipedia articles, which contain lots of named entities, we add NER feature to help detecting them.

2 Attention-Based Decoder

We employ an attention-based GRU decoder to decode the sentence and answer information to generate questions. At decoding time step tt, the GRU decoder reads the previous word embedding wt−1w_{t-1} and context vector ct−1c_{t-1} to compute the new hidden state sts_{t}. We use a linear layer with the last backward encoder hidden state \reflectbox{\vec{\reflectbox{hh}}}_{1} to initialize the decoder GRU hidden state. The context vector ctc_{t} for current time step tt is computed through the concatenate attention mechanism (Luong et al., 2015), which matches the current decoder state sts_{t} with each encoder hidden state hih_{i} to get an importance score. The importance scores are then normalized to get the current context vector by weighted sum:

We then combine the previous word embedding wt−1w_{t-1}, the current context vector ctc_{t}, and the decoder state sts_{t} to get the readout state rtr_{t}. The readout state is passed through a maxout hidden layer (Goodfellow et al., 2013) to predict the next word with a softmax layer over the decoder vocabulary:

where rtr_{t} is a 2d2d-dimensional vector.

3 Copy Mechanism

To deal with the rare and unknown words problem, Gulcehre et al. (2016) propose using pointing mechanism to copy rare words from source sentence. We apply this pointing method in our NQG system. When decoding word tt, the copy switch takes current decoder state sts_{t} and context vector ctc_{t} as input and generates the probability pp of copying a word from source sentence:

where σ\sigma is sigmoid function. We reuse the attention probability in equation 4 to decide which word to copy.

Experiments and Results

We use the SQuAD dataset as our training data. SQuAD is composed of more than 100K questions posed by crowd workers on 536 Wikipedia articles. We extract sentence-answer-question triples to build the training, development and test setsWe re-distribute the processed data split and PCFG-Trans baseline code at http://res.qyzhou.me . Since the test set is not publicly available, we randomly halve the development set to construct the new development and test sets. The extracted training, development and test sets contain 86,635, 8,965 and 8,964 triples respectively. We introduce the implementation details in the appendix.

We conduct several experiments and ablation tests as follows:

The rule-based system1 modified on the code released by Heilman (2011). We modified the code so that it can generate question based on a given word span.

We implement a seq2seq with attention as the baseline method.

We extend the s2s+att with our feature-rich encoder to build the NQG system.

Based on NQG, we incorporate copy mechanism to deal with rare words problem.

Based on NQG+, we initialize the word embedding matrix with pre-trained GloVe (Pennington et al., 2014) vectors.

Based on NQG+, we make the encoder and decoder share the same embedding matrix.

Based on NQG+, we use both pre-train word embedding and STshare methods, to further improve the performance.

Ablation test, the answer position indicator is removed from NQG model.

Ablation test, the POS tag feature is removed from NQG model.

Ablation test, the NER feature is removed from NQG model.

Ablation test, the word case feature is removed from NQG model.

We report BLEU-4 score Papineni et al. (2002) as the evaluation metric of our NQG system.

Table 1 shows the BLEU-4 scores of different settings. We report the beam search results on both development and test sets. Our NQG framework outperforms the PCFG-Trans and s2s+att baselines by a large margin. This shows that the lexical features and answer position indicator can benefit the question generation. With the help of copy mechanism, NQG+ has a 2.05 BLEU improvement since it solves the rare words problem. The extended version, NQG++, has 1.11 BLEU score gain over NQG+, which shows that initializing with pre-trained word vectors and sharing them between encoder and decoder help learn better word representation.

We evaluate the PCFG-Trans baseline and NQG++ with human judges. The rating scheme is, Good (3) - The question is meaningful and matches the sentence and answer very well; Borderline (2) - The question matches the sentence and answer, more or less; Bad (1) - The question either does not make sense or matches the sentence and answer. We provide more detailed rating examples in the supplementary material. Three human raters labeled 200 questions sampled from the test set to judge if the generated question matches the given sentence and answer span. The inter-rater aggreement is measured with Fleiss’ kappa Fleiss (1971).

Table 2 reports the human judge results. The kappa scores show a moderate agreement between the human raters. Our NQG++ outperforms the PCFG-Trans baseline by 0.76 score, which shows that the questions generated by NQG++ are more related to the given sentence and answer span.

Ablation Test

The answer position indicator, as expected, plays a crucial role in answer focused question generation as shown in the NQG−-Answer ablation test. Without it, the performance drops terribly since the decoder has no information about the answer subsequence.

Ablation tests, NQG−-Case, NQG−-POS and NQG−-NER, show that word case, POS and NER tag features contributes to question generation.

Case Study

Table 3 provides three examples generated by NQG++. The words with underline are the target answers. These three examples are with different question types, namely WHEN, WHAT and WHO respectively. It can be observed that the decoder can ‘copy’ spans from input sentences to generate the questions. Besides the underlined words , other meaningful spans can also be used as answer to generate correct answer focused questions.

Type of Generated Questions

Following Wang and Jiang (2016), we classify the questions into different types, i.e., WHAT, HOW, WHO, WHEN, WHICH, WHERE, WHY and OTHER.We treat questions ‘what country’, ‘what place’ and so on as WHERE type questions. Similarly, questions containing ‘what time’, ‘what year’ and so forth are counted as WHEN type questions. We evaluate the precision and recall of each question types. Figure 2 provides the precision and recall metrics of different question types. The precision and recall of a question type TT are defined as:

For the majority question types, WHAT, HOW, WHO and WHEN types, our NQG++ model performs well for both precision and recall. For type WHICH, it can be observed that neither precision nor recall are acceptable. Two reasons may cause this: a) some WHICH-type questions can be asked in other manners, e.g., ‘which team’ can be replaced with ‘who’; b) WHICH-type questions account for about 7.2% in training data, which may not be sufficient to learn to generate this type of questions. The same reason can also affect the precision and recall of WHY-type questions.

Conclusion and Future Work

In this paper we conduct a preliminary study of natural language question generation with neural network models. We propose to apply neural encoder-decoder model to generate answer focused questions based on natural language sentences. The proposed approach uses a feature-rich encoder to encode answer position, POS and NER tag information. Experiments show the effectiveness of our NQG method. In future work, we would like to investigate whether the automatically generated questions can help to improve question answering systems.

References

Appendix A Implementation Details

We use the same vocabulary for both encoder and decoder. The vocabulary is collected from the training data and we keep the top 20,000 frequent words. We set the word embedding size to 300 and all GRU hidden state sizes to 512. The lexical and answer position features are embedded to 32-dimensional vectors. We use dropout (Srivastava et al., 2014) with probability p=0.5p=0.5. During testing, we use beam search with beam size 12.

A.2 Lexical Feature Annotation

We use Stanford CoreNLP v3.7.0 (Manning et al., 2014) to annotate POS and NER tags in sentences with its default configuration and pre-trained models.

A.3 Model Training

We initialize model parameters randomly using a Gaussian distribution with Xavier scheme (Glorot and Bengio, 2010). We use a combination of Adam (Kingma and Ba, 2015) and simple SGD as our the optimizing algorithms. The training is separated into two phases, the first phase is optimizing the loss function with Adam and the second is with simple SGD. For the Adam optimizer, we set the learning rate α=0.001\alpha=0.001, two momentum parameters β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999 respectively, and ϵ=10−8\epsilon=10^{-8}. We use Adam optimizer until the BLEU score on the development set drops for six consecutive tests (we test the BLEU score on the development set for every 1,000 batches). Then we switch to a simple SGD optimizer with initial learning rate α=0.5\alpha=0.5 and halve it if the BLEU score on the development set drops for twelve consecutive tests. We also apply gradient clipping (Pascanu et al., 2013) with range $$ for both Adam and SGD phases. To both speed up the training and converge quickly, we use mini-batch size 64 by grid search.

Appendix B Human Evaluation Examples

We evaluate the PCFG-Trans baseline and NQG++ with human judges. The rating scheme is provided in Table 4.

The human judges are asked to label the generated questions if they match the given sentence and answer span according to the rating scheme and examples. We provide some example questions with different scores in Table 5. For the first score 3 example, the question makes sense and the target answer “reason” can be used to answer it given the input sentence. For the second score 2 example, the question is inadequate for answering the sentence since the answer is about prime number. However, given the sentence, a reasonable person will give the targeted answer of the question. For the third score 1 example, the question is totally wrong given the sentence and answer.