Improving Question Generation with Sentence-level Semantic Matching and Answer Position Inferring

Xiyao Ma, Qile Zhu, Yanlin Zhou, Xiaolin Li, Dapeng Wu

Introduction

Question Generation (QG), an inverse problem of Question Answering (QA), aims to generate a semantically relevant question given a context and a corresponding answer. It has huge potential in education scenario (?), dialogue system, and question answering (?). A bunch of models using sequence-to-sequence (seq-to-seq) models (?) with the attention mechanism (?) have been proposed for neural question generation (?; ?).

Enriched linguistic features with Part-Of-Speech (POS) tags, relative position information, and paragraph context are incorporated in the embedding layers (?; ?; ?). Copy mechanism (?) is exploited to enhance the output quality of decoders (?; ?).

However, checking the questions generated by the strong baseline models NQG++ (?) and Pointer-generator (?) originally solving text summarization, the modern question generation models face the two main issues as follows: (1) Wrong keywords and question words: The model may ask questions with wrong keywords and wrong question words, as shown in the examples in Table 1. (2) Poor copy mechanism: The model copies the context words semantically irrelevant to the answer (?), as illustrated in the examples in Table 2.

Generally, the decoder with parameters θd\theta_{d} in seq-to-seq models (?; ?; ?) is trained by maximizing the generation probability p(yt∣y<t,z;θd)p(y_{t}|y_{<t},z;\theta_{d}) of the reference question word yty_{t}, given the previous generated words conditioned on the encoded context zz. However, the decoder may focus on local word semantics while ignoring the global question semantics during generation, resulting in above-mentioned issues. Meanwhile, the answer position-aware features are not exploited well by the copy mechanism, resulting in copying answer-irrelevant context words from input.

To alleviate these issues, we claim that learning the sentence-level semantics and answer position-awareness in a Multi-Task Learning (MTL) fashion results in a better performance as shown in Table 1 and 2. To do so, we first propose sentence-level semantic matching module for learning global semantics from both the encoder and decoder simultaneously. Then, answer position inferring module is introduced to enforce the model with the copy mechanism (?) to emphasize the relevant context words with the answer position-awareness. Furthermore, we propose answer-aware gated fusion mechanism for improved answer-aware sentence vector for decoder.

We further conduct extensive experiments on SQuAD (?) and MS MARCO (?) dataset to show the superiority of our proposed model. The experimental results show that our model not only outperforms the SOTA models on main metrics, auxiliary metrics, and human judgments, but also improves different models due to its generality. Our contributions are three-fold:

We analyze the questions generated by strong baselines and find two issues: wrong keywords and wrong question words and copying answer-irrelevant context words. We identify that lacking whole question semantics and expoiting answer position-awareness not weel are the key root causes.

To address the issues, we propose neural question generation model with sentence-level semantic matching, answer position inferring, and gated fusion.

We conduct extensive experiments to demonstrate the superiority of our proposed model for improving question generation performance in terms of the main metrics, auxiliary machine comprehension metrics, and human judgments. Besides, our work can improve current models significantly due to its generality.

Proposed Model

In this section, we describe the details of our proposed models, starting with an overview of question generation problem. Then, we illustrate our backbone seq-to-seq model with gated fusion for improved answer-aware sentence vector for generation. Finally, we illustrate sentence-level semantic matching and answer position inferring to alleviate the issues we discussed in the previous section.

In a question generation problem, a sentence X={xt}t=1MX=\{x_{t}\}^{M}_{t=1} containing an answe AA, a contiguous span of the sentence, is given to generate a question Y={yt}t=1NY=\{y_{t}\}^{N}_{t=1} matching with the sentence XX and the answer AA semantically.

Seq-to-seq model with Answer-aware Gated Fusion

Answer-aware Gated Fusion:

Instead of passing the last hidden state hmh_{m} of the encoder to the decoder as the initial hidden state, we propose gated fusion to provide an improved answer-aware sentence vector zz for the decoder.

Similar to the gates in LSTM, we use two information flow gates computed by SigmoidSigmoid functions to control the information flow of sentence vector and answer vector:

Decoder:

Meanwhile, the attention mechanism (?) is exploited by attending the current decoder state sts_{t} to the encoder context H=[h1,h2,...,hm]H=[h_{1},h_{2},...,h_{m}]. The context vector ctc_{t} is computed with normalized attention vector αt\alpha_{t} by weighted-sum:

Question word yty_{t} is generated from vocabulary VV with SoftmaxSoftmax function:

where ff is realized by a two-layer feed-forward network.

Copy Mechanism / Pointer-generator:

Generally, when generating the question word yty_{t}, a copy switch gcopyg_{copy} is computed to decide whether the generated word is generated from vocab or copied from source sentence, given the current decoder hidden state sts_{t} and context vector ctc_{t}:

where WcW^{c}, UcU^{c}, and bcb^{c} are learnable weights and biases.

The final word distribution is obtained by combining the probability of generate mode and the probability of copy mode:

where θ\theta, θ1\theta_{1}, and θ2\theta_{2} are the parameters of neural network.

We use the negative log likelihood for the seq-to-seq loss:

where θs2s\theta_{s2s} is the parameters of the seq-to-seq model, and N is the number of data in the train dataset.

Sentence-level Semantic Matching

Existing models, especially the decoders, generate question words given the generated and partial question words without considering the global whole question semantic, prone to wrong question words or keywords. Meanwhile, we observe that there exist different reference questions targeting the different answers in the same sentence in SQuAD and MARCO datasets. For example, we have <sentence,answer1,question1><sentence,answer1,question1> and <sentence,answer2,question2><sentence,answer2,question2>. However, the baseline model is prone to generating generic questions in this case. To overcome this problem, we propose the sentence-level semantic matching module to learn the sentence-level semantics from both the encoder and decoder sides in a multi-task learning way.

Generally, we have the improved answer-aware sentence vector zz obtained by our gated fusion. Regarding the decoder, a uni-directional LSTM, as an encoder for question, we take the last hidden state sns_{n} as the question vector.

where [zi,sni][z^{i},s_{n}^{i}] is the concatenation of the sentence vector snis_{n}^{i} and the question vector snis_{n}^{i}.

We jointly train the classifiers with seq-to-seq model in a multi-task learning fashion by minimizing the cross entropy loss:

where θsm\theta_{sm} is the parameters of the two classifiers, and psmpsm is the output probability of the classifiers.

Answer Position Inferring

Another issue of the baseline model is that it copies the answer-irrelevant words from the input sentence. One potential reason is that the model does not learn the answer position features well, and the attention matrix is not signified by the context words relevant to the answer. To address the issue, we leverage answer position inferring module to steer the model to learn the answer position-awareness, still in a Multi-Task Learning fashion.

Then, two two-layer bidirectional LSTMs are used to capture the interactions among the sentence words conditioned on the question (?). The answer starting index and end index are predicted by the output layer with SoftmaxSoftmax function:

where Wp1W_{p^{1}} and Wp2W_{p^{2}} are trainable weights, and ff function is a trainable multi-layer perception (MLP) network.

We compute the loss with the negative log likelihood of the ground truth answer starting index yi1y_{i}^{1} and ending index yi2y_{i}^{2} with the predicted distribution:

where θap\theta_{ap} is the parameters to be updated of the answer position inferring module, and yi1,yi2y_{i}^{1},y_{i}^{2} are the ground truth answer position labels.

To jointly train the generation model with the proposed modules in a multi-task Learning approach, we minimize the total loss during the training:

where α\alpha and β\beta control the magnitude of the sentence-level semantic matching loss and the answer position inferring loss, and θall\theta_{all} are all the parameters of our model. By minimizing the above loss function, our model is expected to discover and utilize the sentence-level and answer position-aware semantics of the question and sentence.

Experiments and Results

In this section, we conduct extensive experiments on the SQuAD and MS MARCO dataset, demonstrating the superiority of our proposed model compared with existing approaches.

SQuAD V1.1 dataset contains 536 Wikipedia articles and more than 100K questions posed about the articles (?). The answer is also given with corresponding questions as the sub-span of the sentence. Following the baseline (?), we use the training dataset (86635) to train our model, and we split the dev dataset into dev (8965) and test dataset (8964) with a ratio of 50%-50% for evaluation.

MS MARCO contains more than one million queries along with answers either generated by human or selected from passages (?). We select a subset of MS MARCO, where the answers are sub-spans of the passages. We split them into train set (86039), dev set (9480), and test set (7921) for model training and evaluation purpose.

We report automatic evaluation with BLEU-1, BLEU-2, BLEU-3, BLEU-4 (?), METEOR (?), and ROUGE-L (?) as the main metrics.

Baselines

In the experiments, we have several baselines for comparisons:

NQG++ (?): It is a baseline for Neural Question Generation task. It uses enriched semantic and lexical features in the encoder embedding layer of the seq-to-seq model. Attention mechanism and copy mechanismare also used.

Feature-enriched Pointer-generator (?): It is a seq-to-seq model with attention mechanism and copy mechanism. The copy mechanism is realized differently from NQG++. We add enriched features used in NQG++ in the embedding layer.

Answer-focused (?): It is a SOTA model on QG that uses an additional vocabulary for question word generation with relative answer position information instead of BIO used in NQG++.

Gated Self-attention (?). It is also a SOTA model on QG that leverages paragraph as input with gated self-attention above RNN in the encoder. Meanwhile, an improved maxout pointer is introduced.

Results and Analysis

We report the main metrics of different models on SQuAD and MS MARCO dataset in Table 3.

Answer-focused model (?) improves the performance by using separate vocabulary for question word generation along with answer relative position. The Gated Self-Attention model (?) emphasizes the intra-attention among the sentence with improved maxout pointer.

Different from the models above, our work aims to improve the model by learning the sentence-level semantic-matching features on both the encoder and decoder sides. The result shows that our model outperforms the two SOTA models on the main metrics.

Auxiliary Metrics

Although the main metrics can reflect the similarity between the generated question and the references, it has its limits on reflecting the semantics of generated question (?).

Alternatively, considering that machine comprehension takes the article and the corresponding question as the input to find the answer in the passages, we adopt the machine comprehension metrics (?) to evaluate the quality of the questions generated by different models (?).

We show the performances of BiDAF (?) pre-trained by AllenNLP (?) in terms of Exact Match (EM) and F1 metrics on reference questions, questions generated by baseline, and questions generated by our model in Table 4.

Our model outperforms NQG++ and Pointer-generator on EM and F1 significantly, since our model generates more answer-relevant questions by discovering sentence-level semantics and answer position features.

Sentence-level Semantic Matching Analysis

To analyze the quality of our model on generating the right question words and keywords, we randomly sample 200 questions generated by NQG++, Pointer-generator, and our model, respectively. Generally, the generated question is claimed to have the right question words if it has the same question words to the reference question. For example, we have a generated question ”what place …” and a reference question ”where …”, and we claim that the model generate a question with the right question words. In addition, we choose the words with most semantics importance as the keywords, which indicate the sentence topic and content. We report the number of the questions with right question words and keywords by different models in Table 5.

The main reason that our model outperforms the existing model is that learning the sentence-level semantics helps to capture the key semantics and results in better performance on generating the semantic-matching keywords.

Answer Position Inferring Analysis

We also conduct the similar experiment on evaluating the copy mechanisms in different models in terms of precision and recall used in (?). Given one generated question G and reference question R, we definite precision and recall as:

As reported in Table 6, the improvement of Precision and Recall indicates that answer position inferring can help copy OOV words from the input sentence.

Model Generality

To show the effectiveness and generality of our work, we evaluate the validness of our work by applying it to current representative models without revising the models. As shown in Table 7, our work can improve existing models by more than 2% on QG tasks due to its effectiveness and generality.

Human Evaluation

We also conduct human evaluation to examine the quality of the questions generated by the models and reference questions by scoring them on a scale of 1 to 5 in terms of semantics matching, fluency, and syntactically correctness. As reported in Table 8, our model generates questions with higher scores on the three metrics than the two baseline models, indicating the superiority of our proposed model by utilizing the sentence-level semantics and answer position-awareness.

Case Study

In this section, we present some examples of questions generated by our model.

Furthermore, we present a pair of examples, which have the same input sentence in Table 9. Different from that NQG++ generate similar and non-semantic-matching questions, our model can ask different and more semantic-matching questions than baselines, targeting the different answers.

Implementation Details

Followed NQG++ (?), we conduct our experiments on the preprocessed data provided by (?). We use 1 layer LSTM as the RNN cell for both the encoder and the decoder, and a bidirectional LSTM is used for the encoder. The hidden size of the encoder and decoder are 512. We use a 300-dimension pre-trained Glove vector as the word embedding (?). As same as NQG++ (?), the dimensions of lexical features and answer position are 16. We use Adam (?) Optimizer for model training with an initial learning rate as 0.001, and we halve it when the validation score does not improve. During the training of Sentence-level Semantic Matching module, we sample the negative sentences and questions from nearby data samples in the same batch, due to the preprocessed data (?) lacking of the information about which data samples are from the same passage. We compute our total loss function with α\alpha of 1 and β\beta of 2. Models are trained for 20 epochs with mini-batch of size 32. We choose model achieving the best performance on the dev dataset.

Related Work

Question generation tasks can be categorized into two classes: one is the rule-based method, meaning manually design lexical rules or templates to convert context into questions without deep understanding on the context semantic (?; ?). The other one is neural network based methods, which adopt seq-to-seq (?) or an encoder-decoder (?) framework to generate question words from scratches (?; ?). Our work focuses on the second category.

(?) firstly proposes to generate question with a seq-to-seq model given a context automatically. However, the model does not take the answer into consideration. Then (?) proposes to use a feature-enriched encoder to encode the input sentence by concatenating word embedding with lexical features as the encoder input, and answer position are involved in informing the model where the answer is. It is shown that it brings considerable improvements to the model. With the success of reinforcement learning, (?) propose to combine supervised learning and reinforcement learning together for question generation by using policy gradient after training the model in supervised learning way. The reward term in the policy gradient loss function can be perplexity and the BLEU scores (?). To tackle the issue that question words do not match with the answer type, (?) introduce a vocabulary only to generate question words. (?) propose to use paragraph as the input for providing more semantic information with an improved maxout pointer for copying words from the input.

Different from existing methods focusing on utilizing more informative features and improving the copy mechanism, we point out that incapability of capturing sentence-level semantics and exploiting answer-aware features are the main reasons, and we alleviate the problem by proposing two modules which can be integrated with any base models named sentence-level semantic matching and answer position inferring in Multi-Task Learning fashion.

Conclusion

In this paper, we observe two issues with the widely used baseline model on question generation. We point out the root cause is that existing models neither consider the whole question semantics nor exploit the answer position-aware features well. To address the issue, we propose the neural question generation model with sentence-level semantic matching, answer position inferring, and gated fusion. Extensive experimental results show that our work improves existing models significantly and outperforms the SOTA models on SQuAD and MARCO datasets.

Acknowledgement

This research was supported by CBL industry and agency members and by the IUCRC Program of the National Science Foundation under Grant No. CNS-1747783.

References