Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering
Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, Wei Xu
Introduction
Question answering (QA) with neural network, i.e. neural QA, is an active research direction along the road towards the long-term AI goal of building general dialogue agents [\citenameWeston et al.2016]. Unlike conventional methods, neural QA does not rely on feature engineering and is (at least nearly) end-to-end trainable. It reduces the requirement for domain specific knowledge significantly and makes domain adaption easier. Therefore, it has attracted intensive attention in recent years.
Resolving QA problem requires several fundamental abilities including reasoning, memorization, etc. Various neural methods have been proposed to improve such abilities, including neural tensor networks [\citenameSocher et al.2013], recursive networks [\citenameIyyer et al.2014], convolution neural networks [\citenameYih et al.2014, \citenameDong et al.2015, \citenameYin et al.2015], attention models [\citenameHermann et al.2015, \citenameYin et al.2015, \citenameSantos et al.2016], and memories [\citenameGraves et al.2014, \citenameWeston et al.2015, \citenameKumar et al.2016, \citenameBordes et al.2015, \citenameSukhbaatar et al.2015], etc. These methods achieve promising results on various datasets, which demonstrates the high potential of neural QA. However, we believe there are still two major challenges for neural QA:
System development and/or evaluation on real-world data: Although several high quality and well-designed QA datasets have been proposed in recent years, there are still problems about using them to develop and/or evaluate QA system under real-world settings due to data size and the way they are created. For example, bAbI [\citenameWeston et al.2016] and the 30M Factoid Question-Answer Corpus [\citenameSerban et al.2016] are artificially synthesized; the TREC datasets [\citenameHarman and Voorhees2006], Free917 [\citenameCai and Yates2013] and WebQuestions [\citenameBerant et al.2013] are human generated but only have few thousands of questions; SimpleQuestions [\citenameBordes et al.2015] and the CNN and Daily Mail news datasets [\citenameHermann et al.2015] are large but generated under controlled conditions. Thus, a new large-scale real-world QA dataset is needed.
In this work, we address the above two challenges by a new dataset and a new neural QA model. Our contributions are two-fold:
We propose a new large-scale real-world factoid QA dataset WebQA with more than 42k questions and 566k evidences, where an evidence is a piece of text that contains relevant information to answer the question. On one hand, our dataset is an order of magnitude larger than existing real-world QA datasets [\citenameHarman and Voorhees2006, \citenameCai and Yates2013, \citenameBerant et al.2013], which are generally insufficient to train an end-to-end QA system. On the other hand, all questions in our dataset are asked by real-world users in daily life, which is significantly more close to real-world settings than those generated under controlled conditions [\citenameBordes et al.2015, \citenameHermann et al.2015]. Besides, as we also provide multiple human annotated evidences for each question, the dataset can be used in research such as evidence ranking and answer sentence selection as well.
Experimental results show that our model outperforms baselines with a large margin on the WebQA dataset, indicating that it is effective. Furthermore, our model even achieves an F1 score of 70.97% on character-based input, which is comparable with the 74.69% F1 score on word-based input, demonstrating that our model is robust.
Factoid QA as Sequence Labeling
In this work, we focus on open-domain factoid QA. Taking Figure 1 as an example, we formalize the problem as follows: given each question Q, we have one or more evidences E, and the task is to produce the answer A, where an evidence is a piece of text of any length that contains relevant information to answer the question. The advantage of this formalization is that evidences can be retrieved from web or unstructured knowledge base, which can improve system coverage significantly.
Formally, we formalize QA as a sequence labeling problem as follows: suppose we have a vocabulary of size , given question and evidence , where and are one-hot vectors of dimension , and and are the number of words in the question and evidence respectively. The problem is to find the label sequence which maximizes the conditional probability under parameter
In this work, we model by a neural network composed of LSTMs and CRF.
Recurrent Sequence Labeling Model
Figure 2 shows the structure of our model. The model consists of three components: (1) question LSTM for computing question representation; (2) evidence LSTMs for evidence analysis; and (3) a CRF layer for sequence labeling. The question LSTM in a form of a single layer LSTM equipped with a single time attention takes the question as input and generates the question representation . The three-layer evidence LSTMs takes the evidence, question representation and optional features as input and produces “features” for the CRF layer. The CRF layer takes the “features” as input and produces the label sequence. The details will be given in the following sections.
2 Long Short-Term Memory (LSTM)
Following [\citenameGraves2013], we define as a function mapping its input , previous state and output to current state and output :
3 Question LSTM
The question LSTM consists of a single-layer LSTM Multi-layer LSTMs can be used but no gain was observed. and a single-time attention model. The question is fed into the LSTM to produce a sequence of vector representations
4 Evidence LSTMs
The three-layer evidence LSTMs processes evidence to produce “features” for the CRF layer.
The first LSTM layer takes evidence , question representation and optional features as input. We find the following two simple common word indicator features are effective:
Question-Evidence common word feature (q-e.comm): for each word in the evidence, the feature has value 1 when the word also occurs in the question, otherwise 0. The intuition is that words occurring in questions tend not to be part of the answers for factoid questions.
Evidence-Evidence common word feature (e-e.comm): for each word in the evidence, the feature has value 1 when the word occurs in another evidence, otherwise 0. The intuition is that words shared by two or more evidences are more likely to be part of the answers.
Although counterintuitive, we found non-binary e-e.comm feature values does not work well. Because the more evidences we considered, the more words tend to get non-zero feature values, and the less discriminative the feature is.
The second LSTM layer stacks on top of the first LSTM layer, but processes its output in a reverse order. The third LSTM layer stacks upon the first and second LSTM layers with cross layer links, and its output serves as features for CRF layer.
Formally, the computations are defined as follows
5 Sequence Labeling
Following [\citenameHuang et al.2015, \citenameZhou and Xu2015], we use CRF on top of evidence LSTMs for sequence labeling. The probability of a label sequence given question and evidence is computed as
Training
We use a minibatch stochastic gradient descent (SGD) [\citenameLecun et al.1998] algorithm with rmsprop [\citenameTieleman and Hinton2012] to minimize the objective function. The initial learning rate is 0.001, batch size is 120, and . We also apply dropout [\citenameHinton et al.2012] to the output of all the LSTM layers. The dropout rate is 0.05. All these hyper-parameters are determined empirically via grid search on validation set.
WebQA Dataset
In order to train and evaluate open-domain factoid QA system for real-world questions, we build a new Chinese QA dataset named as WebQA. The dataset consists of tuples of (question, evidences, answer), which is similar to example in Figure 1. All the questions, evidences and answers are collected from web. Table 1 shows some statistics of the dataset.
The questions and answers are mainly collected from a large community QA website Baidu Zhidao http://zhidao.baidu.com and a small portion are from hand collected web documents. Therefore, all these questions are indeed asked by real-world users in daily life instead of under controlled conditions. All the questions are of single-entity factoid type, which means (1) each question is a factoid question and (2) its answer involves only one entity (but may have multiple words). The question in Figure 1 is a positive example, while the question “Who are the children of Albert Enistein?” is a counter example because the answer involves three persons. The type and correctness of all the question answer pairs are verified by at least two annotators.
All the evidences are retrieved from Internet by using a search engine with questions as queries. We download web pages returned in the first 3 result pages and take all the text pieces which have no more than 5 sentences and include at least one question word as candidate evidences. As evidence retrieval is beyond the scope of this work, we simply use TF-IDF values to re-rank these candidates.
For each question in the training set, we provide the top 10 ranked evidences to annotate (“Annotated Evidence” in Table 1). An evidence is annotated as positive if the question can be answered by just reading the evidence without any other prior knowledge, otherwise negative. Only evidences whose annotations are agreed by at least two annotators are retained. We also provide trivial negative evidences (“Retrieved Evidence” in Table 1), i.e. evidences that do not contain golden standard answers.
For each question in the validation and test sets, we provide one major positive evidence, and maybe an additional positive one to compute features. Both of them are annotated. Raw retrieved evidences are also provided for evaluation purpose (“Retrieved Evidence” in Table 1).
The dataset will be released on the project page http://idl.baidu.com/WebQA.html.
Evaluation on WebQA Dataset
We compare our model with two sets of baselines:
MemN2N [\citenameSukhbaatar et al.2015] is an end-to-end trainable version of memory networks [\citenameWeston et al.2015]. It encodes question and evidence with a bag-of-word method and stores the representations of evidences in an external memory. A recurrent attention model is used to retrieve relevant information from the memory to answer the question.
Attentive and Impatient Readers [\citenameHermann et al.2015] use bidirectional LSTMs to encode question and evidence, and do classification over a large vocabulary based on these two encodings. The simpler Attentive Reader uses a similar way as our work to compute attention for the evidence. And the more complex Impatient Reader computes attention after processing each question word.
The key difference between our model and the two readers is that they produce answer by doing classification over a large vocabulary, which is computationally expensive and has difficulties in handling unseen words. However, as our model uses an end-to-end trainable sequence labeling technique, it avoids both of the two problems by its nature.
2 Evaluation Method
The performance is measured with precision (P), recall (R) and F1-measure (F1) Measures such MAP and MRR are often also used for evaluating QA system. However, as our model gives only conditional probabilities which are not directly comparable for different answers, we will not include these measures in this work.
where is the list of correctly answered questions, is the list of produced answers, and is the list of all questions As the baselines will produce exactly one answer for each question, P, R and F1 will be identical for them..
As WebQA is collected from web, the same answer may be expressed in different surface forms in the golden standard answer and the evidence, e.g. “北京 (Beijing)” v.s. “北京市 (Beijing province)”. Therefore, we use two ways to count correctly answered questions, which are referred to as “strict” and “fuzzy” in the tables:
Strict matching: A question is counted if and only if the produced answer is identical to the golden standard answer;
Fuzzy matching: A question is counted if and only if the produced answer is a synonym The synonyms will also be released. of the golden standard answer;
And we also consider two evaluation settings:
Annotated evidence: Each question has one major annotated evidence and maybe another annotated evidence for computing q-e.comm and e-e.comm features (Section 3.4);
Retrieved evidence: Each question is provided with at most 20 automatically retrieved evidences (see Section 5 for details). All the evidences will be processed by our model independently and answers are voted by frequency to decide the final result. Note that a large amount of the evidences are negative and our model should not produce any answer for them.
3 Model Settings
If not specified, the following hyper-parameters will be used in the reset of this section: LSTM layer width (Section 3.2), word embedding dimension (Section 3.3), feature embedding dimension (Section 3.3). The word embeddings are initialized with pre-trained embeddings using a 5-gram neural language model [\citenameBengio et al.2003] and is fixed during training.
We will show that injecting noise data is important for improving performance on retrieved evidence setting in Section 6.5. In the following experiments, 20% of the training evidences will be negative ones randomly selected on the fly, of which 25% are annotated negative evidences and 75% are retrieved trivial negative evidences (Section 5). The percentages are determined empirically. Intuitively, we provide the noise data to teach the model learning to recognize unreliable evidence.
For each evidence, we will randomly sample another evidence from the rest evidences of the question and compare them to compute the e-e.comm feature (Section 3.4). We will develop more powerful models to process multiple evidences in a more principle way in the future.
As the answer for each question in our WebQA dataset only involves one entity (Section 5), we distinguish label Os before and after the first B in the label sequence explicitly to discourage our model to produce multiple answers for a question. For example, the golden labels for the example evidence in Figure 1 will became “Einstein/O1 married/O1 his/O1 first/O1 wife/O1 Mileva/B Marić/I in/O2 1903/O2”, where we use “O1” and “O2” to denote label Os before and after the first B All the words in a negative evidence will get label “O1”.. “Fuzzy matching” is also used for computing golden standard labels for training set.
For each setting, we will run three trials with different random seeds and report the average performance in the following sections.
4 Comparison with Baselines
As the baselines can only predict one-word answers, we only do experiments on the one-word answer subset of WebQA, i.e. only questions with one-word answers are retained for training, validation and test. As shown in Table 2, our model achieves significant higher F1 scores than all the baselines.
The main reason for the relative low performance of MemN2N is that it uses a bag-of-word method to encode question and evidence such that higher order information like word order is absent to the model. We think its performance can be improved by designing more complex encoding methods [\citenameHill et al.2016] and leave it as a future work.
The Attentive and Impatient Readers only have access to the fixed length representations when doing classification. However, our model has access to the outputs of all the time steps of the evidence LSTMs, and scores the label sequence as a whole. Therefore, our model achieves better performance.
5 Evaluation on the Entire WebQA Dataset
In this section, we evaluate our model on the entire WebQA dataset. The evaluation results are shown in Table 3. Although producing multi-word answers is harder, our model achieves comparable results with the one-word answer subset (Table 2), demonstrating that our model is effective for both single-word and multi-word word settings.
“Noise” in Table 3 means whether we inject noise data or not (Section 6.3). As all evidences are positive under the annotated evidence setting, the ability for recognizing unreliable evidence will be useless. Therefore, the performance of our model with and without noise is comparable under the annotated evidence setting. However, the ability is important to improve the performance under the retrieved evidence setting because a large amount of the retrieved evidences are negative ones. As a result, we observe significant improvement by injecting noise data for this setting.
6 Effect of Word Embedding
As stated in Section 6.3, the word embedding is initialized with LM embedding and kept fixed in training. We evaluate different initialization and optimization methods in this section. The evaluation results are shown in Table 4. The second row shows the results when the embedding is optimized jointly during training. The performance drops significantly. Detailed analysis reveals that the trainable embedding enlarge trainable parameter number and the model gets over fitting easily. The model acts like a context independent entity tagger to some extend, which is not desired. For example, the model will try to find any location name in the evidence when the word “在哪 (where)” occurs in the question. In contrary, pre-trained fixed embedding forces the model to pay more attention to the latent syntactic regularities. And it also carries basic priors such as “梨 (pear)” is fruit and “李世石 (Lee Sedol)” is a person, thus the model will generalize better to test data with fixed embedding. The third row shows the result when the embedding is randomly initialized and jointly optimized. The performance drops significantly further, suggesting that pre-trained embedding indeed carries meaningful priors.
7 Effect of q-e.comm and e-e.comm Features
As shown in Table 6, both the q-e.comm and e-e.comm features are effective, and the q-e.comm feature contributes more to the overall performance. The reason is that the interaction between question and evidence is limited and q-e.comm feature with value 1, i.e. the corresponding word also occurs in the question, is a strong indication that the word may not be part of the answer.
8 Effect of Question Representations
9 Effect of Evidence LSTMs Structures
We investigate the effect of evidence LSTMs layer number, layer width and cross layer links in this section. The results are shown in Figure 7. For fair comparison, we do not use cross layer links in Figure 7 (a) (dotted lines in Figure 2), and highlight the results with cross layer links (layer width 64) with circle and square for retrieved and annotated evidence settings respectively. We can conclude that: (1) generally the deeper and wider the model is, the better the performance is; (2) cross layer links are effective as they make the third evidence LSTM layer see information in both directions.
10 Word-based v.s. Character-based Input
Our model achieves fuzzy matching F1 scores of 69.78% and 70.97% on character-based input in annotated and retrieved evidence settings respectively (Table 7), which are only 3.72 and 3.72 points lower than the corresponding scores on word-based input respectively. The performance is promising, demonstrating that our model is robust and effective.
Conclusion and Future Work
In this work, we build a new human annotated real-world QA dataset WebQA for developing and evaluating QA system on real-world QA data. We also propose a new end-to-end recurrent sequence labeling model for QA. Experimental results show that our model outperforms baselines significantly.
There are several future directions we plan to pursue. First, multi-entity factoid and non-factoid QA are also interesting topics. Second, we plan to extend our model to multi-evidence cases. Finally, inspired by Residual Network [\citenameHe et al.2016], we will investigate deeper and wider models in the future.