Hierarchical Graph Network for Multi-hop Question Answering

Yuwei Fang, Siqi Sun, Zhe Gan, Rohit Pillai, Shuohang Wang, Jingjing Liu

Introduction

In contrast to one-hop question answering Rajpurkar et al. (2016); Trischler et al. (2016); Lai et al. (2017) where answers can be derived from a single paragraph Wang and Jiang (2017); Seo et al. (2017); Liu et al. (2018); Devlin et al. (2019), many recent studies on question answering focus on multi-hop reasoning across multiple documents or paragraphs. Popular tasks include WikiHop Welbl et al. (2018), ComplexWebQuestions Talmor and Berant (2018), and HotpotQA Yang et al. (2018).

An example from HotpotQA is illustrated in Figure 1. In order to correctly answer the question (“The director of the romantic comedy ‘Big Stone Gap’ is based in what New York city”), the model is required to first identify P1 as a relevant paragraph, whose title contains the keywords that appear in the question (“Big Stone Gap”). S1, the first sentence of P1, is then chosen by the model as a supporting fact that leads to the next-hop paragraph P2. Lastly, from P2, the span “Greenwich Village, New York City” is selected as the predicted answer.

Most existing studies use a retriever to find paragraphs that contain the right answer to the question (P1 and P2 in this case). A Machine Reading Comprehension (MRC) model is then applied to the selected paragraphs for answer prediction Nishida et al. (2019); Min et al. (2019b). However, even after successfully identifying a reasoning chain through multiple paragraphs, it still remains a critical challenge how to aggregate evidence from scattered sources on different granularity levels (e.g., paragraphs, sentences, entities) for joint answer and supporting facts prediction.

To better leverage fine-grained evidences, some studies apply entity graphs through query-guided multi-hop reasoning. Depending on the characteristics of the dataset, answers can be selected either from entities in the constructed entity graph Song et al. (2018); Dhingra et al. (2018); De Cao et al. (2019); Tu et al. (2019); Ding et al. (2019), or from spans in documents by fusing entity representations back into token-level document representation Xiao et al. (2019). However, the constructed graph is mostly used for answer prediction only, while insufficient for finding supporting facts. Also, reasoning through a simple entity graph Ding et al. (2019) or paragraph-entity hybrid graph Tu et al. (2019) lacks the ability to support complicated questions that require multi-hop reasoning.

Intuitively, given a question that requires multiple hops through a set of documents to reach the right answer, a model needs to: (ii) identify paragraphs relevant to the question; (iiii) determine strong supporting evidence in those paragraphs; and (iiiiii) pinpoint the right answer following the garnered evidence. To this end, Graph Neural Network with its inherent message passing mechanism that can pass on multi-hop information through graph propagation, has great potential of effectively predicting both supporting facts and answer simultaneously for complex multi-hop questions.

Motivated by this, we propose a Hierarchical Graph Network (HGN) for multi-hop question answering, which empowers joint answer/evidence prediction via multi-level fine-grained graphs in a hierarchical framework. Instead of only using entities as nodes, for each question we construct a hierarchical graph to capture clues from sources with different levels of granularity. Specifically, four types of graph node are introduced: questions, paragraphs, sentences and entities (see Figure 2). To obtain contextualized representations for these hierarchical nodes, large-scale pre-trained language models such as BERT Devlin et al. (2019) and RoBERTa Liu et al. (2019) are used for contextual encoding. These initial representations are then passed through a Graph Neural Network for graph propagation. The updated node representations are then exploited for different sub-tasks (e.g., paragraph selection, supporting facts prediction, entity prediction). Since answers may not be entities in the graph, a span prediction module is also introduced for final answer prediction.

The main contributions of this paper are three-fold: (ii) We propose a Hierarchical Graph Network (HGN) for multi-hop question answering, where heterogeneous nodes are woven into an integral hierarchical graph. (iiii) Nodes from different granularity levels mutually enhance each other for different sub-tasks, providing effective supervision signals for both supporting facts extraction and answer prediction. (iiiiii) On the HotpotQA benchmark, the proposed model achieves new state of the art in both Distractor and Fullwiki settings.

Related Work

Multi-hop question answering requires a model to aggregate scattered pieces of evidence across multiple documents to predict the right answer. WikiHop Welbl et al. (2018) and HotpotQA Yang et al. (2018) are two recent datasets designed for this purpose. Existing work on HotpotQA Distractor setting focuses on converting the multi-hop reasoning task into single-hop sub-problems. Specifically, QFE Nishida et al. (2019) regards evidence extraction as a query-focused summarization task, and reformulates the query in each hop. DecompRC Min et al. (2019b) decomposes a compositional question into simpler sub-questions and leverages single-hop MRC models to answer the sub-questions. A neural modular network is also proposed in Jiang and Bansal (2019b), where neural modules are dynamically assembled for more interpretable multi-hop reasoning. Recent studies Chen and Durrett (2019); Min et al. (2019a); Jiang and Bansal (2019a) have also studied the multi-hop reasoning behaviors that models have learned in the task.

Graph Neural Network

Recent studies on multi-hop QA also build graphs based on entities and reasoning over the constructed graph using graph neural networks Kipf and Welling (2017); Veličković et al. (2018). MHQA-GRN Song et al. (2018) and Coref-GRN Dhingra et al. (2018) construct an entity graph based on co-reference resolution or sliding windows. Entity-GCN De Cao et al. (2019) considers three different types of edges that connect different entities in the entity graph. HDE-Graph Tu et al. (2019) enriches information in the entity graph by adding document nodes and creating interactions among documents, entities and answer candidates. Cognitive Graph QA Ding et al. (2019) employs an MRC model to predict answer spans and possible next-hop spans, and then organizes them into a cognitive graph. DFGN Xiao et al. (2019) constructs a dynamic entity graph, where in each reasoning step irrelevant entities are softly masked out and a fusion module is designed to improve the interaction between the entity graph and documents.

More recently, SAE Tu et al. (2020) defines three types of edge in the sentence graph based on the named entities and noun phrases appearing in the question and sentences. C2F Reader Shao et al. (2020) uses graph attention or self-attention on entity graph, and argues that this graph may not be necessary for multi-hop reasoning. Asai et al. (2020) proposes a new graph-based recurrent method to find evidence documents as reasoning paths, which is more focused on information retrieval. Different from the above methods, our proposed model constructs a hierarchical graph, effectively exploring relations on different granularities and employing different nodes to perform different tasks.

Hierarchical Coarse-to-Fine Modeling

Previous work on hierarchical modeling for question answering is mainly based on a coarse-to-fine framework. Choi et al. (2017) proposes to use reinforcement learning to first select relevant sentences and then produce answers from those sentences. Min et al. (2018) investigates the minimal context required to answer a question, and observes that most questions can be answered with a small set of sentences. Swayamdipta et al. (2018) constructs lightweight models and combines them into a cascade structure to extract the answer. Zhong et al. (2019) proposes to use hierarchies of co-attention and self-attention to combine information from evidence across multiple documents. Different from the above methods, our proposed model organizes different granularities in a hierarchical manner and leverages graph neural network to obtain the representations for different downstream tasks.

Hierarchical Graph Network

As illustrated in Figure 2, the proposed Hierarchical Graph Network (HGN) consists of four main components: (ii) Graph Construction Module (Sec. 3.1), through which a hierarchical graph is constructed to connect clues from different sources; (iiii) Context Encoding Module (Sec. 3.2), where initial representations of graph nodes are obtained via a RoBERTa-based encoder; (iiiiii) Graph Reasoning Module (Sec. 3.3), where graph-attention-based message passing algorithm is applied to jointly update node representations; and (iviv) Multi-task Prediction Module (Sec. 3.4), where multiple sub-tasks, including paragraph selection, supporting facts prediction, entity prediction, and answer span extraction, are performed simultaneously.

The hierarchical graph is constructed in two steps: (ii) identifying relevant multi-hop paragraphs; and (iiii) adding edges representing connections between sentences/entities within the selected paragraphs.

We first retrieve paragraphs whose titles match any phrases in the question (title matching). In addition, we train a paragraph ranker based on a pre-trained RoBERTa encoder, followed by a binary classification layer, to rank the probabilities of whether the input paragraphs contain the ground-truth supporting facts. If multiple paragraphs are found by title matching, only two paragraphs with the highest ranking scores are selected. If title matching returns no results, we further search for paragraphs that contain entities appearing in the question. If this also fails, the paragraph ranker will select the paragraph with the highest ranking score. The number of selected paragraphs in the first-hop is at most 2.

Once the first-hop paragraphs are identified, the next step is to find facts and entities within the paragraphs that can lead to other relevant paragraphs (i.e,, the second hop). Instead of relying on entity linking, which could be noisy, we use hyperlinks (provided by Wikipedia) in the first-hop paragraphs to discover second-hop paragraphs. Once the links are selected, we add edges between the sentences containing these links (source) and the paragraphs that the hyperlinks refer to (target), as illustrated by the dashed orange line in Figure 2. In order to allow information flow from both directions, the edges are considered as bidirectional.

Through this two-hop selection process, we are able to obtain several candidate paragraphs. In order to reduce introduced noise during inference, we use the paragraph ranker to select paragraphs with top-NN ranking scores in each step.

Nodes and Edges

Paragraphs are comprised of sentences, and each sentence contains multiple entities. This graph is naturally encoded in a hierarchical structure, and also motivates how we construct the hierarchical graph. For each paragraph node, we add edges between the node and all the sentences in the paragraph. For each sentence node, we extract all the entities in the sentence and add edges between the sentence node and these entity nodes. Optionally, edges between paragraphs and edges between sentences can also be included in the final graph.

Each type of these nodes captures semantics from different information sources. Thus, the hierarchical graph effectively exploits the structural information across all different granularity levels to learn fine-grained representations, which can locate supporting facts and answers more accurately than simpler graphs with homogeneous nodes.

An example hierarchical graph is illustrated in Figure 2. We define different types of edges as follows: (ii) edges between question node and paragraph nodes; (iiii) edges between question node and its corresponding entity nodes (entities appearing in the question, not shown for simplicity); (iiiiii) edges between paragraph nodes and their corresponding sentence nodes (sentences within the paragraph); (iviv) edges between sentence nodes and their linked paragraph nodes (linked through hyperlinks); (vv) edges between sentence nodes and their corresponding entity nodes (entities appearing in the sentences); (vivi) edges between paragraph nodes; and (viivii) edges between sentence nodes that appear in the same paragraph. Note that a sentence is only connected to its previous and next neighboring sentence. The final graph consists of these seven types of edges as well as four types of nodes, which link the question to paragraphs, sentences, and entities in a hierarchical way.

2 Context Encoding

3 Graph Reasoning

For graph propagation, we use Graph Attention Network (GAT) Veličković et al. (2018) to perform message passing over the hierarchical graph. Specifically, GAT takes all the nodes as input, and updates node feature hi′\mathbf{h}_{i}^{\prime} through its neighbors Ni\mathcal{N}_{i} in the graph. Formally,

The graph information will further contribute to the context information for answer span extraction. We merge the context representation M\mathbf{M} and the graph representation H′\mathbf{H}^{\prime} via a gated attention mechanism:

4 Multi-task Prediction

After graph reasoning, the updated node representations are used for different sub-tasks: (ii) paragraph selection based on paragraph nodes; (iiii) supporting facts prediction based on sentence nodes; and (iiiiii) answer prediction based on entity nodes and context representation G\mathbf{G}. Since the answers may not reside in entity nodes, the loss for entity node only serves as a regularization term.

In our HGN model, all three tasks are jointly performed through multi-task learning. The final objective is defined as:

where λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3}, and λ4\lambda_{4} are hyper-parameters, and each loss function is a cross-entropy loss, calculated over the logits (described below).

For both paragraph selection (Lpara\mathcal{L}_{para}) and supporting facts prediction (Lsent\mathcal{L}_{sent}), we use a two-layer MLP as the binary classifier:

We treat entity prediction (Lentity\mathcal{L}_{entity}) as a multi-class classification problem. Candidate entities include all entities in the question and those that match the titles in the context. If the ground-truth answer does not exist among the entity nodes, the entity loss is zero. Specifically,

The entity loss will only serve as a regularization term, and the final answer prediction will only rely on the answer span extraction module as follows.

The logits of every position being the start and end of the ground-truth span are computed by a two-layer MLP on top of G\mathbf{G} in Eqn.(4):

Following previous work Xiao et al. (2019), we also need to identify the answer type, which includes the types of span, entity, yes and no. We use a 3-way two-layer MLP for answer-type classification based on the first hidden representation of G\mathbf{G}:

During decoding, we first use this to determine the answer type. If it is “yes” or “no”, we directly return it as the answer. Overall, the final cross-entropy loss (Ljoint\mathcal{L}_{joint}) used for training is defined over all the aforementioned logits: osent,opara,oentity,ostart,oend,otype\mathbf{o}_{sent},\mathbf{o}_{para},\mathbf{o}_{entity},\mathbf{o}_{start},\mathbf{o}_{end},\mathbf{o}_{type}.

Experiments

In this section, we describe experiments comparing HGN with state-of-the-art approaches and provide detailed analysis on the model and results.

We use HotpotQA dataset Yang et al. (2018) for evaluation, a popular benchmark for multi-hop QA. Specifically, two sub-tasks are included in this dataset: (ii) Answer prediction; and (iiii) Supporting facts prediction. For each sub-task, exact match (EM) and partial match (F1) are used to evaluate model performance, and a joint EM and F1 score is used to measure the final performance, which encourages the model to take both answer and evidence prediction into consideration.

There are two settings in HotpotQA: Distractor and Fullwiki setting. In the Distractor setting, for each question, two gold paragraphs with ground-truth answers and supporting facts are provided, along with 8 ‘distractor’ paragraphs that were collected via a bi-gram TF-IDF retriever Chen et al. (2017). The Fullwiki setting is more challenging, which contains the same training questions as in the Distractor setting, but does not provide relevant paragraphs for test set. To obtain the right answer and supporting facts, the entire Wikipedia can be used to find relevant documents. Implementation details can be found in Appendix B.

2 Experimental Results

Table 1 and 2 summarize results on the hidden test set of HotpotQA. In Distractor setting, HGN outperforms both published and unpublished work on every metric by a significant margin, achieving a Joint EM/F1 score of 47.11/74.21 with an absolute improvement of 2.44/1.48 over previous state of the art. In Fullwiki setting, HGN achieves state-of-the-art results on Joint EM/F1 with 2.57/1.08 improvement, despite using an inferior retriever; when using the same retriever as in SemanticRetrievalMRS Yixin Nie (2019), our method outperforms by a significant margin, demonstrating the effectiveness of our multi-hop reasoning approach. In the following sub-sections, we provide a detailed analysis on the sources of performance gain on the dev set. Additional ablation study on paragraph selection is provided in Appendix D.

Effectiveness of Hierarchical Graph

As described in Section 3.1, we construct our graph with four types of nodes and seven types of edges. For ablation study, we build the graph step by step. First, we only consider edges from question to paragraphs, and from paragraphs to sentences, i.e., only edge type (ii), (iiiiii) and (iviv) are considered. We call this the PS Graph. Based on this, entity nodes and edges related to each entity node (corresponding to edge type (iiii) and (vv)) are added. We call this the PSE Graph. Lastly, edge types (vivi) and (viivii) are added, resulting in the final hierarchical graph.

As shown in Table 4, the use of PS Graph improves the joint F1 score over the plain RoBERTa model by 2.812.81 points. By further adding entity nodes, the Joint F1 increases by 0.300.30 points. This indicates that the addition of entity nodes is helpful, but may also bring in noise, thus only leading to limited performance improvement. By including edges among sentences and paragraphs, our final hierarchical graph provides an additional improvement of 0.240.24 points. We hypothesize that this is due to the explicit connection between sentences that leads to better representations.

Effectiveness of Pre-trained Language Model

To verify the effects of pre-trained language models, we compare HGN with prior state-of-the-art methods using the same pre-trained language models. Results in Table 5 show that our HGN variants outperform DFGN, EPS and SAE, indicating the performance gain comes from better model design.

3 Analysis

In this section, we provide an in-depth error analysis on the proposed model. HotpotQA provides two reasoning types: “bridge” and “comparison”. “Bridge” questions require the identification of a bridge entity that leads to the answer, while “comparison” questions compare two entities to infer the answer, which could be yes, no or a span of text. For analysis, we further split “comparison” questions into “comp-yn” and “comp-span”. Table 6 indicates that “comp-yn” questions are the easiest, on which our model achieves 88.5 joint F1 score. HGN performs similarly on “bridge” and “comp-span” with 74 joint F1 score, indicating that there is still room for further improvement.

To provide a more in-depth understanding of our model’s weaknesses (and provide insights for future work), we randomly sample 100 examples in the dev set with the answer F1 as 0. After carefully analyzing each example, we observe that these errors can be roughly grouped into six categories: (ii) Annotation: the annotation provided in the dataset is not correct; (iiii) Multiple Answers: questions may have multiple correct answers, but only one answer is provided in the dataset; (iii)iii) Discrete Reasoning: this type of error often appears in “comparison” questions, where discrete reasoning is required to answer the question correctly; (iviv) Commonsense & External Knowledge: to answer this type of question, commonsense or external knowledge is required; (vv) Multi-hop: the model fails to perform multi-hop reasoning, and finds the final answer from wrong paragraphs; (vivi) MRC: model correctly finds the supporting paragraphs and sentences, but predicts the wrong answer span.

Note that these error types are not mutually exclusive, but we aim to classify each example into only one type, in the order presented above. For example, if an error is classified as ‘Commonsense & External Knowledge’ type, it cannot be classified as ‘Multi-hop’ or ‘MRC’ error. Table 3 shows examples from each category (the corresponding paragraphs are omitted due to space limit).

We observed that a lot of errors are due to the fact that some questions have multiple answers with the same meaning, such as “a body of water vs. creek”, “EPA vs. Environmental Protection Agency”, and “American-born vs. U.S. born”. In these examples, the former is the ground-truth answer, and the latter is our model’s prediction. Secondly, for questions that require commonsense or discrete reasoning (e.g., “second” means “Code#02”Please refer to Row 4 in Table 3 for more context., “which band has more members”, or “who was born earlier”), our model just randomly picks an entity as answer, as it is incapable of performing this type of reasoning. The majority of the errors are from either multi-hop reasoning or MRC model’s span selection, which indicates that there is still room for further improvement. Additional examples are provided in Appendix F.

4 Generalizability Discussion

The hierarchical graph can be applied to different multi-hop QA datasets, though in this paper mainly tailored for HotpotQA. Here we use Wikipedia hyperlinks to connect sentences and paragraphs. An alternative way is to use an entity linking system to make it more generalizable. For each sentence node, if its entities exist in a paragraph, an edge can be added to connect the sentence and paragraph nodes. In our experiments, we restrict the number of multi-hops to two for the HotpotQA task, which can be increased to accommodate other datasets. The maximum number of paragraphs is set to four for HotpotQA, as we observe that using more documents within a maximum sequence length does not help much (see Table 9 in the Appendix). To generalize to other datasets that need to consume longer documents, we can either: (ii) use sliding-window-based method to chunk a long sequence into short ones; or (iiii) replace the BERT-based backbone with other transformer-based models that are capable of dealing with long sequences Beltagy et al. (2020); Zaheer et al. (2020); Wang et al. (2020).

Conclusion

In this paper, we propose a new approach, Hierarchical Graph Network (HGN), for multi-hop question answering. To capture clues from different granularity levels, our HGN model weaves heterogeneous nodes into a single unified graph. Experiments with detailed analysis demonstrate the effectiveness of our proposed model, which achieves state-of-the-art performances on the HotpotQA benchmark. Currently, in the Fullwiki setting, an off-the-shelf paragraph retriever is adopted for selecting relevant context from large corpus of text. Future work includes investigating the interaction and joint training between HGN and paragraph retriever for performance improvement.

References

Appendix A Datasets

There are two benchmark settings in HotpotQA: Distractor and Fullwiki setting. They both have 90k training samples and 7.4k development samples. In the Distractor setting, there are 2 gold paragraphs and 8 distractors. However, 2 gold paragraphs may not be available in the Fullwiki Setting. Therefore, the Fullwiki setting is more challenge which requires to search the entire Wikipedia to find relevant documents. For both settings, there are 90K hidden test samples. More details about the dataset can be found in Yang et al. (2018).

Appendix B Implementation Details

Our implementation is based on the Transformer library Wolf et al. (2019). To construct the proposed hierarchical graph, we use spacyhttps://spacy.io to extract entities from both questions and sentences. The numbers of entities, sentences and paragraphs in one graph are limited to 60, 40 and 4, respectively. Since HotpotQA only requires two-hop reasoning, up to two paragraphs are connected to each question. Our paragraph ranking model is a binary classifier based on the RoBERTa-large model. For the Fullwiki setting, we leverage the retrieved paragraphs and the paragraph ranker provided by Yixin Nie (2019). We finetune on the training set for 8 epochs, with batch size as 8, learning rate as 1e-5, λ1\lambda_{1} as 1, λ2\lambda_{2} as 5, λ3\lambda_{3} as 1, λ4\lambda_{4} as 1, LSTM dropout rate as 0.3 and GNN dropout rate as 0.3. We search hyperparameters for learning rate from {1e-5, 2e-5, 3e-5} , λ2\lambda_{2} from {1, 3, 5} and dropout rate from {0.1, 0.3, 0.5}.

Appendix C Computing Resources

We conduct experiments on 4 Quadro RTX 8000 GPUs. The parameters of each component in HGN are summarized in Table 7. The computation bottleneck is mainly from RoBERTa. The best model of HGN took around 12 hours for training, which is almost the same as the RoBERTa-large baseline.

Appendix D Effectiveness of Paragraph Selection

The proposed HGN relies on effective paragraph selection to find relevant multi-hop paragraphs. Table 8 shows the performance of paragraph selection on the dev set of HotpotQA. In DFGN, paragraphs are selected based on a threshold to maintain high recall (98.27%), leading to a low precision (60.28%). Compared to both threshold-based and pure Top-NN-based paragraph selection, our two-step paragraph selection process is more accurate, achieving 94.53% precision and 94.53% recall. Besides these two top-ranked paragraphs, we also include two other paragraphs with the next highest ranking scores, to obtain a higher coverage on potential answers. Table 9 summarizes the results on the dev set in the Distractor setting, using our paragraph selection approach for both DFGN and the plain BERT-base model. Note that the original DFGN does not finetune BERT, leading to much worse performance. In order to provide a fair comparison, we modify their released code to allow finetuning of BERT. Results show that our paragraph selection method outperforms the threshold-based one in both models.

Appendix E Case Study

We provide two example questions for case study. To answer the question in Figure 3 (left), QQ needs to be linked with P1P1. Subsequently, the sentence S4S4 within P1P1 is connected to P2P2 through the hyperlink (“John Surtees”) in S4S4. A plain BERT model without using the constructed graph missed S7S7 as additional supporting facts, while our HGN discovers and utilizes both pieces of evidence as the connections among S4S4, P2P2 and S7S7 are explicitly encoded in our hierarchical graph.

For the question in Figure 3 (right), the inference chain is Q→P1→S1→S2→P2→S3Q\rightarrow P1\rightarrow S1\rightarrow S2\rightarrow P2\rightarrow S3. The plain BERT model infers the evidence sentences S2S2 and S3S3 correctly. However, it fails to predict S1S1 as the supporting facts, while HGN succeeds, potentially due to the explicit connections between sentences in the constructed graph.

Appendix F Additional Examples for Error Analysis

Below, we provide additional examples for error analysis, where “Q” denotes question, “A” denotes answer provided with dataset and “P” denotes the prediction of proposed model. A full list of all the 100 examples is provided in Table 10 and 11.

Category: Annotation ID: 5ae2e0fd55429928c4239524 Q: What actor was also a president that Richard Darman worked with when they were in office? A: George H. W. Bush P: Ronald Reagan

ID: 5ab43b755542991779162c21 Q: What sports club based in Hamburg Germany had a Persian born football player who played for eight seasons? A: Mehdi Mahdavikia P: Hamburger SV

ID: 5a72e28f5542992359bc31ba Q: Which technique did the director at Pzena Investment Management outline? A: outlined by Joel Greenblatt P: Magic formula investing

ID: 5a7e71ab55429949594199bc Q: Perfect Imperfection is a 2016 Chinese romantic drama film starring a south Korean actor best known for his roles in what 2016 television drama? A: Reunited Worlds P: Cinderella and Four Knights

ID: 5a7a18b05542990783324e53 Q: What year was the independent regional brewery founded that currently operates in Hasting’s oldest pub? A: since 1864 P: 1698

Category: Multiple Answers ID: 5a8c9641554299585d9e36f5 Q: Which season of Alias does the English actor, who was born 25 June 1961, appear? A: three P: third season

ID: 5ae6179b5542992663a4f25b Q: Which Hong Kong actor born on 19 August 1946 starred in The Sentimental Swordsman A: Tommy Tam Fu-Wing P: Ti LungAlias of the true answer, Tommy Tam Fu-Wing

ID: 5abec66b5542997ec76fd360 Q: What do Josef Veltjens and Hermann Goering have in common? A: A veteran World War I fighter pilot ace P: German

ID: 5a85d6d95542996432c570fb Q: What is one element of House dance where the dancer ripples his or her torso back and forth? A: the jack P: Jacking

ID: 5a79c9395542994bb94570a2 Q: Which two occupations does Ronnie Dunn and Annie Lennox have in common? A: singer, songwriter P: singer-songwriter

Category: Discrete Reasoning ID: 5a8ec3205542995a26add506 Q: Does Dashboard Confessional have more members than World Party? A: yes P: no

ID: 5abfd83f5542997ec76fd45c Q: Which genus has more species, Quesnelia or Honeysuckle? A: Honeysuckle P: Honeysuckles

ID: 5ac44b47554299194317396c Q: Which became a Cathedral first St Chad’s Cathedral, Birmingham or Chelmsford Cathedral? A: Metropolitan Cathedral Church and Basilica of Saint Chad P: St Chad’s

ID: 5ac2455e55429951e9e68512 Q: Were both Life magazine and Strictly Slots magazine published monthly in 1998? A: yes P: no

ID: 5a7d26bd554299452d57bb28 Q: Who was born earlier, Johnny Lujack or Jim Kelly? A: Jim Kelly P: John Christopher Lujack

Category: Commonsense & External Knowledge ID: 5ac275e755429921a00aaf81 Q: From what nation is the football player who was named Man of the Match at the 2001 Intercontinental Cup? A: Ghana P: Ghanaian

ID: 5ac02d345542992a796decc0 Q: Where are Abbey Clancy and Peter Crouch from? A: England P: English

ID: 5ab2beba554299166977408f Q: Who is the father of the Prince in which William Joseph Weaver is most famous for painting a full length portrait of? A: George III P: Queen Victoria

ID: 5a8dab16554299068b959d89 Q: What type of elevation does Aldgate railway station, Adelaide and Aldgate, South Australia have in common? A: Hills P: kilometres

ID: 5a82edae55429966c78a6a9f Q: Swiss music duo Double released their best known single ”The Captain of Her Heart” in what year? A: 1986 P: 1985

Category: Multi-hop ID: 5a7a46605542994f819ef1ad Q: What year did Roy Rogers and his third wife star in a film directed by Frank McDonald? A: 1945 P: 1946

ID: 5a84f7255542991dd0999e33 Q: Which country borders the Central African Republic and is south of Libya and east of Niger? A: Republic of Chad P: Sudan

ID: 5a77152355429966f1a36c2e Q: What was the Roud Folk Song Index of the nursery rhyme inspiring What Are Little Girls Made Of? A: 821 P: 326

ID: 5a7e7c725542991319bc94be Q: In what year did Farda Amiga win a race at the Saratoga Race course? A: (foaled February 1, 1999) P: 1872

ID: 5ae21ef35542994d89d5b35d Q: What college teamdid the point guard that led the way for Philedlphia 76ers in the 2017-18 season play basketball in? A: Washington Huskies P: University of Kansas

Category: MRC ID: 5ae5cf625542996de7b71a22 Q: What sports team included both of the brothers Case McCoy and Colt McCoy during different years? A: University of Texas Longhorns P: Washington Redskins

ID: 5a8fa4a5554299458435d6a3 Q: What is name of the business unit led by Tina Sharkey at a web portal which is originally known as America Online? A: Sesame Street P: community programming

ID: 5a8135cc55429903bc27b943 Q: In the USA, gun powder is used in conjunction with this to start the Boomershot. A: Anvil firing P: an explosive fireball

ID: 5a84bb825542991dd0999dbe Q: Who beacme a star as a comic book character created by Gerry Conway and Bob Oksner? A: Megalyn Echikunwoke P: Stephen Amell

ID: 5a75f1a755429976ec32bcb1 Q: Which actress played a character that dated Mark Brendanawicz? A: Rashida Jones P: Amy Poehler