Compositional Questions Do Not Necessitate Multi-hop Reasoning

Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, Luke Zettlemoyer

Introduction

Multi-hop reading comprehension (RC) requires reading and aggregating information over multiple pieces of textual evidence (Welbl et al., 2017; Yang et al., 2018; Talmor and Berant, 2018). In this work, we argue that it can be difficult to construct large multi-hop RC datasets. This is because multi-hop reasoning is a characteristic of both the question and the provided evidence; even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant. For example, the question in Figure 1 is compositional: a plausible solution is to find “What animal’s habitat was the Réserve Naturelle Lomako Yokokala established to protect?”, and then answer “What is the former name of that animal?”. However, when considering the evidence paragraphs, the question is solvable in a single hop by finding the only paragraph that describes an animal.

Our analysis is centered on HotpotQA Yang et al. (2018), a dataset of mostly compositional questions. In its RC setting, each question is paired with two gold paragraphs, which should be needed to answer the question, and eight distractor paragraphs, which provide irrelevant evidence or incorrect answers. We show that single-hop reasoning can solve much more of this dataset than previously thought. First, we design a single-hop QA model based on BERT (Devlin et al., 2018), which, despite having no ability to reason across paragraphs, achieves performance competitive with the state of the art. Next, we present an evaluation demonstrating that humans can solve over 80% of questions when we withhold one of the gold paragraphs.

To better understand these results, we present a detailed analysis of why single-hop reasoning works so well. We show that questions include redundant facts which can be ignored when computing the answer, and that the fine-grained entity types present in the provided paragraphs in the RC setting often provide a strong signal for answering the question, e.g., there is only one animal in the given paragraphs in Figure 1, allowing one to immediately locate the answer using one hop.

This analysis shows that more carefully chosen distractor paragraphs would induce questions that require multi-hop reasoning. We thus explore an alternative method for collecting distractors based on adversarial paragraph selection. Although this appears to mitigate the problem, a single-hop model re-trained on these distractors can recover most of the original single-hop accuracy, indicating that these distractors are still insufficient. Another method is to consider very large distractor sets such as all of Wikipedia or the entire Web, as done in open-domain HotpotQA and ComplexWebQuestions Talmor and Berant (2018). However, this introduces additional computational challenges and/or the need for retrieval systems. Finding a small set of distractors that induce multi-hop reasoning remains an open challenge that is worthy of follow up work.

Related Work

Large-scale RC datasets (Hermann et al., 2015; Rajpurkar et al., 2016; Joshi et al., 2017) have enabled rapid advances in neural QA models (Seo et al., 2017; Xiong et al., 2018; Yu et al., 2018; Devlin et al., 2018). To foster research on reasoning across multiple pieces of text, multi-hop QA has been introduced Kočiskỳ et al. (2018); Talmor and Berant (2018); Yang et al. (2018). These datasets contain compositional or “complex” questions. We demonstrate that these questions do not necessitate multi-hop reasoning.

Existing multi-hop QA datasets are constructed using knowledge bases, e.g., WikiHop Welbl et al. (2017) and ComplexWebQuestions Talmor and Berant (2018), or using crowd workers, e.g., HotpotQA Yang et al. (2018). WikiHop questions are posed as triples of a relation and a head entity, and the task is to determine the tail entity of the relationship. ComplexWebQuestions consists of open-domain compositional questions, which are constructed by increasing the complexity of SPARQL queries from WebQuestions (Berant et al., 2013). We focus on HotpotQA, which consists of multi-hop questions written to require reasoning over two paragraphs from Wikipedia.

Parallel research from Chen and Durrett (2019) presents similar findings on HotpotQA. Our work differs because we conduct human analysis to understand why questions are solvable using single-hop reasoning. Moreover, we show that selecting distractor paragraphs is difficult using current retrieval methods.

Single-paragraph QA

This section shows the performance of a single-hop model on HotpotQA.

Our model, single-paragraph BERT, scores and answers each paragraph independently (Figure 2). We then select the answer from the paragraph with the best score, similar to Clark and Gardner (2018).Full details in Appendix B. Code available at https://github.com/shmsw25/single-hop-rc.

The model receives a question Q=[q1,..,qm]Q=[q_{1},..,q_{m}] and a single paragraph P=[p1,...,pn]P=[p_{1},...,p_{n}] as input. Following Devlin et al. (2018), S=[q1,...,qm,[SEP],p1,...,pn]S=[q_{1},...,q_{m},{\tt[SEP]},p_{1},...,p_{n}], where [SEP]{\tt[SEP]} is a special token, is fed into BERT:

2 Model Results

HotpotQA has two settings: a distractor setting and an open-domain setting.

The HotpotQA distractor setting pairs the two paragraphs the question was written for (gold paragraphs) with eight spurious paragraphs selected using TF-IDF similarity with the question (distractors). Our single-paragraph BERT model achieves 67.08 F1, comparable to the state-of-the-art (Table 1).Results as of March 4th, 2019. This indicates the majority of HotpotQA questions are answerable in the distractor setting using a single-hop model.

Open-domain Setting

The HotpotQA open-domain setting (Fullwiki) does not provide a set of paragraphs—all of Wikipedia is considered. We follow Chen et al. (2017) and retrieve paragraphs using bigram TF-IDF similarity with the question.

Compositional Questions Are Not Always Multi-hop

This section provides a human analysis of HotpotQA to understand what phenomena enable single-hop answer solutions. HotpotQA contains two question types, Bridge and Comparison, which we evaluate separately.

Bridge questions consist of two paragraphs linked by an entity Yang et al. (2018), e.g., Figure 1. We first investigate single-hop human performance on HotpotQA bridge questions using a human study consisting of NLP graduate students. Humans see the paragraph that contains the answer span and the eight distractor paragraphs, but do not see the other gold paragraph. As a baseline, we show a different set of people the same questions in their standard ten paragraph form.

On a sample of 200 bridge questions from the validation set, human accuracy shows marginal degradation when using only one hop: humans obtain 87.37 F1 using all ten paragraphs and 82.06 F1 when using only nine (where they only see a single gold paragraph). This indicates humans, just like models, are capable of solving bridge questions using only one hop.

Next, we manually categorize what enables single-hop answers for 100 bridge validation examples (taking into account the distractor paragraphs), and place questions into four categories (Table 2).

27%27\% of questions require multi-hop reasoning. The first example of Table 2 requires locating the university where “Ralph Hefferline” was a psychology professor, and multiple universities are provided as distractors. Therefore, the answer cannot be determined in one hop.It is possible that a single-hop model can do well by randomly guessing between two or three well-typed options, but we do not evaluate that strategy here.

Weak Distractors

35%35\% of questions allow single-hop answers in the distractor setting, mostly by entity type matching. Consider the question in the second row of Table 2: in the ten provided paragraphs, only one actress has a government position. Thus, the question is answerable without considering the film “Kiss and Tell.” These examples may become multi-hop in the open-domain setting, e.g., there are numerous actresses with a government position on Wikipedia.

To further investigate entity type matching, we reduce the question to the first five tokens starting from the wh-word, following Sugawara et al. (2018). Although most of these reduced questions appear void of critical information, the F1 score of single-paragraph BERT only degrades about 15 F1 from 67.08 to 52.13.

Redundant Evidence

26%26\% of questions are compositional but are solvable using only part of the question. For instance, in the third example of Table 2 there is only a single founder of “Kaiser Ventures.” Thus, one can ignore the condition on “American industrialist” and “father of modern American shipbuilding.” This category differs from the weak distractors category because its questions are single-hop regardless of the distractors.

Non-compositional Single-hop

8%8\% of questions are non-compositional and single-hop. In the last example of Table 2, one sentence contains all of the information needed to answer correctly.

2 Categorizing Comparison Questions

Comparison questions require quantitative or logical comparisons between two quantities or events. We create rules (Appendix C) to group comparison questions into three categories: questions which require multi-hop reasoning (multi-hop), may require multi-hop reasoning (context-dependent), and require single-hop reasoning (single-hop).

Many comparison questions are multi-hop or context-dependent multi-hop, and single-paragraph BERT achieves near chance accuracy on these types of questions (Table 3).Comparison questions test mainly binary relationships. This shows that most comparison questions are not solvable by our single-hop model.

Can We Find Better Distractors?

In Section 4.1, we identify that 35% of bridge examples are solvable using single-hop reasoning due to weak distractor paragraphs. Here, we attempt to automatically correct these examples by choosing new distractor paragraphs which are likely to trick our single-paragraph model.

We report the F1 score of single-paragraph BERT on these new distractors in Table 4: the accuracy declines from 67.08 F1 to 46.84 F1. However, when the same procedure is done on the training set and the model is re-trained, the accuracy increases to 60.10 F1 on the adversarial distractors.

Type Distractors

We also experiment with filtering the initial list of 50 paragraph to ones whose entity type (e.g., person) matches that of the gold paragraphs. This can help to eliminate the entity type bias described in Section 4.1. As shown in Table 4, the original model’s accuracy degrades significantly (drops to 40.73 F1). However, similar to the previous setup, the model trained on the adversarially selected distractors can recover most of its original accuracy (increases to 58.42 F1).

These results show that single-paragraph BERT can struggle when the distribution of the distractors changes (e.g., using adversarial selection rather than only TF-IDF). Moreover, the model can somewhat recover its original accuracy when re-trained on distractors from the new distribution.

Conclusions

In summary, we demonstrate that question compositionality is not a sufficient condition for multi-hop reasoning. Instead, future datasets must carefully consider what evidence they provide in order to ensure multi-hop reasoning is required. There are at least two different ways to achieve this.

Our single-hop model struggles in the open-domain setting. We largely attribute this to the insufficiencies of standard TF-IDF retrieval for multi-hop questions. For example, we fail to retrieve the paragraph about “Bonobo apes” in Figure 1, because the question does not contain terms about “Bonobo apes.” Table 5 shows that the model achieves 39.12 F1 given 500 retrieved paragraphs, but achieves 53.12 F1 when additional two gold paragraphs are given, demonstrating the significant effect of failure to retrieve gold paragraphs. In this context, we suggest that future work can explore better retrieval methods for multi-hop questions.

Retrieving Strong Distractors

Another way to ensure multi-hop reasoning is to select strong distractor paragraphs. For example, we found 35%35\% of bridge questions are currently single-hop but may become multi-hop when combined with stronger distractors (Section 4.1). However, as we demonstrate in Section 5, selecting strong distractors for RC questions is non-trivial. We suspect this is also due to the insufficiencies of standard TF-IDF retrieval for multi-hop questions. In particular, Table 5 shows that single-paragraph BERT achieves 53.12 F1 even when using 500 distractors (rather than eight), indicating that 500 distractors are still insufficient. In this end, future multi-hop RC datasets can develop improved methods for distractor collection.

Acknowledgements

This research was supported by ONR (N00014-18-1-2826, N00014-17-S-B001), NSF (IIS-1616112, IIS-1252835, IIS-1562364), ARO (W911NF-16-1-0121), an Allen Distinguished Investigator Award, Samsung GRO and gifts from Allen Institute for AI, Google, and Amazon.

The authors would like to thank Shi Feng, Nikhil Kandpal, Victor Zhong, the members of AllenNLP and UW NLP, and the anonymous reviewers for their valuable feedback.

References

Appendix A Example Distractor Question

We present the full example from Figure 1 below. Paragraphs 1 and 5 are the two gold paragraphs.

What is the former name of the animal whose habitat the Réserve Naturelle Lomako Yokokala was established to protect?

Answer

(Gold Paragraph) Paragraph 1

The bonobo (or ; “Pan paniscus”), formerly called the pygmy chimpanzee and less often, the dwarf or gracile chimpanzee, is an endangered great ape and one of the two species making up the genus “Pan”; the other is “Pan troglodytes”, or the common chimpanzee. Although the name “chimpanzee” is sometimes used to refer to both species together, it is usually understood as referring to the common chimpanzee, whereas “Pan paniscus” is usually referred to as the bonobo.

Paragraph 2

The Carriére des Nerviens Regional Nature Reserve (in French “Réserve naturelle régionale de la carriére des Nerviens”) is a protected area in the Nord-Pas-de-Calais region of northern France. It was established on 25 May 2009 to protect a site containing rare plants and covers just over 3 ha. It is located in the municipalities of Bavay and Saint-Waast in the Nord department.

Paragraph 3

Céreste (Occitan: “Ceirésta”) is a commune in the Alpes-de-Haute-Provence department in southeastern France. It is known for its rich fossil beds in fine layers of “Calcaire de Campagne Calavon” limestone, which are now protected by the Parc naturel régional du Luberon and the Réserve naturelle géologique du Luberon.

Paragraph 4

The Grand Cote National Wildlife Refuge (French: “Réserve Naturelle Faunique Nationale du Grand- Cote”) was established in 1989 as part of the North American Waterfowl Management Plan. It is a 6000 acre reserve located in Avoyelles Parish, near Marksville, Louisiana, in the United States.

(Gold Paragraph) Paragraph 5

The Lomako Forest Reserve is found in Democratic Republic of the Congo. It was established in 1991 especially to protect the habitat of the Bonobo apes. This site covers 3,601.88 km2.

Paragraph 6

Guadeloupe National Park (French: “Parc national de la Guadeloupe”) is a national park in Guadeloupe, an overseas department of France located in the Leeward Islands of the eastern Caribbean region. The Grand Cul-de-Sac Marin Nature Reserve (French: “Réserve Naturelle du Grand Cul-de-Sac Marin”) is a marine protected area adjacent to the park and administered in conjunction with it. Together, these protected areas comprise the Guadeloupe Archipelago (French: “l’Archipel de la Guadeloupe”) biosphere reserve.

Paragraph 7

La Désirade National Nature Reserve (French: “Réserve naturelle nationale de La Désirade”) is a reserve in Désirade Island in Guadeloupe. Established under the Ministerial Decree No. 2011-853 of 19 July 2011 for its special geological features it has an area of 62 ha. The reserve represents the geological heritage of the Caribbean tectonic plate, with a wide spectrum of rock formations, the outcrops of volcanic activity being remnants of the sea level oscillations. It is one of thirty three geosites of Guadeloupe.

Paragraph 8

La Tortue ou l’Ecalle or Ile Tortue is a small rocky islet off the northeastern coast of Saint Barthélemy in the Caribbean. Its highest point is 35 m above sea level. Referencing tortoises, it forms part of the Réserve naturelle nationale de Saint-Barthélemy with several of the other northern islets of St Barts.

Paragraph 9

Nature Reserve of Saint Bartholomew (Réserve Naturelle de Saint-Barthélemy) is a nature reserve of Saint Barthélemy (RNN 132), French West Indies, an overseas collectivity of France.

Paragraph 10

Ile Fourchue, also known as Ile Fourche is an island between Saint-Barthélemy and Saint Martin, belonging to the Collectivity of Saint Barthélemy. The island is privately owned. The only inhabitants are some goats. The highest point is 103 meter above sea level. It is situated within Réserve naturelle nationale de Saint-Barthélemy.

Appendix B Full Model Details

Single-paragraph BERT is a pipeline which first retrieves a single paragraph using a classifier and then selects the associated answer. Formally, the model receives a question Q=[q1,..,qm]Q=[q_{1},..,q_{m}] and a single paragraph P=[p1,...,pn]P=[p_{1},...,p_{n}] as input. The question and paragraph are merged into a single sequence, S=[q1,...,qm,[SEP],p1,...,pn]S=[q_{1},...,q_{m},{\tt[SEP]},p_{1},...,p_{n}], where [SEP]{\tt[SEP]} is a special token indicating the boundary. The sequence is fed into BERT-base:

A candidate answer span is then computed separately from the classifier. We define

We use PyTorch (Paszke et al., 2017) based on Hugging Face’s implementation.https://github.com/huggingface/pytorch-pretrained-BERT We use Adam (Kingma and Ba, 2015) with learning rate 5×10−55\times 10^{-5}. We lowercase the input and set the maximum sequence length ∣S∣|{S}| to 300300. If a sequence is longer than 300300, we split it into multiple sequences and treat them as different examples.

Appendix C Categorizing Comparison Questions

This section describes how we categorize comparison questions. We first identify ten question operations that sufficiently cover comparison questions (Table 6). Next, for each question, we extract the two entities under comparison using the Spacyhttps://spacy.io/ NER tagger on the question and the two HotpotQA supporting facts. Using these extracted entities, we identity the suitable question operation following Algorithm 1.

Based on the identified operation, questions are classified into multi-hop, context-dependent multi-hop, or single-hop. First, numerical questions are always multi-hop (e.g., first example of Table 6). Next, the operations And, Or, Is equal, and Not equal are context-dependent multi-hop. For instance, in the second example of Table 6, if “Hot Rod” is not a magazine, one can immediately answer No. Finally, the operations Which is true and Intersection are single-hop because they can be answered using one paragraph regardless of the context. For instance, in the third example of Table 6, if Henry Roth’s paragraph explains he is from England, one can answer Henry Roth, otherwise, the answer is Robert Erskine Childers.