Investigating the Successes and Failures of BERT for Passage Re-Ranking

Harshith Padigela, Hamed Zamani, W. Bruce Croft

Introduction

Recent developments in deep learning and the availability of large-scale datasets have led to significant improvements in various computer vision and natural language processing tasks. In information retrieval (IR), the lack of publicly available large-scale datasets for many tasks, such as ad-hoc retrieval, has restricted observing substantial improvements over traditional methods (Lin, 2019; Zamani et al., 2018). A number of approaches, such as weak supervision (Dehghani et al., 2017; Zamani and Croft, 2018), have been recently proposed to enable deep neural models to learn from limited training data. More recently, Microsoft has released MS MARCO v2 (Nguyen et al., 2016), a large dataset for the passage re-ranking task, to foster the neural information retrieval research.

In this paper, we first show that a simple neural model that uses bidirectional encoder representations from Transformers (BERT) (Devlin et al., 2018) for question and passage representations performs surprisingly well compared to state-of-the-art retrieval models, including traditional term-matching models, conventional feature-based learning to rank models, and recent neural ranking models. This has been also discovered by other researchers, such as (Nogueira and Cho, 2019) in parallel with this study. Looking at the leaderboard of the MS MARCO passage re-ranking task shows the effectiveness of the BERT representations for retrieval.The leaderboard is available at http://www.msmarco.org/leaders.aspx.

We believe that understanding the performance of effective neural IR models, e.g., BERT, is important. It could potentially provide guidelines for the IR researchers for further development of neural IR models. Given this motivation, this paper mainly analyzes the results obtained by BERT for passage re-ranking and studies the reasons behind its success. To do so, we compare the results obtained by both BM25 and BERT, and highlight their differences. We choose BM25 as our basis for comparison, due to its effectiveness and more importantly its simplicity and explainable behavior, which makes the analysis easier.

In more detail, this paper studies the following hypotheses:

H1: BM25 is more biased towards higher query term frequency compared to BERT.

H2: Bias towards higher query term frequency hurts the BM25 performance.

H3: BERT retrieves documents with more novel words.

H4: BERT’s improvement over BM25 is higher for longer queries.

In addition we also identify the query types for which BERT does and does not perform well. Our experiments provide interesting insights into the performance of this model.

BERT

Representations learned using language modelling (Mikolov et al., 2013; Pennington et al., 2014) have shown to be useful in many downstream natural language tasks (Collobert et al., 2011). There exist two primary approaches for using these pre-trained representations: (1) feature-based models and (2) fine-tuning (Devlin et al., 2018). In the feature-based approach, task-specific architectures are designed on top of the pre-trained feature representations. While in the fine-tuning approach, minimal task specific parameters are added, which will be fine-tuned in addition to the pre-trained representations for the downstream task. BERT (Devlin et al., 2018), which falls into the latter category, is a multi-layer bidirectional transformer encoder utilizing the transformer units described in (Vaswani et al., 2017). The BERT model uses bidirectional self-attention to capture interaction between the input tokens and is pre-trained on the masked language modelling task (Devlin et al., 2018).

Pre-trained BERT models, fine-tuned using a single additional layer have been shown to achieve state-of-the-art results in a wide range of natural language tasks, including machine reading comprehension (MRC) and natural language inference (NLI) (Devlin et al., 2018). In this paper, we also use the same setting, by adding a single layer on top of the BERT’s large pre-trained model, and fine-tuning it using a pointwise setting with a maximum likelihood objective. This is also similar to the setting used in (Nogueira and Cho, 2019).

Empirical Analysis

We consider the MS MARCO dataset for passage re-ranking (Nguyen et al., 2016) in our analysis. The MS MARCO dataset is generated using queries sampled from the Bing’s query logs and corresponding relevant passages marked by human assessors. It is notable that the relevance judgments provided by the MS MARCO dataset are different from the traditional TREC-style relevance judgments. They utilized the information that the human assessors provided for the machine reading comprehension data. This means that a marked passage is a true positive/relevant, however an unmarked passage may not be a true negative. For every query, a set of top 1000 candidate passages are extracted using BM25 for re-ranking. Since the original relevant documents were picked from a set of 10 candidate documents chosen by the Bing’s ranking stack, the relevant passages might not be present in the 1000 passages chosen by BM25.

The training set consists of approximately 400 million tuples of query, relevant passage, and non-relevant passage. The development set contains 6,980 queries with their corresponding set of 1,000 candidate passages. On average, each query has one relevant passage. Around 1242 queries have no marked relevant passages. We primarily focus our analysis only on the 5738 queries in the development set that have at least one relevant passage.

2. Experimental Setup

We use BERT and BM25 for our analysis and comparison. We used the BERT large model trained on MS MARCO. The training setup is similar to the one described in (Nogueira and Cho, 2019). For BM25 relevance matching we indexed all the passages using Elasticsearch (Gormley and Tong, 2015) with the default analyzer and default parameters of b=0.75b=0.75 and k1=1.2k1=1.2. Since most queries have only 1 relevant document, we use mean reciprocal rank of the top 10 retrieved passages (MRR@10) as our main retrieval metric, which is also suggested by MS MARCO (Nguyen et al., 2016).Due to the incomplete judgments, recall-oriented metrics such as mean average precision (MAP), are not suitable for this dataset.

3. Results and Discussion

The performance of various models on the entire development and evaluation (test) sets of MS MARCO are shown in Table 1. We can see that the BERT model which was originally trained on the masked language modelling (MLM) task (Devlin et al., 2018) and further fine-tuned using a pointwise training on the MS MARCO data, outperforms the existing traditional retrieval models and recent neural ranking models by a large margin. In order to understand these improvements, we look into the BERT’s and the BM25’s performances on the development set.Note that the evaluation set is not publicly accessible.

In the following, we study a set of hypotheses and provide empirical evidence to either validate or invalidate them.

Hypothesis I: BM25 is more biased towards higher query term frequency compared to BERT. We hypothesize that in many queries the top results from BM25 were just the passages that contain multiple repetitions of query words without actually conveying any useful information, which is not the case for BERT. To validate this hypothesis, we calculate the fraction of query tokens (FQT) as follows: For each query, we take the top kk results, remove stopwords and punctuations, and calculate the fraction of query tokens in the remaining tokens. If d1,d2,⋯ ,dkd_{1},d_{2},\cdots,d_{k} are the set of results for a query qq without stopwords and punctuation, then,

where N(di,q)N(d_{i},q) denotes the number of occurrences of query tokens qq in the document did_{i}. We limit kk to a maximum of 10.

We find that the FQT average across queries is 0.2 for BM25 and 0.147 for BERT. In 95.96% of the queries BM25 has a higher FQT value than BERT. These results validate our first hypothesis, saying that BM25 has a higher bias towards query term frequency in document matches, compared to BERT. An example can be seen in Table 5 Query 5

Hypothesis II: Bias towards higher query term frequency hurts the BM25 performance. We hypothesize that the bias towards query term frequency affects the BM25 performance significantly, compared to BERT.

To investigate this, we see how MRR changes across different ranges of FQT. The FQT range of is split into 5 buckets and the average MRR value and the number of queries in each bucket is shown in Table 2. As FQT value increases, we can see that the MRR value decreases in both BERT and BM25. Because of BM25’s bias towards high FQT (validated by Hypothesis I and also evident by the number of queries), we can see the decrease in MRR (as we go from the lowest to highest FQT buckets) is more prominent for BM25 (34.5%) than BERT (16.7%). The signed t-test for measuring the difference between two pairs of data, applied on difference between FQT values of BM25 and BERT yields a p-value of 0.0, indicating statistically significant difference between the FQT values.

Hypothesis III: BERT retrieves documents with more novel words. Since the recent neural models trained on the language modeling task have been shown to capture semantic similarities, we hypothesize that BERT can retrieve results with more novel words, compared to BM25. To validate this, we calculate the fraction of novel terms (FNT) as follows. Let d1,d2,⋯ ,dkd_{1},d_{2},\cdots,d_{k} be the results for a query qq, which are stripped of stopwords and punctuation. Then

where U(di)U(d_{i}) gives the number of unique terms in document did_{i} and N′(di,q)N^{\prime}(d_{i},q) gives the number of unique terms in document did_{i} which are not present in the query qq. We limit kk to a maximum of 10. We find that the FNT average across queries is 0.88 for BM25 and 0.9 for BERT. In 85.85%85.85\% of queries BERT has a higher FNT value than BM25. The signed t-test on the difference between FNT values of BERT and BM25 yields a p-value of 0.0, indicating statistically significant difference between the FNT values. This validates our hypothesis that BERT retrieves documents with more novel words than BM25.

Hypothesis IV: BERT′s improvement over BM25 is higher for longer queries. Since the BERT model is designed to learn context-aware word representations, we hypothesize that its improvements for longer queries, which generally provide richer context, are more significant. To validate this hypothesis, we calculate the average MRR per query length for both BERT and BM25, shown in Table 3. We can see that BERT performs significantly better than BM25 across all query lengths. But as the query length increases from 2 to 10, the performance of both BM25 and BERT generally decreases, and this decrease is more prominent for BERT (39%) compared to BM25 (33%), indicating its higher sensitivity to query length than BM25. The MRR difference between BERT and BM25 also decreases from 0.29 to 0.16 as query length increases from 2 to 10, which indicates that our fourth hypothesis is incorrect and that BERT’s improvement is lower for longer queries. Interestingly, BERT performs surprisingly well for very short queries compared to longer ones. The reason might be that BERT is not successful at capturing the query context properly for long queries. This can be also observed from the examples, such as Queries 5, 5 in Table 5.

4. Result Analysis

We conduct various analyses to understand the similarities and differences between BM25 and BERT. We discuss them below.

Per Query Analysis. We analyze the per query performance of BERT compared to BM25. Figure 1 plots Δ\DeltaMRR per query (i.e., MRRBERT−MRRBM25\text{MRR}_{\text{BERT}}-\text{MRR}_{\text{BM25}}), which are sorted in descending order. As depicted in the figure, in 3257 questions (57% out of 5738) BERT has a better performance compared to BM25 and in 690 questions (12%) BM25 performs better than BERT. For 525525 (9%) queries Δ\DeltaMRR is equal to 11, meaning that a relevant answer is retrieved by BERT as the first ranked passage, however, no relevant answer is retrieved by BM25 in the top 10 result list. For 1791 queries, both BERT and BM25 perform similarly.

This experiments show that BERT not only performs better than BM25 (on average), but it also performs more accurately for substantially more queries.

Similarity between the BERT’s and the BM25’s result lists. To measure the similarity between the results of BERT and BM25, we calculate the following metric, MUR - matches upto result. MUR(i,q)(i,q) for a query qq measures the number of matches in the top ii results of BERT and BM25. We can see the average MUR for each i∈i\in in Table 4 which indicates the low extent of similarity between BERT and BM25. The number of matches increases linearly with ii with slope of about 0.330.33 and intercept around -0.210.21, indicating a consistent linear relationship between BERT and BM25.

Comparison using query starting ngrams: Here we look at the most frequent bigrams with which the queries start. The idea is that looking at the starting ngrams can help us understand the type of queries. We extract the most frequent 15 bigrams and compute the average MRR using BERT for each of them. The result is shown in Figure 2. We can see that the bigrams corresponding to numeric type questions, such as “how much” and “how long”, as well as location type questions like “what county / where is” and entity type questions, such as “what type” have a low MRR. This is consistent with our observations in the previous experiment (see Table 6).

Semantic similarity: Being trained on a language modeling task, we expect BERT to capture various semantic relationships. While in some cases these help in arriving at the right answer sometimes they can also lead to incorrect answers. We will discuss two such examples below. In question 5 of Table 5, BERT captures similarity between the word “confident” in query and “confidence” in the passage, which helps it arrive at the right answer. This can be seen by visualizing the attention values between query and document words as shown in Figure 3. Similarly in Example 5 of Table 5, the question asks for another name for word “reaper”, which in this context means synonyms for the word “reaper”. However, BERT relates name to a character name reaper (see attention map 4). This leads to an incorrect answer.

Conclusions and Future Work

BERT performs surprisingly well for a passage re-ranking task. In this paper, we provide empirical analysis to understand the performance of BERT and how its results are different from a typical retrieval model, e.g., BM25. We showed that BM25 is more biased towards high query term frequency and this bias hurts its performance. We demonstrated that, as expected, BERT retrieves passages with more novel words. Surprisingly, we found out that BERT is failing at capturing the query context for long queries. Our analysis also suggested that BERT is relatively successful in answering abbreviation answer type questions and relatively poor at numerical and entity type questions.

Although BERT substantially outperforms state-of-the-art models for passage retrieval, it is still far away from a perfect retrieval performance. We believe that future work investigating the relevance preferences captured by BERT across various query types and a better encoding of query context for longer queries could help in developing even better models.

Acknowledgements

This work was supported in part by the Center for Intelligent Information Retrieval and in part by NSF IIS-1715095. Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect those of the sponsor.

References