Overview of the TREC 2020 deep learning track

Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos

Introduction

Deep learning methods, where a computational model learns an intricate representation of a large-scale dataset, yielded dramatic performance improvements in speech recognition and computer vision (LeCun et al., 2015). When we have seen such improvements, a common factor is the availability of large-scale training data (Deng et al., 2009; Bellemare et al., 2013). For ad hoc ranking in information retrieval, which is a core problem in the field, we did not initially see dramatic improvements in performance from deep learning methods. This led to questions about whether deep learning methods were helping at all (Yang et al., 2019a). If large training data sets are a factor, one explanation for this could be that the training sets were too small.

The TREC Deep Learning Track, and associated MS MARCO leaderboards (Bajaj et al., 2016), have introduced human-labeled training sets that were previously unavailable. The main goal is to study information retrieval in the large training data regime, to see which retrieval methods work best.

The two tasks, document retrieval and passage retrieval, each have hundreds of thousands of human-labeled training queries. The training labels are sparse, with often only one positive example per query. Unlike the MS MARCO leaderboards, which evaluate using the same kind of sparse labels, the evaluation at TREC uses much more comprehensive relevance labeling. Each year of TREC evaluation evaluates on a new set of test queries, where participants submit before the test labels have even been generated, so the TREC results are the gold standard for avoiding multiple testing and overfitting. However, the comprehensive relevance labeling also generates a reusable test collections, allowing reuse of the dataset in future studies, although people should be careful to avoid overfitting and overiteration.

The main goals of the Deep Learning Track in 2020 have been: 1) To provide large reusable training datasets with associated large scale click dataset for training deep learning and traditional ranking methods in a large training data regime, 2) To construct reusable test collections for evaluating quality of deep learning and traditional ranking methods, 3) To perform a rigorous blind single-shot evaluation, where test labels don’t even exist until after all runs are submitted, to compare different ranking methods, and 4) To study this in both a traditional TREC setup with end-to-end retrieval and in a re-ranking setup that matches how some models may be deployed in practice.

Task description

The track has two tasks: Document retrieval and passage retrieval. Participants were allowed to submit up to three runs per task, although this was not strictly enforced. Submissions to both tasks used the same set of 200200 test queries. In the pooling and judging process, NIST chose a subset of the queries for judging, based on budget constraints and with the goal of finding a sufficiently comprehensive set of relevance judgments to make the test collection reusable. This led to a judged test set of 4545 queries for document retrieval and 5454 queries for passage retrieval. The document queries are not a subset of the passage queries.

When submitting each run, participants indicated what external data, pretrained models and other resources were used, as well as information on what style of model was used. Below we provide more detailed information about the document retrieval and passage retrieval tasks, as well as the datasets provided as part of these tasks.

The first task focuses on document retrieval, with two subtasks: (i) Full retrieval and (ii) top-100100 reranking.

In the full retrieval subtask, the runs are expected to rank documents based on their relevance to the query, where documents can be retrieved from the full document collection provided. This subtask models the end-to-end retrieval scenario.

In the reranking subtask, participants were provided with an initial ranking of 100100 documents, giving all participants the same starting point. This is a common scenario in many real-world retrieval systems that employ a telescoping architecture (Matveeva et al., 2006; Wang et al., 2011). The reranking subtask allows participants to focus on learning an effective relevance estimator, without the need for implementing an end-to-end retrieval system. It also makes the reranking runs more comparable, because they all rerank the same set of 100 candidates.

The initial top-100100 rankings were retrieved using Indri (Strohman et al., 2005) on the full corpus with Krovetz stemming and stopwords eliminated.

Judgments are on a four-point scale: {etaremune}[start=3]

Perfectly relevant: Document is dedicated to the query, it is worthy of being a top result in a search engine.

Highly relevant: The content of this document provides substantial information on the query.

Relevant: Document provides some information relevant to the query, which may be minimal.

Irrelevant: Document does not provide any useful information about the query. For metrics that binarize the judgment scale, we map document judgment levels 3,2,1 to relevant and map document judgment level 0 to irrelevant.

2 Passage retrieval task

Similar to the document retrieval task, the passage retrieval task includes (i) a full retrieval and (ii) a top-10001000 reranking tasks.

In the full retrieval subtask, given a query, the participants were expected to retrieve a ranked list of passages from the full collection based on their estimated likelihood of containing an answer to the question. Participants could submit up to 10001000 passages per query for this end-to-end retrieval task.

In the top-10001000 reranking subtask, 10001000 passages per query were provided to participants, giving all participants the same starting point. The sets of 1000 were generated based on BM25 retrieval with no stemming as applied to the full collection. Participants were expected to rerank the 1000 passages based on their estimated likelihood of containing an answer to the query. In this subtask, we can compare different reranking methods based on the same initial set of 10001000 candidates, with the same rationale as described for the document reranking subtask.

Judgments are on a four-point scale: {etaremune}[start=3]

Perfectly relevant: The passage is dedicated to the query and contains the exact answer.

Highly relevant: The passage has some answer for the query, but the answer may be a bit unclear, or hidden amongst extraneous information.

Related: The passage seems related to the query but does not answer it.

Irrelevant: The passage has nothing to do with the query. For metrics that binarize the judgment scale, we map passage judgment levels 3,2 to relevant and map document judgment levels 1,0 to irrelevant.

Datasets

Both tasks have large training sets based on human relevance assessments, derived from MS MARCO. These are sparse, with no negative labels and often only one positive label per query, analogous to some real-world training data such as click logs.

In the case of passage retrieval, the positive label indicates that the passage contains an answer to a query. In the case of document retrieval, we transferred the passage-level label to the corresponding source document that contained the passage. We do this under the assumption that a document with a relevant passage is a relevant document, although we note that our document snapshot was generated at a different time from the passage dataset, so there can be some mismatch. Despite this, machine learning models trained with these labels seem to benefit from using the labels, when evaluated using NIST’s non-sparse, non-transferred labels. This suggests the transferred document labels are meaningful for our TREC task.

This year for the document retrieval task, we also release a large scale click dataset, The ORCAS data, constructed from the logs of a major search engine (Craswell et al., 2020). The data could be used in a variety of ways, for example as additional training data (almost 50 times larger than the main training set) or as a document field in addition to title, URL and body text fields available in the original training data.

For each task there is a corresponding MS MARCO leaderboard, using the same corpus and sparse training data, but using sparse data for evaluation as well, instead of the NIST test sets. We analyze the agreement between the two types of test in Section 4.

Table 1 and Table 2 provide descriptive statistics for the dataset derived from MS MARCO and the ORCAS dataset, respectively. More details about the datasets—including directions for download—is available on the TREC 2020 Deep Learning Track websitehttps://microsoft.github.io/TREC-2020-Deep-Learning. Interested readers are also encouraged to refer to (Bajaj et al., 2016) for details on the original MS MARCO dataset.

Results and analysis

The TREC 2020 Deep Learning Track had 25 participating groups, with a total of 123 runs submitted across both tasks.

Based run submission surveys, we manually classify each run into one of three categories:

nnlm: if the run employs large scale pre-trained neural language models, such as BERT (Devlin et al., 2018) or XLNet (Yang et al., 2019b)

nn: if the run employs some form of neural network based approach—e.g., Duet (Mitra et al., 2017; Mitra and Craswell, 2019) or using word embeddings (Joulin et al., 2016)—but does not fall into the “nnlm” category

trad: if the run exclusively uses traditional IR methods like BM25 (Robertson et al., 2009) and RM3 (Abdul-Jaleel et al., 2004).

We placed 70 (57%57\%) runs in the “nnlm” category, 13 (10%10\%) in the “nn” category, and the remaining 40 (33%33\%) in the “trad” category. In 2019, 33 (44%44\%) runs were in the “nnlm” category, 20 (27%27\%) in the “nn” category, and the remaining 22 (29%29\%) in the “trad” category. While there was a significant increase in the total number of runs submitted compared to last year, we observed a significant reduction in the fraction of runs in the “nn” category.

We further categorize runs based on subtask:

rerank: if the run reranks the provided top-kk candidates, or

fullrank: if the run employs their own phase 1 retrieval system.

We find that only 37 (30%30\%) submissions fall under the “rerank” category—while the remaining 86 (70%70\%) are “fullrank”. Table 3 breaks down the submissions by category and task.

Overall results

Our main metric in both tasks is Normalized Discounted Cumulative Gain (NDCG)—specifically, NDCG@10, since it makes use of our 4-level judgments and focuses on the first results that users will see. To get a picture of the ranking quality outside the top-10 we also report Average Precision (AP), although this binarizes the judgments. For comparison to the MS MARCO leaderboard, which often only has one relevant judgment per query, we report the Reciprocal Rank (RR) of the first relevant document on the NIST judgments, and also using the sparse leaderboard judgments.

Some of our evaluation is concerned with the quality of the top-kk results, where k=100k=100 for the document task and k=1000k=1000 for the passage task. We want to consider the quality of the top-kk set without considering how they are ranked, so we can see whether improving the set-based quality is correlated with an improvement in NDCG@10. Although we could use Recall@kk as a metric here, it binarizes the judgments, so we instead use Normalized Cumulative Gain (NCG@kk) (Rosset et al., 2018). NCG is not supported in trec_eval. For trec_eval metrics that are correlated, see Recall@kk and NDCG@kk.

The overall results are presented in Table 4 for document retrieval and Table 5 for passage retrieval. These tables include multiple metrics and run categories, which we now use in our analysis.

Neural vs. traditional methods.

The first question we investigated as part of the track is which ranking methods work best in the large-data regime. We summarize NDCG@10 results by run type in Figure 1.

For document retrieval runs (Figure 1(a)) the best “trad” run is outperformed by “nn” and “nnlm” runs by several percentage points, with “nnlm” also having an advantage over “nn”. We saw a similar pattern in our 2019 results. This year we encouraged submission of a variety of “trad” runs from different participating groups, to give “trad” more chances to outperform other run types. The best performing run of each category is indicated, with the best “nnlm” and “nn” models outperforming the best “trad” model by 23%23\% and 11%11\% respectively.

For passage retrieval runs (Figure 1(b)) the gap between the best “nnlm” and “nn” runs and the best “trad” run is larger, at 42%42\% and 17%17\% respectively. One explanation for this could be that vocabulary mismatch between queries and relevant results is greater in short text, so neural methods that can overcome such mismatch have a relatively greater advantage in passage retrieval. Another explanation could be that there is already a public leaderboard, albeit without test labels from NIST, for the passage task. (We did not launch the document ranking leaderboard until after our 2020 TREC submission deadline.) In passage ranking, some TREC participants may have submitted neural models multiple times to the public leaderboard, so are relatively more experienced working with the passage dataset than the document dataset.

In query-level win-loss analysis for the document retrieval task (Figure 2) the best “nnlm” model outperforms the best “trad” run on 38 out of the 45 test queries (i.e., 84%84\%). Passage retrieval shows a similar pattern in Figure 3. Similar to last year’s data, neither task has a large class of queries where the “nnlm” model performs worse.

End-to-end retrieval vs. reranking.

Our datasets include top-kk candidate result lists, with 100 candidates per query for document retrieval and 1000 candidates per query for passage retrieval. Runs that simply rerank the provided candidates are “rerank” runs, whereas runs that perform end-to-end retrieval against the corpus, with millions of potential results, are “fullrank” runs. We would expect that a “fullrank” run should be able to find a greater number of relevant candidates than we provided, achieving higher NCG@kk. A multi-stage “fullrank” run should also be able to optimize the stages jointly, such that early stages produce candidates that later stages are good at handling.

According to Figure 4, “fullrank” did not achieve much better NDCG@10 performance than “rerank” runs. In fact, for the passage retrieval task, the top two runs are of type “rerank”. While it was possible for “fullrank” to achieve better NCG@kk, it was also possible to make NCG@kk worse, and achieving significantly higher NCG@kk does not seem necessary to achieve good NDCG@10.

Specifically, for the document retrieval task, the best “fullrank” run achieves 5%5\% higher NDCG@10 over the best “rerank’ run; whereas for the passage retrieval task, the best “fullrank” run performs slightly worse (0.3%0.3\% lower NDCG@10) compared to the best “rerank’ run.

Similar to our observations from Deep Learning Track 2019, we are not yet seeing a strong advantage of “fullrank” over “rerank”. However, we hope that as the body of literature on neural methods for phase 1 retrieval (e.g., (Boytsov et al., 2016; Zamani et al., 2018; Mitra et al., 2019; Nogueira et al., 2019)) grows, we would see a larger number of runs with deep learning as an ingredient for phase 1 in future editions of this TREC track.

Effect of ORCAS data

Based on the descriptions provided, ORCAS data seems to have been used by six of the runs (ndrm3-orc-full, ndrm3-orc-re, uogTrBaseL17, uogTrBaseQL17o, uogTr31oR, relemb_mlm_0_2). Most runs seem to be make use of the ORCAS data as a field, with some runs using the data as an additional training dataset as well. Most runs used the ORCAS data for the document retrieval task, with relemb_mlm_0_2 being the only run using the ORCAS data for the passage retrieval task.

This year it was not necessary to use ORCAS data to achieve the highest NDCG@10. However, when we compare the performance of the runs that use the ORCAS dataset with those that do not use the dataset within the same group, we observe that usage of the ORCAS dataset always led to an improved performance in terms of NDCG@10, with maximum increase being around 0.05130.0513 in terms of NDCG@10. This suggests that the ORCAS dataset is providing additional information that is not available in the training data. This could also imply that even though the training dataset provided as part of the track is very large, deep models are still in need of more training data.

NIST labels vs. Sparse MS MARCO labels.

Our baseline human labels from MS MARCO often have one known positive result per query. We use these labels for training, but they are also available for test queries. Although our official evaluation uses NDCG@10 with NIST labels, we now compare this with reciprocal rank (RR) using MS MARCO labels. Our goal is to understand how changing the labeling scheme and metric affects the overall results of the track, but if there is any disagreement we believe the NDCG results are more valid, since they evaluate the ranking more comprehensively and a ranker that can only perform well on labels with exactly the same distribution as the training set is not robust enough for use in real-world applications, where real users will have opinions that are not necessarily identical to the preferences encoded in sparse training labels.

Figure 5 shows the agreement between the results using MS MARCO and NIST labels for the document retrieval and passage retrieval tasks. While the agreement between the evaluation setup based on MS MARCO and TREC seems reasonable for both tasks, agreements for the document ranking task seems to be lower (Kendall correlation of 0.460.46) than agreements for the passage task (Kendall correlation of 0.690.69). This value is also lower than the correlation we observed for the document retrieval task for last year.

In Table 6 we show how the agreement between the two evaluation setups varies across task and run type. Agreement on which are the best neural network runs is high, but correlation for document trad runs is close to zero.

One explanation for this low correlation could be use of the ORCAS dataset. ORCAS was mainly used in the document retrieval task, and could bring search results more in line with Bing’s results, since Bing’s results are what may be clicked. Since MS MARCO sparse labels were also generated based on top results from Bing, we would expect to see some correlation between ORCAS runs and MS MARCO labels (and Bing results). By contrast, NIST judges had no information about what results were retrieved or clicked in Bing, so may have somewhat less correlation with Bing’s results and users.

In Figure 6 we compare the results from the two evaluation setups when the runs are split based on the usage of the ORCAS dataset. Our results suggest that runs that use the ORCAS dataset did perform somewhat better based on the MS MARCO evaluation setup. While the similarities between the ORCAS dataset and the MS MARCO labels seem to be one reason for the mismatch between the two evaluation results, it is not enough to fully explain the 0.030.03 correlation in Table6. Removing the ORCAS “trad” runs only increases the correlation to 0.130.13. In the future we plan to further analyze the possible reasons for this poor correlation, which could also be related to 1) the different metrics used in the two evaluation setups (RR vs. NDCG@10), 2) the different sensitivity of the datasets due to the different number of queries and number of documents labelled per query), or 3) difference in relevance labels provided by NIST assessors vs. labels derived from clicks.

Conclusion

The TREC 2020 Deep Learning Track has provided two large training datasets, for a document retrieval task and a passage retrieval task, generating two ad hoc test collections with good reusability. The main document and passage training datasets in 2020 were the same as those in 2019. In addition, as part of the 2020 track, we have also released a large click dataset, the ORCAS dataset, which was generated using the logs of the Bing search engine.

For both tasks, in the presence of large training data, this year’s non-neural network runs were outperformed by neural network runs. While usage of the ORCAS dataset seems to help improve the performance of the systems, it was not necessary to use ORCAS data to achieve the highest NDCG@10.

We compared reranking approaches to end-to-end retrieval approaches, and in this year’s track there was not a huge difference, with some runs performing well in both regimes. This is another result that would be interesting to track in future years, since we would expect that end-to-end retrieval should perform better if it can recall documents that are unavailable in a reranking subtask.

This year the number of runs submitted for both tasks have increased compared to last year. In particular, number of non-neural runs have increased. Hence, test collections generated as part of this year’s track may be more reusable compared to last year since these test collections may be fairer towards evaluating the quality of unseen non-neural runs. We note that the number of “nn” runs also seems to be smaller this year. We will continue to encourage a variety of approaches in submission, to avoid converging too quickly on one type of run, and to diversify the judging pools.

Similar to last year, in this year’s track we have two types of evaluation label for each task. Our official labels are more comprehensive, covering a large number of results per query, and labeled on a four point scale at NIST. We compare this to the MS MARCO labels, which usually only have one positive result per query. While there was a strong correlation between the evaluation results obtained using the two datasets for the passage retrieval task, the correlation for the document retrieval task was lower. Part of this low correlation seems to be related to the usage of the ORCAS dataset (which is generated using similar dataset as the one used to generate the MS MARCO labels) by some runs, and evaluation results based on MS MARCO data favoring these runs. However, our results suggest that while the ORCAS dataset could be one reason for the low correlation, there might be other reasons causing this reduced correlation, which we plan to explore as future work.

References