Tevatron: An Efficient and Flexible Toolkit for Dense Retrieval

Luyu Gao, Xueguang Ma, Jimmy Lin, Jamie Callan

Introduction

Dense retrieval’s popularity in the research community has greatly grown in the past years (Karpukhin et al., 2020; Xiong et al., 2021; Lin et al., 2021; Qu et al., 2021; Gao and Callan, 2021a). By modeling relevance with query-document vector products, dense retrievers can carry out efficient and effective semantic search.

While the idea of vector-based search is not new, the adoption of deep pre-trained language models as encoder (Devlin et al., 2019) has substantially boosted the effectiveness of dense retrieval (Karpukhin et al., 2020). Meanwhile, similar to other research that relies on deep learning, the success of dense retrieval will not be possible without large data. Many of recent research works are based on their own software with specialized support only for specific datasets and models (Karpukhin et al., 2020; Xiong et al., 2021). We however believe flexible generalization across models and datasets is critical. Tevatron provides researchers with access to the latest state-of-the-art models and makes it easy for them to start a new research problem on a new dataset.

In our past research on dense retrieval (Gao and Callan, 2021a, b; Ma et al., 2021), we have run into several engineering challenges specific to dense systems. For example, in terms of resources, large corpora and training sets require large CPU memory; accelerator (GPU/TPU) memory usage also grows with model size. While orthogonal to actual research, these engineering problems slow down and constrain researchers, especially those with limited hardware resources. With Tevatron, we aim at providing a unified solution to common engineering problems.

Tevatron incorporates several popular widely-used open-source packages, including datasets (Lhoest et al., 2021), transformers (Wolf et al., 2020) and FAISS (Johnson et al., 2019) respectively as backbone for our data management, neural network modeling and embedding-based retrieval components.

To accommodate different research needs, we select two deep learning frameworks for Tevatron, Pytorch (Paszke et al., 2019) and JAX(Bradbury et al., 2018). Pytorch’s eager execution patterns and intuitive object-oriented design have gained its massive user base in the research community. On the other hand, JAX, backed by just-in-time (JIT) XLA compilation, offers smooth transitions across hardware stacks with optimized performance.

The rest of the paper is organized as follows. Section 2 gives an overview of Tevatron. Section 3 demonstrates Tevatron usage and command-line interface. Section 4 shows the experimental results of running Tevatron with various models and datasets.

Toolkit Overview

Tevatronhttp://tevatron.ai is packaged as a Python module available on the Python Package Index. Tevatron can be installed via pip, as follows:

In this section, we give an overview of the core components of Tevatron. We demonstrate how these components respectively support the full pipeline of data preparation, training, encoding, and search. Code and documentation of Tevatron are available at its website, tevatron.ai.

Having data ready to use is a critical preliminary step before training or encoding starts. Data access overhead and constraints could directly affect training/encoding performance. In Tevatron, we adopt the following core design: 1) text data are pre-tokenized before training or encoding happens, 2) keep tokenized data memory-mapped instead of lazy-loaded or in-memory. The former avoids overheads when running sub-word/piece level tokenizers and also reduces data traffic compared to raw text. The latter allows random data access in the training/encoding loop without consuming a large amount of physical memory.

Tevatron defines two basic raw input format templates for IR and QA context. As shown in Fig. 1, for the IR dataset (e.g. MS MARCO (Bajaj et al., 2018)), we organize a training instance into an anchor query, a list of positive target texts, and a list of negative target texts. The positive targets are usually human judged and the negative texts are usually non-relevant texts from top results of a baseline retrieval system such as BM25.

The second format (not shown due to space limits) has an additional answers field for QA tasks (e.g. Natural Question (Kwiatkowski et al., 2019)), since the positive passages for QA dataset are usually judged by answer exact match (Chen et al., 2017; Karpukhin et al., 2020).

Users can pass raw data file pointer and processing specifications to Tevatron’s dataset class (HFTrainDataset for training, HFQueryDataset and HFCorpusDataset for encoding) which will perform fast parallel data formatting and tokenization. Processed data is internally represented as a datasets.Dataset object and is stored in Apache Arrow format which can be memory-mapped and randomly accessed by offset.

For researchers who are focusing on building new models, we make a collection of popular open-access datasets self-contained within the Tevatron toolkit. For instance, with a single line of command, one can load the training set of MS-MARCO. Under the hood, Tevatron will first download the raw data set we hosted through Huggingfacehttps://huggingface.co/tevatron. Then it will run the corresponding pre-defined pre-processing script to format and tokenize the downloaded data.

2. Dense Retrieval Model

Tevatron’s model class DenseModel is a Pytorch nn.Module subclass that defines the deep neural encoder of the dense retriever. Functionally, it interfaces the underlying Transformer models and provides methods for text encoding and loss computation. Thanks to duck typing in python, DenseModel class support models in the Huggingface transformer library that return standard base model output. This means new Transformers models can be loaded into Tevatron as soon as they are available in the transformer library. On the other hand, this helps Tevatron avoid maintaining the transformer codes and reduce code reduplication. Internally, DenseModel wraps a transformer module and optionally a pooler module which controls how mapping from transformer output tensor to final representations. DenseModel class also handles loss computation during training. It implements a contrastive loss with in-batch negatives and can perform negative sharing across devices using parallel collective defined in the NCCL library.

Tevatron has a sub-package tevax that implements core functionality for JAX. Following JAX’s functional nature (Bradbury et al., 2018), tevax is designed with a different philosophy. We define loss functions that can be composed with other JAX transformations. In practice, they can be combined with Flax models in the transformer library for dense retriever training. Two classes TiedParams and DualParams for parameter managing are registered as Pytrees that JAX can differentiate through. A RetrieverTrainState class manages parameters and model transformations. With JAX as backend, tevax makes it possible for a single piece of code to run on a single GPU, multiple GPUs, or TPU systems.

3. Trainer

To complete the dense retriever training setup, we introduce a DenseTrainer which implements miscellaneous training utilities. It controls basic setups such as batch size and the number of training epochs. When running on multiple GPUs, the trainer will properly set up distributed training and wrap models for gradient reduction. During training, it will asynchronously load training data to overlap computation and I/O operations. At each training step, the trainer turns a batch of loaded data into tensors and passes them to the model. In this way, the trainer glues the data sets and models together.

DenseTrainer is a subclass of Trainer in the transformers library. It inherits a collection of advanced utilities including mixed-precision training and optimizer state sharding. It is also possible to further subclass DenseTrainer to create unique training behaviors. Concretely in Tevatron, we implement a subclass GCTrainer which uses gradient caching to support large batch training on memory-limited devices (Luyu Gao and Callan, 2021).

By combining data processor, dense retrieval model, and trainer all together, Tevatron abstracts the training loop of dense retrieval model into the code block shown in Fig. 2.

4. Retriever

The retriever classes in Tevatron build dense retrieval index from text embeddings and execute search over the index. We use FAISS library (Johnson et al., 2019) as our retriever’s backend. It implements several efficient indices in C++ and exposes them through Python interfaces. For users who want the best performance, Tevatron provides a simple class BaseFaissIPRetriever which wraps a flat faiss.IndexFlatIP index for exact search. Those who want to trade-off between efficiency and effectiveness can use the more powerful FaissRetriever class. FaissRetriever takes an additional index_spec string argument in its initialization method and use faiss.index_factory method to flexibly build the specified index. Users can take advantage of this interface to build approximate search indices like HNSW (Malkov and Yashunin, 2020) or PQ (Jégou et al., 2011).

Toolkit Usage

On top of the various core components, Tevatron provides a set of command-line interfaces (CLI) to drive the dense retrieval pipeline. With the flexible design in data and neural model support, one could conduct research of various types without writing code. In this section, we give an example using Tevatron CLI to run the previously discussed components to learn the model and perform open domain retrieval on Natural Questions (Kwiatkowski et al., 2019).

With Tevatron, we are able to replicate the training of DPR model for NQ dataset (see details in section 4.1) by a single command:

As introduced in section 2.1, Tevatron will automatically handle the downloading and pre-processing of our self-contained train data Tevatron/wikipedia-nq. Then the preprocessed dataset and initialized DenseModel will be fed into Trainer class as shown in figure 2. Since the above command enables the grad_cache option, it will uses GCTrainer during training. Here we also enable mix precision training (Micikevicius et al., 2018) via the --fp16 option to improve efficiency.

2. Encoding

Besides training data, Tevatron also self-contains corresponding corpus data for each dataset. Again, we simplifies corpus encoding process into a single command:

As encoding the entire corpus within a single process may cost large RAM usage and a long time, Tevatron support encoding the corpus by sharding. For example, the above command encodes the first 1/20 split of the entire corpus. Users can easily run multiple processes for multiple shards in parallel to speed up the encoding process.

3. Retrieval

By taking query and corpus embeddings, we can run retrieval with following command:

where --batch_size controls the number of queries passed to the FAISS index each search call and -1 will pass all queries in one call. Larger batches typically run faster due to better memory access patterns and hardware utilization. The results will be saved in a text file with each line stores query_id passage_id score.

Experiments

In this section, we demonstrate the system effectiveness and efficiency of Tevatron by running experiments on two common-use collections for QA and IR tasks, Wikipedia and MS MARCO.

The DPR work by Karpukhin et al. (Karpukhin et al., 2020) is among the first works that show text retrieval using learned dense representations outperforms traditional text retrieval using heuristic sparse representations (e.g. BM25) on open-domain question-answering tasks.

We evaluate the effectiveness of Tevatron by replicating the retrieval results on QA tasks (Kwiatkowski et al., 2019; Joshi et al., 2017; Rajpurkar et al., 2016; Voorhees and Tice, 2000; Berant et al., 2013) reported in original DPR work (Karpukhin et al., 2020). We compare the models trained under the "Single" setting defined in the original work where each model is trained by the corresponding individual dataset. Following the similar hyperparameters setting, we train the models with a learning rate of 1e-5 for 40 epochs with batch size 128. In Table 1, except having slightly lower accuracy than the numbers in DPR paper on SQuAD, Tevatron gives even a bit higher top-kk accuracy on all other four datasets. Overall, all top-kk accuracy results obtained via the Tevatron pipeline are at the same level of accuracy as original work. Therefore, we conclude that this is a successful replication, proving that the Tevatron pipeline is effective.

We demonstrate the efficiency of Tevatron by comparing it with the original DPR repo https://github.com/facebookresearch/DPR To be clear, the efficiency results are based on the master branch on 2022-02-12 on three dimensions: RAM usage, GPU memory usage, and training time. The experiments are conducted on a machine with NVIDIA A100 GPUs. In both DPR-repo and Tevatron-default settings, we train the dense retriever model on 4 GPUs in distributed data-parallel mode of Pytorch. By comparing the first two rows in Table 2, we see that Tevatron is more efficient on all three dimensions than the original codebase. Concretely, Tevatron costs 3/4 less RAM, 12G less GPU memory, being 1/4 faster than training using DPR repo. This means, given the same resources, Tevatron has the potential to support larger training data, larger batch size, and faster training.

The gradient cache feature of Tevatron can further improve the GPU memory efficiency (Luyu Gao and Callan, 2021). DPR training requires batch size 128 to get the level of retrieval accuracy as reported above. With the original DPR repo, users cannot train a model with enough batch size if the GPU resource is limited, which will result in a drop in retrieval accuracy. Tevatron provides users the option to train dense retrievers using limited GPU resources but keeps the same amount of batch size for each optimization step. To illustrate this, we conduct experiments with Tevatron-GradCache on a single GPU. Tevatron-GradCache trains dense retrievers by splitting the batch of size 128 into sub-batches of size 32. Via conducting two round forward steps described in the gradient cache work (Luyu Gao and Callan, 2021), the model update step of Tevatron-GradCache is mathematically equivalent to Tevatron-default. In the experiment, it costs only 4G RAM and 15G GPU memory, to train a DPR model on NQ dataset with desired batch size. By reducing the sub-batch size, Tevatron-GradCache can save more GPU memory.

We also evaluated the performance of the training dense retriever model using the Jax backend of Tevatron on a V3-8 TPU VM. Back in the days when DPR (Karpukhin et al., 2020) first came out, it cost around a day to train dense retriever on NQ dataset using the initial DPR repo with 8×\times Nvidia V100-large GPUsThis is recorded by DPR authors in the GitHub page.. Now it is exciting to see that, such training can be done within one hour with Tevatron.

2. Supervised IR

To further show the flexibility of the Tevatron toolkit across model architectures and accelerator platforms, we train multiple dense retrieval baselines on MS MARCO passage ranking task with different Transformer backbones from HuggingFace hubhttps://huggingface.co/models. The experiments are conducted on both GPU and TPU platforms. The models are trained with a learning rate of 5e-6 with batch size 64 for 3 epochs using our self-contained dataset Tevatron/msmarco-passage.

In Table 3, we show MRR@10 for each model and training time on 4×\times A100 GPU and V3-8 TPU. The model backbones we choose varying across:

model size: row(1), row(2) and row(4) are models in BERT family with different size {distil, base, large}.

model type: row(4), row(5) are models in same level of size with different backbone structure {bert, roberta}.

model parameters: row(2), row(3) are same model backbone but the later one is further pre-trained from row(2), i.e. {original, fine-tuned}.

For all the variants of model initialization, Tevatron can train dense retrievers effectively and efficiently on different platforms with the Tevatron CLI commands.

Finally, we evaluated two models trained with hard negative mining(Gao and Callan, 2021b) mined with Tevatron retriever. We craft the augmented training data by combining the hard negative passages mined using the first round dense retriever model with the original training dataset. Then we retrain the models using the hard negative augmented data. The last two rows in Table 3 show that by augmenting training data with hard negative, we can further improve the effectiveness of the dense retriever model. We also demonstrate Tevatron can replicate the state-of-the-art co-Condenser retriever(Gao and Callan, 2021b) on MS MARCO passage ranking.

3. Cross-lingual Retrieval

The success of dense retrieval also drives research in multilingual retrieval (Asai et al., 2021; Zhang et al., 2021; Clark et al., 2020). We additionally show our Tevatron toolkit can generalize to multilingual retrieval tasks by replicating the dense retrieval baseline reported in the XOR-Retrieve task (Asai et al., 2021).

We train dense retriever with Tevatron that encode seven languages queries and English corpus into the same embedding space. Such a method can conduct retrieval in a single stage without addition need for translation. Results in Table 4 show that the baseline model replicated with Tevatron gaining on average 6 points over original baseline results on seven languages.

Conclusion

This paper introduces Tevatron, an efficient and flexible toolkit for training and running dense retrievers with Transformers. The toolkit has a modularized design for easy research exploration and a set of command-line interfaces for fast development and evaluation. Our experiments show that Tevatron can be used to train dense retrieval models effectively and efficiently. The flexible and generalizable functionalities provide IR community convenience in future dense retrieval research.

Acknowledge

We would like to thank Google’s TPU Research Cloud (TRC) for access to Cloud TPUs and Compute Canada for access to GPU clusters.

References