SciFive: a text-to-text transformer model for biomedical literature

Long N. Phan, James T. Anibal, Hieu Tran, Shaurya Chanana, Erol Bahadroglu, Alec Peltekian, Grégoire Altan-Bonnet

Introduction

Biomedical literature is widely accessible to the scientific community through databases such as Pubmed, PMC, and ScienceDirect. Within seconds, researchers can access millions of journal articles relating to an input query. Text generation tasks such as document summarization and question answering can allow researchers to quickly obtain important information from a large collection of papers, yet current methods generally underperform in these areas. Thus, new NLP methods are needed to parse the increasingly immense amounts of information.

The introduction of the transformer (Vaswani et al., 2017) marked a significant achievement for natural language processing. This is demonstrated by the success of transformer-based architectures such as BERT (Devlin et al., 2018), which, at the time of publication, achieved state-of-the-art (SOTA) results on common NLP tasks. Furthermore, the BERT model has been extended for domain-specific tasks in NLP. Domain-specific language (i.e. biomedical language) is often challenging for NLP models because of the significant differences in vocabulary compared to standard langauge corpora such as Wikipedia. To solve this problem, BERT models have been pre-trained for domain-specific tasks. With this approach, SOTA results were achieved in areas such as clinical notes, biomedical literature, and general scientific literature.

Approach

BERT (Devlin et al., 2018) is not a unified transfer learning method because BERT-style models can only produce a single prediction for a given input. These models are simply not designed for text generation tasks such as question-answering or summarization. The text-to-text transfer transformer (T5) model proposed by Raffel et al. (2019) overcomes this limitation by outputting a string of text for each input, allowing for both question-answering, summarization and other tasks where a single output is generally insufficient. In this report, we introduce SciFive, a pretrained, domain-specific adaptation of the T5 model that is intended for tasks relating to biomedical literature. We here outline two primary contributions of our work.

(1) Our model achieves SOTA results on a variety of common classification tasks in biomedical NLP, including named entity recognition (NER) and relation-extraction (RE).

(2) Second, our model can be extended to tasks requiring extended outputs and achieves superior results on BioAsq question-answering challenges when compared to BioBERT, the current SOTA method to the best of our knowledge (Lee et al., 2019)

Unlabeled Dataset

In this section we will describe our biomedical unlabeled datasets which are used in the transfer learning pre-training stage. These large datasets overcome the drawback of overfitting when building a language model in the biomedical domain. (Ruder, 2017). For SciFive, we use two different corpora of biomedical language in order to generalize our model within the domain.

PubMed Abstract https://pubmed.ncbi.nlm.nih.gov: The PubMed database contains more than 32 millions citations and abstracts of biomedical literature. For the purpose of model pre-training, we use only the abstracts.

PubMed Central (PMC) https://www.ncbi.nlm.nih.gov/pmc: PMC is a corpus of free full-text articles in the domain of biomedical and life sciences. We hypothesize that training the language model with full-text articles can improve the learning in biomedical context while still containing a generalized representation of natural language overall.

Methods

Here, we describe our approach to implementing the SciFive model, which retains the original structure and parameters of the T5 model (Raffel et al., 2019).

The text-to-text transfer transformer (T5) model (Raffel et al., 2019) is highly similar to the transformer-based encoder-decoder model introduced by Vaswani et al. (2017). Each encoder block consists of a self-attention layer and a feed-forward neural network. Each decoder block consists of a self-attention layer, an encoder-decoder attention layer, and a feedforward neural network. There are, however, minor differences between T5 and the transformer-based encoder-decoder model. For example, layer normalization is applied between the components of each encoder block and each decoder block. Compared to BERT (Devlin et al., 2018), the addition of the decoder block allows T5 to generate outputs that are sequences of text. T5 is pre-trained with self-supervision through a learning objective called span-based language masking. (Raffel et al., 2019).

2 SciFive

SciFive follows the sequence-to-sequence encoder-decoder architecture proposed by Vaswani et al. (2017) and the T5 framework https://github.com/google-research/text-to-text-transfer-transformer released by Raffel et al. (2019). The original T5 work implemented five different model sizes - Small, Base, Large, 3B, and 11B. Due to limited computing resources, we will use only the base and large model for this study. The base and large models have 220 million parameters and 770 million parameters respectively.

We first initialized SciFive with the pre-trained weights from the base T5 model. We then re-trained SciFive on various combinations of the C4 corpus (Dodge et al., 2021), a corpus of PubMed abstracts, and a corpus of PMC full-text articles. We trained SciFive for extra 200k steps to optimize the pre-trained weights from T5 in the context of biomedical literature. We also trained a large version of the SciFive model, using 1.2 millions steps (200k additional steps compared to the regular model). With the provided TPU v2-8 on Google Colab, we used the self-supervised training setting recommended by Raffel et al. (2019) with a batch size of 128 for the base model and 64 for the large model. We used a learning rate of 0.001 and sequence length 1024 tokens for both input and target as we noticed that unlabeled biomedical text during self-supervised training is long. For the purpose of generalization of biomedical text, we train SciFive on various combinations of biomedical corpus as describe in Table 1.

3 Input/Output Representation

Consistent with the original T5 model Raffel et al. (2019), SciFive converts all of the biomedical tasks into a text-to-text format. During self-supervised training, a text input sequence is given and the model will try to learn a target input going through a learning objective called span-based mask language modeling. Spans of text are randomly masked and the target sequence is predicted as a concatenation of the same sentinel tokens and the real masked spans. An illustration of span-based mask learning objective is in Figure 1.

During supervised training, a sequence of text for both input and target is given to the model for the purpose of learning to generate text. For example, when performing Named-entity recognition (NER), we generate the target sequence by prepending and appending a special token to the named entities in a sentence. The target sequence for Question Answering task is the text corresponding to the answer for a given question (the question text is the input).

4 Vocabulary

For every pre-trained language model (LMs), vocabulary plays a crucial role, as these models attempt to derive effective contextualized word vector representations from the training corpus. For SciFive, we use the Sentence Piece model (Kudo and Richardson, 2018) as a base vocabulary model. Sentence Piece is used in all of our SciFive models because it extracts sub-words that contain the semantic meaning of a sequence. This overcomes the drawbacks of word-level tokenization and eliminates the need for an immense vocabulary set.

5 Multi-Task Learning

SciFive is trained with a maximum likelihood objective using "teacher forcing" (Raffel et al., 2019) for all tasks, thereby enabling multi-task learning. During supervised fine-tuning, a task-specific token is prepended to the input sequence. In one example, we leverage this type of training for the Named-entity recognition task. We believe that this strategy will boost performance for biomedical NER by using the attention of each named entity across all the tasks. Figure 2 illustrates multi-task learning for our NER tasks.

6 Fine-Tuning SciFive

We fine-tuned SciFive on five categories of biomedical NLP tasks.

(1) Named entity recognition (NER) involves predicting a predefined category that describes a proper noun. For example “Lupus” may be classified as “Disease".

(2) Relation Extraction (RE) involves identifying relationships within text (i.e. gene-disease).

(3) Natural Language Inference involves determining the validity of a hypothesis (i.e., True, False).

(4) Document Classification involves assigning a document to a category based on the text.

(5) Question answering involves generating an answer if given a question and a sequence of text containing the answer to that question.

We fine-tuned in both multi-tasking and single-task learning using the final checkpoints of our SciFive model, 200k steps for both base and large models. Similar to the setting during self-supervised training on TPU v2-8, we choose the batch size of 128 and 64 for the base and large respectively with learning rate 0.001. The input and output specification setting for each task is described in Table 2.

Results

We tested SciFive on 7 NER tasks, 5 RE asks, 1 inference task, 1 document classification task, and 3 question answering tasks. We then compared the SciFive results with the current SOTA on these tasks.

We describe here the datasets and the preprocessing techniques we used. In most cases, we use the same preprocessing procedure as the current baseline models (i.e. BioBERT from Lee et al. (2019) and BlueBERT from Peng et al. (2019)).

We tested SciFive on 7 datasets commonly used for biomedical NER: NCBI disease (Doğan et al., 2014), BC5CDR disease (Li et al., 2016), BC5CDR chemical (Li et al., 2016), BC4CHEMD (Krallinger et al., 2015), BC2GM (Smith et al., 2008), JNLPBA (Collier and Kim, 2004), and Species800 Pafilis et al. (2013). We follow the processing pipeline and the train/valid/test split similiar to Lee et al. (2019). For all NER tasks, we evaluate the performance of SciFive based on precision (P), recall (R), and F-1 score (F).

1.2 Relation Extraction

We tested SciFive on 2 RE tasks: CHEMPROT (Islamaj Doğan et al., 2019) and DDI (Herrero-Zazo et al., 2013). We follow the same preprocessing technique as Peng et al. (2019). We also evaluate the F1-scores of each class in the two relation extraction corpus.

1.3 Natural Language Inference

To assess the NLI capabilities of SciFive, we use the MedNLI datasets from MIMIC-III (Romanov and Shivade, 2018) with the same preprocessing technique and training/testing sets.

1.4 Document Classification

We use SciFive to classify documents from the HoC dataset (Baker et al., 2015), evaluating the F1 score on the sample average in the same manner as Zhang et al. (2017).

1.5 Question Answering

Question Answering (QA) is perhaps the most important component of our assessment, as we expect a text-to-text model to vastly outperform BERT-like models in this area. We test SciFive on the factoid questions from the BioASQ 4b, 5b, and 6b challenges Tsatsaronis et al. (2015). To preprocess the BioASQ data, we use the same approach as Lee et al. (2019).

Using the same approach as the original T5, (Raffel et al., 2019), SciFive converts all problems into a text-to-text format. Therefore, we cannot use the same evaluation procedure as BioBERT. (Lee et al., 2019). BioBERT determines the final answer for a question by taking the highest scoring answer across all the snippets of text corresponding to that question. Our model outputs a sequence of text, not a probability distribution, so we cannot determine our "best" answer in the same way as BioBERT. This key difference prevents us from evaluating strict accuracy as done by Lee et al. (2019), so we evaluate only the lenient accuracy for each task. For a single question, SciFive answers questions using a sequence of text rather than probabilities for the start and end of the answer. SciFive uses each piece of context to answer that question individually. If SciFive answers correctly using one or more of the contextual snippets, we say SciFive has answered the question correctly according to the lenient accuracy metric.

To evaluate our results, we rely on an expert assessment. SciFive outputs full-sentence answers that often do not correspond to the exact BioASQ answer provided for a given question, but, in many cases, these answers are still scientifically correct. For a meaningful assessment of Q/A results, the scientific accuracy must be considered rather than the phrasing of the answer. Table 4 shows several examples of SciFive answers compared to BioBERT answers. It can be easily seen from these examples that SciFive provides clearer, more complete answers than BioBERT.

2 Experimental Results

In Table 5, we show the results of SciFive compared to the SOTA approaches. For NER, RE, NLI, and documentation classification, we compare the F1 scores obtained by SciFive to the F1 scores obtained by the SOTA method pre-BioBERT, BioBERT Lee et al. (2019), BlueBERT Peng et al. (2019), BERT Devlin et al. (2018), and T5 Raffel et al. (2019). For the BioASQ tasks (Table 3), we compare the lenient accuracy of base SciFive only with base T5 and base BioBERT due to the time required for thorough expert assessment. It should be noted that BioBERT was the winner of these BioASQ challenges. We achieved SOTA results on 3/7 NER tasks, 2/2 RE tasks, 1/1 NLI tasks, and 3/3 question answering tasks (Table 5). We also achieved a near-SOTA result on the HoC document classification task. Based on these results, we emphasize the following point: SciFive (both base and large model) competitive results on classification tasks while also providing SOTA results on text generation tasks such as question-answering. This is a significant improvement over BERT-based models, which demonstrates weaker performances on question-answering tasks.

Discussion

We used SciFive to explore the role of text generation models in broad-spectrum biomedical NLP, achieving SOTA results on a variety of tasks. This is particularly true for question answering, where SciFive achieved SOTA results. Both T5 and SciFive significantly outperformed BioBERT, highlighting the value of text generation models in biomedical NLP. However, question answering is relatively simplistic compared to other text generation tasks. To fully examine the potential of text generation models in the context of domain-specific literature, SciFive will be applied to tasks such as document summarization and abstract generation.

From our results, it can be seen that the SOTA results are split between the various versions of SciFive. While we expected the Pubmed+PMC model to have the best performances given the mixture of abstracts and full text articles, our results show that further study is needed to understand the optimal nature of biomedical corpora.

Conclusion

In this manuscript, we introduce SciFive, a domain-specific text-to-text model trained specifically for tasks involving biomedical literature. SciFive is effective for NER, RE, NLI, and question answering tasks, achieving SOTA or near-SOTA results in all cases. This outcome supports our conclusion that text-to-text (text generation) models are highly versatile and broadly applicable within domain-specific contexts. These models can be used for common tasks and tasks which require a longer sequence of text as an output (i.e. question answering). Our results suggest the need for further study of domain-specific text generation models applied to more difficult tasks such as a document summarization and abstract generation.

Funding

This work has been supported by the Cancer Research Training Award (CRTA) through the National Cancer Institute to JTA and EB. This research was supported by the Intramural Research Program of the NIH.

References