Contextualized Representations Using Textual Encyclopedic Knowledge
Mandar Joshi, Kenton Lee, Yi Luan, Kristina Toutanova
Introduction
Current self-supervised representations, trained at large scale from document-level contexts, are known to encode linguistic Tenney et al. (2019) and factual Petroni et al. (2019) knowledge into their parameters. Yet, even large pretrained representations are unable to capture and preserve all factual knowledge they have “read” during pretraining due to the long tail of entity and event-specific information Logan et al. (2019). For open-domain tasks, where the input consists of only a question or a statement out of context, as in open-domain QA or factuality prediction, previous work has retrieved and used text that may contain or entail the needed answer to build representations for the task Chen et al. (2017); Guu et al. (2020); Lewis et al. (2020); Oh et al. (2017); Kadowaki et al. (2019).
On the other hand, when relevant text is provided as input, such as in reading comprehension tasks Rajpurkar et al. (2016), relation extraction, syntactic analysis, etc., which can be cast as tasks of labeling spans in the input text, prior work has not focused on drawing background information from external text sources. Instead, most research has explored architectures to integrate background from structured knowledge bases to form input text representations Bauer et al. (2018); Mihaylov and Frank (2018); Yang et al. (2019); Zhang et al. (2019); Peters et al. (2019). A notable exception is Weissenborn et al. (2017), with a specialized architecture which uses textual entity descriptions.
We posit that representations should be able to directly integrate textual background knowledge since a wider scope of information is more readily available in textual form. Our method represents input texts by jointly encoding them with dynamically retrieved sentences from the Wikipedia pages of entities they mention. We term these representations TEK-enriched, for Textual Encyclopedic Knowledge (Figure 2 shows an illustration), and use them for reading comprehension (RC) by contextualizing questions and passages together with retrieved Wikipedia background sentences. Such background knowledge can help reason about the relationships between questions and passages. Figure 1 shows an example question from the TriviaQA dataset Joshi et al. (2017) asking for the pen-name of a gossip columnist. Encoding relevant background knowledge (pseudonymous byline of a gossip column published in the Daily Express) helps ground the vague reference to the William Hickey column in the given document context.
Using text as background knowledge allows us to directly reuse powerful pretrained BERT-style encoders Devlin et al. (2019). We show that an off-the-shelf RoBERTa Liu et al. (2019b) model can be directly finetuned on minimally structured TEK-enriched inputs, which are formatted to allow the encoder to distinguish between the original passages and background sentences. This method considerably improves on current state-of-the art methods which only consider context from a single input document (Section 4). The improvement comes without an increase in the length of the input window for the Transformer Vaswani et al. (2017).
Although existing pretrained models provide a good starting point for task-specific TEK-enriched representations, there is still a mismatch between the type of input seen during pretraining (single document segments) and the type of input the model is asked to represent for downstream tasks (document text with background Wikipedia sentences from multiple pages). We show that the Transformer model can be substantially improved by reducing this mismatch via self-supervised masked language model (MLM) Devlin et al. (2019) pretraining on TEK-augmented input texts.
Our approach records considerable improvements over state of the art base (12-layer) and large (24-layer) Transformer models for in-domain and out-of-domain document-level extractive question answering (QA), for tasks where factual knowledge about entities is important and well-covered by the background collection. On TriviaQA, we see improvements of 1.6 to 3.1 F1, respectively, over comparable RoBERTa models which do not integrate background information. On MRQA Fisch et al. (2019), a large collection of diverse QA datasets, we see consistent gains in-domain along with large improvements out-of-domain on BioASQ (2.1 to 4.2 F1), TextbookQA (1.6 to 2.0 F1), and DuoRC (1.1 to 2.0 F1).
TEK-enriched Representations
We follow recent work on pretraining bidirectional Transformer representations on unlabeled text, and finetuning them for downstream tasks Devlin et al. (2019). Subsequent approaches have shown significant improvements over BERT by improving the training example generation, the masking strategy, the pretraining objectives, and the optimization methods Liu et al. (2019b); Joshi et al. (2020). We build on these improvements to train TEK-enriched representations and use them for extractive QA.
Our approach seeks to contextualize input text jointly with relevant textual encyclopedic background knowledge retrieved dynamically from multiple documents. We define a retrieval function, , which takes as input and retrieves a list of text spans from the corpus . In our implementation, each of the text spans is a sentence. The encoder then represents by jointly encoding with using such that the output representations of are cognizant of the information present in (see Figure 2). We use a deep Transformer encoder operating over the input sequence [CLS] [SEP] [SEP]for .
We refer to inputs generically as contexts. These could be either contiguous word sequences from documents (passages), or, for the QA application, question-passage pairs, which we refer to as RC-contexts. For a fixed Transformer input length limit (which is necessary for computational efficiency), there is a trade-off between the length of the document context (the length of ) and the amount of background knowledge (the length of ). Section 5 explores this trade-off and shows that for an encoder input limit of 512, the values of for the length of and for the length of provide an effective compromise.
We use a simple implementation of the background retrieval function , using an entity linker for finetuning (Section 2.1) and Wikipedia hyperlinks for pretraining (Section 2.2), and a way to score the relevance of individual sentences using ngram overlap.
The input for the extractive QA task consists of the question and a candidate passage . We use the following retrieval function to obtain relevant background .
We detect entity mentions in using a proprietary Wikipedia-based entity linker,We also report results on publicly available linkers showing that our method is robust to the exact choice of the linker (Section 5). and form a candidate pool of background segments as the union of the sentences in the Wikipedia pages of the detected entities. These sentences are then ranked based on their number of overlapping ngrams with the question (equally weighted unigrams, bigrams, and trigrams). To form the input for the Transformer encoder, each background sentence is minimally structured as by prepending the name of the entity whose page it belongs to along with a separator ‘:’ token. Each sentence is followed by [SEP]. Appendix A shows an example of an RC-context with background knowledge segments.
QA Model
Following BERT, our QA model architecture consists of two independent linear classifiers for predicting the answer span boundary (start and end) on top of the output representations of . We assume that the answer, if present, is contained only in the given passage, , and do not consider potential mentions of the answer in the background . For instances which do not contain the answer, we set the answer span to be the special token [CLS]. We use a fixed Transformer input window size of , and use a sliding window with a stride of tokens to handle longer documents. Our TEK-enriched representations use document passages of length while baselines use longer passages of length .
2 TEK-enriched Pretraining
Standard pretraining uses contiguous document-level natural language inputs. Since TEK-augmented inputs are formatted as natural language sequences, off-the-shelf pretrained models can be used as a starting point for creating TEK-enriched representations. As one of our approaches, we use a standard single-document pretraining model.
While the input format is the same, there is a mismatch between contiguous document segments and TEK-augmented inputs sourced from multiple documents. We propose an additional pretraining stage—starting from the RoBERTa parameters, we resume pretraining using an MLM objective on TEK-augmented document text , which encourages the model to integrate the knowledge from multiple background segments.
In pretraining, is a contiguous block of text from Wikipedia. The retrieval function returns where each is a sentence from the Wikipedia page of some entity hyperlinked from a span in . We use high-precision Wikipedia hyperlinks instead of an entity linker for pretraining. The background candidate sentences are ranked by their ngram overlap with . The top ranking sentences in up to tokens are used. If no entities are found in , is constructed from the context following from the same document.
Training Objective
We continue pretraining a deep Transformer using the MLM objective Devlin et al. (2019) after initializing the parameters with pretrained RoBERTa weights. Following improvements in SpanBERT Joshi et al. (2020), we mask spans with lengths sampled from a geometric distribution in the entire input ( and ). We use a single segment ID, and remove the next sentence prediction objective which has been shown to not improve performance Joshi et al. (2020); Liu et al. (2019b) for multiple tasks including QA. We evaluate two methods building textual-knowledge enriched representations for QA differing in the pretraining approach used:
TEKPF
TEKF
TEKF replaces the first specialized pretraining stage in TEKPF with 200K steps for standard single-document-context pretraining for a fair comparison with TEKPF, but follows the same finetuning regimen.
Experimental Setup
We perform experiments on TriviaQA and MRQA, two large extractive question answering benchmarks (see Table 1 for dataset statistics).
TriviaQA Joshi et al. (2017) contains trivia questions paired with evidence collected via entity linking and web search. The dataset is distantly supervised in that the answers are contained in the evidence but the context may not support answering the questions. We experiment with both the Wikipedia and Web tasks.
MRQA
The MRQA shared task Fisch et al. (2019) consists of several widely used QA datasets unified into a common format aimed at evaluating out-of-domain generalization. The data consists of a training set, in-domain and out-of-domain dev sets, and a private out-of-domain test set. The training and the in-domain dev sets consist of modified versions of corresponding sets from SQuAD Rajpurkar et al. (2016), NewsQA Trischler et al. (2017), SearchQA Dunn et al. (2017), TriviaQA Web Joshi et al. (2017), HotpotQA Yang et al. (2018) and Natural Questions Kwiatkowski et al. (2019). The out-of-domain test evaluation, including access to questions and passages, is only available through Codalab. Due to the complexity of our system which involves entity linking and retrieval, we perform development and model selection on the in-domain dev set and treat the out-of-domain dev set as the test set. The out-of-domain set we evaluate on has examples from BioASQ Tsatsaronis et al. (2015), DROP Dua et al. (2019), DuoRC Saha et al. (2018), RACE Lai et al. (2017), RelationExtraction Levy et al. (2017), and TextbookQA Kembhavi et al. (2017).
1 Baselines
We compare TEKPF and TEKF with two baselines, RoBERTa and RoBERTa++. Both use the same architecture as our approach, but use only original RC-contexts for finetuning and inference, and use standard single-document-context RoBERTa pretraining. TEKPF and TEKF use and , while both baselines use and .
We finetune the model on QA data without knowledge augmentation starting from the same RoBERTa checkpoint that is used as an initializer for TEK-augmented pretraining.
RoBERTa++
For a fair evaluation of the new TEK-augmented pretraining method while controlling for the number of pretraining steps and other hyperparameters, we extend RoBERTa’s pretraining for an additional 200K steps on single contiguous blocks of text (without background information). We use the same masking and other hyperparameters as in TEK-augmented pretraining. This pretrained checkpoint is also used to initialize parameters for our TEKF approach.
The implementation details of all models, including hyperpameters, can be found in Appendix B.
Results
Table 2 compares our approaches with baselines and previous work. The 12-layer variant of our RoBERTa baseline outperforms or matches the performance of several previous systems including ELMo-based ones Wang et al. (2018); Lewis (2018) which are specialized for this task. We also see that RoBERTa++ outperforms RoBERTa, indicating that there is still room for improvement by simply pretraining for more steps on task-domain relevant text. Furthermore, the 12-layer and 24-layer variants of our TEKF approach considerably improve over a comparable RoBERTa++ baseline for both Wikipedia (1.9 and 1.1 F1 respectively) and Web (1.6 and 1.0 F1 respectively) indicating that TEK representations are useful even without additional TEK-pretraining. The base variant of our best model TEKPF, which uses TEK-pretrained TEK-enriched representations records even bigger gains of 3.1 F1 and 2.0 F1 on Wikipedia and Web respectively over a comparable 12-layer RoBERTa++ baseline. The 24-layer models show similar trends with improvements of 1.6 and 1.7 F1 over RoBERTa++.
MRQA
Table 3 shows in-domain and out-of-domain evaluation on MRQA. As in the case of TriviaQA, the 12-layer variants of our RoBERTa baselines are competitive with previous work, which includes D-Net Li et al. (2019) and Delphi Longpre et al. (2019), the top two systems of the MRQA shared task, while the 24-layer variants considerably outperform the current state of the art across all datasets. RoBERTa++ again performs better than RoBERTa on all datasets except DROP and RACE. DROP is designed to test arithmetic reasoning, while RACE contains (often fictional and thus not groundable to Wikipedia) passages from English exams for middle and high school students in China. The performance drop after further pretraining on Wikipedia could be a result of multiple factors including the difference in style of required reasoning or content; we leave further investigation of this phenomenon for future work. The base variants of TEKF and TEKPF outperform both baselines on all other datasets. Comparing the base variant of our full TEKPF approach to RoBERTa++, we observe an overall improvement of 1.6 F1 with strong gains on BioASQ (4.2 F1), DuoRC (2.0 F1), and TextbookQA (2.0 F1). The 24-layer variants of TEKPF show similar trends with improvements of 2.1 F1 on BioASQ, 1.1 F1 on DuoRC, and 1.6 F1 on TextbookQA. Our large models see a reduction in the average gain mostly due to drop in performance on DROP. Like in the case of TriviaQA, TEK-pretraining generally improves performance even further where TEK-finetuning is useful (with the exception of DuoRC which sees a small loss of 0.24 F1 due to TEK-pretraining for the large modelsAccording to the Wilcoxon signed rank test of statistical significance, the large TEKPF is significantly better than TEKF on BioASQ and TextbookQA -value , and is not significantly different from it for DuoRC.), with the biggest gains seen on BioASQ.
Takeaways
Both TEKPF and TEKF record strong gains on benchmarks that focus on factual reasoning outperforming the RoBERTa-based baselines that use only RC-contexts. The success of TEKF underscores the advantage of textual encyclopedic knowledge in that it improves current models even without additional TEK-pretraining. Finally, TEK-pretraining further improves the model’s ability to use the retrieved background knowledge for the downstream RC task.
Ablation Studies
We also compare the two pretraining setups for models which do not use background knowledge to form representations for the finetuning tasks. Table 4 shows results for all four combinations of the pretraining and finetuning method variables, using 12-layer base models on the development sets of TriviaQA and MRQA (in-domain). Comparing rows 1 and 3, we see marginal gains across all datasets for TEK pretraining indicating that pretraining with encyclopedic knowledge does not hurt QA performance even when such information is not available during finetuning and inference. While previous work Liu et al. (2019b); Joshi et al. (2020) has shown that pretraining with single contiguous chunks of text clearly outperforms BERT’s bi-sequence pipeline,BERT randomly samples the second sequence from a different document in the corpus with a probability of 0.5. our results suggest that using background sentences from other documents during pretraining has no adverse effect on the downstream tasks we consider.
Trade-off between Document Context and Knowledge
Our approach uses a part of the Transformer window for textual knowledge, instead of additional context from the same document. Having established the usefulness of the background knowledge even without tailored pretraining, we now consider the trade-off between neighboring context and retrieved knowledge (Table 5). We first compare using a shorter window of tokens for RC-contexts with using tokens for RC-contexts (the first two rows). Using longer document context results in consistent gains, some of which our TEK-enriched representations need to sacrifice. We then consider the trade-off for varying values of context length and background length (rows 2-5). The partitioning of tokens for context and for background outperforms other configurations. This suggests that relevant encyclopedic knowledge from outside of the current document is more useful than long-distance neighboring text from the same document for these benchmarks.
Choice of the Entity Linker
Table 6 compares the performance of TEKPF when used with publicly available entity linkers, Google Cloud Natural Language API (abbreviated as GC)https://cloud.google.com/natural-language/docs/basics#entity analysis and TagMe Ferragina and Scaiella (2010). Using TagMe results in a minor drop of around 0.3 F1 from TEKPF across benchmarks while still maintaining major gains over RoBERTa++. The results indicate that the choice of entity linker can make a difference but our method is robust and performs well with multiple linkers.
Discussion
When are TEK-enriched representations most useful for question answering? The strongest gains we have seen are on TriviaQA, BioASQ, and TextbookQA. All three datasets involve questions targeting the long tail of factual information, which has sizable coverage in Wikipedia, the encyclopedic collection we use. We hypothesize that enriching representations with encyclopedic knowledge could be particularly useful when factual information that might be difficult to “memorize” during pretraining is important. Current pretraining methods are able to store a significant amount of world knowledge into model parameters Petroni et al. (2019); this might enable the model to make correct predictions even from contexts with complex phrasing or partial information. TEK-enriched representations complement this strength via dynamic retrieval of factual knowledge. Unlike structured KBs which have been used prominently in previous work, encyclopedic text is more likely to be available for a variety of domains (e.g., biomedical and legal). Improvements on the science-based BioASQ and TextbookQA datasets further suggest that Wikipedia can be used as a bridge corpus for more effective domain adaptation for QA.
For 75% of the examples in the TriviaQA Wikipedia development set where our approach outperforms the context-only baselines, the answer string is mentioned in the background text. A qualitative analysis of these examples indicates that the retrieved background information typically falls into two categories – (a) where the background helps disambiguate between multiple answer candidates by providing partial pieces of information missing from the original context, and (b) where the background sentences help by providing a redundant but more direct phrasing of the information need compared to the original context. Figure 3 provides examples of each category.
Even when the retrieved background contains the answer string, our model uses the background only to refine representations of the candidate answers in the original document context; possible answer positions in the background are not considered in our model formulation. This highlights the strength of an encoder with full cross-attention between RC-contexts and background knowledge. The encoder is able to build representations for, and consider possible answers in all document passages, while integrating knowledge from multiple pieces of external textual evidence.
The exact form of background knowledge is dependent on the retrieval function. Our results have shown that contextualizing the input with textual background knowledge, especially after suitable pretraining, improves state of the art methods even with simple entity linking and ngram-match retrieval functions. We hypothesize that more sophisticated retrieval methods could further significantly improve performance (for example, by prioritizing for more complementary information).
Related Work
Many NLP tasks require the use of multiple kinds of background knowledge Fillmore (1976); Minsky (1986). Earlier work Ratinov and Roth (2009); Nakashole and Mitchell (2015) combined features over the given task data with hand-engineered features over knowledge repositories. Other forms of external knowledge include relational knowledge between word or entity pairs, typically integrated via embeddings from structured knowledge graphs (KGs) Yang and Mitchell (2017); Bauer et al. (2018); Mihaylov and Frank (2018); Wang and Jiang (2019) or via word pair embeddings trained from text Joshi et al. (2019). Weissenborn et al. (2017) used a specialized architecture to integrate background knowledge from ConceptNet and Wikipedia entity descriptions. For open-domain QA, recent works Sun et al. (2019); Xiong et al. (2019) jointly reasoned over text and KGs, via specialized graph-based architectures for defining the flow of information between them. These methods did not take advantage of large scale unlabeled text to pre-train deep contextualized representations which have the capacity to encode even more knowledge in their parameters.
Most relevant to ours is work building upon these powerful pretrained representations, and further integrating external knowledge. Recent work focuses on refining pretrained contextualized representations using entity or triple embeddings from structured KGs Peters et al. (2019); Yang et al. (2019); Zhang et al. (2019). The KG embeddings are trained separately (often to predict links in the KG), and knowledge from KG is fused with deep Transformer representations via special-purpose architectures. Some of these prior works also pre-train the knowledge fusion layers from unlabeled text through self-supervised objectives Zhang et al. (2019); Peters et al. (2019). Instead of separately encoding structured KBs, and then attending to their single-vector embeddings, we explore directly using wider-coverage textual encyclopedic background knowledge. This enables direct application of a pretrained deep Transformer (RoBERTa) for jointly contextualizing input text and background knowledge. We showed background knowledge integration can be further improved by additional knowledge-augmented self-supervised pretraining.
Liu et al. (2019a) augment text with relevant triples from a structured KB. They process triples as word sequences using BERT with a special-purpose attention masking strategy. This allows the model to partially re-use BERT for encoding and integrating the structured knowledge. Our work uses wider-coverage textual sources instead and shows the power of additional knowledge-tailored self-supervised pretraining.
Question Answering
For open-domain QA, where documents known to answer the question are not given as input (e.g. OpenBookQA Mihaylov et al. (2018)), methods exploring retrieval of relevant textual knowledge are a necessity. Recent work in these areas has focused on improving the evidence retrieval components Lee et al. (2019); Banerjee et al. (2019); Guu et al. (2020), and has used Wikidata triples with textual descriptions of Wikipedia entities as a source of evidence Min et al. (2019). Other approaches use pseudo-relevance feedback (PRF) Xu and Croft (1996) style multi-step retrieval of passages by query reformulation Buck et al. (2018); Nogueira and Cho (2017), entity linking Das et al. (2019b), and more complex reader-retriever interaction Das et al. (2019a). When multiple candidate contexts are retrieved for open-domain QA, they are sometimes jointly contextualized using a specialized architecture Min et al. (2019). We are the first to explore pretraining of representations which can integrate background from multiple documents, and hypothesize that these representations could be further improved by more sophisticated retrieval approaches.
Conclusion
We presented a method to build text representations by jointly contextualizing the input with dynamically retrieved textual encyclopedic knowledge. We showed consistent improvements, in- and out-of-domain, across multiple reading comprehension benchmarks that require factual reasoning and knowledge well represented in the background collection.
References
Appendices
Table 7 shows pretraining (left) and QA finetuning (right) examples which encode contexts with background sentences from Wikipedia.
B Implementation
We implemented all models in TensorFlow Abadi et al. (2015). For pretraining, we used the 12-layer RoBERTa-base (125M parameters) and 24-layer RoBERTa-large (355M parameters) configurations, and initialized the parameters from their respective checkpoints. In TEK-augmented pretraining, we further pretrained the models for 200K steps with a batch size of and BERT’s triangular learning rate schedule with a warmup of steps on TEK-augmented contexts. We used a peak learning rate of for base and for large models. All models were trained and evaluated on Google Cloud TPUs. We apply the following finetuning hyperparameters to all methods, including the baselines. For each method, we chose the best model based on dev set performance measured using F1.
We follow the input preprocessing of Clark and Gardner (2018). The input to our model is the concatenation of the first four 400-token passages selected by their linear passage ranker. For training, we define the gold span to be the first occurrence of the gold answer(s) in the context Joshi et al. (2017); Talmor and Berant (2019). We choose learning rates from {1e-5, 2e-5} and finetune for 5 epochs with a batch size of 32.
MRQAhttps://github.com/mrqa/MRQA-Shared-Task-2019
We choose learning rates from {1e-5, 2e-5} and number of epochs from {2, 3, 5} with a batch size of 32. For both benchmarks, especially for large models, we found higher learning rates to perform sub-optimally on the development sets. Table 8 reports best performing hyperparamter configurations for each benchmark.