A Frustratingly Easy Approach for Entity and Relation Extraction

Zexuan Zhong, Danqi Chen

Introduction

Extracting entities and their relations from unstructured text is a fundamental problem in information extraction. This problem can be decomposed into two subtasks: named entity recognition Sang and De Meulder (2003); Ratinov and Roth (2009) and relation extraction Zelenko et al. (2002); Bunescu and Mooney (2005). Early work employed a pipelined approach, training one model to extract entities Florian et al. (2004, 2006), and another model to classify relations between them Zhou et al. (2005); Kambhatla (2004); Chan and Roth (2011). More recently, however, end-to-end evaluations have been dominated by systems that model these two tasks jointly Li and Ji (2014); Miwa and Bansal (2016); Katiyar and Cardie (2017); Zhang et al. (2017a); Li et al. (2019); Luan et al. (2018, 2019); Wadden et al. (2019); Lin et al. (2020); Wang and Lu (2020). There has been a long held belief that joint models can better capture the interactions between entities and relations and help mitigate error propagation issues.

In this work, we re-examine this problem and present a simple approach which learns two encoders built on top of deep pre-trained language models Devlin et al. (2019); Beltagy et al. (2019); Lan et al. (2020). The two models — which we refer them as to the entity model and relation model throughout the paper — are trained independently and the relation model only relies on the entity model to provide input features. Our entity model builds on span-level representations and our relation model builds on contextual representations specific to a given pair of spans. Despite its simplicity, we find this pipelined approach to be extremely effective: using the same pre-trained encoders, our model outperforms all previous joint models on three standard benchmarks: ACE04, ACE05 and SciERC, advancing the previous state-of-the-art by 1.7%–2.8% absolute in relation F1.

To better understand the effectiveness of this approach, we carry out a series of careful analyses. We observe that, (1) the contextual representations for the entity and relation models essentially capture distinct information, so sharing their representations hurts performance; (2) it is crucial to fuse the entity information (both boundary and type) at the input layer of the relation model; (3) leveraging cross-sentence information is useful in both tasks. Hence, we expect that this simple model will serve as a very strong baseline in end-to-end relation extraction and make us rethink the value of joint modeling of entities and relations.

Finally, one possible shortcoming of our approach is that we need to run our relation model once for every pair of entities. To alleviate this issue, we present a novel and efficient alternative by approximating and batching the computations for different groups of entity pairs at inference time. This approximation achieves an 8-16×\times speedup with only a slight reduction in accuracy (e.g., 1.0%1.0\% F1 drop on ACE05), which makes our model fast and accurate to use in practice. Our final system is called PURE (the Princeton University Relation Extraction system) and we make our code and models publicly available for the research community.

We summarize our contributions as follows:

We present a simple and effective approach for end-to-end relation extraction, which learns two independent encoders for entity recognition and relation extraction. Our model establishes the new state-of-the-art on three standard benchmarks and surpasses all previous joint models.

We conduct careful analyses to understand why our approach performs so well and how different factors impact the final performance. We conclude that it is more effective to learn distinct contextual representations for entities and relations than to learn them jointly.

To speed up the inference time of our model, we also propose a novel efficient approximation, which achieves a large runtime improvement with only a small accuracy drop.

Related Work

Traditionally, extracting relations between entities in text has been studied as two separate tasks: named entity recognition and relation extraction. In the last several years, there has been a surge of interest in developing models for joint extraction of entities and relations Li and Ji (2014); Miwa and Sasaki (2014); Miwa and Bansal (2016). We group existing joint models into two categories: structured prediction and multi-task learning:

Structured prediction approaches cast the two tasks into one unified framework, although it can be formulated in various ways. Li and Ji (2014) propose an action-based system which identifies new entities as well as links to previous entities, Zhang et al. (2017a); Wang and Lu (2020) adopt a table-filling approach proposed in Miwa and Sasaki (2014); Katiyar and Cardie (2017) and Zheng et al. (2017) employ sequence tagging-based approaches; Sun et al. (2019) and Fu et al. (2019) propose graph-based approaches to jointly predict entity and relation types; and, Li et al. (2019) convert the task into a multi-turn question answering problem. All of these approaches need to tackle a global optimization problem and perform joint decoding at inference time, using beam search or reinforcement learning.

Multi-task learning

This family of models essentially builds two separate models for entity recognition and relation extraction and optimizes them together through parameter sharing. Miwa and Bansal (2016) propose to use a sequence tagging model for entity prediction and a tree-based LSTM model for relation extraction. The two models share one LSTM layer for contextualized word representations and they find sharing parameters improves performance (slightly) for both models. The approach of Bekoulis et al. (2018) is similar except that they model relation classification as a multi-label head selection problem. Note that these approaches still perform pipelined decoding: entities are first extracted and the relation model is applied on the predicted entities.

The closest work to ours is DYGIE and DYGIE++ Luan et al. (2019); Wadden et al. (2019), which builds on recent span-based models for coreference resolution Lee et al. (2017) and semantic role labeling He et al. (2018). The key idea of their approaches is to learn shared span representations between the two tasks and update span representations through dynamic graph propagation layers. A more recent work Lin et al. (2020) further extends DYGIE++ by incorporating global features based on cross-substask and cross-instance constraints.This is an orthogonal contribution to ours and we will explore it for future work. Our approach is much simpler and we will detail the differences in Section 3.2 and explain why our model performs better.

Method

In this section, we first formally define the problem of end-to-end relation extraction in Section 3.1 and then detail our approach in Section 3.2. Finally, we present our approximation solution in Section 3.3, which considerably improves the efficiency of our approach during inference.

The input of the problem is a sentence XX consisting of nn tokens x1,x2,…,xnx_{1},x_{2},\dots,x_{n}. Let S={s1,s2,…,sm}S=\{s_{1},s_{2},\dots,s_{m}\} be all the possible spans in XX of up to length LL and \textscSTART(i)\textsc{START}(i) and \textscEND(i)\textsc{END}(i) denote start and end indices of sis_{i}. Optionally, we can incorporate cross-sentence context to build better contextual representations (Section 3.2). The problem can be decomposed into two sub-tasks:

Let E\mathcal{E} denote a set of pre-defined entity types. The named entity recognition task is, for each span si∈Ss_{i}\in S, to predict an entity type ye(si)∈Ey_{e}(s_{i})\in\mathcal{E} or ye(si)=ϵy_{e}(s_{i})=\epsilon representing span sis_{i} is not an entity. The output of the task is Ye={(si,e)\mathchar58si∈S,e∈E}Y_{e}=\{(s_{i},e)\mathrel{\mathop{\mathchar 58\relax}}s_{i}\in S,e\in\mathcal{E}\}.

Relation extraction

Let R\mathcal{R} denote a set of pre-defined relation types. The task is, for every pair of spans si∈S,sj∈Ss_{i}\in S,s_{j}\in S, to predict a relation type yr(si,sj)∈Ry_{r}(s_{i},s_{j})\in\mathcal{R}, or there is no relation between them: yr(si,sj)=ϵy_{r}(s_{i},s_{j})=\epsilon. The output of the task is Yr={(si,sj,r)\mathchar58si,sj∈S,r∈R}Y_{r}=\{(s_{i},s_{j},r)\mathrel{\mathop{\mathchar 58\relax}}s_{i},s_{j}\in S,r\in\mathcal{R}\}.

2 Our Approach

As shown in Figure 1, our approach consists of an entity model and a relation model. The entity model first takes the input sentence and predicts an entity type (or ϵ\epsilon) for each single span. We then process every pair of candidate entities independently in the relation model by inserting extra marker tokens to highlight the subject and object and their types. We will detail each component below, and finally summarize the differences between our approach and DYGIE++ Wadden et al. (2019).

Our entity model is a standard span-based model following prior work Lee et al. (2017); Luan et al. (2018, 2019); Wadden et al. (2019). We first use a pre-trained language model (e.g., BERT) to obtain contextualized representations xt\mathbf{x}_{t} for each input token xtx_{t}. Given a span si∈Ss_{i}\in S, the span representation he(si)\mathbf{h}_{e}(s_{i}) is defined as:

Relation model

The relation model aims to take a pair of spans si,sjs_{i},s_{j} (a subject and an object) as input and predicts a relation type or ϵ\epsilon. Previous approaches Luan et al. (2018, 2019); Wadden et al. (2019) re-use the span representations he(si),he(sj)\mathbf{h}_{e}(s_{i}),\mathbf{h}_{e}(s_{j}) to predict the relationship between sis_{i} and sjs_{j}. We hypothesize that these representations only capture contextual information around each individual entity and might fail to capture the dependencies between the pair of spans. We also argue that sharing the contextual representations between different pairs of spans may be suboptimal. For instance, the words is a in Figure 1 are crucial in understanding the relationship between MORPA and parser but not for MORPA and text-to-speech.

Our relation model instead processes each pair of spans independently and inserts typed markers at the input layer to highlight the subject and object and their types. Specifically, given an input sentence XX and a pair of subject-object spans si,sjs_{i},s_{j}, where sis_{i}, sjs_{j} have a type of ei,ej∈E∪{ϵ}e_{i},e_{j}\in\mathcal{E}\cup\{\epsilon\} respectively. We define text markers as ⟨\textscS:ei⟩\langle\textsc{S:}e_{i}\rangle, ⟨\textsc/S:ei⟩\langle\textsc{/S:}e_{i}\rangle, ⟨\textscO:ej⟩\langle\textsc{O:}e_{j}\rangle, and ⟨\textsc/O:ej⟩\langle\textsc{/O:}e_{j}\rangle, and insert them into the input sentence before and after the subject and object spans (Figure 1 (b)).Our final model indeed only considers ei,ej≠ϵe_{i},e_{j}\neq\epsilon. We have explored strategies using spans which are predicted as ϵ\epsilon for the relation model but didn’t find improvement. See Section 5.3 for more discussion. Let X^\widehat{X} denote this modified sequence with text markers inserted:

We apply a second pre-trained encoder on X^\widehat{X} and denote the output representations by x^t\mathbf{\widehat{x}}_{t}. We concatenate the output representations of two start positions and obtain the span-pair representation:

where \textscSTART(i)^\widehat{\textsc{START}(i)} and \textscSTART(j)^\widehat{\textsc{START}(j)} are the indices of ⟨\textscS:ei⟩\langle\textsc{S:}e_{i}\rangle and ⟨\textscO:ej⟩\langle\textsc{O:}e_{j}\rangle in X^\widehat{X}. Finally, the representation hr(si,sj)\mathbf{h}_{r}(s_{i},s_{j}) will be fed into a feedforward network to predict the probability distribution of the relation type r∈R∪{ϵ}r\in\mathcal{R}\cup\{\epsilon\}: Pr(r∣si,sj)P_{r}(r|s_{i},s_{j}).

This idea of using additional markers to highlight the subject and object is not entirely new as it has been studied recently in relation classification Zhang et al. (2019); Soares et al. (2019); Peters et al. (2019). However, most relation classification tasks (e.g., TACRED Zhang et al. (2017b)) only focus on a given pair of subject and object in an input sentence and its effectiveness has not been evaluated in the end-to-end setting in which we need to classify the relationships between multiple entity mentions. We observed a large improvement in our experiments (Section 5.1) and this strengthens our hypothesis that modeling the relationship between different entity pairs in one sentence require different contextual representations. Furthermore, Zhang et al. (2019); Soares et al. (2019) only consider untyped markers (e.g., ⟨\textscS⟩\langle\textsc{S}\rangle, ⟨\textsc/S⟩\langle\textsc{/S}\rangle) and previous end-to-end models (e.g., Wadden et al. (2019)) only inject the entity type information into the relation model through auxiliary losses. We find that injecting type information at the input layer is very helpful in distinguishing entity types — for example, whether “Disney” refers to a person or an organization— before trying to understand the relations.

Cross-sentence context

Cross-sentence information can be used to help predict entity types and relations, especially for pronominal mentions. Luan et al. (2019); Wadden et al. (2019) employ a propagation mechanism to incorporate cross-sentence context. Wadden et al. (2019) also add a 3-sentence context window which is shown to improve performance. We also evaluate the importance of leveraging cross-sentence context in our approach. As we expect that pre-trained language models to be able to capture long-range dependencies, we simply incorporate cross-sentence context by extending the sentence to a fixed window size WW for both the entity and relation model. Specifically, given an input sentence with nn words, we augment the input with (W−n)/2(W-n)/2 words from the left context and right context respectively.

Training & inference

For both entity model and relation model, we fine-tune the two pre-trained language models using task-specific losses. We use cross-entropy loss for both models:

where ei∗e_{i}^{*} represents the gold entity type of sis_{i} and ri,j∗r_{i,j}^{*} represents the gold relation type of span pair si,sjs_{i},s_{j} in the training data. For training the relation model, we only consider the gold entities SG⊂SS_{G}\subset S in the training set and use the gold entity labels as the input of the relation model. We considered training on predicted entities as well as all spans SS (with pruning), but none of them led to meaningful improvements compared to this simple pipelined training (see more discussion in Section 5.3). During inference, we first predict the entities by taking ye(si)=arg max⁡e∈E∪{ϵ}Pe(e∣si)y_{e}(s_{i})=\operatorname*{arg\,max}_{e\in\mathcal{E}\cup\{\epsilon\}}P_{e}(e|s_{i}). Denote Spred={si\mathchar58ye(si)≠ϵ}S_{\text{pred}}=\{s_{i}\mathrel{\mathop{\mathchar 58\relax}}y_{e}(s_{i})\neq\epsilon\}, we enumerate all the spans si,sj∈Spreds_{i},s_{j}\in S_{\text{pred}} and use ye(si),ye(sj)y_{e}(s_{i}),y_{e}(s_{j}) to construct the input for the relation model Pr(r∣si,sj)P_{r}(r\mid s_{i},s_{j}).

Differences from DYGIE++

Our approach differs from DYGIE++ Luan et al. (2019); Wadden et al. (2019) in the following ways: (1) We use separate encoders for the entity and relation models, without any multi-task learning. The predicted entity types are used directly to construct the input for the relation model. (2) The contextual representations in the relation model are specific to each pair of spans by using the text markers. (3) We only incorporate cross-sentence information by extending the input with additional context (as they did) and we do not employ any graph propagation layers and beam search.They also incorporated coreferences and event prediction in their framework. We focus on entity and relation extraction in this paper and we leave these extensions to future work. As a result, our model is much simpler. As we will show in the experiments (Section 4), it also achieves large gains in all the benchmarks, using the same pre-trained encoders.

3 Efficient Batch Computations

One possible shortcoming of our approach is that we need to run our relation model once for every pair of entities. To alleviate this issue, we propose a novel and efficient alternative to our relation model. The key problem is that we would like to re-use computations for different pairs of spans in the same sentence. This is impossible in our original model because we must insert the entity markers for each pair of spans independently. To this end, we propose an approximation model by making two major changes to the original relation model. First, instead of directly inserting entity markers into the original sentence, we tie the position embeddings of the markers with the start and end tokens of the corresponding span:

where \textscp(⋅)\textsc{p}(\cdot) denotes the position id of a token. As the example shown in Figure 1, if we want to classify the relationship between MORPA and parser, the first entity marker ⟨S\langle S\mathchar58\textscMethod⟩\mathrel{\mathop{\mathchar 58\relax}}\textsc{Method}\rangle will share the position embedding with the token mor. By doing this, the position embeddings of the original tokens will not be changed.

Second, we add a constraint to the attention layers. We enforce the text tokens to only attend to text tokens and not attend to the marker tokens while an entity marker token can attend to all the text tokens and all the 4 marker tokens associated with the same span pair. These two modifications allow us to re-use the computations of all text tokens, because the representations of text tokens are independent of the entity marker tokens. Thus, we can batch multiple pairs of spans from the same sentence in one run of the relation model. In practice, we add all marker tokens to the end of the sentence to form an input that batches a set of span pairs (Figure 1(c)). This leads to a large speedup at inference time and only a small drop in performance (Section 4.3).

Experiments

We evaluate our approach on three popular end-to-end relation extraction datasets: ACE05catalog.ldc.upenn.edu/LDC2006T06, ACE04catalog.ldc.upenn.edu/LDC2005T09, and SciERC Luan et al. (2018). Table 2 shows the data statistics of each dataset. The ACE05 and ACE04 datasets are collected from a variety of domains, such as newswire and online forums. The SciERC dataset is collected from 500 AI paper abstracts and defines scientific terms and relations specially for scientific knowledge graph construction. We follow previous work and use the same preprocessing procedure and splits for all datasets. See Appendix A for more details.

Evaluation metrics

We follow the standard evaluation protocol and use micro F1 measure as the evaluation metric. For named entity recognition, a predicted entity is considered as a correct prediction if its span boundaries and the predicted entity type are both correct. For relation extraction, we adopt two evaluation metrics: (1) boundaries evaluation (Rel): a predicted relation is considered as a correct prediction if the boundaries of two spans are correct and the predicted relation type is correct; (2) strict evaluation (Rel+): in addition to what is required in the boundaries evaluation, predicted entity types also must be correct. More discussion of the evaluation settings can be found in Bekoulis et al. (2018); Taillé et al. (2020).

Implementation details

We use bert-base-uncased Devlin et al. (2019) and albert-xxlarge-v1 Lan et al. (2020) as the base encoders for ACE04 and ACE05, for a fair comparison with previous work and an investigation of small vs large pre-trained models. As detailed in Table 1, some previous work used BERT-large models. We are not able to do a comprehensive study of all the pre-trained models and our BERT-base results are generally higher than most published results using larger models. We also use scibert-scivocab-uncased Beltagy et al. (2019) as the base encoder for SciERC, as this in-domain pre-trained model is shown to be more effective than BERT Wadden et al. (2019). We use a context window size of W=300W=300 for the entity model and W=100W=100 for the relation model in our default setting using cross-sentence contextWe use a context window size W=100W=100 for the ALBERT entity models to reduce GPU memory usage. and the effect of different context sizes is provided in Section 5.4. We consider spans up to L=8L=8 words. For all the experiments, we report the averaged F1 scores of 5 runs. More implementation details can be found in Appendix B.

2 Main Results

Table 1 compares our approach PURE to all the previous results. We report the F1 scores in both single-sentence and cross-sentence settings. As is shown, our single-sentence models achieve strong performance and incorporating cross-sentence context further improves the results considerably. Our BERT-base (or SciBERT) models achieve similar or better results compared to all the previous work including models built on top of larger pre-trained LMs, and our results are further improved by using a larger encoder ALBERT.

For entity recognition, our best model achieves an absolute F1 improvement of +1.4%+1.4\%, +1.7%+1.7\%, +1.4%+1.4\% on ACE05, ACE04, and SciERC respectively. This shows that cross-sentence information is useful for the entity model and pre-trained Transformer encoders are able to capture long-range dependencies from a large context. For relation extraction, our approach outperforms the best previous methods by an absolute F1 of +1.8%+1.8\%, +2.8%+2.8\%, +1.7%+1.7\% on ACE05, ACE04, and SciERC respectively. We also obtained a 4.3%4.3\% higher relation F1 on ACE05 compared to DYGIE++ Wadden et al. (2019) using the same BERT-base pre-trained model. Compared to the previous best approaches using either global features Lin et al. (2020) or complex neural models (e.g., MT-RNNs) Wang and Lu (2020), our approach is much simpler and achieves large improvements on all the datasets. Such improvements demonstrate the effectiveness of learning representations for entities and relations of different entity pairs, as well as early fusion of entity information in the relation model. We also noticed that compared to the previous state-of-the-art model Wang and Lu (2020) based on ALBERT, our model achieves a similar entity F1 (89.5 vs 89.7) but a substantially better relation F1 (67.6 vs 69.0) without using context. This clearly demonstrates the superiority of our relation model. Finally, we also compare our model to a joint model (similar to DYGIE++) of different data sizes to test the generality of our results. As shown in Appendix C, our findings are robust to data sizes.

3 Batch Computations and Speedup

In Section 3.3, we proposed an efficient approximation solution for the relation model, which enables us to re-use the computations of text tokens and batch multiple span pairs in one input sentence. We evaluate this approximation model on ACE05 and SciERC. Table 3 shows the relation F1 scores and the inference speed of the full relation model and the approximation model. On both datasets, our approximation model significantly improves the efficiency of the inference process.Note that we only applied this batch computation trick at inference time, because we observed that training with batch computation leads to a slightly (and consistently) worse result. We hypothesize that this is due to the impact of increased batch sizes. We still modified the position embedding and attention masks during training (without batching the instances though). For example, we obtain a 11.9×11.9\times speedup on ACE05 and a 8.7×8.7\times speedup on SciERC in the single-sentence setting. By re-using a large part of computations, we are able to make predictions on the full ACE05 test set (2k sentences) in less than 1010 seconds on a single GPU. On the other hand, this approximation only leads to a small performance drop and the relaion F1 measure decreases by only 1.0%1.0\% and 1.2%1.2\% on ACE05 and SciERC respectively in the single-sentence setting. Considering the accuracy and efficiency of this approximation model, we expect it to be very effective to use in practice.

Analysis

Despite its simple design and training paradigm, we have shown that our approach outperforms all previous joint models. In this section, we aim to take a deeper look and understand what contributes to its final performance.

Our key observation is that it is crucial to build different contextual representations for different pairs of spans and an early fusion of entity type information can further improve performance. To validate this, we experiment the following variants on both ACE05 and SciERC:

Text: We use the span representations defined in the entity model (Section 3.2) and concatenate the hidden representations for the subject and the object, as well as their element-wise multiplication: [he(si),he(sj),he(si)⊙he(sj)][\mathbf{h}_{e}(s_{i}),\mathbf{h}_{e}(s_{j}),\mathbf{h}_{e}(s_{i})\odot\mathbf{h}_{e}(s_{j})]. This is similar to the relation model in Luan et al. (2018, 2019).

Markers: We use untyped entity types (⟨\textscS⟩\langle\textsc{S}\rangle, ⟨\textsc/S⟩\langle\textsc{/S}\rangle, ⟨\textscO⟩\langle\textsc{O}\rangle, ⟨\textsc/O⟩\langle\textsc{/O}\rangle) at the input layer and concatenate the representations of two spans’ starting points.

MarkersELoss: We also consider a variant which uses untyped markers but add another FFNN to predict the entity types of subject and object through auxiliary losses. This is similar to how the entity information is used in multi-task learning Luan et al. (2019); Wadden et al. (2019).

TypedMarkers: This is our final model described in Section 3.2 with typed entity markers.

Table 4 summarizes the results of all the variants using either gold entities or predicted entities from the entity model. As is shown, different input representations make a clear difference and the variants of using marker tokens are significantly better than standard text representations and this suggests the importance of learning different representations with respect to different pairs of spans. Compared to Text, TypedMarkers improved the F1 scores dramatically by +5.0%+5.0\% and +7.4%+7.4\% absolute when gold entities are given. With the predicted entities, the improvement is reduced as expected while it remains large enough. Finally, entity type is useful in improving the relation performance and an early fusion of entity information is particularly effective (TypedMarkers vs MarkersEType and MarkersELoss). We also find that MarkersEType to perform even better than MarkersEloss which suggests that using entity types directly as features is better than using them to provide training signals through auxiliary losses.

2 Modeling Entity-Relation Interactions

One main argument for joint models is that modeling the interactions between the two tasks can contribute to each other. In this section, we aim to validate if it is the case in our approach. We first study whether sharing the two representation encoders can improve performance or not. We train the entity and relation models together by jointly optimizing Le+Lr\mathcal{L}_{e}+\mathcal{L}_{r} (Table 5). We find that simply sharing the encoders hurts both the entity and relation F1. We think this is because the two tasks have different input formats and require different features for predicting entity types and relations, thus using separate encoders indeed learns better task-specific features. We also explore whether the relation information can improve the entity performance. To do so, we add an auxiliary loss to our entity model, which concatenates the two span representations as well as their element-wise multiplication (see the Text variant in Section 5.1) and predicts the relation type between the two spans (r∈Rr\in\mathcal{R} or ϵ\epsilon). Through joint training with this auxiliary relation loss, we observe a negligible improvement (<0.1%<0.1\%) on averaged entity F1 over 55 runs on the ACE05 development set. To summarize, (1) entity information is clearly important in predicting relations (Section 5.1). However, we don’t find that relation information to improve our entity model substantiallyMiwa and Bansal (2016) observed a slight improvement on entity F1 by sharing the parameters (80.8 →\rightarrow 81.8 F1) on the ACE05 development data. Wadden et al. (2019) observed that their relation propagation layers improved the entity F1 slightly on SciERC but it hurts performance on ACE05.; (2) simply sharing the encoders does not provide benefits to our approach.

3 Mitigating Error Propagation

A well-known drawback of pipeline training is the error propagation issue. In our final model, we use gold entities (and their types) to train the relation model and the predicted entities during inference and this may lead to a discrepancy between training and testing. In the following, we describe several attempts we made to address this issue.

We first study whether using predicted entities — instead of gold entities — during training can mitigate this issue. We adopt a 10-way jackknifing method, which is a standard technique in many NLP tasks such as dependency parsing Agić and Schluter (2017). Specifically, we divide the data into 1010 folds and predict the entities in the kk-th fold using an entity model trained on the remainder. As shown in Table 6, we find that jackknifing strategy hurts the final relation performance surprisingly. We hypothesize that it is because it introduced additional noise during training.

Second, we consider using more pairs of spans for the relation model at both training and testing time. The main reason is that in the current pipeline approach, if a gold entity is missed out by the entity model during inference, the relation model will not be able to predict any relations associated with that entity. Following the beam search strategy used in the previous work Luan et al. (2019); Wadden et al. (2019), we consider using λn\lambda n (λ=0.4\lambda=0.4 and nn is the sentence length)This pruning strategy achieves a recall of 96.7%96.7\% of gold relations on the development set of ACE05. top spans scored by the entity model. We explored several different strategies for encoding the top-scoring spans for the relation model: (1) typed markers: the same as our main model except that we now have markers e.g., ⟨\textscS:ϵ⟩\langle\textsc{S:}\epsilon\rangle, ⟨\textsc/S:ϵ⟩\langle\textsc{/S:}\epsilon\rangle as input tokens; (2) untyped markers: in this case, the relation model is unaware of a span is an entity or not; (3) untyped markers trained with an auxiliary entity loss (e∈Ee\in\mathcal{E} or ϵ\epsilon). As Table 6 shows, none of these changes led to significant improvements and using untyped markers is especially worse because the relation model struggles to identify whether a span is an entity or not.

In sum, we do not find any of these attempts improved performance significantly and our simple pipelined training turns out to be a surprisingly effective strategy. We do not argue that this error propagation issue does not exist or cannot be solved, while we will need to explore better solutions to address this issue.

4 Effect of Cross-sentence Context

In Table 1, we demonstrated the improvements from using cross-sentence context on both the entity and relation performance. We explore the effect of different context sizes WW in Figure 2. We find that using cross-sentence context clearly improves both entity and relation F1. However, we find the relation performance doesn not further increase from W=100W=100 to W=300W=300. In our final models, we use W=300W=300 for the entity model and W=100W=100 for the relation model.

Conclusion

In this paper, we present a simple and effective approach for end-to-end relation extraction. Our model learns two encoders for entity recognition and relation extraction independently and our experiments show that it outperforms previous state-of-the-art on three standard benchmarks considerably. We conduct extensive analyses to undertand the superior performance of our approach and validate the importance of learning distinct contextual representations for entities and relations and using entity information as input features for the relation model. We also propose an efficient approximation, obtaining a large speedup at inference time with a small reduction in accuracy. We hope that this simple model will serve as a very strong baseline and make us rethink the value of joint training in end-to-end relation extraction.

Acknowledgements

We thank Yi Luan for the help with the datasets and evaluation. We thank Howard Chen, Ameet Deshpande, Dan Friedman, Karthik Narasimhan, and the anonymous reviewers for their helpful comments and feedback. This work is supported in part by a Graduate Fellowship at Princeton University.

References

Appendix A Datasets

We use ACE04, ACE05, and SciERC datasets in our experiments. Table 2 shows the data statistics of each dataset.

The ACE04 and ACE05 datasets are collected from a variety of domains, such as newswire and online forums. We follow Luan et al. (2019)’s preprocessing stepsWe use the script provided by Luan et al. (2019): https://github.com/luanyi/DyGIE/tree/master/preprocessing. and split ACE04 into 5 folds and ACE05 into train, development, and test sets.

The SciERC dataset is collected from 12 AI conference/workshop proceedings in four AI communities Luan et al. (2018). SciERC includes annotations for scientific entities, their relations, and coreference clusters. We ignore the coreference annotations in our experiments. We use the processed dataset which is downloaded from the project websitehttp://nlp.cs.washington.edu/sciIE/ of Luan et al. (2018).

Appendix B Implementation Details

We implement our models based on HuggingFace’s Transformers library Wolf et al. (2019). For the entity model, we follow Wadden et al. (2019) and set the width embedding size as dF=150d_{F}=150 and use a 2-layer FFNN with 150150 hidden units and ReLU activations to predict the probability distribution of entity types:

For the relation model, we use a linear classifier on top of the span pair representation to predict the probability distribution of relation types:

For our approximation model (Section 4.3), we batch candidate pairs by adding 44 markers for each pair to the end of the sentence, until the total number of tokens exceeds 250250. We train our models with Adam optimizer of a linear scheduler with a warmup ratio of 0.10.1. For all the experiments, we train the entity model for 100100 epochs, and a learning rate of 1e-5 for weights in pre-trained LMs, 5e-4 for others and a batch size of 16. We train the relation model for 1010 epochs with a learning rate of 2e-5 and a batch size of 32.

Appendix C Performance with Varying Data Sizes

We compare our pipeline model to a joint model with 10%, 25%, 50%, 100% of training data on the ACE05 dataset. Here, our goal is to understand whether our finding still holds when the training data is smaller (and hence it is expected to have more errors in entity predictions).

Our baseline of joint model is our reimplementation of DYGIE++ Wadden et al. (2019), without using propagation layers (the encoders are shared for the entity and relation model and no input marker is used; the top scoring 0.4n0.4n entities are considered in beam pruning). As shown in Table 7, we find that our model achieves even larger gains in relation F1 over the joint model, when the number of training examples is reduced. This further highlights the importance of explicitly encoding entity boundaries and type features in data-scarce scenarios.