Two are Better than One: Joint Entity and Relation Extraction with Table-Sequence Encoders

Jue Wang, Wei Lu

Introduction

Named Entity Recognition (NER, Florian et al. 2006, 2010) and Relation Extraction (RE, Zhao and Grishman 2005; Jiang and Zhai 2007; Sun et al. 2011; Plank and Moschitti 2013) are two fundamental tasks in Information Extraction (IE). Both tasks aim to extract structured information from unstructured texts. One typical approach is to first identify entity mentions, and next perform classification between every two mentions to extract relations, forming a pipeline Zelenko et al. (2002); Chan and Roth (2011). An alternative and more recent approach is to perform these two tasks jointly Li and Ji (2014); Miwa and Sasaki (2014); Miwa and Bansal (2016), which mitigates the error propagation issue associated with the pipeline approach and leverages the interaction between tasks, resulting in improved performance.

Among several joint approaches, one popular idea is to cast NER and RE as a table filling problem Miwa and Sasaki (2014); Gupta et al. (2016); Zhang et al. (2017). Typically, a two-dimensional (2D) table is formed where each entry captures the interaction between two individual words within a sentence. NER is then regarded as a sequence labeling problem where tags are assigned along the diagonal entries of the table. RE is regarded as the problem of labeling other entries within the table. Such an approach allows NER and RE to be performed using a single model, enabling the potentially useful interaction between these two tasks. One example The exact settings for table filling may be different for different papers. Here we fill the entire table (rather than the lower half of the table), and assign relation tags to cells involving two complete entity spans (rather than part of such spans). We also preserve the direction of the relations. is illustrated in Figure 1.

Unfortunately, there are limitations with the existing joint methods. First, these methods typically suffer from feature confusion as they use a single representation for the two tasks – NER and RE. As a result, features extracted for one task may coincide or conflict with those for the other, thus confusing the learning model. Second, these methods underutilize the table structure as they usually convert it to a sequence and then use a sequence labeling approach to fill the table. However, crucial structural information (e.g., the 4 entries at the bottom-left corner of Figure 1 share the same label) in the 2D table might be lost during such conversions.

In this paper, we present a novel approach to address the above limitations. Instead of predicting entities and relations with a single representation, we focus on learning two types of representations, namely sequence representations and table representations, for NER and RE respectively. On one hand, the two separate representations can be used to capture task-specific information. On the other hand, we design a mechanism to allow them to interact with each other, in order to take advantage of the inherent association underlying the NER and RE tasks. In addition, we employ neural network architectures that can better capture the structural information within the 2D table representation. As we will see, such structural information (in particular the context of neighboring entries in the table) is essential in achieving better performance.

The recent prevalence of BERT Devlin et al. (2019) has led to great performance gains on various NLP tasks. However, we believe that the previous use of BERT, i.e., employing the contextualized word embeddings, does not fully exploit its potential. One important observation here is that the pairwise self-attention weights maintained by BERT carry knowledge of word-word interactions. Our model can effectively use such knowledge, which helps to better learn table representations. To the best of our knowledge, this is the first work to use the attention weights of BERT for learning table representations.

We summarize our contributions as follows:

We propose to learn two separate encoders – a table encoder and a sequence encoder. They interact with each other, and can capture task-specific information for the NER and RE tasks;

We propose to use multidimensional recurrent neural networks to better exploit the structural information of the table representation;

We effectively leverage the word-word interaction information carried in the attention weights from BERT, which further improves the performance.

Our proposed method achieves the state-of-the-art performance on four datasets, namely ACE04, ACE05, CoNLL04, and ADE. We also conduct further experiments to confirm the effectiveness of our proposed approach.

Related Work

NER and RE can be tackled by using separate models. By assuming gold entity mentions are given as inputs, RE can be regarded as a classification task. Such models include kernel methods Zelenko et al. (2002), RNNs Zhang and Wang (2015), recursive neural networks Socher et al. (2012), CNNs Zeng et al. (2014), and Transformer models Verga et al. (2018); Wang et al. (2019). Another branch is to detect cross-sentence level relations Peng et al. (2017); Gupta et al. (2019), and even document-level relations Yao et al. (2019); Nan et al. (2020). However, entities are usually not directly available in practice, so these approaches may require an additional entity recognizer to form a pipeline.

Joint learning has been shown effective since it can alleviate the error propagation issue and benefit from exploiting the interrelation between NER and RE. Many studies address the joint problem through a cascade approach, i.e., performing NER first followed by RE. Miwa and Bansal (2016) use bi-LSTM Graves et al. (2013) and tree-LSTM Tai et al. (2015) for the joint task. Bekoulis et al. (2018a, b) formulate it as a head selection problem. Nguyen and Verspoor (2019) apply biaffine attention Dozat and Manning (2017) for RE. Luan et al. (2019), Dixit and Al (2019), and Wadden et al. (2019) use span representations to predict relations.

Miwa and Sasaki (2014) tackle joint NER and RE as from a table filling perspective, where the entry at row ii and column jj of the table corresponds to the pair of ii-th and jj-th word of the input sentence. The diagonal of the table is filled with the entity tags and the rest with the relation tags indicating possible relations between word pairs. Similarly, Gupta et al. (2016) employ a bi-RNN structure to label each word pair. Zhang et al. (2017) propose a global optimization method to fill the table. Tran and Kavuluru (2019) investigate CNNs on this task.

Recent work Luan et al. (2019); Dixit and Al (2019); Wadden et al. (2019); Li et al. (2019); Eberts and Ulges (2019) usually leverages pre-trained language models such as ELMo Peters et al. (2018), BERT Devlin et al. (2019), RoBERTa Liu et al. (2019), and ALBERT Lan et al. (2019). However, none of them use pre-trained attention weights, which convey rich relational information between words. We believe it can be useful for learning better table representations for RE.

Problem Formulation

In this section, we formally formulate the NER and RE tasks. We regard NER as a sequence labeling problem, where the gold entity tags yNER\boldsymbol{y}^{\text{NER}} are in the standard BIO (Begin, Inside, Outside) scheme Sang and Veenstra (1999); Ratinov and Roth (2009). For the RE task, we mainly follow the work of Miwa and Sasaki (2014) to formulate it as a table filling problem. Formally, given an input sentence x=[xi]1≤i≤N\boldsymbol{x}=[x_{i}]_{1\leq i\leq N}, we maintain a tag table yRE=[yi,jRE]1≤i,j≤N\boldsymbol{y}^{\text{RE}}=[y^{\text{RE}}_{i,j}]_{1\leq i,j\leq N}. Suppose there is a relation with type rr pointing from mention xib,..,xiex_{i^{b}},..,x_{i^{e}} to mention xjb,..,xjex_{j^{b}},..,x_{j^{e}}, we have yi,jRE=r→y^{\text{RE}}_{i,j}=\overrightarrow{r} and yj,iRE=r←y^{\text{RE}}_{j,i}=\overleftarrow{r} for all i∈[ib,ie]∧j∈[jb,je]i\in[i^{b},i^{e}]\wedge j\in[j^{b},j^{e}]. We use ⊥\bot for word pairs with no relation. An example was given earlier in Figure 1.

Model

We describe the model in this section. The model consists of two types of interconnected encoders, a table encoder for table representation and a sequence encoder for sequence representation, as shown in Figure 2. Collectively, we call them table-sequence encoders. Figure 3 presents the details of each layer of the two encoders, and how they interact with each other. In each layer, the table encoder uses the sequence representation to construct the table representation; and then the sequence encoder uses the table representation to contextualize the sequence representation. With multiple layers, we incrementally improve the quality of both representations.

where each word is represented as an HH dimensional vector.

2 Table Encoder

The table encoder, shown in the left part of Figure 3, is a neural network used to learn a table representation, an N×NN\times N table of vectors, where the vector at row ii and column jj corresponds to the ii-th and jj-th word of the input sentence.

Next, we use the Multi-Dimensional Recurrent Neural Networks (MD-RNN, Graves et al. 2007) with Gated Recurrent Unit (GRU, Cho et al. 2014) to contextualize Xl\boldsymbol{X}_{l}. We iteratively compute the hidden states of each cell to form the contextualized table representation Tl\boldsymbol{T}_{l}, where:

We provide the multi-dimensional adaptations of GRU in Appendix A to avoid excessive formulas here.

Generally, it exploits the context along layer, row, and column dimensions. That is, it does not consider only the cells at neighbouring rows and columns, but also those of the previous layer.

The time complexity of the naive implementation (i.e., two for-loops) for each layer is O(N×N)O(N\times N) for a sentence with length NN. However, antidiagonal entries We define antidiagonal entries to be entries at position (i,j)(i,j) such that i+j=N+1+Δi+j=N+1+\Delta, where Δ∈[−N+1,N−1]\Delta\in[-N+1,N-1] is the offset to the main antidiagonal entries. can be calculated at the same time as they do not depend on each other. Therefore, we can optimize it through parallelization and reduce the effective time complexity to O(N)O(N).

The above illustration describes a unidirectional RNN, corresponding to Figure 4(a). Intuitively, we would prefer the network to have access to the surrounding context in all directions. However, this could not be done by one single RNN. For the case of 1D sequence modeling, this problem is resolved by introducing bidirectional RNNs. Graves et al. (2007) discussed quaddirectional RNNs to access the context from four directions for modeling 2D data. Therefore, similar to 2D-RNN, we also need to consider RNNs in four directionsIn our scenario, there is an additional layer dimension. However, as the model always traverses from the first layer to the last layer, only one direction shall be considered for the layer dimension.. We visualize them in Figure 4.

Empirically, we found the setting only considering cases (a) and (c) in Figure 4 achieves no worse performance than considering four cases altogether. Therefore, to reduce the amount of computation, we use such a setting as default. The final table representation is then the concatenation of the hidden states of the two RNNs:

3 Sequence Encoder

The sequence encoder is used to learn the sequence representation – a sequence of vectors, where the ii-th vector corresponds to the ii-th word of the input sentence. The architecture is similar to Transformer Vaswani et al. (2017), shown in the right portion of Figure 3. However, we replace the scaled dot-product attention with our proposed table-guided attention. Here, we mainly illustrate why and how the table representation can be used to compute attention weights.

First of all, given Q\boldsymbol{Q} (queries), K\boldsymbol{K} (keys) and V\boldsymbol{V} (values), a generalized form of attention is defined in Figure 5. For each query, the output is a weighted sum of the values, where the weight assigned to each value is determined by the relevance (given by score function ff) of the query with all the keys.

For each query QiQ_{i} and key KjK_{j}, Bahdanau et al. (2015) define ff in the form of:

where UU is a learnable vector and gg is the function to map each query-key pair to a vector. Specifically, they define g(Qi,Kj)=tanh⁡(QiW0+KjW1)g(Q_{i},K_{j})=\tanh(Q_{i}W_{0}+K_{j}W_{1}), where W0,W1W_{0},W_{1} are learnable parameters.

Our attention mechanism is essentially a self-attention mechanism, where the queries, keys and values are exactly the same. In our case, they are essentially sequence representation Sl−1S_{l-1} of the previous layer (i.e., Q=K=V=Sl−1\boldsymbol{Q}=\boldsymbol{K}=\boldsymbol{V}=\boldsymbol{S}_{l-1}). The attention weights (i.e., the output from the function ff in Figure 5) are essentially constructed from both queries and keys (which are the same in our case). On the other hand, we also notice the table representation Tl\boldsymbol{T}_{l} is also constructed from Sl−1\boldsymbol{S}_{l-1}. So we can consider Tl\boldsymbol{T}_{l} to be a function of queries and keys, such that Tl,i,j=g(Sl−1,i,Sl−1,j)=g(Qi,Kj)T_{l,i,j}=g(S_{l-1,i},S_{l-1,j})=g(Q_{i},K_{j}). Then we put back this gg function to Equation 7, and get the proposed table-guided attention, whose score function is:

We show the advantages of using this table-guided attention: (1) we do not have to calculate gg function since Tl\boldsymbol{T}_{l} is already obtained from the table encoder; (2) Tl\boldsymbol{T}_{l} is contextualized along the row, column, and layer dimensions, which corresponds to queries, keys, and queries and keys in the previous layer, respectively. Such contextual information allows the network to better capture more difficult word-word dependencies; (3) it allows the table encoder to participate in the sequence representation learning process, thereby forming the bidirectional interaction between the two encoders.

The table-guided attention can be extended to have multiple heads Vaswani et al. (2017), where each head is an attention with independent parameters. We concatenate their outputs and use a fully-connected layer to get the final attention outputs.

The remaining parts are similar to Transformer. For layer ll, we use position-wise feedforward neural networks (FFNN) after self-attention, and wrap attention and FFNN with a residual connection He et al. (2016) and layer normalization (Ba et al. 2016), to get the output sequence representation:

4 Exploit Pre-trained Attention Weights

In this section, we describe the dashed lines in Figures 2 and 3, which we ignored in the previous discussions. Essentially, they exploit information in the form of attention weights from a pre-trained language model such as BERT.

We keep the rest unchanged. We believe this simple yet novel use of the attention weights allows us to effectively incorporate the useful word-word interaction information captured by pre-trained models such as BERT into our table-sequence encoders for improved performance.

Training and Evaluation

We use SL\boldsymbol{S}_{L} and TL\boldsymbol{T}_{L} to predict the probability distribution of the entity and relation tags:

where YNER{\boldsymbol{Y}}^{\text{NER}} and YRE{\boldsymbol{Y}}^{\text{RE}} are random variables of the predicted tags, and Pθ{P}_{\theta} is the estimated probability function with θ\theta being our model parameters.

For training, both NER and RE adopt the prevalent cross-entropy loss. Given the input text x\boldsymbol{x} and its gold tag sequence yNER\boldsymbol{y}^{\text{NER}} and tag table yRE\boldsymbol{y}^{\text{RE}}, we then calculate the following two losses:

The goal is to minimize both losses LNER+LRE\mathcal{L}^{\text{NER}}+\mathcal{L}^{\text{RE}}.

During evaluation, the prediction of relations relies on the prediction of entities, so we first predict the entities, and then look up the relation probability table Pθ(YRE){P}_{\theta}({\boldsymbol{Y}}^{\text{RE}}) to see if there exists a valid relation between predicted entities.

Specifically, we predict the entity tag of each word by choosing the class with the highest probability:

The whole tag sequence can be transformed into entities with their boundaries and types.

Relations on entities are mapped to relation classes with highest probabilities on words of the entities. We also consider the two directed tags for each relation. Therefore, for two entity spans (ib,ie)(i^{b},i^{e}) and (jb,je)(j^{b},j^{e}), their relation is given by:

where the no-relation type ⊥\bot has no direction, so if r→=⊥\overrightarrow{r}=\bot, we have r←=⊥\overleftarrow{r}=\bot as well.

Experiments

We evaluate our model on four datasets, namely ACE04 Doddington et al. (2004), ACE05 Walker et al. (2006), CoNLL04 Roth and tau Yih (2004) and ADE Gurulingappa et al. (2012). More details could be found in Appendix B.

Following the established line of work, we use the F1 measure to evaluate the performance of NER and RE. For NER, an entity prediction is correct if and only if its type and boundaries both match with those of a gold entity. Follow Li and Ji (2014); Miwa and Bansal (2016), we use head spans for entities in ACE. And we keep the full mention boundary for other corpora. For RE, a relation prediction is considered correct if its relation type and the boundaries of the two entities match with those in the gold data. We also report the strict relation F1 (denoted RE+), where a relation prediction is considered correct if its relation type as well as the boundaries and types of the two entities all match with those in the gold data. Relations are asymmetric, so the order of the two entities in a relation matters.

2 Model Setup

We tune hyperparameters based on results on the development set of ACE05 and use the same setting for other datasets. GloVe vectors Pennington et al. (2014) are used to initialize word embeddings. We also use the BERT variant – ALBERT as the default pre-trained language model. Both pre-trained word embeddings and language model are fixed without fine-tuning. In addition, we stack three encoding layers (L=3L=3) with independent parameters including the GRU cell in each layer. For the table encoder, we use two separate MD-RNNs with the directions of “layer+row+col+” and “layer+row-col-” respectively. For the sequence encoder, we use eight attention heads to attend to different representation subspaces. We report the averaged F1 scores of 5 runs for our models. For each run, we keep the model that achieves the highest averaged entity F1 and relation F1 on the development set, and evaluate and report its score on the test set. Other hyperparameters could be found in Appendix C.

3 Comparison with Other Models

Table 1 presents the comparison of our model with previous methods on four datasets. Our NER performance is increased by 1.2, 0.9, 1.2/0.6 and 0.4 absolute F1 points over the previous best results. Besides, we observe even stronger performance gains in the RE task, which are 3.6, 4.2, 2.1/2.5 (RE+) and 0.9 (RE+) absolute F1 points, respectively. This indicates the effectiveness of our model for jointly extracting entities and their relations. Since our reported numbers are the average of 5 runs, we can consider our model to be achieving new state-of-the-art results.

4 Comparison of Pre-trained Models

In this section, we evaluate our method with different pre-trained language models, including ELMo, BERT, RoBERTa and ALBERT, with and without attention weights, to see their individual contribution to the final performance.

The overall results reported in Table 2 confirm the importance of leveraging the attention weights, which bring improvements for both NER and RE tasks. This allows the system using vanilla BERT to obtain results no worse than RoBERTa and ALBERT in relation extraction.

5 Ablation Study

We design several additional experiments to understand the effectiveness of components in our system. The experiments are conducted on ACE05.

We also compare different table filling settings, which are included in Appendix E.

We first focus on the understanding of the necessity of modeling the bidirectional interaction between the two encoders. Results are presented in Table 3. “RE (gold)” is presented so as to compare with settings that do not predict entities, where the gold entity spans are used in the evaluation.

We first try optimizing the NER and RE objectives separately, corresponding to “w/o Relation Loss” and “w/o Entity Loss”. Compared with learning with a joint objective, the results of these two settings are slightly worse, which indicates that learning better representations for one task not only is helpful for the corresponding task, but also can be beneficial for the other task.

Next, we investigate the individual sequence and table encoder, corresponding to “w/o Table Encoder” and “w/o Sequence Encoder”. We also try jointly training the two encoders but cut off the interaction between them, which is “w/o Bi-Interaction”. Since no interaction is allowed in the above three settings, the table-guided attention is changed to conventional multi-head scaled dot-product attention, and the table encoding layer always uses the initial sequence representation S0\boldsymbol{S}_{0} to enrich the table representation. The results of these settings are all significantly worse than the default one, which indicates the importance of the bidirectional interaction between sequence and table representation in our table-sequence encoders.

We also experiment the use of the main diagonal entries of the table representation to tag entities, with results reported under “NER on diagonal”. This setup attempts to address NER and RE in the same encoding space, in line with the original intention of Miwa and Sasaki (2014). By exploiting the interrelation between NER and RE, it achieves better performance compared with models without such information. However, it is worse than our default setting. We ascribe this to the potential incompatibility of the desired encoding space of entities and relations. Finally, although it does not directly use the sequence representation, removing the sequence encoder will lead to performance drop for NER, which indicates the sequence encoder can help improve the table encoder by better capturing the structured information within the sequence.

5.2 Encoding Layers

Table 4 shows the effect of the number of encoding layers, which is also the number of bidirectional interactions involved. We conduct one set of experiments with shared parameters for the encoding layers and another set with independent parameters. In general, the performance increases when we gradually enlarge the number of layers LL. Specifically, since the shared model does not introduce more parameters when tuning LL, we consider that our model benefits from the mutual interaction inside table-sequence encoders. Typically, under the same value LL, the non-shared model employs more parameters than the shared one to enhance its modeling capability, leading to better performance. However, when L>3L>3, there is no significant improvement by using non-shared model. We believe that increasing the number of layers may bring the risk of over-fitting, which limits the performance of the network. We choose to adopt the non-shared model with L=3L=3 as our default setting.

5.3 Settings of MD-RNN

Table 5 presents the comparisons of using different dimensions and directions to learn the table representation, based on MD-RNN. Among those settings, “Unidirectional” refers to an MD-RNN with direction “layer+row+col+”; “Bidirectional” uses two MD-RNNs with directions “layer+row+col+” and “layer+row-col-” respectively; “Quaddirectional” uses MD-RNNs in four directions, illustrated in Figure 4. Their results are improved when adding more directions, showing richer contextual information is beneficial. Since the bidirectional model is almost as good as the quaddirectional one, we leave the former as the default setting.

In addition, we are also curious about the contribution of layer, row, and column dimensions for MD-RNNs. We separately removed the layer, row, and column dimension. As we can see, the results are all lower than the original model without removal of any dimension. “Layer-wise only” removed row and col dimensions, and is worse than others as it does not exploit the sentential context.

More experiments with more settings are presented in Appendix D. Specifically, all unidirectional RNNs are consistently worse than others, while bidirectional RNNs are usually on-par with quaddirectional RNNs. Besides, we also tried to use CNNs to implement the table encoder. However, since it is usually difficult for CNNs to learn long-range dependencies, we found the performance was worse than the RNN-based models.

6 Attention Visualization

We visualize the table-guided attention with bertviz Vig (2019)https://github.com/jessevig/bertviz for a better understanding of how the network works. We compare it with pre-trained Transformers (ALBERT) and human-defined ground truth, as presented in Figure 6.

Our discovery is similar to Clark et al. (2019). Most attention heads in the table-guided attention and ALBERT show simple patterns. As shown in the left part of Figure 6, these patterns include attending to the word itself, the next word, the last word, and the punctuation.

The right part of Figure 6 also shows task-related patterns, i.e., entities and relations. For a relation, we connect words from the head entity to the tail entity; For an entity, we connect every two words inside this entity mention. We can find that our proposed table-guided attention has learned more task-related knowledge compared to ALBERT. In fact, not only does it capture the entities and their relations that ALBERT failed to capture, but it also has higher confidence. This indicates that our model has a stronger ability to capture complex patterns other than simple ones.

7 Probing Intermediate States

Figure 7 presents an example picked from the development set of ACE05. The prediction layer after training (a linear layer) is used as a probe to display the intermediate state of the model, so we can interpret how the model improves both representations from stacking multiple layers and thus from the bidirectional interaction. Such probing is valid since we use skip connection between two adjacent encoding layers, so the encoding spaces of the outputs of different encoding layers are consistent and therefore compatible with the prediction layer.

In Figure 7, the model made many wrong predictions in the first layer, which were gradually corrected in the next layers. Therefore, we can see that more layers allow more interaction and thus make the model better at capturing entities or relations, especially difficult ones. More cases are presented in Appendix F.

Conclusion

In this paper, we introduce the novel table-sequence encoders architecture for joint extraction of entities and their relations. It learns two separate encoders rather than one – a sequence encoder and a table encoder where explicit interactions exist between the two encoders. We also introduce a new method to effectively employ useful information captured by the pre-trained language models for such a joint learning task where a table representation is involved. We achieved state-of-the-art F1 scores for both NER and RE tasks across four standard datasets, which confirm the effectiveness of our approach. In the future, we would like to investigate how the table representation may be applied to other tasks. Another direction is to generalize the way in which the table and sequence interact to other types of representations.

Acknowledgements

We would like to thank the anonymous reviewers for their helpful comments and Lidan Shou for his suggestions and support on this work. This work was done during the first author’s remote internship with the StatNLP Group in Singapore University of Technology and Design. This research is supported by Ministry of Education, Singapore, under its Academic Research Fund (AcRF) Tier 2 Programme (MOE AcRF Tier 2 Award No: MOE2017-T2-1-156). Any opinions, findings and conclusions or recommendations expressed in this material are those of the authors and do not reflect the views of the Ministry of Education, Singapore.

References

Appendix A MD-RNN

In this section we present the detailed implementation of MD-RNN with GRU.

Formally, at the time-step layer ll, row ii, and column jj, with the input Xl,i,jX_{l,i,j}, the cell at layer ll, row ii and column jj calculates the gates as follows:

where WW and bb are trainable parameters and please note that they share parameters in different rows and columns but not necessarily in different layers. Besides, ⊙\odot is the element-wise product, and σ\sigma is the sigmoid function.

As in GRU, rr is the reset gate controlling whether to forget previous hidden states, and zz is the update gate, selecting whether the hidden states are to be updated with new hidden states. In addition, we employ a lambda gate λ\lambda, which is used to weight the predecessor cells before passing them through the update gate.

And we found in our preliminary experiments that both of them performed as well as each other, and we choose the former, which saves some computation.

The time complexity of the naive implementation (i.e., two for-loops in each layer) is O(L×N×N)O(L\times N\times N) for a sentence with length NN and the number of encoding layer LL. However, antidiagonal entries can be calculated at the same time because their values do not depend on each other, shown in the same color in Figure 8. Therefore, we can optimize it through parallelization and reduce the effective time complexity to O(L×N)O(L\times N).

Appendix B Data

Table 6 shows the dataset statistics after pre-processing. We keep the same pre-processing and evaluation standards used by most previous works.

The ACE04 and ACE05 corpora are collected from a variety of domains, such as newswire and online forums. We use the same entity and relation types, data splits, and pre-processing as Li and Ji (2014) and Miwa and Bansal (2016) We use the prepocess script provided by Luan et al. (2019): https://github.com/luanyi/DyGIE/tree/master/preprocessing. Specifically, they use head spans for entities but not use the full mention boundary.

The CoNLL04 dataset provides entity and relation labels. We use the same train-test split as Gupta et al. (2016) https://github.com/pgcool/TF-MTRNN/tree/master/data/CoNLL04, and we use the same 20% train set as development set as Eberts and Ulges (2019) http://lavis.cs.hs-rm.de/storage/spert/public/datasets/conll04/. Both micro and macro average F1 are used in previous work, so we will specify this while comparing with other systems.

The ADE dataset is constructed from medical reports that describe the adverse effects arising from drug use. It contains a single relation type “Adverse-Effect” and the two entity types “Adverse-Effect” and “Drug”. Similar to previous work, we filter out instances containing overlapping entities, only accounting for 2.8% of total.

Following prior work, we perform 5-fold cross-validation for ACE04 and 10-fold for ADE. Besides, we use 15% of the training set as the development set. We report the average score of 5 runs for every dataset. For each run, we use the model that achieves the best performance (averaged entity metric score and relation metric score) on the development set, and evaluate and report its score on the test set.

Appendix C Hyperparameters and Pre-trained Language Models

The detailed hyperparameters are present in Table 7. For the word embeddings, we use 100-dimensional GloVe word embeddings trained on 6B tokenshttps://nlp.stanford.edu/projects/glove/ as initialization. We disable updating the word embeddings during training. We set the hidden size to 200, and since we use bidirectional MD-RNNs, the hidden size for each MD-RNN is 100. We use inverse time learning rate decay: lr^=lr/(1+decay_rate×steps/decay_steps)\hat{lr}={lr}/(1+\text{decay\_rate}\times\text{steps}/\text{decay\_steps}), with decay rate 0.05 and decay steps 1000.

Besides, the tested pre-trained language models are shown as follows:

[ELMo] Peters et al. (2018): Character-based pre-trained language model. We use the large checkpoint, with embeddings of dimension 3072.

[BERT] Devlin et al. (2019): Pre-trained Transformer. We use the bert-large-uncased checkpoint, with embeddings of dimension 1024 and attention weight feature of dimension 384 (24 layers ×\times 16 heads).

[RoBERTa] Liu et al. (2019): Pre-trained Transformer. We use the roberta-large checkpoint, with embeddings of dimension 1024 and attention weight feature of dimension 384 (24 layers ×\times 16 heads).

[ALBERT] Lan et al. (2019): A lite version of BERT with shared layer parameters. We use the albert-xxlarge-v1 checkpoint, with embeddings of dimension 4096 and attention weight feature of dimension 768 (12 layers ×\times 64 heads). We by default use this pre-trained model.

We use the implementation provided by Wolf et al. (2019)https://github.com/huggingface/Transformers and Akbik et al. (2019)https://github.com/flairNLP/flair to generate contextualized embeddings and attention weights. Specifically, we generate the contextualized word embedding by averaging all sub-word embeddings in the last four layers; we generate the attention weight feature (if available) by summing all sub-word attention weights for each word, which are then concatenated for all layers and all heads. Both of them are fixed without fine-tuning.

Appendix D Ways to Leverage the Table Context

Table 8 presents the comparisons of different ways to learn the table representation.

Importance of context Setting “layer+row col” does not exploit the table context when learning the table representation, instead, only layer-wise operations are used. As a result, it performs much worse than the ones exploiting the context, confirming the importance to leverage the context information.

Context along row and column Neighbors along both the row and column dimensions are important. setting “layer+row+col ; layer+row-col” and “layer+row col+; layer+row col-” remove the row and column dimensions respectively, and their performance is though better than “layer+row col”, but worse than setting “layer+row+col+; layer+row-col-”.

Multiple dimensions Since in setting “layer+row+col+”, the cell at row ii and column jj only knows the information before the ii-th and jj-th word, causing worse performance than bidirectional (“layer+row+col+; layer+row-col-” and “layer+row+col-; layer+row-col+”) and quaddirectional (“layer+row+col+; layer+row-col-; layer+row+col-; layer+row-col+”) settings. Besides, the quaddirectional model does not show superior performance than bidirectional ones, so we use the latter by default.

Layer dimension Different from the row and column dimensions, the layer dimension does not carry more sentential context information. Instead, it carries the information from previous layers, so the model can reason high-level relations based on low-level dependencies captured by predecessor layers, which may help recognize syntactically and semantically complex relations. Moreover, recurring along the layer dimension can also be viewed as a layer-wise short-cut, serving similarly to high way Srivastava et al. (2015) and residual connection He et al. (2016) and making it possible for the networks to be very deep. By removing it (results under “layer row+col+; layer row-col-”), the performance is harmed.

Other network Our model architecture can be adapted to other table encoders. We try CNN to encode the table representation. For each layer ll, given inputs Xl\boldsymbol{X}_{l}, we have:

We also try different kernel sizes for CNN. However, despite its advantages in training time, its performance is worse than the MD-RNN based ones.

Appendix E Table Filling Formulations

Our table filling formulation does not exactly follow Miwa and Sasaki (2014). Specifically, we fill the entire table instead of only the lower (or higger) triangular part, and we assign relation tags to cells where entity spans intersect instead of where last words intersect. To maintain the ratio of positive instances to negative instances, although the entire table can express directed relations by undirected tags, we still keep the directed relation tags. I.e, if yi,jRE=r→y^{\text{RE}}_{i,j}=\overrightarrow{r} then yj,iRE=r←y^{\text{RE}}_{j,i}=\overleftarrow{r}, and vice versa. Table 9 ablates our formulation (last row), and compares it with the original one Miwa and Sasaki (2014) (first row).

Appendix F Probing Intermediate States

Figure 9 presents examples picked from the development set of ACE05. The prediction layer (a linear layer) after training is used as a probe to display the intermediate state of the model, so we can interpret how the model improves both representations from stacking multiple layers and thus from the bidirectional interaction.

Such probing is valid since for the table encoder, the encoding spaces of different cells are consistent as they are connected through gate mechanism, including cells in different encoding layers; for the sequence encoder, we used residual connection so the encoding spaces of the inputs and outputs are consistent. Therefore, they are all compatible with the prediction layer. Empirically, the intermediate layers did give valid predictions, although they are not directly trained for prediction.

In Figure 9(a), the model made a wrong prediction with the representation learned by the first encoding layer. But after the second encoding layer, this mistake has been corrected by the model. This is also the case that happens most frequently, indicating that two encoding layers are already good enough for most situations. For some more complicated cases, the model needs three encoding layers to determine the final decision, shown in Figure 9(b). Nevertheless, more layers do not always push the prediction towards the correct direction, and Figure 9(c) shows a negative example, where the model made a correct prediction in the second encoding layer, but in the end it decided not to output one relation, resulting in a false-negative error. But we must note that such errors rarely occur, and the more common errors are that entities or relationships are not properly captured at all encoding layers.