OneEE: A One-Stage Framework for Fast Overlapping and Nested Event Extraction

Hu Cao, Jingye Li, Fangfang Su, Fei Li, Hao Fei, Shengqiong Wu, Bobo Li, Liang Zhao, Donghong Ji

Introduction

Event Extraction (EE) is a fundamental yet challenging task in information extraction research Miwa and Bansal (2016); Katiyar and Cardie (2016); Fei et al. (2020b); Li et al. (2021b); Fei et al. (2022a). EE facilitates the development of practical applications such as knowledge graph construction Wei et al. (2019c); Bosselut et al. (2021), biological process analysis Miwa et al. (2013), and financial market surveillance Nuij et al. (2013). The goal of EE is to recognize event triggers as well as the associated arguments from texts. As an example, Figure 1(a) illustrates a Share Reduction event including a trigger “reduced” and a subject argument “Wang Yawei”.

Traditional methods for EE Li et al. (2013); Chen et al. (2015); Nguyen et al. (2016); Liu et al. (2018); Nguyen and Nguyen (2019) regard event extraction as a sequence labeling task, assuming that event mentions do not overlap with each other. However, they neglect complicated irregular EE scenarios (i.e., overlapped and nested EE) Fei et al. (2020a, 2021a). As exemplified in Figure 1(b), there are two overlapped events, Investment, and Share Transfer, which share the same trigger word “acquired” and the argument words “Guangzhou Securities”. Figure 1(c) illustrates an example of nested events where the event Gene Expression is the Theme argument of another event Positive Regulation.

Prior studies for overlapped and nested EE Yang et al. (2019); Li et al. (2020) employ pipeline-based methods that extract event triggers and arguments in several successive stages. Recently, the state-of-the-art model Sheng et al. Sheng et al. (2021) also uses such a method that consecutively performs event type detection, trigger extraction, and argument extraction. The main problem with such a method is that the latter stage relies on the former stage, which inherently brings the error propagation problem.

To address the above issue, we present a novel tagging scheme that transforms overlapping and nested EE into word-word relation recognition. As shown in Figure 2, we design two types of relations, including the span relation (S-*) and role relation (R-*). S-* handles trigger and argument identification, denoting whether two words are the head-tail boundary of a trigger (T) or argument (A). R-* addresses argument role classification, indicating whether the argument plays the “*” role in the event.

Based on this scheme, we further propose a one-stage event extraction model, OneEE, which mainly includes three parts. First, it adopts BERT Devlin et al. (2019) as the encoder to get contextualized word representations. Afterward, an adaptive event fusion layer composed of an attention module and two gate fusion modules are used to obtain event-aware contextual representations for each event type. In the prediction layer, we parallelly predict the span and role relations between each pair of words by calculating distance-aware scores. Finally, event triggers, arguments, and their roles can be decoded out using these relation labels in one stage without error propagation.

We evaluate OneEE on 3 overlapped and nested EE datasets (FewFC Zhou et al. (2021), Genia11 Kim et al. (2011), and Genia13 kim2013Genia), and conduct extensive experiments and analyses. Our contributions can be summarized as follows:

∙\bullet We design a new tagging scheme that casts event extraction as a word-word relation recognition task, providing a novel and simple solution for overlapped and nested EE.

∙\bullet We propose OneEE, a one-stage model that effectively extracts word-word relations in parallel for overlapped and nested EE.

∙\bullet We further present an adaptive event fusion layer to obtain event-aware contextual representations and effectively integrate event information.

∙\bullet OneEE outperforms the SoTA model with regard to both the performance and inference speed.

Related Work

Information extraction is one of the key research track in natural language processing Miwa and Bansal (2016); Fei et al. (2021c), among which the event extraction is the most complicated task Chen et al. (2015); Fei et al. (2022c). Traditional EE (i.e., flat or regular EE) Li et al. (2013); Nguyen et al. (2016); Liu et al. (2018); Sha et al. (2018); Nguyen and Nguyen (2019) formulates EE into a sequence labeling task, assigning each token with a label (e.g., BIO tagging scheme). For example, Nguyen et al. Nguyen et al. (2016) uses two bidirectional RNNs to get richer representation which is then utilized to predict event triggers and argument roles jointly. Liu et al. Liu et al. (2018) jointly extracts multiple event triggers and arguments by introducing attention-based GCN to model the dependency graph information Fei et al. (2021b); Li et al. (2021a); Fei et al. (2022b). However, their underlying assumption that event mentions do not overlap with each other is not always valid. Irregular EE (i.e., overlapped and nested EE) has not received much attention, which is more challenging and realistic.

Existing methods for overlapped and nested EE Yang et al. (2019); Li et al. (2020) perform event extraction in a pipeline manner with several steps. To solve the argument overlap, Yang et al. Yang et al. (2019) adopts multiple sets of binary classifiers where each severs for a role to detect the role-specific argument spans but fails in solving trigger overlap. Except for pipeline methods, the latest attempt dealing with overlapped EE is Sheng et al. Sheng et al. (2021) in a joint framework with cascade decoding. They are the first to simultaneously tackle all the overlapping patterns. Sheng et al. Sheng et al. (2021) sequentially performs type detection, trigger extraction, and argument extraction, where the overlapped targets are extracted separately conditioned on the specific former prediction. Nevertheless, most of the multi-stage methods suffer from error propagation.

2 Tagging-based Information Extraction

Tagging scheme in the field of information extraction has been extensively investigated. Traditional sequence labeling approaches tagging each token once (e.g., BIO) is hard to tackle irregular information extraction (e.g., overlapped NER). Several researchers Zheng et al. (2017) extend the BIO label scheme to adapt to more complex scenarios. However, they suffer from the label ambiguity problem due to limited flexibility. Recently, the grid tagging scheme is used in a lot of information extraction tasks, such as opinion mining Wu et al. (2020), relation extraction Wang et al. (2020), and named entity recognition Wang et al. (2021), due to its characteristic of presenting relations between word pairs. For example, TPLinker Wang et al. (2020) realizes one-stage joint relation extraction without a gap between training and inference by tagging token pairs with link labels. Inspired by these works, we design our tagging scheme to address overlapping and nested EE, which predicts relations between trigger or argument words parallelly in one stage.

Also it is noteworthy explicitly that this work inherits the recent success of the idea of word-word relation detection, as in Li et al. Li et al. (2022b). Li et al. Li et al. (2022b) propose to unify all the NER (including the flat, nested and discontinuous mentions) with a word-word modeling based on the grid tagging scheme. This work however differs from Li et al. Li et al. (2022b) in two folds. First, we extend the idea of the word-word tagging from NER to EE successfully, where we re-design two relation types for the nested and overlapped events. Second, from the modeling perspective, we devise an adaptive event fusion layer to fully support the one-stage (end-to-end) complex event detection, which greatly helps avoid error propagation.

Problem Formulation

The goal of event extraction includes extracting event triggers and their arguments. We can formalize overlapping and nested EE as follows: given an input sentence consisting of N tokens or words X={x1,x2,…,xN}X=\{x_{1},x_{2},\dots,x_{N}\} and event type e∈Ee\in\mathcal{E}, the task aims to extract the span relations S\mathcal{S} and the role relations R\mathcal{R} between each token pair (xi,xj)(x_{i},x_{j}), where E\mathcal{E} denotes the event type collection, S\mathcal{S} and R\mathcal{R} are pre-defined tags. These relations can be explained below, and we also give an example as demonstrated in Figure 2 for better understanding.

S\mathcal{S}: the span relation indicates that xix_{i} and xjx_{j} are the starting and ending token of the extracted trigger span S-T or argument span S-A, where 1≤i≤j≤N1\leq i\leq j\leq N.

R\mathcal{R}: the role relation indicates that the argument with xjx_{j} acts the certain role R-* of the event with the trigger containing xix_{i}, where 1≤i,j≤N1\leq i,j\leq N. * indicates the role type.

NONE, indicating that the word pair does not have any relation defined in this paper.

Framework

The architecture of our model is illustrated in Figure 3, which mainly consists of three components. First, the widely-used pre-trained language model, BERT Devlin et al. (2019), is used as the encoder to yield contextualized word representations from the input sentences. Then, an adaptive event fusion layer consisting of an attention module and two gate modules is used to integrate the target event type embedding into contextual representations. Afterward, a prediction layer is employed to jointly extract the span relations and the role relations between word pairs.

2 Adaptive Event Fusion Layer

Since the goal of our framework is to predict the relations between word pairs for the target event type ete_{t}, it is important to generate event-aware representations. Therefore, to fuse the event information and contextual information provided by the encoder, we design an adaptive fusion layer. As shown in Figure 3, it consists of an attention module, modeling the interaction among events and obtaining the global event information, and two gate fusion modules for integrating the global and target event information with contextualized word representations.

Motivated by the self-attention in Transformer Vaswani et al. (2017); Wei et al. (2019b), we first introduce an attention mechanism, of which input consists of queries, keys, and values. The output is computed as a weighted sum of the values, where the weight assigned to each value is the dot product of the query with the corresponding key. The attention mechanism can be formulated as:

where dh\sqrt{d_{h}} is a scaling factor, Q\bm{Q}, K\bm{K} and V\bm{V} are query, key and value tensors, represented by Eq. 4.

Gate Fusion Mechanism

We design a gate fusion mechanism to integrate two kinds of features and filter the unnecessary information. The gate vector g\bm{g} is produced by a fully-connection layer with the sigmoid function, which can adaptively control the flow of the input:

where p\bm{p} and q\bm{q} are input vectors, represented by Eq. 5 and Eq. 6. σ(⋅)\sigma(\cdot) is a sigmoid activation function, ⊙\odot and [;][;] denote element-wise product and concatenation operations, respectively. Wg\bm{W}_{g} and bg\bm{b}_{g} are trainable parameters.

where Eg\bm{E}^{g} is the output of the attention mechanism, Wq\bm{W}_{q}, Wk\bm{W}_{k} and Wv\bm{W}_{v} are learnable parameters.

To encode global event information into word representations, we adopt a gate module to fuse the contextual word representations and global event representations. After that, we employ another gate mechanism to integrate the target event type embedding and the output of the last gate module. the overall process can be formulated as:

3 Joint Prediction Layer

After the adaptive event fusion layer, we obtain the event-aware word representations Vt\bm{V}^{t}, which are used to jointly predict the span and role relations between each pair of words. For each word pair (wi,wj)(w_{i},w_{j}), we calculate a score to measure the possibility of them for the relation s∈Ss\in\mathcal{S} and r∈Rr\in\mathcal{R}.

To integrate relative distance information and word pair representations, we introduce a distance-aware score function. For two vectors pi\bm{p}_{i} and pj\bm{p}_{j} from a sequence of representations, we combine them with corresponding position embeddings from Su et al. Su et al. (2021), and then calculate the score by the dot product of them:

where Ri\bm{R}_{i} and Rj\bm{R}_{j} are position embeddings of pi\bm{p}_{i} and pj\bm{p}_{j}, Rj−i=Ri⊤Rj\bm{R}_{j-i}=\bm{R}_{i}^{\top}\bm{R}_{j}. Thus, we can obtain the span score cijsc^{s}_{ij} and the role score cijrc^{r}_{ij} of the word pair (wi,wj)(w_{i},w_{j}) for target event type tt:

where Ws1\bm{W}_{s1}, Ws2\bm{W}_{s2}, Wr1\bm{W}_{r1} and Wr2\bm{W}_{r2} denote parameters. vit\bm{v}^{t}_{i} and vjt\bm{v}^{t}_{j} are from Eq. 6.

4 Training Details

For the score cij⋆c^{\star}_{ij}, where ⋆\star denotes the relation ss or rr, our training target is to minimize a variant of circle loss Sun et al. (2020) which extends softmax cross-entropy loss to figure out multi-label classification problem. In addition, we introduce a threshold score δ\delta, noting that the scores of the pairs with relation are larger than δ\delta, and the other pairs are less than it. The loss function can be formulated as:

where Ω⋆\Omega^{\star} denotes the pair set of relation ⋆\star, δ\delta is set to zero.

Finally, we enumerate all event types in the selected event type set E′\mathcal{E}^{\prime} and get the total loss:

where S′\mathcal{S}^{\prime} is a subset sampled from S\mathcal{S}, we detail the sampling strategy in the appendix.

5 Inference

During the inference period, our model is able to extract all events by parallelly injecting their event type embeddings to the adaptive event fusion layer. As shown in Figure 4, once all the tags of a certain event type are predicted by our model in one stage, the overall decoding process can be summarized as four steps: First, we get starting and ending indices of the trigger or argument. Second, we obtain the trigger and argument spans.Note that if two pairs with the same span relation clash in the boundaries, the pair with higher score will be selected. Third, we match the trigger and arguments according to the R-* relations. Finally, the event type is assigned to this event structure. Specially, we repeat the above four steps for each event type.

As shown in Table 1, we follow previous work Sheng et al. (2021), adopting FewFC Zhou et al. (2021), a Chinese financial event extraction benchmark for overlapped EE. FewFC annotates 10 event types and 18 argument role classes with about 22% sentences containing overlapped events. We also experiment on two biomedical datasets for nested EE, namely Genia11 Kim et al. (2011) and Genia13 kim2013Genia, with around 18% sentences containing nested events. Genia11 annotates 9 event types and 10 argument role classes while the figures for Genia13 are 13 and 7. We split the train/dev/test as 8.0:1.0:1.0 for both of them.

2 Implementation Details

We employ the Chinese Bert-base model for FewFC and BioBERT Lee et al. (2020) for Genia11 and Genia13. We adopt AdamW Loshchilov and Hutter (2019) optimizer with the learning rate of 2e−52e-5 for BERT and 1e−31e-3 for the other modules. The batch size is 8 and the hidden size dhd_{h} is 768. We train our model with 20 epochs on FewFC and Genia11 and 30 epochs on Genia13. All the hyper-parameters are tuned on the development set. All the event type embeddings are trained from scratch.

3 Evaluation Metrics

For evaluation, we follow the traditional criteria of previous work Chen et al. (2015); Du and Cardie (2020); Sheng et al. (2021). 1) Trigger Identification (TI): A trigger is correctly identified if the predicted trigger span matches with a golden label; 2) Trigger Classification (TC): A trigger is correctly classified if it is correctly identified and assigned to the right type; 3) Argument Identification (AI): An argument is correctly identified if its event type is correctly recognized and the predicted argument span matches with a golden label; 4) Argument Classification (AC): An argument is correctly classified if it is correctly identified and the predicted role matches any of the golden labels. We report Precision (P), Recall (R), and F measure (F1) for each of the four metrics.

4 Baselines

These methods cast the EE task into a sequence labeling task by assigning each token a label. BERT-softmax uses BERT to get feature representations for classifying triggers and arguments. BERT-CRF adds the CRF layer on BERT to capture label dependencies. BERT-CRF-joint extends the BIO tagging scheme to joint labels of type and role as B/I/O-type-role, inspired by joint extraction of entity and relation Zheng et al. (2017). All these methods are incapable to solve the overlapping problem due to label conflicts.

Multi-stage Methods for Overlapped and Nested EE

These methods perform EE in several stages. PLMEE Yang et al. (2019) solves the argument overlap problem by extracting role-specific argument according to the trigger predicted by the trigger extractor in a pipeline manner. CasEE Sheng et al. (2021) sequentially performs type&trigger&argument extractions, where the overlapped targets are separately extracted conditioned on former predictions and all subtasks are jointly learned.

Table 4.5 reports the result of all methods on the overlapped EE dataset, FewFC, while Table 6.1 reports the results of the nested EE datasets, Genia11 and Genia13. We can observe that:

1) Our method significantly outperforms all other methods and achieves the state-of-the-art F1 score on all three datasets.

2) In comparison with sequence labeling methods, our model achieves better recall and F1 scores. Specifically, our model outperforms BERT-CRF-joint by 11.7% and 6.3% in recall and the F1 score of AC on the FewFC dataset and achieves a substantial improvement of 4.4% in F1 score of AC on two Genia datasets averagely. It shows the effectiveness of our model on overlapped and nested EE since the sequence labeling methods can only solve flat EE.

3) In comparison with multi-stage methods, our model also improves the performance on the F1 score considerably. Our model outperforms the state-of-the-art model, CasEE, by 2.1% in the F1 score of TC on three datasets averagely. We consider this is because that the event feature is well learned by our adaptive event fusion module. Especially, our model improves 3.4% on AI and 1.6% on AC over CasEE on an average of three datasets. The results reveal the superiority of our one-stage framework which elegantly realizes overlapped and nested event extraction without error propagation.

To evaluate the effectiveness of our proposed model in recognizing overlapping and nested event mentions, we further report the results on sentences containing at least one overlapping event in FewFC and sentences containing at least one nested event in Genia11, respectively.

Figure 5 illustrates the results of TC and AC on overlapping and nested sentences in testing. It shows that our method outperforms other methods on overlapping and nested sentences. The reasons are mainly two-fold: 1) We solve all the overlapping patterns while BERT-CRF-joint could not handle overlapped and nested EE and PLMEE only solve the argument overlap. 2) Our one-stage model outperforms CasEE because we effectively learn event-aware representations and extract word-word relations in parallel, while CasEE performs in three sequential steps with error propagation.

In this section, we investigate the effect of position embeddings for the prediction layer of OneEE. We divide the arguments in the test set of FewFC into 6 groups according to their distance from corresponding triggers and report the recall scores of the model with and without position embeddings. As shown in Figure 6, the AC recall declines as the distance between trigger and argument in an event go up. This indicates that it is more difficult for the model to detect roles correctly if the distance is longer in an event. Furthermore, the model with position embeddings outperforms another one without position embeddings, revealing that the relative distance information is beneficial for event extraction.

Table 5 lists the stage numbers, parameter numbers, and inference speeds of two baselines and our model. For a fair comparison, all of these models are implemented using PyTorch and tested using the NVIDIA RTX 3090 GPU, where the batch size is set as 1. As seen, PLMEE has 2 times as many parameters as the other two models, due to the utilization of two BERT-based modules for each stage. Moreover, the inference speed of our model is about 3 times faster than that of PLMEE Yang et al. (2019) and 0.3 times faster than that of CasEE Sheng et al. (2021), which verifies the efficiency of our model. Last but not least, when the batch size is set as 8, the inference speed of our model is 9.4 times as fast as that of PLMEE, which also demonstrates the advantage of our model, that is, it supports parallel inference. In one word, our model leverages fewer parameters but achieves better performance and faster inference speed.

In this section, we investigate the effect of the role strategies for AC performance. As shown in Figure 7, we introduce 4 different strategies to predict the role relation between trigger and argument: the role labels only exist in 1) trigger and argument head pairs (TH-AH), 2) trigger word and argument head pairs (TW-AH), 3) trigger head and argument word pairs (TH-AW), and 4) trigger and argument word pairs (TW-AW). The results of our model with 4 strategies are demonstrated in Figure 8. We can learn that TW-AW achieves the best results against all other strategies on both FewFC and Genia11 datasets. It is largely due to that its labels are denser than other strategies.

We further investigate the effect of the event number for EE, and the results are shown in Figure 9. We can observe that BERT-CRF-joint, PLMEE, and CasEE achieve similar performances on single-event sentences, while CasEE outperforms PLMEE and BERT-CRF-joint on the sentences with multiple events. Most importantly, our system achieves the best results against all other baselines for different event numbers, indicating the advances of our proposed method.

In this paper, we propose a novel one-stage framework based on word-word relation recognition to address overlapped and nested EE concurrently. The relations between word pairs are pre-defined as the word-word relations within a trigger or argument and cross a trigger-argument pair. Moreover, we propose an efficient model that consists of an adaptive event fusion layer for integrating the target event representation, and a distance-aware prediction layer for identifying all kinds of relations jointly. Experimental results show that our proposed model achieves new SoTA results on three datasets and faster speed than the SoTA model. Through ablation studies, we find that the adaptive event fusion layer and distance-aware prediction layer are effective in improving the model performance. In future work, we will extend our method to other structured prediction tasks, such as structured sentiment analysis and overlapped entity relation extraction.

This work is supported by the National Natural Science Foundation of China (No. 62176187), the National Key Research and Development Program of China (No. 2017YFC1200500), the Research Foundation of Ministry of Education of China (No. 18JZD015), the Youth Fund for Humanities and Social Science Research of Ministry of Education of China (No. 22YJCZH064), the General Project of Natural Science Foundation of Hubei Province (No.2021CFB385). L Zhao would like to thank the support from Center for Artificial Intelligence (C4AI-USP), the Sao Paulo Research Foundation (FAPESP grant #2019/07665-4), the IBM Corporation, and China Branch of BRICS Institute of Future Networks.

Appendix A Parallel Training with Sampling

We parallelly inject multiple target event type embeddings at the adaptive event fusion layer during training period, which results in huge computation resources. To this end, we use a subset E′\mathcal{E}^{\prime} to replace E\mathcal{E} for each sample, where the number of E′\mathcal{E}^{\prime} is KK. It consists of one positive event type (the event type annotated in the sample) and K−1K-1 negative event types selected randomly from the event types that does not appear in the sample. In other words, we inject KK different event type embeddings into the gate module of Eq. 6 simultaneously. If there is no positive event type in the sample, we will select KK negative event types.

Appendix B Decoding for Nested EE

In the manuscript, we have already shown the decoding process of our model for overlapped EE in Section 4.5. Due to page limitation, we show an example of nested in Figure 10(a). We also demonstrate its decoding process in Figure 10(b), which is the same as the overlapped EE decoding.

Appendix C Analysis of the Event Sampling Number

To further analyze the effect of sampling number KK and the sampling strategy, we also evaluate our model with positive and negative sampling and random sampling and compare them with different sampling numbers. Figure 11 shows the TC F1 change trend as the number of sampling increases. As seen, both two models with 6 event type samplings achieve the best performance, compared with the other sampling numbers. Specifically, our model with one positive sampling and K−1K-1 negative samplings outperforms the model with KK randomly selected samplings when KK is less than 7, which demonstrates that our sampling strategy is helpful for the model training.