Event Extraction by Answering (Almost) Natural Questions

Xinya Du, Claire Cardie

Introduction

Event extraction is a long-studied and challenging task in Information Extraction (IE) Sundheim (1992); Grishman and Sundheim (1996); Riloff (1996). Its goal is to extract structured information — “what is happening” and the persons/objects that are involved — from unstructured text. The task is illustrated via an example in Figure 1, which depicts an ownership transfer event (the event type), triggered by the word “sale" (the event trigger) and accompanied by its extracted arguments — text spans denoting entities that fill a set of (semantic) roles associated with the event type (e.g., buyer, seller and artifact for ownership transfer events).

Recent successful approaches to event extraction have benefited from dense features extracted by neural models Chen et al. (2015); Nguyen et al. (2016); Liu et al. (2018) as well as contextualized lexical representations from pretrained language models Zhang et al. (2019b); Wadden et al. (2019). The approaches, however, exhibit two key weaknesses. First, they rely heavily on entity information for argument extraction. In particular, event argument extraction generally consists of two steps – first identifying entities and their general semantic class with trained models Wadden et al. (2019) or a parser Sha et al. (2018), then assigning argument roles (or no role) to each entity. Although joint models Yang and Mitchell (2016); Nguyen and Nguyen (2019); Zhang et al. (2019a); Lin et al. (2020) have been proposed to mitigate this issue, error propagation Li et al. (2013) still occurs during event argument extraction.

A second weakness of neural approaches to event extraction is their inability to exploit the similarities of related argument roles across event types. For example, the ACE 2005 Doddington et al. (2004) Conflict.Attack events and Justice.Execute events have target and person argument roles, respectively. Both roles, however, refer to a human being (who) is affected by an action. Ignoring the similarity can hurt performance, especially for argument roles with few/no examples at training time (e.g., similar to the zero-shot setting in Levy et al. (2017)).

In this paper, we propose a new paradigm for the event extraction task – formulating it as a question answering (QA)/machine reading comprehension (MRC) task (Contribution 1). The general framework is illustrated in Figure 2. Using BERT Devlin et al. (2019) as the base model for obtaining contextualized representations from the input sequences, we develop two BERT-based QA models – one for event trigger detection and the other for argument extraction. For each, we design one or more Question Templates that map the input sentence into the standard BERT input format. Thus, trigger detection becomes a request to identify “the action” or the “verb” in the input sentence and determine its event type; and argument extraction becomes a sequence of requests to identify the event’s arguments, each of which is a text span in the input sentence. Details will be explained in Section 2. To the best of our knowledge, this is the first attempt to cast event extraction as a QA task.

Treating event extraction as QA overcomes the weaknesses in existing methods identified above (Contribution 2): (1) Our approach requires no entity annotation (gold or predicted entity information) and no entity recognition pre-step; event argument extraction is performed as an end-to-end task; (2) The question answering paradigm naturally permits the transfer of argument extraction knowledge across semantically related argument roles. We propose rule-based question generation strategies (including incorporating descriptions in annotation guidelines) for templates creation, and conduct extensive experiments to evaluate our framework on the Automatic Content Extraction (ACE) event extraction task and show empirically that the performance on both trigger and argument extraction outperform prior methods (Section 3.2). Finally, we show that our framework extends to the zero-shot setting – it is able to extract event arguments for unseen roles (Contribution 3).

Methodology

In this section, we first provide an overview for the framework (Figure 2), then go deeper into details of its components: question generation strategies for template creation, as well as training and inference of QA models.

Our QA framework for event extraction relies on two sets of Question Templates that map an input sentence to a suitable input sequence for two instances of a standard pre-trained bidirectional transformer (BERT Devlin et al. (2019)). The first of these, BERT_QA_Trigger (green box in Figure 2), extracts from the input sentence the event trigger which is a single token, and its type (one of a fixed set of pre-defined event types). The second QA model, BERT_QA_Arg (orange box in Figure 2), is applied to the input sequence, the extracted event trigger and its event type to iteratively identify candidate event arguments (spans of text) in the input sentence. Finally, a dynamic threshold is applied to the extracted candidate arguments, and only the arguments with probability above the threshold are retained.

The input sequences for the two QA models share a standard BERT-style format:

where [CLS] is BERT’s special classification token, [SEP] is the special token to denote separation, and is the tokenized input sentence. We provide details on how to obtain the in Section 2.2. Details on the QA models and the inference process will be explained in Section 2.3.

2 Question Generation Strategies

For our QA-based framework for event extraction to be easily moved from one domain to the other, we concentrated on developing question generation strategies that not only worked well for the task, but can be quickly and easily implemented. For event trigger detection, we experiment with a set of four fixed templates – “what is the trigger”, “trigger”, “action”, “verb”. Basically, we use the fixed literal phrase as the question. For example, if we choose the “action” template, the input sequence for the example sentence in Figures 1 and 2 is instantiated as:

[CLS] action [SEP] As part of the 11-billion-dollar sale … [SEP]

As for event argument extraction, we design three templates with argument role name, basic argument based question and annotation guideline based question, respectively:

Template 1 (Role Name) For this template, is simply instantiated with the argument role name (e.g., artifact, agent, place).

Template 2 (Type + Role) Instead of directly using the argument role name () as the question, we first determine the argument role’s general semantic type — one of person, place, other; and construct the associated “WH" word question – who for person, where for place and what for all other cases, of the following form:

Examples are shown in Table 1 for the arguments of event type Movement.Transport. By adding the WH word, more semantic information is included as compared to Template 1.

Template 3 (Incorporating Annotation Guidelines) To incorporate even more semantic information and make the question more natural sounding, we utilize the descriptions of each argument role provided in the ACE annotation guidelines for events Linguistic Data Consortium (2005) for generating the questions.

+ “in ” Finally, for each template type, it is possible to encode the trigger information by adding “in ” at the end of the question (where is instantiated with the real trigger token obtained from the trigger detection phase). For example, the Template 2 question incorporating trigger information would be:

To help better understand all the strategies above, Table 1 presents an example for argument roles of event type Movement.Transport. We see that the annotation guideline based questions are more natural and encode more semantics about a given argument role, than the simple Type + Role question “what is the artifact?”.

3 Question Answering Models

We use BERT Devlin et al. (2019) as the base model for getting contextualized representations for the input sequences for both BERT_QA_Trigger and BERT_QA_Arg. After the instantiation with question templates the sequences are of format [CLS] [SEP] [SEP].

Then we get the contextualized representations of each token for trigger detection and argument extraction with BERTTr\texttt{BERT}_{Tr} and BERTArg\texttt{BERT}_{Arg}, respectively. For the input sequence (e1,e2,...,eN)(e_{1},e_{2},...,e_{N}) prepared for trigger detection, we have:

For the input sequence (a1,a2,...,aM)(a_{1},a_{2},...,a_{M}) prepared for argument span extraction, we have:

The output layer of each QA model, however, differs: BERT_QA_Trigger predicts the event type for each token in sentence (or None if it is not an event trigger), while BERT_QA_Arg predicts the start and end offsets for the argument span with a different decoding strategy.

At test time, for trigger detection, to obtain the type for each token e1,e2,...,eNe_{1},e_{2},...,e_{N}, we simply apply argmax to PtrP_{tr}.

To train the models (BERT_QA_Trigger and BERT_QA_Arg), we minimize the negative log-likelihood loss for both models, parameters are updated during the training process. In particular, the loss for the argument extraction model is the sum of two parts: the start token loss and end end token loss. For the training examples with no argument span (no answer case), we minimize the start and end probability of the first token of the sequence ([CLS]).

At test time, predicting the argument spans is more complex – for each argument role, there can be several or no spans to be extracted. After the output layer, we have the probability of each token ai∈(a1,a2,...,aM)a_{i}\in(a_{1},a_{2},...,a_{M}) being the start (Ps(i)P_{s}(i)) and end (Pe(i)P_{e}(i)) of the argument span.

Firstly, we run an algorithm to harvest all valid argument spans candidates for each argument role (Algorithm 1). Basically, we:

Enumerate all the possible combinations of start offset (startstart) and end offset (endend) of the argument spans (line 1–2);

Eliminate the spans not satisfying the constraints: start and end token must be within the sentence; the length of the span should be shorter than a maximum length constraint; Argument spans should have larger probability than the probability of “no argument” (which is stored at the [CLS] token) (line 3–5);

Calculate the relative no answer score (no_ans_scoreno\_ans\_score) for the candidate span and add the candidate to list (line 6–8).

Then we run another algorithm to filter out candidate arguments that should not be included (Algorithm 2). More specifically, we obtain a probability threshold (best_threshbest\_thresh) that helps achieve best evaluation results on the dev set (line 1–9) and keep only those arguments with no_ans_scoreno\_ans\_score smaller than the threshold (line 10–13). With the dynamic threshold for determining the number of arguments to be extracted for each roleEach role has a separate threshold., we avoid adding a (hard) hyperparameter for this purpose.

Another easier way to get final argument predictions is to directly include all the candidates with no_ans_score<0no\_ans\_score<0, which does not require tuning the dynamic threshold best_threshbest\_thresh.

Experiments

We conduct experiments on the ACE 2005 corpus Doddington et al. (2004), it contains documents crawled between year 2003 and 2005 from a variety of areas such as newswire (nw), weblogs (wl), broadcast conversations (bc) and broadcast news (bn). The part that we use for evaluation is fully annotated with 5,272 event triggers and 9,612 arguments. We use the same data split and pre-processing step as in the prior works Zhang et al. (2019b); Wadden et al. (2019).

As for evaluation, we adopt the same criteria defined in Li et al. (2013): An event trigger is correctly identified (ID) if its offsets match those of a gold-standard trigger; and it is correctly classified if its event type (33 in total) also matches the type of the gold-standard trigger. An event argument is correctly identified (ID) if its offsets and event type match those of any of the reference argument mentions in the document; and it is correctly classified if its semantic role (22 in total) is also correct. Though our framework does not involve the trigger/argument identification step and tackles the identification + classification in an end-to-end way, we still report the trigger/argument identification’s results to compare to prior work. It could be seen as a more lenient evaluation metric, as compared to the final trigger detection and argument extraction metric (ID + Classification), which requires both the offsets and the type to be correct. All the aforementioned elements are evaluated using precision (denoted as P), recall (denoted as R) and F1 scores (denoted as F1).

2 Results

We compare our framework’s performance to a number of prior competitive models: dbRNN Sha et al. (2018) is an LSTM-based framework that leverages the dependency graph information to extract event triggers and argument roles. Joint3EE Nguyen and Nguyen (2019) is a multi-task model that performs entity recognition, trigger detection and argument role assignment by shared BiGRU hidden representations. GAIL Zhang et al. (2019b) is an ELMo-based model that utilizes a generative adversarial network to help the model focus on harder-to-detect events. DYGIE++ Wadden et al. (2019) is a BERT-based framework that models text spans and captures within-sentence and cross-sentence context. OneIE Lin et al. (2020) is a joint neural model for extraction with global features.Slightly different from our and Wadden et al. (2019)’s data pre-processing, OneIE skips lines before the ¡text¿ tag (e.g., headline, datetime).

In Table 2, we present the comparison of models’ performance on trigger detection. We also implement a BERT fine-tuning baseline and it reaches nearly same performance as its counterpart in DYGIE++. We observe that our BERT_QA_Trigger model with the best trigger questioning strategy reaches comparable (better) performance with the baseline models.Note that OneIE is concurrent to our work and reports better performance. On trigger detection, it reaches 74.7 F1 as compare to our 72.39. On argument extraction (affected by trigger detection), it reaches 56.8 as compared to our 53.31.

Table 3 shows the comparison between our model and baseline systems on argument extraction. Notice that the performance of argument extraction is directly affected by trigger detection. Because argument extraction correctness requires the trigger to which the argument refers to be correctly identified and classified. We observe, (1) Our BERT_QA_Arg model with the best argument question generation strategy (annotation guideline based questions) outperforms prior work significantly, although it uses no entity recognition resources; (2) Drop of F1 performance from argument identification (correct offset) to argument ID + classification (both correct offset and argument role) is only around 1%, while the gap is around 3% for prior models which rely on entity recognition and a multi-step process for argument extraction. This once again demonstrates the benefit of our new formulation for the task as question answering.

To better understand how the dynamic threshold is affecting our framework’s performance. We conduct an ablation study on this (Table 3) and find that the threshold increases the precision and the general F1 substantially. The last row in the table shows the test time ensemble performance of the predictions from BERT_QA_Arg trained with template 2 question, and another BERT_QA_Arg trained with template 3 question (the two relatively better questioning strategies). The ensemble system outperforms the non-ensemble system in both precision and recall, demonstrating the benefit from both templates.

Evaluation on Unseen Argument Roles

To verify how our formulation provides advantages for extracting arguments with unseen argument roles (similar to the zero-shot relation extraction setting in Levy et al. (2017)), we conduct another experiment, where we keep 80% of the argument roles (16 roles) seen at training time, and 20% (6 roles) only seen at test time. Specifically, the unseen roles are “Vehicle, Artifact, Target, Victim, Recipient, Buyer”. Notice that during training, we use the subset of sentences from the training set, which are known to contain arguments of seen roles as positive examples. At test time, we evaluate the models on the subset of sentences from the test set, which contains arguments of unseen roles.We omit the trigger detection phase in this evaluation.

Table 5 presents the results. Random NE is our random baseline that selects a named entity in the sentence, it has a reasonable performance of near 25%. Prior models such as GAIL are not capable of handling the unseen roles. ZSTE Huang et al. (2018) is a framework for zero-shot transfer learning of event extraction with AMR. It maps each parsed candidate span to a specific type in a target event ontology. Its argument extraction results are affected by AMR performance and their reported F1 is around 20-30% in their evaluation setting.

Using our QA-based framework, as we leverage more semantic information and naturalness into the question (from question template 1 to 2, to 3), both the precision and recall increase substantially.

Further Analysis

To investigate how the question generation strategies affect the performance of event extraction, we perform experiments on trigger and argument extractions with different strategies, respectively.

In Table 6, we try different fixed questions for trigger detection. By “leaving empty”, we mean instantiating the question with empty string.In this case, the model degrades to a token classification model, which matches our BERT FineTune baseline’s performance. There’s no substantial gap between different alternatives. By using “verb” as the question, our BERT_QA_Trigger model achieves best performance (measured by F1 score). The QA model also encodes the semantic interactions between the fixed question (“verb”) and the sentence, this explains why BERT_QA_Trigger is better than BERT FineTune in trigger detection.

The comparison between different question generation strategies for argument extraction is even more interesting. In Table 4, we present the results in two settings: event argument extraction with predicted triggers (the same setting as in Table 3), and with gold triggers. In summary, we find that:

Adding “in ” after the question consistently improves the performance. It serves as an indicator for what/where the trigger is in the input sentence. Without adding the “in ”, for each template (1, 2 & 3), the F1 of models’ predictions drop around 3 percent when given predicted triggers, and more when given gold triggers.

Our template 3 questioning strategy which is most natural achieves the best performance. As we mentioned earlier, template 3 questions are based on descriptions for argument roles in the annotation guideline, thus encoding more semantic information about the role name. And this corresponds to the accuracy of models’ predictions – template 3 is more effective than templates 1&2 in both with “in ” and without “in ” settings. What’s more, we observe that template 2 (adding a WH_word to form the questions) achieves better performance than the template 1 (directly using argument role name).

2 Error Analysis

We further conduct error analysis and provide a number of representative examples. Table 7 summarizes error statistics for trigger detection and argument extraction.

For event triggers, the majority of the errors relate to missing or spurious predictions, and only 8.29% involve misclassified event types (e.g., an Elect event is mistaken for a Start-Position event). For event arguments, on the sentences that come with at least one event in gold data, our framework extracts more arguments only around 14% of the cases. Most of the time (54.37%), our framework extracts fewer arguments than it should; this corresponds to the results in Table 3, where the precision of our models are higher. In around 30% of the cases, our framework extracts the same number of arguments as in the gold data, almost half of which match exactly the gold arguments.

After examining the example predictions, we find that reasons for errors can be mainly divided into the following categories:

More complex sentence structures. In the following example, the input sentence has multiple clauses, each with trigger and arguments (such as when triggers are partial or elided). Our model is capable of also extracting “Tom” as another Entity of the Contact.Meet event in the first example:

[She]Entity visited the store and [Tom]Entity did too.

But in the second example, when there is a higher-order event expressed spanning events in nested clauses, our model did not extract the entire Victim correctly, which shows the difficulty of handling complex clause structures.

Canadian authorities arrested two Vancouver-area men on Friday and charged them in the deaths of [329 passengers and crew members of an Air-India Boeing 747 that blew up over the Irish Sea in 1985, en route from Canada to London]Victim.

Lack of reasoning with document-level context. In the sentence “MCI must now seize additional assets owned by Ebbers, to secure the loan.” There is a Transfer-Money event triggered by loan, with MCI as the Giver and Ebbers, the Recipient. In the previous paragraph, it’s mentioned that “Ebbers failed to make repayment of certain amount of money on the loan from MCI.” Without this context, it is hard to determine that Ebbers should be the recipient of the loan.

Lack of knowledge to obtain exact boundary of the argument span. For example, in “Negotiations between Washington and Pyongyang on their nuclear dispute have been set for April 23 in Beijing …”, for the Entity role, two argument spans should be extracted (“Washington” and “Pyongyang”). While our framework predicts the entire “Washington and Pyongyang” as the argument span. Although there’s an overlap between the prediction and gold-data, the model gets no credit for it.

Data and lexical sparsity. In the following two examples, our model fails to detect the triggers of type End-Position. “Minister Tony Blair said ousting Saddam Hussein now was key to solving similar crises.” “There’s no indication if Erdogan would purge officials who opposed letting in the troops.” It’s partially due to they not being seen during training as triggers. “ousting” is a rare word and is not in the tokenizers’ vocabulary. Purely inferring from the sentence context is hard to make the correct prediction.

Related Work

Most event extraction research has focused on the 2005 Automatic Content Extraction (ACE) sentence-level event task Walker et al. (2006). In recent years, continuous representations from convolutional neural networks Nguyen and Grishman (2015); Chen et al. (2015) and recurrent neural networks Nguyen et al. (2016) have been proved to help substantially for pipeline-based classifiers by automatically extracting features. To mitigate the effect of error propagation, joint models have been proposed for event extraction. Yang and Mitchell (2016) consider structural dependencies between events and entities, which requires heavy feature engineering to capture discriminative information. Nguyen and Nguyen (2019) propose a multitask model that performs entity recognition, trigger detection and argument role prediction by sharing BiGRU hidden representations. Zhang et al. (2019a) utilize a neural transition-based extraction framework Zhang and Clark (2011), which requires specially designed transition actions. It still requires recognizing entities during decoding, though entity recognition and argument role prediction are done jointly.

These methods generally perform trigger detection →\rightarrow entity recognition →\rightarrow argument role assignment during decoding. Different from the works above, our framework completely bypasses the entity recognition stage (thus no annotation resources for NER needed), and directly tackles event argument extraction. Also related to our work includes DYGIE++ Wadden et al. (2019) – it models the entity/argument spans (with start and end offset) instead of labeling with the BIO scheme. Different from our work, its learned span representations are later used to predict the entity/argument type. While our QA model directly extracts the spans for certain argument role types. Contextualized representations produced by pre-trained language models Peters et al. (2018); Devlin et al. (2019) have been shown to be helpful for event extraction Zhang et al. (2019b); Wadden et al. (2019) and question answering Rajpurkar et al. (2016). The attention mechanism helps capture relationships between tokens in the question and input sequence tokens. We use BERT in our framework for capturing these semantic relationships.

Machine Reading Comprehension (MRC)

Span-based MRC tasks involve extracting a span from a paragraph Rajpurkar et al. (2016) or multiple paragraphs Joshi et al. (2017); Kwiatkowski et al. (2019). Recently, there have been explorations on formulating NLP tasks as a question answering problem. McCann et al. (2018) proposes natural language decathlon challenge (decaNLP), which consists of ten tasks (e.g., machine translation, summarization, question answering). They cast all tasks as question answering over a context and propose a general model for this. In the information extraction literature, Levy et al. (2017) propose the zero-shot relation extraction task and reduce the task to answering crowd-sourced reading comprehension questions. Li et al. (2019) casts entity-relation extraction as a multi-turn question answering task. Their questions lack diversity and naturalness. For example for the PART-WHOLE relation, the template question is “find Y that belongs to X”, where X is instantiated with the pre-given entity. The follow-up work for named entity recognition from Li et al. (2020) propose better query strategies incorporating synonyms and examples. Different from the works above, we focus on the more complex event extraction task, which involves both trigger detection and argument extraction. Our generated questions for extracting event arguments are somewhat more natural (incorporating descriptions from annotation guidelines) and leverage trigger information.

Question Generation

To generate question templates 2&3 (Type + Role question and annotation guideline based question) which are more natural, we draw insights from the literature of automatic rule-based question generation Heilman and Smith (2010). Heilman (2011) propose to use linguistically motivated rules for WH word (question phrase) selection. In their more general case of question generation from sentences, answer phrases can be noun phrases, prepositional phrases, or subordinate clauses. Complicated rules are designed with help from the superTagger Ciaramita and Altun (2006). In our case, event arguments are mostly noun phrases and the rules are simpler – “who” for person, “where” for place and “what” for all other types of entities. We sample around 10 examples from the development set to determine the entity type of each argument role. In the future, it will be interesting to investigate how to utilize machine learning-based question generation methods Du et al. (2017). They would be more beneficial for the setting where the schema/ontology contains a large number of argument types.

Conclusion

In this paper, we introduce a new paradigm for event extraction based on question answering. We investigate how the question generation strategies affect the performance of our framework on both trigger detection and argument span extraction, and find that more natural questions lead to better performance. Our framework outperforms prior works on the ACE 2005 benchmark, and is capable of extracting event arguments of roles not seen at training time. For future work, it would be interesting to try incorporating broader context (e.g., paragraph/document-level context Ji and Grishman (2008); Huang and Riloff (2011); Du and Cardie (2020) in our methods to improve the accuracy of the predictions.

Acknowledgments

We thank the anonymous reviewers and Heng Ji for helpful suggestions. This research is based on work supported in part by DARPA LwLL Grant FA8750-19-2-0039.

References

Appendix A Questions Based on Annotation Guidelines

Questions based on annotation guidelines for each argument role.