Unified Structure Generation for Universal Information Extraction

Yaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, Hua Wu

Introduction

Information extraction (IE) aims to identify and structure user-specified information from unstructured texts (Andersen et al., 1992; Grishman, 2019). IE tasks are highly diversified due to its varying targets (entity, relation, event, sentiment, etc.), heterogeneous structures (spans, triplets, records, etc.), and demand-specific schemas (Grishman and Sundheim, 1996; Mitchell et al., 2005; Ji and Grishman, 2011).

Currently, most IE approaches are task-specialized, which leads to dedicated architectures, isolated models, and specialized knowledge sources for different IE task. These task-specialized solutions greatly hinder the rapid architecture development, effective knowledge sharing, and quick cross-domain adaptation of IE systems. First, it is very complicated to develop dedicated architectures for a large amount of IE tasks/settings/scenarios. Second, learning isolated models severely restricts the knowledge sharing between related tasks and settings. Finally, it is costly and time-consuming to construct data sets and knowledge sources specialized for different IE tasks. Therefore, it will be of great benefit to develop a universal IE architecture that can uniformly model different IE tasks, adaptively predict heterogeneous structures and effectively learn from various resources, which we referred to as Universal IE.

Fundamentally, all IE tasks can be modeled as text-to-structure transformations, with different tasks correspond to different structures. For example, as shown in Figure 1, an entity is a named span structure, an event is a schema-defined record structure. These text-to-structure transformations in IE can be further decomposed into several atomic transformation operations: 1) Spotting, which locates the desirable spans concerning to given specific semantic types (Kripke and Munitz, 1971; Chen and Yuille, 2004). For example, locating span “Steve” as a Person entity and locating “excited” as a sentiment expression. 2) Associating, which connects spans by assigning them with semantic roles in pre-defined schemas (Onyshkevych, 1994; Milward and Thomas, 2000). For example, associating “Steve” and “Apple” by assigning them as the Arg1 and the Arg2 of a Work-for relation. In this way, different IE tasks can be decomposed into a sequence of atomic text-to-structure transformations, and all IE models share the same underlying spotting and associating abilities. For example, entity extraction can be viewed as spotting mention spans of corresponding entity types, while event detection can be reformulated as spotting triggers spans with event types. And the spotting abilities can be shared between these two tasks.

Based on the above observations, we propose UIE, a unified text-to-structure generation architecture that can universally model different IE tasks, adaptively generate targeted structures, and collaboratively learn general IE abilities from different knowledge sources. Specifically, to model heterogeneous IE structures, we design a structural extraction language (SEL) that can effectively encode different IE structures into a uniform representation, so that various IE tasks can be universally modeled in the same text-to-structure generation framework. To adaptively generate targeted structures for different IE tasks, we propose structural schema instructor (SSI), a schema-based prompt mechanism which controls what to spot, what to associate, and what to generate in UIE. To learn common IE abilities for UIE, we pre-train UIE on large-scale, heterogeneous datasets mined from easily accessible web sources. The large-scale pre-trained UIE model provides a solid foundation for knowledge sharing and quick adaptation to new IE settings, and significantly boosts the IE performance in all supervised, low-resource, and few-shot settings.

We conduct experiments on 13 datasets of 4 main IE tasks (entity/relation/event/sentiment extraction and their unification), and supervised, low-resource, and few-shot settings. Experiment results show that UIE achieves significant improvements in all settings. On supervised settings, UIE achieved 1.42% F1 scores improvements over the state-of-the-art, task-specialized architectures on all datasets. On few-shot and low-resource settings, UIE exhibits strong on-demand adaptation ability: it outperforms baselines dramatically by a large margin. These results verified the effectiveness, universality, and transferability of UIE across different IE tasks, settings, and scenarios.

The main contributions of this paper are:

1) We propose UIE, a unified text-to-structure generation architecture that can universally model different IE tasks, adaptively generate targeted structures, and collaboratively learn general IE abilities from different knowledge sources.

2) We design a unified structure generation network, which encodes heterogeneous IE structures into a uniform representation via a structural extraction language, and controls the UIE model which to spot, which to associate, and which to generate via structural schema instructor mechanism.

3) We pre-train a large-scale text-to-structure generation model via a unified pre-training algorithm. To the best of our knowledge, this is the first text-to-structure pre-trained extraction model, which can benefit future IE studies.

Unified Structure Generation for Universal Information Extraction

Information extraction tasks can be formulated as text-to-structure problems, where different IE tasks correspond to different structures. This paper aims to uniformly model the text-to-structure transformations of different IE tasks via a single framework, i.e., different structure transformations will share the same underlying operations and different transformation abilities in a universal model. Formally, given a specific pre-defined schema ss and texts xx, a universal IE model needs to generate a structure that contains the desirable structural information in the text xx indicated by the schema ss.

Generally, there are two main challenges here. Firstly, due to the diversity of IE tasks, there are many different target structures to extract, e.g., entity, relation, event, etc. Secondly, IE tasks are often demand-specific which are defined using different schemas, therefore we need to adaptively control the extraction process.

In this section, we describe how to jointly formulate, learn, and conduct various IE tasks in a unified text-to-structure generation architecture, named UIE. Specifically, we first design structured extraction language (SEL) to uniformly encode heterogeneous extraction structures, i.e., encode entity, relation, event into a unified representation. Then we describe structural schema instructor (SSI), a schema-based prompt mechanism that controls the UIE model which to spot, which to associate, and which to generate for different extraction settings. The details are as follows.

This section describes how to encode heterogeneous IE structures into a uniform representation. Based on the above discussions, IE structure generation can be decomposed into two atomic operations:

Spotting indicates locating target information pieces from the sentence, e.g., the entity and the trigger word in the event.

Associating indicates connecting different information pieces based on the desirable associations, e.g., the relation between entity pair or the role between event and its argument.

Then different IE structures can be represented as a combination of atomic structure generation operations.

Concretely, we design a unified structured extraction language (SEL), which encodes different IE structures via the spotting-associating structure. As shown in Figure 2(a), each SEL expression contains three types of semantic units: 1) SpotName represents there is a specific information piece with the type of spot name existing in the source text; 2) AssoName indicates there exists a specific information piece in the source text that is with the AssoName association to its upper-level Spotted information in the structure; 3) InfoSpan represents the text span corresponding to the specific spotting or associating information piece in the source text. Furthermore, “:” in the SEL indicates the mapping from InfoSpan to its spotting or associating names, and the two structure indicators “(” and “)” are used to form the hierarchical structure between the extracted information.

Using SEL, Figure 2(b) shows how to represent entity, relation, and event structures. There are three entities and each entity is represented as a spotting structure such as “person:Steve”, “organization:Apple”, and “time:1997”; one relation which is represented as an association structure between “Steve” and “Apple” with association name work for; and one event which is represented as an association structure, where the trigger is a spotting structure “start-position:became”, and its arguments are associated with the trigger: Steve as employee, Apple as employer, 1997 as time.

We can see that, SEL have the advantages that: 1) uniformly encodes varying IE structures, therefore different IE tasks can be modeled as the same text-to-structure generation process; 2) efficiently represents all extraction results of a sentence in the same structure, thus can perform joint extraction naturally; 3) the output structure of generation is very compact, which greatly reduce the complexity of decoding.

For example, the two different tasks entity recognition and event detection can be revisited using the same “(SpotName: InfoSpan)” grammar. While both relation extraction and event extraction can be formulated using the grammar “(SpotName: InfoSpan (AssoName: InfoSpan), …)”, even they are with totally different binary “entity-relation-entity” and N-ary “event-arguments” structures. Such a unified structured extraction language enables UIE to learn from and adapt to different IE tasks without designing task-specialized architectures, because these IE tasks are all universally formulated as the transformation from texts to SEL representations.

2 Structural Schema Instructor for Controllable IE Structure Generation

Using SEL, UIE can uniformly generate different IE structures. However, because different IE tasks have different schemas, one challenge here is how to adaptively control which information we want to generate during extraction. For example, given a sentence “Steve became CEO of Apple in 1997.”, an entity recognition system will generate “((person: Steve) (organization: Apple) (Time: 1997))”, and an event extraction system will generate “(start position: became (employee: Steve) (employer: Apple))”. To this end, we propose structural schema instructor (SSI), a schema-based prompt mechanism that controls which kinds of information need to be spotted and associated.

Figure 3 shows the overall framework of UIE. Formally, UIE takes the given structural schema instructor (ss) and the text sequence (xx) as input, and generates the linearized SEL (yy) which contains the extracted information from xx based on schema ss:

where x=[x1,...,x∣x∣]x=[x_{1},...,x_{|x|}] is the text sequence, s=[s1,...,s∣s∣]s=[s_{1},...,s_{|s|}] is the structural schema instructor, and y=[y1,...,y∣y∣]y=[y_{1},...,y_{|y|}] is a SEL sequence that can be easily converted into the extracted information record.

To describe the extraction target of a task, the structural schema instructor constructs a schema-based prompt and uses it as a prefix during generation.

Specifically, corresponding to the spotting-association structure, the structural schema instructor contains three types of token segments: 1) SpotName: the targeted spotting name in the specific information extraction task, such as “person“ in the NER task; 2) AssoName: the targeted association name, such as “work for” in the relation extraction task; 3) Special Symbols ([spot], [asso], [text]) which are added before each SpotName, AssoName, and input text sequence. All tokens in SSI are concatenated and put before the original text sequences. As shown in Figure 3, the entire input for UIE is in the form of:

For example, the SSI “[spot] person [spot] company [asso] work for [text]” indicates extracting records of the relation schema “the person works for the company” from the sentence. Given the SSI ss, UIE first encodes the text xx, then generates the target record yy in linearized SEL using an encoder-decoder-style architecture.

We found that the schema-based prompt can: 1) effectively guide the SEL generation of UIE, so that the general IE ability can be transferred to new IE tasks; 2) adaptively control which to spot, which to associate, and which to generate, so that semantic knowledge across different labels and tasks can be better shared.

2.2 Structure Generation with UIE

Given SSI ss and text xx as input, UIE extracts targeted information by generating a linearized SEL. We formulate this text-to-SEL generation process using an encoder-decoder-style architecture. Given the raw text sequence xx and the schema instructor ss, UIE first compute the hidden representation H=[s1,...,s∣s∣,x1,...,x∣x∣]\mathbf{H}=[\mathbf{s}_{1},...,\mathbf{s}_{|s|},\mathbf{x}_{1},...,\mathbf{x}_{|x|}] of each token:

where Encoder(⋅)\text{Encoder}(\cdot) is a Transformer encoder. Then UIE will decode the input text into a linearized SEL in an auto-regressive style. At the step ii of decoding, UIE generates the ii-th token yiy_{i} in the SEL sequence and the decoder state hid\mathbf{h}_{i}^{d} as following:

Decoder(⋅)\text{Decoder}(\cdot) is a transformer decoder, which predicts the conditional probability p(yi∣y<i,x,s)p(y_{i}|y_{<i},x,s) of token yiy_{i}. Finally, Decoder(⋅)\text{Decoder}(\cdot) finishes prediction when outputting the end symbol , then we convert the predicted SEL expression into the extracted information record.

Compared with previous IE studies which treat labels as specific symbols, the text-to-structure generation paradigm treats labels as natural language tokens. By verbalizing and generating labels and structures, our method can effectively transfer knowledge from pre-trained language models such as BART (Lewis et al., 2020), T5 (Raffel et al., 2020), and related tasks can easily share knowledge because their labels have similar semantics (e.g., location and place) and share common label-text associations (e.g., victim for different event types).

Pre-training and Fine-tuning for UIE

In this section, we describe: 1) how to pre-train a large-scale UIE model which captures common IE abilities for different IE tasks; 2) how to adapt UIE to different IE tasks in different settings via quick fine-tuning. Specifically, we first collect several large-scale datasets from the Web, including structured (e.g., knowledge bases), unstructured (e.g., raw texts), and parallel (e.g., Wikipedia-Wikidata links) data, then we uniformly pre-train our UIE model on these heterogeneous datasets. Finally, we adapt the pre-trained UIE model to the specific downstream IE tasks via on-demand fine-tuning. We found that the pre-trained UIE model provides a solid foundation for capturing, sharing, and transferring knowledge between different IE tasks, and new IE tasks can be effectively solved because UIE learns general IE ability.

UIE needs to encode the text, map text to structure, and decode valid structure. Therefore, we collect a large-scale pre-training corpus from easily accessible web data sources (more details are in the appendix):

Dpair\mathcal{D}_{\text{pair}} is the text-structure parallel data, where each instance is a parallel pair (token sequence xx, structured record yy). We collect large-scale parallel text-structure pairs by aligning Wikidata with English Wikipedia. Dpair\mathcal{D}_{\text{pair}} is used to pre-train the text-to-structure transformation ability of UIE.

Drecord\mathcal{D}_{\text{record}} is the structure dataset where each instance is structured record yy. We collect structured records from ConceptNet Speer et al. (2017) and Wikidata. Drecord\mathcal{D}_{\text{record}} is used to pre-train the structure decoding ability of UIE.

Dtext\mathcal{D}_{\text{text}} is the unstructured text dataset, and we use all plain texts in English Wikipedia. Dtext\mathcal{D}_{\text{text}} is used to pre-train the semantic encoding ability of UIE.

2 Pre-training

We pre-train UIE using three sequence generation tasks with above mentioned pre-training datasets.

To capture the fundamental text-to-structure mapping ability, we pre-train UIE using Dpair={(x,y)}\mathcal{D}_{\text{pair}}=\{(x,y)\}. Specifically, for each parallel pair (xx, yy), we extract the spot type ss+s_{s+} and the associating type sa+s_{a+} in the record yy as the positive schema s+=ss+∪sa+s_{+}=s_{\text{s+}}\cup s_{\text{a+}}. However, we found that if we only feed UIE with a positive schema, it will only simply remember the triplet in the pre-training data. To learn general mapping ability, we also automatically construct negative schemas for each pair, i.e., we first sample negative spots ss-s_{\text{s-}} and negative association set sa-s_{\text{a-}}, then concatenate meta-schema smeta=s+∪ss-∪sa-s_{\text{meta}}=s_{+}\cup s_{\text{s-}}\cup s_{\text{a-}}, and construct the final extraction target. For example, person and work for is the positive schema in the record “((person: Steve (work for: Apple)))”, and we sample vehicle and located in as the negative schema to construct meta-schema. Finally, the objective of text-to-structure pre-training is:

where θe\theta_{e} and θd\theta_{d} are the parameter of encoder and decoder, respectively.

To pre-train the ability of generating valid structures defined by SEL and schemas, we pre-train UIE on Drecord\mathcal{D}_{\text{record}}. We pre-train UIE decoder as an structured language model, where each record in Drecord\mathcal{D}_{\text{record}} is an expression of SEL:

By pre-training for structure generation, the decoder can capture the regularity of SEL and the interactions between different labels.

During text-to-structure pre-training, we continually pre-train UIE also with the masked language model tasks (Raffel et al., 2020) on Dtext\mathcal{D}_{\text{text}} to retrofit semantic representations of UIE. Specifically, we add span corruption based mask language modeling objective in the pre-training stage:

where x′x^{\prime} is the corrupted source text and x′′x^{\prime\prime} is corrupted target spans. We found this pre-training can effectively alleviate the catastrophic forgetting of token semantics especially on SpotName and AssoName tokens.

We initialize UIE-base and UIE-large with T5-v1.1-base and T5-v1.1-large (Raffel et al., 2020), and the model architectures are shown in Table 7. The final objective is the combine of the above tasks:

For implementation, we uniformly represent all pre-training data as triplets. For text data (xx) in Dtext\mathcal{D}_{\text{text}}, we build a triplet (None, x′x^{\prime}, x′′x^{\prime\prime}) where x′x^{\prime} is the corrupted source text and x′′x^{\prime\prime} is corrupted spans. For text-record data (xx, yy) in Dpair\mathcal{D}_{\text{pair}}, we construct (ss, xx, yy) by sampling the meta-schema ss for each text-record pair. For record data (yy) in Drecord\mathcal{D}_{\text{record}}, we take (None, None, yy) as the input triplet. We randomly pack instances for different tasks in one batch, and details are shown in Algorithm 1 in the appendix.

3 On-Demand Fine-tuning

Given the pre-trained UIE model, we can quickly adapt it to different IE tasks and settings through model fine-tuning. Given a labeled corpus Dtask={(s,x,y)}\mathcal{D}_{\text{task}}=\{(s,x,y)\}, we fine-tune the UIE model using teacher-forcing cross-entropy loss:

To alleviate the exposure bias (Ranzato et al., 2016; Zhang et al., 2020) of the auto-regressive model during decoding, we also design a Rejection Mechanism for effective fine-tuning. Specifically, given an instance (ss, xx, yy), we first encode yy using SEL language, then we randomly insert several [null] unit with negative SpotName and AssoName: (SpotName, [null]) and (AssoName, [null]) into the ground-truth SEL with the probability of pϵp_{\epsilon}. For example, in Table 1, facilityfacility is the negative spot in the schema prompt, i.e., there is no facility entity in the sentence “Steve became CEO of Apple in 1997”. Therefore, we randomly inject the noise of “(facility: [null])” into the target record during model learning. In this way, the UIE can effectively learn to reject misleading generation by generating [null] token.

Experiments

To verify the effectiveness of UIE, we conducted experiments on different IE tasks and settings.

We conduct experiments on 13 IE benchmarks across 4 well-representative IE tasks (including entity extraction, relation extraction, event extraction, structured sentiment extraction) and their combinations (e.g., joint entity-relation extraction). The used datasets includes ACE04 (Mitchell et al., 2005), ACE05 (Walker et al., 2006); CoNLL03 (Tjong Kim Sang and De Meulder, 2003), CoNLL04 (Roth and Yih, 2004), SciERC (Luan et al., 2018), NYT (Riedel et al., 2010), CASIE (Satyapanich et al., 2020), SemEval-14 (Pontiki et al., 2014), SemEval-15 (Pontiki et al., 2015), SemEval-16 (Pontiki et al., 2016), see Table 8 for detail. We employ the end-to-end setting for all extraction tasks, which takes the raw text as input and directly generates the target structure.

We use the same evaluation metrics as all previous methods, and details of metrics are shown in the appendix. For each fine-tuning experiment, we report the average performance on 3 random seeds. Because UIE only generates text spans, we map spans to offsets by finding the first matched offsets that are not already matched in the same SEL hierarchical level (details in appendix). We found this simple heuristic rule is very effective (<0.5% error offsets) and more complicated mapping approaches (such as attention-weight guided span mapping) are left as the future work.

2 Experiments on Supervised Settings

UIE provides a universal backbone for IE tasks. This section assesses the UIE performance in supervised settings. We compare UIE with the state-of-the-art, task-specific supervised models. For a fair comparison, we only compare the state-of-the-art without leveraging additional dataset-specific knowledge or larger-scale contexts. These extensions are good complementary of UIE, and can be left for further improvement. Table 2 shows the performance of UIE on the 13 IE datasets across 4 tasks. We can observe that:

1) By modeling IE as text-to-structure generation and encoding with an effective SEL language, UIE provides an effective universal architecture for IE. The UIE model achieves state-of-the-art performance on nearly all datasets and tasks, even without pre-training (SEL). 2) The large-scale pre-trained model provides a solid foundation for universal IE. Compared with baselines, the pre-trained model achieves the performance of the state-of-the-art in most datasets and improves 1.42% F1 on average. 3) By universally modeling IE tasks and pre-training using large-scale datasets, UIE can effectively capture, share, and transfer IE abilities. Pre-training improves all tasks at the same time, especially events and sentiment knowledge rarely appear in the pre-train dataset. It proves that SEL is a unified and cross-task transferable structured representation for IE, which allows UIE to share learned capabilities and information across different and various information extraction tasks.

3 Experiments on Low-resource Settings

To verify the quick adaptation ability of UIE, we conducted low-resource experiments on six different partitions of the original training sets (1/5/10-shot, 1/5/10% ratio) across 4 tasks. For the few-shot experiments, we sample 1/5/10 sentences for each entity/relation/event/sentiment type in the training set. To avoid the influence of random sampling, we repeated each experiment 10 times with different samples and reported their averaged results as previous works (Huang et al., 2021).

We compare UIE with the following pre-trained model: 1) T5-v1.1-base is an initial model of UIE-base; 2) Fine-tuned T5-base is fine-tuned with sequence generation tasks such as summarization, which have been shown effective in many low-resource NLP tasks Paolini et al. (2021); 3) UIE-base w/o SSI is the distant supervised version of UIE without SSI in the pre-training stage, which is used to verify the necessity of SSI when adapting UIE in low-resource settings. Table 3 shows the performance of 4 IE tasks under 6 low-resource settings. We observe that: 1) By guiding the generation using schema-based prompts, SSI is an effective way for adaptively controlling which to extract. Compared with the UIE model w/o SSI, UIE equipped with SSI achieves improvements of 4.16 and 3.30 on average for n-shot and n-ratio experiments. 2) Our pre-training algorithms can learn general IE ability rather than capture task-specific information. Even the pre-training of UIE didn’t include event and sentiment knowledge, UIE still achieved significantly better performance on these tasks compared to the baseline with only a small number of samples.

4 Ablations on Pre-training Tasks

To investigate the effect of different pre-training tasks, Table 4 shows ablation experiment results of UIE-base on four downstream tasks. We can see that: (1) The pre-training of SEL (LRecord\mathcal{L}_{\text{Record}}) and sequence-to-structure mapping (LPair\mathcal{L}_{\text{Pair}}) is crucial for UIE, and such a structure generation pre-training is especially useful for small-scale datasets. In small datasets CoNLL04 and 16res, adding structure generation pre-training (from T5-v1.1-base to UIE-base w/o LText\mathcal{L}_{\text{Text}}), the performance significantly increases from 72.12 to 75.70 and 72.03 to 74.28. (2) Retrofitting semantic using the mask language model task (LText\mathcal{L}_{\text{Text}}) is more important for the complex extraction task. In the tasks with more semantic types such as event extraction (33 types), the performance drops significantly after removing the LText\mathcal{L}_{\text{Text}} task, e.g., 72.63→\rightarrow70.89 and 57.27→\rightarrow54.16. (3) The mapping pre-training with LPair\mathcal{L}_{\text{Pair}} enables the model to learn the ability of extraction. After ablating LPair\mathcal{L}_{\text{Pair}}, the extraction ability of UIE is significantly decreased, i.e., the performance on the relation (-0.90), event (-1.43/-1.48), and sentiment (-0.46) tasks all see large decline.

5 Effects of Rejection Noise

This section investigates the effect of the proposed rejection noise. Table 5 shows the results of the different pre-trained models on the development set of CoNLL 03 under the 10-shot setting. The mis-generated label has a negative influence on the precision of the proposed generation method leading to a large number of error extraction results. The proposed rejection noise is useful for the generation method, which leads to improvements of 13.16 precision (P) on average.

Related Work

Building and pre-training universal models of NLP tasks has attracted a lot of attention in recent years, e.g., contextualized representation (Devlin et al., 2019; Liu et al., 2019), text generation (Lewis et al., 2020; Raffel et al., 2020), multi-modal (Li et al., 2021b; Cho et al., 2021), and multi-lingual (Conneau et al., 2020; Xue et al., 2021). This paper proposes and pre-trains the first universal model for information extraction.

IE is a long-researched area and many classical neural architectures have been proposed, such as sequence tagging (Lample et al., 2016; Zheng et al., 2017; Lin et al., 2019), span classification (Sohrab and Miwa, 2018; Lin et al., 2018; Wadden et al., 2019), and MRC (Levy et al., 2017; Li et al., 2020; Du and Cardie, 2020). And several task-specific pre-training techniques are proposed on these architectures (Mengge et al., 2020; Wang et al., 2021b; Qin et al., 2021). More relevant to our work are generation-based IE methods, which generate text spans via tagging (Straková et al., 2019; Ma et al., 2019), index pointer (Ren et al., 2021; Yan et al., 2021b) or copy mechanism (Zeng et al., 2018), and these methods usually employ specific classifiers to represent labels. The generation can be enhanced using label templates (Li et al., 2021a; Liu et al., 2021; Cui et al., 2021), schema (Lu et al., 2021; Ahmad et al., 2021), and augmented language methods (Paolini et al., 2021).

Compared with previous IE studies which focus on developing more effective task-specialized models, this paper aims to universally model various IE tasks in an unified text-to-structure framework, which can greatly benefit the rapid development, effective knowledge sharing, and quick adaptation of IE systems.

Conclusion

In this paper, we propose a unified text-to-structure generation framework – UIE, which can universally model different IE tasks, adaptively generate targeted structures, and unfiedly learn general IE abilities from different knowledge sources. Experimental results show that UIE achieves very competitive performance in both supervised and low-resource settings, which verified its universality, effectiveness, and transferability. A large-scale pre-trained text-to-structure model is also released, which will benefit future studies. For future work, we want to extend UIE to KB-aware IE tasks such as entity linking (Cao et al., 2021), and document-aware IE tasks such as co-reference (Lee et al., 2017; Lu et al., 2022).

Acknowledgements

We sincerely thank the reviewers for their insightful comments and valuable suggestions. This research work is supported by the National Natural Science Foundation of China under Grants no. U1936207, 62122077 and 62106251, the Project of the Chinese Language Committee under Grant no. YB2003C002.

References

Appendix A Experiment Details

This section describes the details of experiments, including pre-training and fine-tuning on downstream tasks.

We use the 20210401 version of Wikipediahttps://www.wikipedia.org/ and Wikidatahttps://www.wikidata.org/ dump and ConceptNethttps://conceptnet.io/ to construct the pre-train dataset.

For Wikidata and Wikipedia, we use them to collect the tuples Tw={<Th,eh,r,et,X>}\mathcal{T}_{w}=\{<T_{h},e_{h},r,e_{t},X>\}, where ThT_{h} is head entity type, ehe_{h} is head entity, rr is relation, ete_{t} is tail entity, XX is sentence, and the Tw\mathcal{T}_{w} can be used to construct Dpair\mathcal{D}_{\text{pair}}, Drecord\mathcal{D}_{\text{record}} and Dtext\mathcal{D}_{\text{text}}. Firstly, we construct entity type dictionary L\mathcal{L} and relation dictionary P\mathcal{P} from Wikidata. Wikidata has more than 40M entity items and each item has its corresponding properties which indicate the association between entities. For type dictionary L\mathcal{L}, we regard each item as an entity, use the “instance of” and “subclass of” property values as its corresponding entity types and consider other properties as the relation of the entity with others. To learn general knowledge, all entity types will be retained except those whose instances are < 5. For the type whose name is longer than 3 tokens, we use its headwords as the final type for simplicity, e.g.,“state award of the Republic of Moldova” is converted to “state award”. For relation dictionary P\mathcal{P}, Wikidata has more than 9K kinds of propertieshttps://www.wikidata.org/wiki/Wikidata:List_of_properties, we filter out the properties of external-id, URL, and math types. In this way, we obtain a collection of 31K types and retained 1535 properties which can serve as a solid foundation for universal IE. Secondly, we collect the mentions of each entity by using its anchor texts in Wikipedia and the top 3 frequent noun phrase occurrences of its entry page (Li et al., 2010). Then for each mention, we identify its entity types by linking it to its Wikidata item’s types. For each Wikipedia page, we split the text into sentencesnltk.tokenize.punkt and filter out sentences that have no entities. Thirdly, we regard each entity as a head entity and find the associated entities according to its properties. The associated entity will set as as tail entity, and the property value will set as association type. If a head entity has no type, ThT_{h} will be blank or has no associated tail entity, rr and ete_{t} will be blank. To this end, given a sentence, we can construct instances based on the collected tuples Tw\mathcal{T}_{w} by setting ehe_{h} and ete_{t} as InfoSpan, and assigning ThT_{h} as SpotName, rr as AssoName. Finally, from Wikipedia and Wikidata, we construct Dpair\mathcal{D}_{\text{pair}}, Drecord\mathcal{D}_{\text{record}} and Dtext\mathcal{D}_{\text{text}} with 65M instances, respectively. And we keep 50K as the development dataset.

To add common sense knowledge to structured extraction language (SEL), we extract the tuples Tc\mathcal{T}_{c} from ConceptNet. ConceptNet contains 48 associations and has no context or entity types. So we leave the ThT_{h}, TtT_{t} XX blank and finally construct 1M instances.

We first initialize UIE-base and UIE-large with T5-v1.1-base and T5-v1.1-large checkpoints (Raffel et al., 2020), and the model architectures are shown in Table 7. We employ Adam optimizer Kingma and Ba (2015) as the optimizer with learning rate=1e-4, and use linear scheduling with a warming up proportion 6%. For negative spots and associations in the LPair\mathcal{L}_{\text{Pair}}, we randomly select negative spots and associations up to 10 for each instance, respectively. For LText\mathcal{L}_{\text{Text}}, we set the corruption rate to 15% and the average corrupting span length to 3, following Raffel et al. (2020). We truncate the concatenated overall length of schema prompt ss and raw text xx, as well as the length of SEL expression yy, together to 128 during pre-training. We train our base model and large model for both 500K steps with batch size 512 on 8 NVIDIA A100 GPUs.

The detailed pre-training process in a python-like style is shown in Algorithm 1. In each batch of pre-training processes for UIE, we construct a batch of triplets (ss, xx, yy) containing text-record pairs, text instances, and record instances. In practice, since 8 GPUs could only run the large model with an overall batch of 128 (batch=16 on each GPU), we update the model parameters after accumulating 4 gradients.

A.2 Details of Downstream Tasks

We conduct downstream tasks on 4 IE tasks, 13 datasets, and the detailed statistic of each dataset is shown in Table 8.

We conduct entity extraction experiments on three entity datasets: ACE04https://catalog.ldc.upenn.edu/LDC2005T09 (Mitchell et al., 2005), ACE05-Enthttps://catalog.ldc.upenn.edu/LDC2006T06 (Walker et al., 2006), and CoNLL03 (Tjong Kim Sang and De Meulder, 2003). For nested entity extraction datasets ACE04 and ACE05-Ent, we follow the pre-processing steps and data split of previous works Li et al. (2020).

We conduct experiments on four wide-used end-to-end relation extraction datasets across several languages and domains: ACE05-Rel (Walker et al., 2006), CoNLL04https://github.com/btaille/sincere (Roth and Yih, 2004), NYThttps://github.com/yubowen-ph/JointER (Riedel et al., 2010), and SciERChttp://nlp.cs.washington.edu/sciIE/ (Luan et al., 2018). We follow the preprocessing steps and data split of previous works (Taillé et al., 2020; Yu et al., 2020; Wadden et al., 2019).

For ACE05-Evt, we follow the same types, data splits, and pre-processing steps as Lin et al. (2020). For CASIE (Satyapanich et al., 2020), we first remove three incomplete annotated documents (999, 10001, 10002), then split the remaining documents into three sets: train/val/test=697/100/200 according to the time order of each document. We employ the state-of-the-art generation-based event extraction method Text2Event (Lu et al., 2021) as the comparable state-of-the-art system.

We conduct sentiment extraction experiments on the sentiment triplet extraction (Xu et al., 2020) of SemEval 14/15/16 aspect sentiment analysis datasets. We employ the pre-processing datasets of the previous work (Yan et al., 2021a)https://github.com/yhcc/BARTABSA.

We use span-based offset Micro-F1 as the primary metric to evaluate the model:

Entity: an entity mention is correct if its offsets and type match a reference entity.

Relation Strict: relation with strict match, a relation is correct if its relation type is correct and the offsets and entity types of the related entity mentions are correct.

Relation Triplet: relation with boundary match, a relation is correct if its relation type is correct and the string of the subject/object are correct.

Event Trigger: an event trigger is correct if its offsets and event type matches a reference trigger.

Event Argument: an event argument is correct if its offsets, role type, and event type match a reference argument mention.

Sentiment Triplet: a correct triplet requires the offsets boundary of the target, the offsets boundary of the opinion span, and the target sentiment polarity to be all correct at the same time.

To make a fair comparison with baseline systems, we mapped the generated string-level extraction results to offset-level for model evaluation. In detail, we reconstructed the offset of predicted entity/trigger mentions by finding the matched utterance in the input sequence one by one. For argument mentions in relation and event tasks, we found the nearest matched utterance to the predicted entity/trigger mention as the predicted offset. This simple heuristic offset strategy achieves high accuracy. Compared to the string level evaluation, the error rate of the reported offset level evaluation is less than 0.5%. More complicated mapping approaches are left as future work.

Table 6 shows the detailed hyper-parameters for downstream tasks.

A.3 Comparison of UIE-base

This section introduces detailed experiment results of UIE-base.

Table 9 shows the performance of UIE-base and the state-of-the-art systems on the four aspect-based sentiment analysis datasets. As shown in Table 9, the proposed SEL and SSI also have strong portability to sentiment triplets extraction, which achieves the competitive performance with the state-of-the-art with task-specific architectures. With the unified pre-training, UIE-base achieves an improvement of 3.24 on average over T5-v1.1-base across four datasets. This verifies the proposed unified pre-training algorithms can learn general IE ability even the sentiment knowledge is rarely in the pre-training stage.

Table 10 shows the performance of SEL-SSI with the T5-v1.1-base for NYT. Due to the high overlapping of NYT and pre-trained data, we didn’t conduct the experiment of UIE on NYT. Even without pre-training, SSI + SEL still achieved the state-of-the-art performance on NYT. This is because of the flexible generation architecture and the universal SEL expression, UIE can naturally handle entity overlap problems.