Learning to Denoise Distantly-Labeled Data for Entity Typing
Yasumasa Onoe, Greg Durrett
Introduction
With the rise of data-hungry neural network models, system designers have turned increasingly to unlabeled and weakly-labeled data in order to scale up model training. For information extraction tasks such as relation extraction and entity typing, distant supervision Mintz et al. (2009) is a powerful approach for adding more data, using a knowledge base Del Corro et al. (2015); Rabinovich and Klein (2017) or heuristics Ratner et al. (2016); Hancock et al. (2018) to automatically label instances. One can treat this data just like any other supervised data, but it is noisy; more effective approaches employ specialized probabilistic models Riedel et al. (2010); Ratner et al. (2018a), capturing its interaction with other supervision Wang and Poon (2018) or breaking down aspects of a task on which it is reliable Ratner et al. (2018b). However, these approaches often require sophisticated probabilistic inference for training of the final model. Ideally, we want a technique that handles distant data just like supervised data, so we can treat our final model and its training procedure as black boxes.
This paper tackles the problem of exploiting weakly-labeled data in a structured setting with a two-stage denoising approach. We can view a distant instance’s label as a noisy version of a true underlying label. We therefore learn a model to turn a noisy label into a more accurate label, then apply it to each distant example and add the resulting denoised examples to the supervised training set. Critically, the denoising model can condition on both the example and its noisy label, allowing it to fully leverage the noisy labels, the structure of the label space, and easily learnable correspondences between the instance and the label.
Concretely, we implement our approach for the task of fine-grained entity typing, where a single entity may be assigned many labels. We learn two denoising functions: a relabeling function takes an entity mention with a noisy set of types and returns a cleaner set of types, closer to what manually labeled data has. A filtering function discards examples which are deemed too noisy to be useful. These functions are learned by taking manually-labeled training data, synthetically adding noise to it, and learning to denoise, similar to a conditional variant of a denoising autoencoder Vincent et al. (2008). Our denoising models embed both entities and labels to make their predictions, mirroring the structure of the final entity typing model itself.
We evaluate our model following Choi et al. (2018). We chiefly focus on their ultra-fine entity typing scenario and use the same two distant supervision sources as them, based on entity linking and head words. On top of an adapted model from Choi et al. (2018) incorporating ELMo Peters et al. (2018), naïvely adding distant data actually hurts performance. However, when our learned denoising model is applied to the data, performance improves, and it improves more than heuristic denoising approaches tailored to this dataset. Our strongest denoising model gives a gain of 3 F1 absolute over the ELMo baseline, and a 4.4 F1 improvement over naive incorporation of distant data. This establishes a new state-of-the-art on the test set, outperforming concurrently published work Xiong et al. (2019) and matching the performance of a BERT model Devlin et al. (2018) on this task. Finally, we show that denoising helps even when the label set is projected onto the OntoNotes label set Hovy et al. (2006); Gillick et al. (2014), outperforming the method of Choi et al. (2018) in that setting as well.
Setup
We consider the task of predicting a structured target associated with an input . Suppose we have high-quality labeled data of (input, target) pairs , and noisily labeled data of (input, target) pairs . For our tasks, is collected through manual annotation and is collected by distant supervision. We use two models to denoise data from : a filtering function disposes of unusable data (e.g., mislabeled examples) and a relabeling function transforms the noisy target labels to look more like true labels. This transformation improves the noisy data so that we can use it to without introducing damaging amounts of noise. In the second stage, a classification model is trained on the augmented data ( combined with denoised ) and predicts given in the inference phase.
The primary task we address here is the fine-grained entity typing task of Choi et al. (2018). Instances in the corpus are assigned types from a vocabulary of more than 10,000 types, which are divided into three classes: general types, fine-grained types, and ultra-fine types. This dataset consists of 6K manually annotated examples and approximately 25M distantly-labeled examples. 5M examples are collected using entity linking (EL) to link mentions to Wikipedia and gather types from information on the linked pages. 20M examples (HEAD) are generated by extracting nominal head words from raw text and treating these as singular type labels.
Figure 1 shows examples from these datasets which illustrate the challenges in automatic annotation using distant supervision. The manually-annotated example in (a) shows how numerous the gold-standard labeled types are. By contrast, the HEAD example (b) shows that simply treating the head word as the type label, while correct in this case, misses many valid types, including more general types. The EL example (c) is incorrectly annotated as region, whereas the correct coarse type is actually person. This error is characteristic of entity linking-based distant supervision since identifying the correct link is a challenging problem in and of itself Milne and Witten (2008): in this case, Gascoyne is also the name of a region in Western Australia. The EL example in (d) has reasonable types; however, human annotators could choose more types (grayed out) to describe the mention more precisely. The average number of types annotated by humans is per example while the two distant supervision techniques combined yields types per example on average.
In summary, distant supervision can (1) produce completely incorrect types, and (2) systematically miss certain types.
Denoising Model
To handle the noisy data, we propose to learn a denoising model as shown in Figure 2. This denoising model consists of filtering and relabeling functions to discard and relabel examples, respectively; these rely on a shared mention encoder and type encoder, which we describe in the following sections. The filtering function is a binary classifier that takes these encoded representations and predicts whether the example is good or bad. The relabeling function predicts a new set of labels for the given example.
We learn these functions in a supervised fashion. Training data for each is created through synthetic noising processes applied to the manually-labeled data, as described in Sections 3.3 and 3.4.
For the entity typing task, each example takes the form , where is the sentence, is the mention span, and is the set of types (either clean or noisy).
This encoder is a function which maps a sentence and mention to a real-valued vector . This allows the filtering and relabeling function to recognize inconsistencies between the given example and the provided types. Note that these inputs and are the same as the inputs for the supervised version of this task; we can therefore share an encoder architecture between our denoising model and our final typing model. We use an encoder following Choi et al. (2018) with a few key differences, which are described in Section 4.
2 Type Encoder
The second component of our model is a module which produces a vector . This is an encoder of an unordered bag of types. Our basic type encoder uses trainable vectors as embeddings for each type and combines these with summing. That is, the noisy types are embedded into type vectors . The final embedding of the type set .
Using trainable type embeddings exposes the denoising model to potential data sparsity issues, as some types appear only a few or zero times in the training data. Therefore, we also assign each type a vector based on its definition in WordNet Miller (1995). Even low-frequent types are therefore assigned a plausible embedding.We found this technique to be more effective than using pretrained vectors from GloVe or ELMo. It gave small improvements on an intrinsic evaluation over not incorporating it; results are omitted due to space constraints.
Let denote the th word of the th type’s most common WordNet definition. Each is embedded using GloVe Pennington et al. (2014). The resulting word embedding vectors are fed into a bi-LSTM Hochreiter and Schmidhuber (1997); Graves and Schmidhuber (2005), and a concatenation of the last hidden states in both directions is used as the definition representation . The final representation of the definitions is the sum over these vectors for each type: .
Our final , the concatenation of the type and definition embedding vectors.
3 Filtering Function
The filtering function is a binary classifier designed to detect examples that are completely mislabeled. Formally, is a function mapping a labeled example to a binary indicator of whether this example should be discarded or not.
In the forward computation, the feature vectors and are computed using the mention and type encoders. The model prediction is defined as , where is a sigmoid function, is a parameter vector, and is a 1-layer highway network Srivastava et al. (2015). We can apply to each distant pair in our distant dataset and discard any example predicted to be erroneous ().
We do not know a priori which examples in the distant data should be discarded, and labeling these is expensive. We therefore construct synthetic training data for based on the manually labeled data . For 30% of the examples in , we replace the gold types for that example with non-overlapping types taken from another example. The intuition for this procedure follows Figure 1: we want to learn to detect examples in the distant data like Gascoyne where heuristics like entity resolution have misfired and given a totally wrong label set.
Formally, for each selected example , we repeatedly draw another example from until we find that does not have any common types with . We then create a positive training example . We create a negative training example using the remaining of examples. is trained on using binary cross-entropy loss.
4 Relabeling Function
We train the relabeling function on another synthetically-noised dataset generated from the manually-labeled data . To mimic the type distribution of the distantly-labeled examples, we take each example and randomly drop each type with a fixed rate independent of other types to produce a new type set . We perform this process for all examples in and create a noised training set , where a single training example is . is trained on with a binary classification loss function over types used in Choi et al. (2018), described in the next section.
Typing Model
Figure 3 outlines the overall architecture of our typing model. The encoder consists of four vectors: a sentence representation , a word-level mention representation , a character-level mention representation , and a headword mention vector . The first three of these were employed by Choi et al. (2018). We have modified the mention encoder with an additional bi-LSTM to better encode long mentions, and additionally used the headword embedding directly in order to focus on the most critical word. These pieces use pretrained contextualized word embeddings (ELMo) Peters et al. (2018) as input.
To obtain a mention representation, we use both word and character information. For the word-level representation, the mention’s contextualized word vectors are fed into a bi-LSTM with hidden dimension is . The concatenated hidden states of both directions are summed by a span attention layer to form the word-level mention representation: .
Second, a character-level representation is computed for the mention. Each character is embedded and then a 1-D convolution Collobert et al. (2011) is applied over the characters of the mention. This gives a character vector .
Finally, we take the contextualized word vector of the headword as a third component of our representation. This can be seen as a residual connection He et al. (2016) specific to the mention head word. We find the headwords in the mention spans by parsing those spans in isolation using the spaCy dependency parser Honnibal and Johnson (2015). Empirically, we found this to be useful on long spans, when the span attention would often focus on incorrect tokens.
We use the same loss function as Choi et al. (2018) for training. This loss partitions the labels in general, fine, and ultra-fine classes, and only treats an instance as an example for types of the class in question if it contains a label for that class. More precisely:
Note that this loss function already partially repairs the noise in distant examples from missing labels: for example, it means that examples from HEAD do not count as negative examples for general types when these are not present. However, we show in the next section that this is not sufficient for denoising.
The settings of hyperparameters in our model largely follows Choi et al. (2018) and recommendations for using the pretrained ELMo-Small model.https://allennlp.org/elmo The word embedding size is . The type embedding size and the type definition embedding size are set to . For most of other model hyperparameters, we use the same settings as Choi et al. (2018): , , . The number of filters in the 1-d convolutional layer is . Dropout is applied with for the pretrained embeddings, and for the mention representations. We limit sentences to 50 words and mention spans to 20 words for computational reasons. The character CNN input is limited to characters; most mentions are short, so this still captures subword information in most cases. The batch size is set to . For all experiments, we use the Adam optimizer Kingma and Ba (2014). The initial learning rate is set to 2e-03. We implement all modelsThe code for experiments is available at https://github.com/yasumasaonoe/DenoiseET using PyTorch. To use ELMo, we consult the AllenNLP source code.
Experiments
We evaluate our approach on the ultra-fine entity typing dataset from Choi et al. (2018). The 6K manually-annotated English examples are equally split into the training, development, and test examples by the authors of the dataset. We generate synthetically-noised data, and , using the 2K training set to train the filtering and relabeling functions, and . We randomly select 1M EL and 1M HEAD examples and use them as the noisy data . Our augmented training data is a combination of the manually-annotated data and .
In addition, we investigate if denoising leads to better performance on another dataset. We use the English OntoNotes dataset Gillick et al. (2014), which is a widely used benchmark for fine-grained entity typing systems. The original training, development, and test splits contain 250K, 2K, and 9K examples respectively. Choi et al. (2018) created an augmented training set that has 3.4M examples. We also construct our own augmented training sets with/without denoising using our noisy data , using the same label mapping from ultra-fine types to OntoNotes types described in Choi et al. (2018).
1 Ultra-Fine Typing Results
We first compare the performance of our approach to several benchmark systems, then break down the improvements in more detail. We use the model architecture described in Section 4 and train it on the different amounts of data: manually labeled only, naive augmentation (adding in the raw distant data), and denoised augmentation. We compare our model to Choi et al. (2018) as well as to BERT Devlin et al. (2018), which we fine-tuned for this task. We adapt our task to BERT by forming an input sequence ”[CLS] sentence [SEP] mention [SEP]” and assign the segment embedding A to the sentence and B to the mention span.We investigated several approaches, including taking the head word piece from the last layer and using that for classification (more closely analogous to what Devlin et al. (2018) did for NER), but found this one to work best. Then, we take the output vector at the position of the [CLS] token (i.e., the first token) as the feature vector v, analogous to the usage for sentence pair classification tasks. The BERT model is fine-tuned on the 2K manually annotated examples. We use the pretrained BERT-Base, uncased modelhttps://github.com/google-research/bert with a step size of 2e-05 and batch size 32.
Table 1 compares the performance of these systems on the development set. Our model with no augmentation already matches the system of Choi et al. (2018) with augmentation, and incorporating ELMo gives further gains on both precision and recall. On top of this model, adding the distantly-annotated data lowers the performance; the loss function-based approach of Choi et al. (2018) does not sufficiently mitigate the noise in this data. However, denoising makes the distantly-annotated data useful, improving recall by a substantial margin especially in the general class. A possible reason for this is that the relabeling function tends to add more general types given finer types. BERT performs similarly to ELMo with denoised distant data. As can be seen in the performance breakdown, BERT gains from improvements in recall in the fine class.
Table 2 shows the performance of all settings on the test set, with the same trend as the performance on the development set. Our approach outperforms the concurrently-published Xiong et al. (2019); however, that work does not use ELMo. Their improved model could be used for both denoising as well as prediction in our setting, and we believe this would stack with our approach.
Our model with ELMo trained on denoised data matches the performance of the BERT model. We experimented with incorporating distant data (raw and denoised) in BERT, but the fragility of BERT made it hard to incorporate: training for longer generally caused performance to go down after a while, so the model cannot exploit large external data as effectively. Devlin et al. (2018) prescribe training with a small batch size and very specific step sizes, and we found the model very sensitive to these hyperparameters, with only 2e-05 giving strong results. The ELMo paradigm of incorporating these as features is much more flexible and modular in this setting. Finally, we note that our approach could use BERT for denoising as well, but this did not work better than our current approach. Adapting BERT to leverage distant data effectively is left for future work.
1.1 Comparing Denoising Models
We now explicitly compare our denoising approach to several baselines. For each denoising method, we create the denoised EL, HEAD, and EL & HEAD dataset and investigate performance on these datasets. Any denoised dataset is combined with the 2K manually-annotated examples and used to train the final model.
These heuristics target the same factors as our filtering and relabeling functions in a non-learned way.
For each type observed in the distant data, we add its synonyms and hypernyms using WordNet Miller (1995). This is motivated by the data construction process in Choi et al. (2018).
We use type pair statistics in the manually labeled training data. For each base type that we observe in a distant example, we add any type which is seen more than 90% of the time the base type occurs. For instance, the type art is given at least 90% of the times the film type is present, so we automatically add art whenever film is observed.
Table 3 compares the results on the development set. We report the performance on each of the EL & HEAD, EL, and HEAD dataset. On top of the baseline Original, adding synonyms and hypernyms by consulting external knowledge does not improve the performance. Expanding labels with the Pair technique results in small gains over Original. Overlap is the most effective heuristic technique. This simple filtering and expansion heuristic improves recall on EL. Filter, our model-based example selector, gives similar improvements to Pair and Overlap on the HEAD setting, where filtering noisy data appears to be somewhat important.One possible reason for this is identifying stray word senses; film can refer to the physical photosensitive object, among other things. Relabel and Overlap both improve performance on both EL and HEAD while other methods do poorly on EL. Combining the two model-based denoising techniques, Filter & Relabel outperforms all the baselines.
2 OntoNotes Results
We compare our different augmentation schemes for deriving data for the OntoNotes standard as well. Table 4 lists the results on the OntoNotes test set following the adaptation setting of Choi et al. (2018). Even on this dataset, denoising significantly improves over naive incorporation of distant data, showing that the denoising approach is not just learning quirks of the ultra-fine dataset. Our augmented set is constructed from 2M seed examples while Choi et al. (2018) have a more complex procedure for deriving augmented data from 25M examples. Ours (total size of 2.1M) is on par with their larger data (total size of 3.4M), despite having 40% fewer examples. In this setting, BERT still performs well but not as well as our model with augmented training data.
One source of our improvements from data augmentation comes from additional data that is able to be used because some OntoNotes type can be derived. This is due to denoising doing a better job of providing correct general types. In the EL setting, this yields 730k usable examples out of 1M (vs 540K for no denoising), and in HEAD, 640K out of 1M (vs. 73K).
3 Analysis of Denoised Labels
To understand what our denoising approach does to the distant data, we analyze the behavior of our filtering and relabeling functions. Table 5 reports the average numbers of types added/deleted by the relabeling function and the ratio of examples discarded by the filtering function.
Overall, the relabeling function tends to add more and delete fewer number of types. The HEAD examples have more general types added than the EL examples since the noisy HEAD labels are typically finer. Fine-grained types are added to both EL and HEAD examples less frequently. Ultra-fine examples are frequently added to both datasets, with more added to EL; the noisy EL labels are mostly extracted from Wikipedia definitions, so those labels often do not include ultra-fine types. The filtering function discards similar numbers of examples for the EL and HEAD data: and respectively.
Figure 4 shows examples of the original noisy labels and the denoised labels produced by the relabeling function. In example (a), taken from the EL data, the original labels, {location, city}, are correct, but human annotators might choose more types for the mention span, Minneapolis. The relabeling function retains the original types about the geography and adds ultra-fine types about administrative units such as {township, municipality}. In example (b), from the HEAD data, the original label, {dollar}, is not so expressive by itself since it is a name of a currency. The labeling function adds coarse types, {object, currency}, as well as specific types such as {medium of exchange, monetary unit}. In another EL example (c), the relabeling function tries to add coarse and fine types but struggles to assign multiple diverse ultra-fine types to the mention span Michelangelo, possibly because some of these types rarely cooccur (painter and poet).
Related Work
Past work on denoising data for entity typing has used multi-instance multi-label learning Yaghoobzadeh and Schütze (2015, 2017); Murty et al. (2018). One view of these approaches is that they delete noisily-introduced labels, but they cannot add them, or filter bad examples. Other work focuses on learning type embeddings Yogatama et al. (2015); Ren et al. (2016a, b); our approach goes beyond this in treating the label set in a structured way. The label set of Choi et al. (2018) is distinct in not being explicitly hierarchical, making past hierarchical approaches difficult to apply.
Denoising techniques for distant supervision have been applied extensively to relation extraction. Here, multi-instance learning and probabilistic graphical modeling approaches have been used Riedel et al. (2010); Hoffmann et al. (2011); Surdeanu et al. (2012); Takamatsu et al. (2012) as well as deep models Lin et al. (2016); Feng et al. (2017); Luo et al. (2017); Lei et al. (2018); Han et al. (2018), though these often focus on incorporating signals from other sources as opposed to manually labeled data.
Conclusion
In this work, we investigated the problem of denoising distant data for entity typing tasks. We trained a filtering function that discards examples from the distantly labeled data that are wholly unusable and a relabeling function that repairs noisy labels for the retained examples. When distant data is processed with our best denoising model, our final trained model achieves state-of-the-art performance on an ultra-fine entity typing task.
Acknowledgments
This work was partially supported by NSF Grant IIS-1814522, NSF Grant SHF-1762299, a Bloomberg Data Science Grant, and an equipment grant from NVIDIA. The authors acknowledge the Texas Advanced Computing Center (TACC) at The University of Texas at Austin for providing HPC resources used to conduct this research. Results presented in this paper were obtained using the Chameleon testbed supported by the National Science Foundation. Thanks as well to the anonymous reviewers for their thoughtful comments, members of the UT TAUR lab and Pengxiang Cheng for helpful discussion, and Eunsol Choi for providing the full datasets and useful resources.