Few-Shot Named Entity Recognition: A Comprehensive Study

Jiaxin Huang, Chunyuan Li, Krishan Subudhi, Damien Jose, Shobana Balakrishnan, Weizhu Chen, Baolin Peng, Jianfeng Gao, Jiawei Han

Introduction

Named Entity Recognition (NER) involves processing unstructured text, locating and classifying named entities (certain occurrences of words or expressions) into particular categories of pre-defined entity types, such as persons, organizations, locations, medical codes, dates and quantities. NER serves as an important first component for tasks such as information extraction Ritter et al. (2012), information retrieval Guo et al. (2009), question answering Mollá et al. (2006), task-oriented dialogues Peng et al. (2020a); Gao et al. (2019) and other language understanding applications Nadeau and Sekine (2007); Shaalan (2014). Deep learning has shown remarkable success in NER in recent years, especially with self-supervised pre-trained language models (PLMs) such as BERT Devlin et al. (2019) and RoBERTa Liu et al. (2019c). State-of-the-art (SoTA) NER models are often initialized with PLM weights, fine-tuned with standard supervised learning. One classical approach is to add a linear classifier on top of representations provided by PLMs, and fine-tune the entire model with a cross-entropy objective on domain labels Devlin et al. (2019). Desipite its simplicity, the approach provides strong results on several benchmarks and is served as baseline in this study.

Unfortunately, even with these PLMs, building NER systems remains a labor-intensive, time-consuming task. It requires rich domain knowledge and expert experience to annotate a large corpus of in-domain labeled tokens to make the models work well. However, this is in contrast to the real-world application scenarios, where only very limited amounts of labeled data are available for new domains. For example, a new customer would prefer to provide very few labeled examples for specific domains in cloud-based NER services. The cost of building NER systems at scale with rich annotations (i.e., hundreds of different enterprise use-cases/domains) can be prohibitively expensive. This draws attentions to a challenging but practical research problem: few-shot NER.

To deal with the challenge of few-shot learning, we focus on improving the generalization ability of PLMs for NER from three complementary directions, shown in Figure 1. Instead of limiting ourselves in making use of limited in-domain labeled tokens with the classical approach, (i)(\textup{\it i}) we create prototypes as the representations for different entity types, and assign labels via the nearest neighbor criterion; (ii)(\textup{\it ii}) we continuously pre-train PLMs using web data with noisy labels that is available in much larger quantities to improve NER accuracy and robustness; (iii)(\textup{\it iii}) we employ unlabeled in-domain tokens to predict their soft labels using self-training, and perform semi-supervised learning in conjunction with the limited labeled data.

Our contributions include: (i)(\textup{\it i}) We present the first systematic study for few-shot NER, a problem that is previously little explored in the literature. Three distinctive schemes and their combinations are investigated. (ii)(\textup{\it ii}) We perform comprehensive comparisons of these schemes on 10 public NER datasets from different domains. (iii)(\textup{\it iii}) Compared with existing methods on few-shot and training-free NER settings , the proposed schemes achieve SoTA performance despite their simplicity. To shed light on future research on few-shot NER, our study suggests that: (i)(\textup{\it i}) Noisy supervised pre-training can significantly improve NER accuracy, and we will release our pre-trained checkpoints. (ii)(\textup{\it ii}) Self-training consistently improves few-shot learning when the ratio of data amounts between unlabeled and labeled data is high. (iii)(\textup{\it iii}) The performance of prototype learning varies on different datasets. It is useful when the number of labeled examples is small, or when new entity types are given in the training-free settings.

Background on Few-shot NER

a sequence labeling task, where the input is a text sequence (e.g., sentence) of length TT, X=[x1,x2,...,xT]{{\bf X}}=[\boldsymbol{x}_{1},\boldsymbol{x}_{2},...,\boldsymbol{x}_{T}], and the output is a corresponding TT-length labeling sequence Y=[y1,y2,...,yT]{\bf Y}=[\boldsymbol{y}_{1},\boldsymbol{y}_{2},...,\boldsymbol{y}_{T}], y∈Y\boldsymbol{y}\in\mathcal{Y} is a one-hot vector indicating the entity type of each token from a pre-defined discrete label space. The training dataset for NER often consists of pair-wise data DL={(Xn,Yn)}n=1N\mathcal{D}^{\mathtt{L}}=\{({{\bf X}}_{n},{\bf Y}_{n})\}_{n=1}^{N}, where NN is the number of training examples. Traditional NER systems are trained in the standard supervised learning paradigms, which usually requires a large number of pairwise examples, i.e., NN is large. In real-world applications, the more favorable scenarios are that only a small number of labeled examples are given for each entity type (NN is small), because expanding labeled data increases annotation cost and decreases customer engagement. This yields a challenging task few-shot NER.

Linear Classifier Fine-tuning.

Following the recent self-supervised PLMs Devlin et al. (2019); Liu et al. (2019c), a typical method for NER is to utilize a Transformer-based backbone network to extract the contextualized representation of each token z=fθ0(x)\boldsymbol{z}=f_{\boldsymbol{\theta}_{0}}(\boldsymbol{x}) . A linear classifier (i.e., a linear layer with parameter θ1={W,b}\boldsymbol{\theta}_{1}=\{{{\bf W}},{\boldsymbol{b}}\} followed by a Softmax layer) is applied to project the representation z\boldsymbol{z} into the label space fθ1(z)=Softmax(Wz+b)f_{\boldsymbol{\theta}_{1}}(\boldsymbol{z})=\text{Softmax}({{\bf W}}\boldsymbol{z}+{\boldsymbol{b}}). In another word, the end-to-end learning objective for linear classifier based NER can be obtained via a function composition y=fθ1∘fθ0(x)\boldsymbol{y}=f_{\boldsymbol{\theta}_{1}}\circ f_{\boldsymbol{\theta}_{0}}(\boldsymbol{x}), with trainable parameters θ={θ0,θ1}\boldsymbol{\theta}=\{\boldsymbol{\theta}_{0},\boldsymbol{\theta}_{1}\}. The pipeline is shown in Figure 2(a). The model is optimized by minimizing the cross-entropy:

In practice, θ1={W,b}\boldsymbol{\theta}_{1}=\{{{\bf W}},{\boldsymbol{b}}\} is always updated, while θ0\boldsymbol{\theta}_{0} can be either frozen Liu et al. (2019a, b); Jie and Lu (2019) or updated Devlin et al. (2019); Yang and Katiyar (2020).

Methods

When only a small number of labeled tokens are available, it renders difficulties for the classical supervised fine-tuning approach: the model tends to over-fit the training examples and shows poor generalization performance on the testing set Fritzler et al. (2019). In this paper, we provide a comprehensive study specifically for limited NER data settings, and explore three orthogonal directions shown in Figure 1: (i)(\textup{\it i}) How to adapt meta-learning such as prototype-based methods for few-shot NER? (ii)(\textup{\it ii}) How to leverage freely-available web data as noisy supervised pre-training data? (iii)(\textup{\it iii}) How to leverage unlabeled in-domain sentences in a semi-supervised manner? Note that these three directions are complementary to each other, can be further used jointly to extrapolate the methodology space in Figure 1.

Meta-learning Ravi and Larochelle (2017) have shown promising results for few-shot image classification Tian et al. (2020) and sentence classification Yu et al. (2018); Geng et al. (2019). It is natural to adapt this idea to few-shot NER. The core idea is to use episodic classification paradigm to simulate few-shot settings during model training. Specifically in each episode, MM entity types (usually M<∣Y∣M<|\mathcal{Y}|) are randomly sampled from DL\mathcal{D}^{\mathtt{L}}, containing a support set S={(Xi,Yi}i=1M×K\mathcal{S}=\{({{\bf X}}_{i},{\bf Y}_{i}\}_{i=1}^{M\times K} (KK sentences per type) and a query set Q={(X^i,Y^i}i=1M×K′\mathcal{Q}=\{(\hat{{{\bf X}}}_{i},\hat{{\bf Y}}_{i}\}_{i=1}^{M\times K^{\prime}} (K′K^{\prime} sentences per type).

We build our method based on prototypical network Snell et al. (2017), which introduces the notion of prototypes, representing entity types as vectors in the same representation space of individual tokens. To construct the prototype for the mm-th entity type cm{\boldsymbol{c}}_{m}, the average of representations is computed for all tokens belonging to this type in the support set S\mathcal{S}:

where Sm\mathcal{S}_{m} is the tokens set of the mm-th type in S\mathcal{S}, and fθ0f_{\boldsymbol{\theta}_{0}} is defined in (2). For an input token x∈Q\boldsymbol{x}\in\mathcal{Q} from the query set, its prediction distribution is computed by a softmax function of the distance between x\boldsymbol{x} and all the entity prototypes. For example, the prediction probability for the mm-th prototype is:

2 Noisy Supervised Pre-training

Generic representations via self-supervised pre-trained language models Devlin et al. (2019); Liu et al. (2019c) have benefited a wide range of NLP applications. These models are pre-trained with the task of randomly masked token prediction on massive corpora, and are agnostic to the downstream tasks. In other words, PLMs treat each token equally, which is not aligned with the goal of NER: identifying named entities as emphasized tokens and assigning labels to them. For example, for a sentence “ Mr. Bush asked Congress to raise to \$ 6 billion ”, PLMs treat to and Congress equally, while NER aims to highlight entities like Congress and lowlight their collocated non-entity words like to.

This intuition inspires us to endow the backbone network an ability to outweigh the representations of entities for NER. Hence, we propose to employ the large-scale noisy web data WiNER\mathtt{WiNER} Ghaddar (2017) for noisy supervised pre-training (NSP). The labels in WiNER\mathtt{WiNER} are automatically annotated on the 2013 English Wikipedia dump by querying anchored strings as well as their coreference mentions in each wiki page to the Freebase. The WiNER\mathtt{WiNER} dataset is of 6.8GB and contains 113 entity types along with over 50 million sentences. Though introducing inevitable noises (e.g., a random subset of 1000 mentions are manually evaluated and the accuracy of automatic annotations reaches 77%, due to the error of identifying coreferences), this automatic annotation procedure is highly scalable and affordable. The label set of WiNER\mathtt{WiNER} covers a wide range of entity types. They are often related but different from entity types in the downstream datasets. For example in Figure 2(c), the entity types Musician and Artist in Wikipedia are more fine-grained than Person in a typical NER dataset. The proposed NSP learns representations to distinguish entities from others. This particularly favors the few-shot settings, preventing over-fitting via the prior knowledge of extracting entities from various contexts in pre-training.

Two pre-training objectives are considered in NSP, respectively: the first one is to use the linear classifier in (2), the other is a prototype-based objective in (4). For the linear classifier, we found that the batch size of 10241024 and learning rate of 1e−41e^{-4} works best, and for the prototype-based approach, we use the episodic training paradigm with M=5M=5 and set learning rate to be 5e−55e^{-5}. For both objectives, we train the whole corpus for 11 epoch and apply the Adam Optimizer Kingma and Ba (2015) with a linearly decaying schedule with warmup at 0.10.1. We empirically compare both objectives in experiments, and found that the linear classifier in (2) improves pre-training more significantly.

3 Self-training

Though manually labeling entities is expensive, it is easy to collect large amounts of unlabeled data in the target domain. Hence, it becomes desired to improve the model performance by effectively leveraging unlabeled data DU\mathcal{D}^{\mathtt{U}} with limited labeled data DL\mathcal{D}^{\mathtt{L}}. We resort to the recent self-training scheme Xie et al. (2020) for semi-supervised learning. The algorithm operates as follows:

Learn teacher model θtea\boldsymbol{\theta}^{\mathtt{tea}} via cross-entropy using (1) with labeled tokens DL\mathcal{D}^{\mathtt{L}}.

Generate soft labels using a teacher model on unlabeled tokens:

Learn a student model θstu\boldsymbol{\theta}^{\mathtt{stu}} via cross-entropy using (1) on labeled and unlabeled tokens:

where λU\lambda_{\mathtt{U}} is the weighting hyper-parameter.

A visual illustration for self-training procedure shown in Figure 2(d). It is optional to iterate from Step 1 to Step 3 multiple times, by initializing θtea\boldsymbol{\theta}^{\mathtt{tea}} in Step 1 with newly learned θstu\boldsymbol{\theta}^{\mathtt{stu}} in Step 3. We only perform self-training once in our experiments for simplicity, which has already shown excellent performance.

Related Work

NER is a long standing problem in NLP. Deep learning has significantly improve the recognition accuracy. Early efforts include exploring various neural architectures Lample et al. (2016) such as Bidrectional LSTMs Chiu and Nichols (2016) and adding CRFs to capture structures Ma and Hovy (2016). Early studies have noticed the importance of reducing the annotation labor, where semi-supervised learning is employed, such as clustering Lin and Wu (2009), and combining supervised objective with unsupervised word representations Turian et al. (2010). PLMs have recently revolutionized NER, where large-scale Transformer-based architectures Peters et al. (2018); Devlin et al. (2019) are used as backbone network to extract informative representations. Contextualized string embedding Akbik et al. (2018) is proposed to capture subword structures and polysemous words in different usage. Masked words and entities are jointly trained for prediction in Yamada et al. (2020) with entity-aware self-attention. These methods are designed for standard supervised learning, and have a limited generalization ability in few-shot settings, as empirically shown in Fritzler et al. (2019).

Prototype-based methods

recently become popular few-shot learning approaches in machine learning community. It was firstly studied in the context of image classification Vinyals et al. (2016); Sung et al. (2018); Zhao et al. (2020), and has recently been adapted to different NLP tasks such as text classification Wang et al. (2018); Geng et al. (2019); Bansal et al. (2020), machine translation Gu et al. (2018) and relation classification Han et al. (2018). The closest related works to ours is Fritzler et al. (2019) which explores prototypical network on few-shot NER, but only utilizes RNNs as the backbone model and does not leverage the power of large-scale Transformer-based architectures for word representations. Our work is similar to Ziyadi et al. (2020); Wiseman and Stratos (2019) in that all of them utilize the nearest neighbor criterion to assign the entity type, but differs in that Ziyadi et al. (2020); Wiseman and Stratos (2019) consider every individual token instance for nearest neighbor comparison, while ours considers prototypes for comparison. Hence, our method is much more scalable when the number of given examples increases.

Supervised pre-training.

In computer vision, it is a de facto standard to transfer ImageNet-supervised pre-trained models to small image datasets to pursue high recognition accuracy Yosinski et al. (2014). The recent work named big transfer Kolesnikov et al. (2019) has achieved SoTA on various vision tasks via pre-training on billions of noisily labeled web images. To gain a stronger transfer learning ability, one may combine supervised and self-supervised methods Li et al. (2020c, b). In NLP, supervised/grounded pre-training have been recently explored for natural language generation (NLG) Keskar et al. (2019); Zellers et al. (2019); Peng et al. (2020b); Gao et al. (2020); Li et al. (2020a). They aim to endow GPT-2 Radford et al. , an ability of enabling high-level semantic controlling in language generation, and are often pre-trained on massive corpus consisting of text sequences associated with prescribed codes such as text style, content description, and task-specific behavior. In contrast to NLG, to our best knowledge, large-scale supervised pre-training has been little studied for natural language understanding (NLU). There are early works showing promising results by transferring from medium-sized datasets to small datasets in some NLU applications; For example, from MNLI to RTE for sentence classification Phang et al. (2018); Clark et al. (2020); An et al. (2020), and from OntoNER to CoNLL for NER Yang and Katiyar (2020). Our work further increases the supervised pre-training at the scale of web data Ghaddar (2017), 1000 orders of magnitude larger than Yang and Katiyar (2020), showing consistent improvements.

Self-training.

Self-training Scudder (1965) is one of the earliest semi-supervised methods, and has recently achieved improved performance for tasks such as ImageNet classification Xie et al. (2020), visual object detection Zoph et al. (2020), neural machine translation He et al. (2019), sentence classification Mukherjee and Awadallah (2020); Du et al. (2020). It is shown via object detection tasks in Zoph et al. (2020) that stronger data augmentation and more labeled data can diminish the value of pre-training, while self-training is always helpful in both low-data and high-data regimes. Our work presents the first study of self-training for NER, and we observe similar phenomenons: it consistently boosts few-shot learning performance across all 10 datasets.

Experiments

In this section, we first compare the performance of different combinations of the three schemes on 10 benchmark datasets with various proportions of training data, and then compare our approaches with SoTA methods proposed for settings of few-shot learning and immediate inference for unseen entity types.

Throughout our experiments, the pre-trained base RoBERTa model is employed as the backbone network. We investigate the following 6 schemes for the comparative study: (i)(\textup{\it i}) LC is the linear classifier fine-tuning method in Section 2, i.e., adding a linear classifier on the backbone, and directly fine-tuning on entire model on the target dataset; (ii)(\textup{\it ii}) P indicates the prototype-based method in Section 3.1; (iii)(\textup{\it iii}) NSP refers to the noisy supervised pre-training in Section 3.2; Depending on the pre-training objective, we have LC+NSP and P+NSP. (iv)(\textup{\it iv}) ST is the self-training approach in Section 3.3, it is combined with linear classifier fine-tuning, denoted as LC+ST; (v)(\textup{\it v}) LC+NSP+ST.

Datasets.

We evaluate our methods on 10 public benchmark datasets, covering a wide range of domains: OntoNotes 5.0 Ralph et al. (2013), WikiGold https://github.com/juand-r/entity-recognition-datasets Balasuriya et al. (2009) on general domain, CoNLL 2003 Sang and Meulder (2003) on news domain, WNUT 2017 Derczynski et al. (2017) on social domain, MIT Moive Liu et al. (2013b) and MIT Restaurant https://groups.csail.mit.edu/sls/downloads/ Liu et al. (2013a) on review domain, SNIPShttps://github.com/snipsco/nlu-benchmark/tree/master/2017-06-custom-intent-engines Coucke et al. (2018), ATIShttps://github.com/yvchen/JointSLU Hakkani-Tür et al. (2016) and Multiwozhttps://github.com/budzianowski/multiwoz Budzianowski et al. (2018) on dialogue domain, and I2B2https://portal.dbmi.hms.harvard.edu/projects/n2c2-2014/ Stubbs and Uzuner (2015) on medical domain. The detailed statistics of these datasets are summarized in Table 1.

For each dataset, we conduct three sets of experiments using various proportions of the training data: 5-shot, 10% and 100%. For 5-shot setting, we sample 5 sentences for each entity type in the training set and repeat each experiment for 10 times. For 10% setting, we down-sample 10 percent of the training set, and for 100% setting, we use the full training set as labeled data. We only study the self-training method in 5-shot and 10% settings, by using the rest of the training set as unlabeled in-domain corpus.

Hyper-parameters.

We have described details for noisy supervised pre-training in Section 3.2. For training on target datasets, we set a fixed set of hyperparameters across all the datasets: For the linear classifier, we set batch size = 1616 for 100% and 10% settings, batch size = 44 for 5-shot setting. For each episode in the prototype-based method, we set the number of sentences per entity type in support and query set (K,K′)(K,K^{\prime}) to be (5,15)(5,15) for 100% and 10% settings, and (2,3)(2,3) for 5-shot setting. For both training objectives, we set learning rate = 5e−55e^{-5} for 100% and 10% settings, and learning rate = 1e−41e^{-4} for 5-shot setting. For all training data sizes, we set training epoch = 1010, and Adam optimizer Kingma and Ba (2015) is used with the same linear decaying schedule as the pre-training stage. For self-training, we set λU=0.5.\lambda_{\mathtt{U}}=0.5.

Evaluation.

We follow the standard protocols for NER tasks to evaluate the performance on the test set Sang and Meulder (2003). Since RoBERTa tokenizes each word into subwords, we generate word-level predictions based on the first word piece of a word. Word-level predictions are then turned into entity-level predictions for evaluation when calculating the f1-score. Two tagging schemas are typically considered to encode chunks of tokens into entities: BIO schema marks the beginning token of an entity as B-X and the consecutive tokens as I-X, and other tokens are marked as O. IO schema uses I-X to mark all tokens inside an entity, thus is more defective as there is no boundary tag. In our study, we use BIO schema by default, but report the performance evaluated by IO schema for fair comparison with some previous studies.

2 Comprehensive Comparison Results

To gain thorough insights and benchmark few-shot NER, we first perform an extensive comparative study on 6 methods across 10 datasets. The results are shown in Table 2. We can draw the following major conclusions: (i)(\textup{\it i}) By comparing column 1 and 2 (or comparing 3 and 4), it clearly shows that noisy supervised pre-training provides better results in most datasets, especially in the 5-shot setting, which demonstrates that NSP endows the model an ability to extract better NER-related features. (ii)(\textup{\it ii}) The comparison between column 1 and 3 provides a head-to-head comparison between linear classifier and prototype-based methods: while the prototype-based method demonstrates better performance than LC on CoNLL, WikiGold, WNUT17 and Multiwoz in the 5-shot learning setting, it falls behind LC on other datasets and in average statistics. It shows that the prototype-based method only yields better results when there is very limited labeled data: the size of both entity types and examples are small. (iii)(\textup{\it iii}) When comparing column 5 with 1 (or comparing column 6 and 2), we observe that using self-training consistently works better than directly fine-tuning with labeled data only, suggesting that ST is a useful technique to leverage in-domain unlabeled data if allowed. (iv)(\textup{\it iv}) Column 6 shows the highest F1-score in most cases, demonstrating the three proposed schemes in this paper are complementary to each other, and can be combined to yield best results in practice.

3 Comparison with SoTA Methods

The current SoTA on few-shot NER includes: (i)(\textup{\it i}) StructShot Yang and Katiyar (2020), which extends the nearest neighbor classification with a decoding process using abstract tag transition distribution. Both the model and the transition distribution are trained from the source dataset OntoNotes. (ii)(\textup{\it ii}) L-TapNet+CDT Hou et al. (2020) is a slot tagging method which constructs an embedding projection space using label name semantics to well separate different classes. It also includes a collapsed dependency transfer mechanism to transfer label dependency information from source domains to target domains. (iii)(\textup{\it iii}) SimBERT is a simple baseline reported in Yang and Katiyar (2020); Hou et al. (2020); it utilizes a nearest neighbor classifier based on the contextualized representation output by the pre-trained BERT, without fine-tuning on few-shot examples. The results reported in the StructShot paper use IO schema instead of BIO schema, thus we report our performance on both for completeness.

For fair comparison, following Yang and Katiyar (2020), we also continuously pre-train our model on OntoNotes after the noisy supervised pre-training stage. For each 5-shot learning task, we repeat the experiments 10 times by re-sampling few-shot examples each time. The results are reported in Table 3. We observe that our proposed methods consistently outperform the StructShot model across all three datasets, even by simply pre-training the model on large-scale noisily tagged datasets like Wikipedia. Our best model outperforms the previous SoTA by 8% F1-score, which demonstrates that using large amounts of unlabeled in-domain corpus is promising for enhancing the few-shot NER performance.

4 Training-free Method Comparison

Some real-world applications require immediate inference on unseen entity types. For example, novel entity types with a few examples are frequently given in an online fashion, but updating model weights θ\boldsymbol{\theta} frequently is prohibitive. One may store some token examples as supports and utilize them for nearest neighbor classification. The setting is referred to as training-free in Wiseman and Stratos (2019); Ziyadi et al. (2020), as the models identify new entities in a completely unseen target domain using only a few supporting examples in this new domain, without updating θ\boldsymbol{\theta} in that target domain. Our prototype-based method is able to perform such immediate inference. Two recent works on training-free NER are: (i)(\textup{\it i}) Neighbor-tagging Wiseman and Stratos (2019) copies token-level labels from weighted nearest neighbors; (ii)(\textup{\it ii}) Example-based NER Ziyadi et al. (2020) is the SoTA on training-free NER, which identifies the starting and ending tokens of unseen entity types.

We observed that our basic prototype-based method, under the training-free setting, does not gain from more given examples. We hypothesize that this is because tokens belonging to the same entity type are not necessarily close to each other, and are often separated in the representation space. Though it is hard to find one single centroid for all tokens in the same type, we assume that there exist local clusters of tokens belonging to the same type. To resolve such issue, we follow Deng et al. (2020) and extend our method to a version called Multi-Prototype, by creating K/5K/5 prototypes for each type given KK examples per type. (e.g., 2 prototypes per class are used for the 10-shot setting). The prediction score for a testing token belonging to a type is computed via averaging the prediction probabilities from all prototypes of the same type.

We compare with previous methods in Table 4. We observe that multi-prototype methods not only benefit from more support examples, but also surpass neighbor tagging methods and example-based NER by a large margin on two out of three datasets. For the MIT Movie dataset, one entity type can span a large chunk with multiple consecutive words in a sentence, which favors the span-based method like Ziyadi et al. (2020). For example, the underlined part in the sentence “what movie does the quote i dont think we are in kansas anymore come from” is annotated as entity type Quote. The proposed methods in this paper can be combined with the span-based approach to specifically tackle this problem, and we leave it as future work. Further, if slightly fine-tuning is allowed, we see that the prototype-based method achieves 0.438 with 5-shot learning in Table 2, better than 0.395 achieved by example-based NER given 500 examples.

Conclusion

We have presented a comprehensive study on few-shot NER. Three foundational methods and their combinations are systematically investigated: prototype-based methods, noisy supervised pre-training and self-training. They are intensively compared on 10 public datasets under various settings. All of them can improve the PLM’s generalization ability when learning from a few labeled examples, among which supervised pre-training and self-training turn out to be particularly effective. The proposed schemes achieve SoTA on both few-shot and training-free settings compared with recently proposed methods. We will release our benchmarks and code for few-shot NER, and hope that it can inspire future research with more advanced methods to tackle this challenging and practical problem.

References