UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition

Wenxuan Zhou, Sheng Zhang, Yu Gu, Muhao Chen, Hoifung Poon

Introduction

Large language models (LLMs) such as ChatGPT (Ouyang et al., 2022; OpenAI, 2023) have demonstrated remarkable generalization capabilities, but they generally require prohibitive cost in training and inference. Moreover, in mission-critical applications such as biomedicine, white-box access to model weights and inference probabilities are often important for explainability and trust. Consequently, instruction-tuning has become a popular approach for distilling LLMs into more cost-efficient and transparent student models. Such student models, as exemplified by Alpaca (Taori et al., 2023) and Vicuna (Chiang et al., 2023), have demonstrated compelling capabilities in imitating ChatGPT. However, upon close inspection, they still trail the teacher LLM by a large margin, especially in targeted downstream applications (Gudibande et al., 2023). Bounded by limited compute, it is unsurprising that generic distillation can only produce a shallow approximation of the original LLM across all possible applications.

In this paper, we instead explore targeted distillation where we train student models using mission-focused instruction tuning for a broad application class such as open information extraction (Etzioni et al., 2008). We show that this can maximally replicate LLM’s capabilities for the given application class, while preserving its generalizability across semantic types and domains. We choose named entity recognition (NER) for our case study, as it is one of the most fundamental tasks in natural language processing (Wu et al., 2017; Perera et al., 2020). Recent studies (Wei et al., 2023; Li et al., 2023) show that when there are abundant annotated examples for an entity type, LLMs still fall behind the state-of-the-art supervised system for that entity type. However, for the vast majority of entity types, there is little annotated data. New entity types constantly emerge, and it is expensive and time-consuming to generate annotated examples, especially in high-value domains such as biomedicine where specialized expertise is required for annotation. Trained on pre-specified entity types and domains, supervised NER models also exhibit limited generalizability for new domains and entity types.

We present a general recipe for targeted distillation from LLMs and demonstrate that for open-domain NER. We show how to use ChatGPT to generate instruction-tuning data for NER from broad-coverage unlabeled web text, and conduct instruction-tuning on LLaMA (Touvron et al., 2023a) to distill the UniversalNER models (UniNER in short).

To facilitate a thorough evaluation, we assemble the largest and most diverse NER benchmark to date (UniversalNER benchmark), comprising 43 datasets across 9 domains such as biomedicine, programming, social media, law, finance. On zero-shot NER, LLaMA and Alpaca perform poorly on this benchmark (close to zero F1). Vicuna performs much better by comparison, but still trails ChatGPT by over 20 absolute points in average F1. By contrast, UniversalNER attains state-of-the-art NER accuracy across tens of thousands of entity types in the UniversalNER benchmark, outperforming Vicuna by over 30 absolute points in average F1. With a tiny fraction of parameters, UniversalNER not only replicates ChatGPT’s capability in recognizing arbitrary entity types, but also outperforms its NER accuracy by 7-9 absolute points in average F1. Remarkably, UniversalNER even outperforms by a large margin state-of-the-art multi-task instruction-tuned systems such as InstructUIE (Wang et al., 2023a), which uses supervised NER examples. We also conduct thorough ablation studies to assess the impact of various distillation components, such as the instruction prompts and negative sampling.

Related Work

While LLMs such as ChatGPT achieve promising results, these models are often black-box and have high computational costs. To address these issues, distilling the task capabilities of LLMs into smaller, more manageable models has emerged as a promising direction. Knowledge distillation (Hinton et al., 2015) often revolves around the transfer of knowledge from larger, more complex models to their smaller counterparts. Recent work (Taori et al., 2023; Chiang et al., 2023; Peng et al., 2023) seeks to distill the general abilities of LLMs with the objective of matching, if not surpassing, the performance of the original LLMs. Particularly, Alpaca (Taori et al., 2023) automates the generation of instructions (Wang et al., 2023c) and distills the knowledge from a teacher LLM. Vicuna (Chiang et al., 2023) adopts the ShareGPT data, which are comprised of real conversations with ChatGPT conducted by users, thereby providing a more authentic context for distillation. Another line of work (Smith et al., 2022; Jung et al., 2023; Hsieh et al., 2023; Gu et al., 2023) focuses on distilling task-level abilities from LLMs. Particularly, Jung et al. (2023) propose an efficient method to distill an order of magnitude smaller model that outperforms GPT-3 on specialized tasks summarization and paraphrasing in certain domains. Hsieh et al. (2022) propose to distill LLMs’ reasoning abilities into smaller models by chain-of-the-thought distillation. However, these studies perform distillation either on certain datasets or domains, while our work focuses on a more general formulation that can be applied to diverse domains.

Instruction tuning.

As an effective method to adapt LMs to perform a variety of tasks, instruction tuning has attracted an increasing number of community efforts: FLAN (Chung et al., 2022), T0 (Sanh et al., 2021), and Tk-Instruct (Wang et al., 2022) convert a large set of existing supervised learning datasets into instruction-following format, and then fine-tune encoder-decoder models, showing strong zero-shot and few-shot performance on NLP benchmarks. Ouyang et al. (2022) crowd-source high-quality instruction data and fine-tune GPT-3 into InstructGPT, enhancing its ability to understand user intention and follow instructions. Recent advancements (Taori et al., 2023; Chiang et al., 2023; Peng et al., 2023) have also led to smaller models that exhibit task-following capabilities, after being fine-tuned on instruction data generated by LLMs, such as ChatGPT or GPT4. However, these smaller models often struggle to generate high-quality responses for a diverse range of tasks (Wang et al., 2023b). A closer examination on targeted benchmarks reveals a substantial gap between these models to ChatGPT (Gudibande et al., 2023). Our proposed method, in contrast, focuses on tuning models to excel at a specific type of tasks. The diversity in our instructing-tuning method comes from task labels (e.g., relation types for relation extraction, entity types for NER), rather than instructions. By focusing on task-level capabilities and using NER as a case study, we demonstrate that it is possible to devise a tuning recipe that not only closes the performance gap but also surpasses ChatGPT. Wang et al. (2023a) also explore instruction-tuning for information extraction tasks. However, their method relies solely on supervised datasets and yields subpar performance when compared to ChatGPT.

Mission-Focused Instruction Tuning

Instruction tuning (Ouyang et al., 2022; Wei et al., 2021) is a method through which pretrained autoregressive language models are finetuned to follow natural language instructions and generate responses. Existing work focuses on tuning models to do diverse tasks (Taori et al., 2023; Chiang et al., 2023). In contrast, we introduce a general recipe for mission-focused instruction tuning, where the pretrained model is tuned for a broad application class such as open information extraction.

In this paper, we conduct a case study on the NER task, as it is one of the fundamental tasks for knowledge extraction from text. The objective is to learn a model f:(X×T)→Yf:(\mathcal{X}\times\mathcal{T})\rightarrow\mathcal{Y}, where X\mathcal{X} represents the set of inputs, T\mathcal{T} denotes a predefined set of entity types, and Y\mathcal{Y} represents the set of entities of a specific type in the given input.

A typical instruction-tuning example is made of three parts, including instruction, input, and output, where the diversity of instruction causes the models to follow a wide range of task instructions. However, for mission-focused instruction tuning, our goal is to tune the model to maximally generalize across semantic types and domains for the targeted application class. Therefore, we focus on increasing the diversity of input rather than instruction.

While earlier work (Jung et al., 2023) employs language models to generate inputs, these models typically assume that the domains of test data are known and prompt LMs to generate data for each domain. This method falls short when applied to distillation for a broad application class, where the distribution of test data is unknown. Consequently, it is challenging to generate inputs from LMs that provide wide coverage of the test domains.

To address this limitation, we propose an alternative: directly sampling inputs from a large corpus across diverse domains, and then using an LLM to generate outputs. In this paper, we sample inputs from the Pile corpus (Gao et al., 2020), which compiles 22 distinct English sub-datasets. We chunk the articles in Pile to passages of a max length of 256 tokens and randomly sample 50K passages as the inputs. Subsequently, we use ChatGPT (gpt-3.5-turbo-0301) to generate entity mentions and their associated types based on the sampled passages. To ensure stability, we set the generation temperature to 0. The specific prompt for constructing the data is shown in Fig. 1. In this prompt, we do not specify the set of entity types of interest, allowing the LLM to generate outputs encompassing a broad coverage of entity types.

Data statistics. After filtering out unparseable outputs and inappropriate entities, including non-English entities and those classified under ’ELSE’ categories, such as None, NA, MISC, and ELSE, our dataset comprises 45,889 input-output pairs, encompassing 240,725 entities and 13,020 distinct entity types. We divide the entity types according to frequency and show the top 10 entity types in each range in Tab. 1. The distribution of these entity types exhibits a heavy tail, where the top 1% of entities account for 74% of total frequencies. We find that the generated data contain entity types from various domains, ranging from the general domain (e.g., person) to the clinical domain (e.g., medical condition). Moreover, we observe variations in granularity among the entity types. E.g., county is the subset of location, and input device is a subset of product. These data characteristics offer extensive coverage of entity types, making them suitable for distilling capabilities from LLMs across various domains.

Definition-based data construction. Besides entity types, we also prompt ChatGPT to generate entity mentions and define their types using short sentences. To do so, we simply change the prompt in Fig. 1 from “extract all entities and identify their entity types” to “extract all entities and concepts, and define their type using a short sentence”. This method generates a much more diverse set of 353,092 entity types and leads to a tuned model that is less sensitive to entity type paraphrasing (Section 5.5), but performs worse on standard NER benchmarks (Section 5.2).

2 Instruction Tuning

After obtaining the data, we apply instruction tuning to smaller models to distill for a broad application class, e.g., diverse entity types in NER. Our template, as shown in Fig. 2, adopts a conversation-style tuning format. In this approach, the language model is presented with a passage Xpassage{\bm{X}}_{\text{passage}} as input. Then, for each entity type ti{\bm{t}}_{i} that appears in the output, we transform it into a natural language query “What describes ti{\bm{t}}_{i}?” Subsequently, we tune the LM to generate a structured output yi{\bm{y}}_{i} in the form of a JSON list containing all entities of ti{\bm{t}}_{i} in the passage. We consider y1,...,yT{\bm{y}}_{1},...,{\bm{y}}_{T} as gold tokens and apply a language modeling objective on these tokens. Our preliminary experiments show that conversation-style tuning is better than traditional NER-style tuning adopted by Wang et al. (2023a); Sun et al. (2023).

Besides one entity type per query, we also consider combining all entity types in a single query, requiring the model to output all entities in a single response. Detailed results and discussions can be found in Section 5.2.

Negative sampling. Our data construction process follows an open-world assumption where we allow the model to generate entity types that have appeared in the passage. However, the generated data do not account for entity types that are not mentioned in the passage, i.e., negative entity types. As a result, it is challenging for us to apply a model trained on this data to a closed-world setting, where one may ask for entity types that do not exist in the passage. To address this potential mismatch, we sample negative entity types from the collection of all entity types that do not appear in the passage as queries and set the expected outputs as empty JSON lists. The sampling of negative entity types is done with a probability proportional to the frequency of entity types in the entire dataset. This approach greatly improves the instruction tuning results, as shown in Section 5.4.

Supervised finetuning. When we have additional human annotations, model performance can be further improved with supervised data. However, a significant challenge arises when training with multiple datasets, as there might be discrepancies in label definitions among these datasets, resulting in label conflicts. For instance, some datasets like ACE (Walker et al., 2006) consider personal pronouns (e.g., she, he) as person, while other datasets like multiNERD (Tedeschi & Navigli, 2022) do not include pronouns.

To address this issue, we propose to use dataset-specific instruction tuning templates to harmonize the discrepancies in label definitions, as illustrated in Fig. 3. Specifically, we augment the input with an additional field denoting the dataset name D{\bm{D}}. By doing so, the model can learn the dataset-specific semantics of labels. During inference, we use the respective dataset name in the prompt for the supervised setting, whereas we omit the dataset field from the prompt in the zero-shot setting.

Universal NER Benchmark

To conduct a comprehensive evaluation of NER models across diverse domains and entity types, we collect the largest NER benchmark to date. This benchmark encompasses 43 NER datasets across 9 domains, including general, biomedical, clinical, STEM, programming, social media, law, finance, and transportation domains. An overview of data distribution is shown in Fig. 4. Detailed dataset statistics are available in Appendix Tab. 6.

Dataset processing. To make the entity types semantically meaningful to LLMs, we conduct a manual inspection of the labels and convert the original labels into natural language formats. For instance, we replace per with person. While we try to collect a broad coverage of NER datasets, we do not use all entity types. This is because some entity types (e.g., Else) are not coming from consistent sources across the different datasets. Their annotations often come from different ontologies for different purposes. The choices of entity types and their annotation guidelines are not optimized for holistic or comprehensive assessments, which renders them suboptimal for use as a “ground truth” to evaluate a universal NER model. Therefore, we remove those labels from the datasets. In addition, some datasets are at the document level and contain very long contexts, which might exceed the input length limit of models. Therefore, we split all instances in document-level datasets into sentence-level ones.

Experiments

This section presents experimental evaluations of UniversalNER. We start by outlining experimental settings (Section 5.1), followed by presenting the results on both distillation and supervised settings (Sections 5.2 and 5.3). Finally, we conduct analysis (Section 5.4) and case study (Section 5.5) to provide deeper insights into the model’s performance.

Model configurations. We train models based on LLaMAWe also train models based on LLaMA 2 (Touvron et al., 2023b). However, no significant difference is observed in our experiments. (Touvron et al., 2023a) following the training schedule of Chiang et al. (2023) for a fair comparison. Considering the large size of certain test datasets, we perform evaluation by sampling up to 200,000 passage-query pairs from each dataset. We use strict entity-level micro-F1F_{1} in evaluation, requiring both the entity type and boundary to exactly match the ground truth.

Compared models. We compare our model (UniNER) against the following models: (1) ChatGPT (gpt-3.5-turbo-0301). We use the prompting template in Ye et al. (2023) for NER. (2) Vicuna (Chiang et al., 2023) is finetuned with ChatGPT conversations, using LLaMA as the base model. (3) InstructUIE (Wang et al., 2023a) is a supervised model finetuned on diverse information extraction datasets, employing a unified natural language generation objective. It adopts Flan-T5 11B (Chung et al., 2022) as the base model.

2 Distillation

We first evaluate the models in a zero-shot setting. We compare the performance of ChatGPT, Vicuna, and our model UniNER, which is distilled from ChatGPT NER annotations on Pile without human-labeled datasets in training. Results are shown in Fig. 5(a).Due to limited space, we only show the average F1F_{1} of all datasets and the average F1F_{1} of each domain. See Appendix Fig. 9 for full results. We observe that our distilled models, namely UniNER-7B and UniNER-13B, outperform ChatGPT in terms of average F1F_{1}. The average F1F_{1} scores of UniNER-7B and UniNER-13B are 41.7% and 43.4%, respectively, compared to 34.9% for ChatGPT. This demonstrates that our proposed targeted distillation from diverse inputs yields models that have superior performance on a broad application class while maintaining a relatively small model size. Additionally, UniNER-13B exhibits better performance compared to UniNER-7B, indicating that fine-tuning on larger models may lead to improved generalization. In terms of domains, both UniNER-7B and UniNER-13B outperform ChatGPT on all domains, showing that the improvements exist across various domains.

We further compare different variations of UniNER, including (1) UniNER-all-in-one, where the extraction of all entity types are combined into one query and response, and (2) UniNER-definition, where queries in instruction tuning data use entity type definitions generated by ChatGPT instead of entity types. Results are shown in Fig. 5(b). We observe that both UniNER-all-in-one and UniNER-definition underperform UniNER-type by 3.3% and 11.8% on average, respectively. The UniNER-definition variant’s decreased performance could be due to its lower consistency with the evaluation datasets, which all adopt words or short phrases as labels instead of sentences. The performance disparity in the UniNER-all-in-one variant can be potentially attributed to the attention distribution and task complexity. When the model is required to handle multiple entity types within a single query, it might disperse its attention across these varied types, possibly resulting in less accurate identification for each individual type. Conversely, by decomposing the task into several simpler ones, each focusing on one entity type at a time, the model might be better equipped to handle the complexity, thus yielding more accurate results.

3 Supervised Finetuning

We study whether our models can be further improved using additional human annotations. We compare the performance of ChatGPT, Vicuna, InstructUIE (Wang et al., 2023a) Please note that the original evaluation script in InstructUIE contains a critical bug. For passages that do not contain any entities, the script adds none as a placeholder entity and takes it into account when calculating F1F_{1}. To rectify this error, we re-evaluated InstructUIE using their released checkpoint., and UniNER.

Out-of-domain evaluation. We first study whether supervised finetuning leads to better generalization on unseen data. We follow InstructUIE to exclude two datasets CrossNER (Liu et al., 2021) and MIT (Liu et al., 2013) for out-of-domain evaluation, and fine-tune our model using training splits of the remaining datasets in the universal NER benchmark. Results are shown in Tab. 3. Notably, without any fine-tuning, instruction-tuned UniNER 7B and 13B already surpass ChatGPT, Vicuna, and the supervised fine-tuned InstructUIE-11B by a large margin. If we train our model from scratch only using the supervised data, it achieves an average F1F_{1} of 57.2%. Continual fine-tuning UniNER-7B using the supervised data achieves the best average F1F_{1} of 60.0%. These findings suggest that the models’ generalization can be further improved with additional human-annotated data.

In-domain evaluation. We then study the performance of UniNER in an in-domain supervised setting, where we fine-tune UniNER-7B using the same training data as InstructUIE (Wang et al., 2023a). Results are shown in Tab. 2. Our UniNER-7B achieves an average F1F_{1} of 84.78% on the 20 datasets, surpassing both BERT-base and InstructUIE-11B by 4.69% and 3.62%, respectively. This experiment demonstrates the effectiveness of our model in the supervised setting.

4 Analysis

Negative sampling strategies. We experiment with different negative sampling strategies in instruction tuning, including (1) no negative sampling, (2) uniform sampling where entity types are randomly sampled with equal probability for each one, and (3) frequency-based sampling where we sample entity types with probabilities proportional to their frequency in the constructed dataset. Results are shown in Tab. 4. Among the approaches tested, frequency-based sampling yielded the best results, outperforming no sampling and uniform sampling by 21.9% and 5.7%, respectively. These findings highlight the crucial role of negative sampling in instruction tuning, with frequency-based sampling emerging as the most effective method for enhancing model performance in our study.

Dataset-specific template. We compare the results of our dataset-specific instruction tuning template and the original template in the supervised setting. As shown in Fig. 6, we find that the data-specific template outperforms the original template on most datasets. To gain deeper insights into the improvements achieved, we further divide the datasets into two categories: those with label (entity type) overlap with other datasets and those without overlap. Our analysis reveals that datasets with label overlap demonstrate more substantial improvements.

To explore this further, we measure F1F_{1} score across all evaluation datasets and calculate the difference. Apart from the long-tail entity types that manifest a high variance in results, we identify two entity types where the dataset-specific template outperforms the original template by over 10%: facility (22.0%) and time (12.4%). Intriguingly, both labels exhibit inconsistencies in their definitions across various datasets. The facility label has been annotated on pronouns (e.g., it, which) as entities in ACE datasets but are excluded in OntoNotes. The time label denotes well-defined time intervals (e.g., Christmas) in MultiNERD, but may encompass any general time expressions (e.g., 3 pm) in OntoNotes. This finding suggests that the improvements provided by the data-specific template are particularly effective in resolving label conflicts.

Evaluation with partial match. While using strict F1F_{1} as an evaluation metric, we notice that it may underestimate the zero-shot learning capabilities of NER models. In particular, strict F1F_{1} penalizes slight misalignments in the boundaries of the extracted entities, which may not necessarily indicate an incorrect understanding of the text. For instance, given the sentence any asian cuisine around and the entity type cuisine, UniNER extracts asian cuisine as the named entity, while the ground truth only labels asian as the correct entity. However, the model’s prediction can still be viewed as correct, even though it is deemed incorrect by strict F1F_{1}. To better estimate the zero-shot abilities, we also consider partial match (Segura-Bedmar et al., 2013) in evaluation. In this context, a prediction that exhibits word overlap with the ground truth is regarded as half correct (counted as 0.5 in true positive) when computing F1F_{1}. Results are shown in Tab. 5. We find that allowing partial match consistently improves the results. Besides, our models is still the best-performing model on average.

5 Case Study

Sensitivity to entity type paraphrasing. One type of entity can be expressed in multiple ways, so it is essential for our model to give consistent predictions given entity types with similar meanings. An example of sensitivity analysis is present in Fig. 7. We observe that UniNER-7B-type sometimes fails to recognize entities with similar semantic meanings. On the other hand, UniNER-7B-definition, despite performing worse on our Universal NER benchmark, exhibits robustness to entity type paraphrasing. It demonstrates that although using definitions may result in lower performance on standard NER benchmarks, it could yield improved performance for less populous entity types.

Recognition of diverse entity types. We present an example in Fig. 8 showcasing the capabilities of UniNER in recognizing various entities. Particularly, we focus on a novel domain of code and assess UniNER’s ability to extract diverse types of entities within the code. Despite minor mistakes (e.g., from_pretrained is not identified as a method), this case study effectively demonstrates our model’s capacity to capture entities of various types.

Conclusion

We present a targeted distillation approach with mission-focused instruction tuning. Using NER as a case study, we train smaller and more efficient models for open-domain NER. The proposed method successfully distills ChatGPT into a smaller model UniversalNER, achieving remarkable NER accuracy across a wide range of domains and entity types without direct supervision. These models not only retain ChatGPT’s capabilities but also surpass it and other state-of-the-art systems in NER performance.

Acknowledgement

Wenxuan Zhou and Muhao Chen were supported by the NSF Grants IIS 2105329 and ITE 2333736.

References

Appendix A Appendix

One type of entity can be expressed in multiple different ways. In this scenario, it is essential for our model to give consistent predictions given entity types with similar meanings. An example of sensitivity analysis is present in Fig. 7. We observe that UniNER-7B-type sometimes fails to recognize entities with similar semantic meanings. On the other hand, UniNER-7B-definition, despite performing worse on our Universal NER benchmark, exhibits robustness to entity type paraphrasing. It demonstrates that although using definitions may result in lower performance on standard NER benchmarks, it could yield improved performance for less populous entity types.

Recognition of diverse entity types

The code snippet provided is for text generation using the Huggingface library:

In the ensuing conversations with UniNER, we present the queries and corresponding results:

User: What describes library in the text? Assistant: ["transformers"] User: What describes local variable in the text? Assistant: ["input_text", "input_ids", "output", "generated_text"] User: What describes global variable in the text? Assistant: [] User: What describes method in the text? Assistant: ["encode", "generate", "decode"] User: What describes class in the text? Assistant:["GPT2Tokenizer","GPT2LMHeadModel"] Despite minor mistakes (e.g., from_pretrained is not identified as a method), this case study effectively demonstrates our model’s capacity to capture entities of various types.

Appendix B Full Evaluation Results

Full results on ChatGPT, UniNER-7B-type, and UniNER-7B-sup+type are shown in Fig. 9.

Appendix C Data Statistics

We show the full dataset statistics in Universal NER in Tab. 6, including the number of instances in train/dev/test data, number of entity types, average number of tokens in input text, and the average number of entities in each instance.