On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey

Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, Haobo Wang

Introduction

The game-changing emergence of Large Language Models (LLMs) instigated a significant paradigm shift in the field of deep learning Zhang et al. 2023a; Guo et al. 2023; Bang et al. 2023. Despite these advancements, a large amount of high-quality data remains the foundation for building robust NLP models Gandhi et al. 2024. To be more specific, here high-quality data typically refers to diverse data that carries rich supervision signals (generally in the form of labels) closely aligned with human intent. However, fulfilling such data reliance with human data can be challenging or even unrealistic sometimes, due to high costs, data scarcity, privacy concerns, etc. Kurakin et al. 2023. Moreover, several studies Hosking et al. 2023; Singh et al. 2023; Gilardi et al. 2023 have highlighted that human-generated data, being inherently susceptible to biases and errors, may not even be optimal for model training or evaluation. These considerations necessitate a more serious inquiry into the question: are there other more effective and scalable methods of data collection that can overcome the current limitations?

Given the recent advancements in LLMs, which demonstrate the capability to generate fluent text on par with human output Hartvigsen et al. 2022; Sahu et al. 2022; Ye et al. 2022a; Tang et al. 2023; Gao et al. 2023a, synthetic data produced by LLMs emerges as a viable alternative or supplement to human-generated data. Specifically, synthetic data is designed to mimic the characteristics and patterns of real-world data Liu et al. 2024. On the one hand, LLMs, through extensive pretraining, have acquired a vast repository of knowledge and demonstrate exceptional linguistic comprehension Kim et al. 2022; Ding et al. 2023a, which forms a foundation for generating faithful data. On the other hand, the profound instruction-following capabilities of LLMs allow better controllability and adaptability over the generation process, facilitating the creation of tailored datasets for specific applications with more flexible process designs Eldan and Li 2023. These two advantages make LLMs highly promising synthetic data generators.

As a pivotal application of LLMs, synthetic data generation holds significant importance for the development of deep learning. As shown in Figure 1, LLMs-driven synthetic data generation Li et al. 2023c; Wang et al. 2021; Seedat et al. 2023 enables the automation of the entire model training and evaluation process with minimal human participation required in the loop Huang et al. 2023, which allows the advantages of deep learning models to be applied across a broader range of applications. Beyond providing a scalable supply of training and testing data, LLM-driven synthetic data generation also may pave the way for developing next-generation LLMs. Insights from TinyStories Eldan and Li 2023 and the Phi series Gunasekar et al. 2023; Li et al. 2023b emphasize that data quality is crucial for effective model learning, while LLMs empower us to actively “design” what the models learn through data manipulation, significantly enhancing the efficacy and controllability of model training. As of June 2024, there are over 300300 datasets on Hugging Face https://huggingface.co that are tagged as “synthetic”, with many mainstream LLMs leveraging high-quality synthetic data for training, including Alpaca Taori et al. 2023, Vicuna Zheng et al. 2023, OpenHermes 2.5, and Openchat 3.5 Wang et al. 2023a.

Though seemingly straightforward, generating synthetic datasets that simultaneously have high correctness and sufficient diversity requires careful process designs and involves a lot of tricks Gandhi et al. 2024, making LLMs-driven synthetic data generation a non-trivial problem. While most existing works generally target data generation for various tasks (e.g., pre-training Gunasekar et al. 2023; Li et al. 2023b; Eldan and Li 2023, fine-tuning Mukherjee et al. 2023; Mitra et al. 2023; Xu et al. 2023a, evaluation Feng et al. 2023; Wei et al. 2024) across different domains (e.g., math Yu et al. 2023a; Luo et al. 2023a, code Luo et al. 2023b; Wei et al. 2023b, instruction Honovich et al. 2023a; Wang et al. 2023d), they share many common ideas. To address the lack of a unified framework in the emerging field of LLM-driven synthetic data generation and develop a general workflow, this survey investigates recent studies and organizes them according to the topics of generation, curation, and evaluation, which are closely related, as shown in Figure 2. Our primary aim is to provide a comprehensive overview of the current state of the field, identify key areas of focus, and highlight the gaps that remain to be addressed. We hope to bring insights to both the academic and industrial communities and drive further development in LLM-driven synthetic data generation.

Preliminaries

In this paper, we investigate the challenge of generating high-quality synthetic data using pre-trained LLMs, denoted as M\mathcal{M}. Rather than creating new datasets from scratch, in more cases, we perform data augmentation with a small number of seed samples or unlabeled inputs, which we denote uniformly as Dsup\mathcal{D}_{\text{sup}}. Although optional for LLMs-driven synthetic data generation, Dsup\mathcal{D}_{\text{sup}} can typically provide valuable supporting information when available. Consequently, the overall generation task can be formulated as:

where Dgen\mathcal{D}_{\text{gen}} represents the final generated dataset, and pp refers to the prompt used for model inference. T\mathcal{T} specifies the generation task, such as rewriting, question answering, annotation, etc. Notably, data annotation as a specialized paradigm of synthetic data generation, has particularly extensive applicability, including RLAIF Bai et al. 2022 and LLMs-based evaluation Chen et al. 2023b; Zheng et al. 2023; Kim et al. 2023, which may involve specific challenges and corresponding solution techniques. Due to page limitations, further details about data annotation can be found in Appendix A.

2 Requirements of 𝒟gen\mathcal{D}_{\text{gen}}

Briefly speaking, our goal is to generate data that closely aligns with evaluation metrics. While the standard of high-quality data may vary across different downstream tasks, there are two general requirements that are considered challenging in most existing literature:

Faithfulness. To provide valid supervision, the generated data must first be logically and grammatically coherent. However, the inherent problems of hallucination fat-tailed knowledge distribution of LLMs can introduce significant noise into the generated results, manifesting as factual errors, incorrect labels, or irrelevant content. These issues become more pronounced when generating long, complex, or domain-specific data.

Diversity. Diversity captures the variation among the generated data, reflecting differences in text length, topic, or even writing style. It is crucial for generating synthetic samples that mimic the diversified nature of real-world data, thereby preventing overfitting and bias during model training or evaluation. Nevertheless, due to the inherent biases of LLMs, uncontrolled generated content often tends to be monotonous, limiting its applicability in downstream tasks.

These two requirements are the focal points of most current research efforts. In the subsequent workflow, we will introduce how different methods address these issues.

Generic Workflow

Existing studies on LLMs-driven synthetic data generation generally incorporate three main topics: generation, curation, and evaluation. Various approaches are employed within these aspects to collaboratively achieve optimal data generation.

In this section, we systematically summarize some common practices for synthetic data generation with LLMs, which can be roughly divided into prompt engineering and multi-step generation. An overall illustration is provided in Figure 3.

One of the greatest advantages of LLMs for synthetic data generation is their instruction-following capability, which contributes to great controllability Wang et al. 2023c; Radford et al. 2019. Therefore, many approaches try to guide LLMs with heuristic prompts to enhance the faithfulness and diversity of the synthetic data Liu et al. 2024.

Empirically, an effective prompt generally contains three key elements: task specification etaske_{\text{task}}, generation conditions econditione_{\text{condition}}, and in-context demonstrations edemoe_{\text{demo}}, which are then collectively wrapped with a template EE into the form of natural instruction:

As shown above, both the generation task T\mathcal{T} and the support dataset D\mathcal{D} will affect the design of pp. Next, we will proceed to detail how each part of the prompt should be appropriately designed to accommodate various scenarios.

In traditional crowdsourced annotation scenarios, the recruited workers are commonly offered a codebook that specifies the necessary contexts, such as task purpose, data explanation, and other background knowledge, so that they can better understand their jobs Gilardi et al. 2023. Similarly, such task specification is crucial for setting the right context for LLMs-driven data generation, which can also include role-play Li et al. 2023c, format clarification, knowledge augmentation Xu et al. 2023b; Sudalairaj et al. 2024, etc. Evidence shows that a simple prologue such as “suppose you are a {xxx}” can significantly improve the LLMs’ performance by setting up a proper scenario for data generation and allowing the LLMs to better take on the roles Li et al. 2023c. More formally, Yoo et al. 2021 defines the task specification with a triplet of text type, label type, and label-token verbalizer. Such a description header is particularly important when extra domain expertise is demanded to address issues like terminology complexities in both context understanding and data generation. Consequently, Xu et al. 2023b leverages external knowledge graphs and LLMs to obtain domain topics for context-informed prompting, which effectively enhances the faithfulness and complexity of generated data.

As mentioned in Section 2.2, a pivotal challenge in using LLMs for synthetic data generation is ensuring sufficient diversity, as directly prompting the LLMs to produce data for certain tasks often results in highly repetitive outputs, even with a high decoding temperature Gandhi et al. 2024; Liu et al. 2024. Addressing this problem, a widely adopted strategy is conditional prompting, which explicitly and concretely communicates to the LLMs the specific type of data desired. The core of conditional prompting involves delineating the targeted data through the formulation of a series of condition-value pairs:

which effectively characterizes the desired attributes and characteristics of the synthetic data. With different combinations of such attributes, we can automatically achieve a degree of “artificially defined” diversity in the generated samples Gunasekar et al. 2023; Li et al. 2023b; Eldan and Li 2023. Conditional prompting not only allows better control over the diversity and coverage of the generated dataset but also refines the content to a narrower, more focused scope that is more likely to align with our specific expectations and requirements Li et al. 2023c. Current research on conditional prompting primarily centers on the following two subjects:

Conditioning Scope. As the backbone of econditione_{\text{condition}}, conditioning scope defined by {c1,⋯ ,cn}\{c_{1},\cdots,c_{n}\} delineates the dimensions that we utilize to characterize our target data. Early studies Gao et al. 2023a; Ye et al. 2022a; Ye et al. 2022b employed a basic output-conditional prompting strategy, utilizing the specific label associated with the classification task as the conditioning variable. The rationale behind this was primarily to maintain class balance and coverage. However, such a strategy is unsuitable for data lacking explicit category labels. Subsequent work by Yu et al. 2023b argues that conditional-prompting with finer-grained attributes (e.g., topics, length, and style Xu et al. 2023b), can lead to more diversified generation due to the vast number of possible attribute combinations, being also applicable to open-ended data. Additionally, Eldan and Li 2023 also condition each generation on the task of incorporating three randomly chosen words into the generated story. This approach was also proven to significantly enhance the diversity of the generated data, shifting the focus from the heuristic features of the output to a more structured and targeted conditioning mechanism by adding “creative randomness” to the prompt Eldan and Li 2023.

Conditioning Values. After defining the conditioning scope, we then need to assign concrete values to each condition. Despite the seemingly straightforward strategy of sampling from the known classes or labels Ye et al. 2022a, there are cases where such an instance pool is unavailable. Addressing this problem, Josifoski et al. 2023 actively retrieves the conditioning instances from external knowledge graphs, while Xu et al. 2023b; Ding et al. 2023b leverage the LLMs to generate diversified instances for conditional prompting. Specifically, Ding et al. 2023b construct a concept tree to delve into different subtopics, ensuring the coverage of sampled conditioning values, which then contributes to more diverse generated data. Moreover, the prompt template EE can also be considered a special type of condition. It has been demonstrated that incorporating templates with a certain level of randomness throughout the generation process can enhance the diversity of the generated contents Meng et al. 2022.

Due to the inherent bias of LLMs, it remains challenging to elicit favorable responses from the LLMs with merely task specification and conditional prompting. In this case, a straightforward yet effective strategy is to provide several demonstrations, which can serve as a form of implicit human guidance. Research has shown that, owing to LLMs’ remarkable in-context learning (ICL) capabilities, a few exemplars can provide them with insights into the patterns exhibited in real-world data, thereby significantly improving the faithfulness of generated data Li et al. 2023c. In the few-shot setting, where labeled samples are available in the support set Dsup\mathcal{D}_{\text{sup}}, these samples can be directly utilized as demonstrations for ICL. However, in scenarios where no ground truth data is available, approaches like Self-Instruct Wang et al. 2023e and Self-Prompting Li et al. 2022 instead leverage ICL with synthetic demonstrations generated by LLMs. This allows the models to learn from their own predictions or other teacher models, even in the absence of labeled data.

However, given the constraint of prompt length and data inconsistency, the quality of in-context samples significantly affects the effectiveness of in-context learning. Sudalairaj et al. 2024 argue that randomly selecting in-context examples from the pool of seed samples, as done in Self-Instruct Wang et al. 2023e, results in a lack of diversity and quality in the generated data. To address this issue, Sudalairaj et al. 2024 opt for selecting examples that concentrate on specific aspects to better stimulate the long tail of knowledge inherent in LLMs. Liu et al. 2022b and Su et al. 2023 prioritize consistent samples as demonstrative examples based on their cosine similarity in the embedding space. Alternatively, Ye et al. 2022b selects the most informative samples using quantified influence scores to steer the generation process. To enhance the informativeness of in-context examples, He et al. 2023 prompts LLMs to provide an explanation for each sample before integrating it into the prompt. This approach not only offers valuable additional information but also aligns well with the subsequent Chain-of-Thought generation.

1.2 Multi-Step Generation

In the previous paragraphs, we have introduced some common prompting strategies, which are typically designed for a specific generation task T\mathcal{T}. However, in most cases, due to the lack of enough reasoning abilities, it is unrealistic to expect the LLMs to generate the entire desired dataset within a single reference, especially when targeting data with complex structures or semantics Cui and Wang 2023. In addressing this problem, a common strategy is multi-step generation, through which the overall generation process is manually decomposed into a chain of simpler sub-tasks T1:k\mathcal{T}_{1:k}, to force the LLMs to produce data in a step-by-step manner as scheduled:

where D0=Dsup\mathcal{D}_{0}=\mathcal{D}_{\text{sup}}. Each intermediate output Di\mathcal{D}_{i} is generated using model Mi\mathcal{M}^{i}, prompted by pip_{i}, for a sub-task Ti\mathcal{T}_{i}. These outputs can then potentially be used in subsequent generations. By manually scheduling the generation procedure, we implicitly align the reasoning paths of LLMs with human prior knowledge. Specifically, there are two common strategies for task decomposition: sample-wise and dataset-wise decomposition, which mainly aim at enhancing the quality of synthetic data at different scales.

A typical use-case of multi-step generation is for addressing the challenges of long-text processing and logical reasoning when dealing with multi-text data such as dialogues and entity-relation triplets. In such cases, a straightforward approach is to divide the sample into smaller chunks and generate only a portion of each sample at a time Li et al. 2022; Ye et al. 2023; Wang et al. 2023e. In this way, D1:k\mathcal{D}_{1:k} can be considered as different parts of Dgen\mathcal{D}_{\text{gen}}:

Notably, as shown in Eq. 4, each iteration of the generation process can be conditioned on the previously generated contents. For example, Ding et al. 2023b prompts the LLMs to alternate between acting as the assistant and the user, replying to each other based on the context, ultimately producing a complete conversation transcript. In this way, the coherence among each internal component Di\mathcal{D}_{i} can be pointedly reinforced with separated instructions, thus making it easier for the model to follow the requirements and generate more faithful data. It should be noted that D1:kD_{1:k} may not necessarily form part of the final DgenD_{\text{gen}}, instead, explicitly outputting some intermediate reasoning steps can also improve the generation of complex data Bai et al. 2022; He et al. 2023. Chain-of-Thought (CoT) prompting stands out as one of the most popular strategies for improving the faithfulness of LLM-generated content Wei et al. 2022. Nevertheless, current research on the exploration of such latent metadata is still insufficient, leaving sample-wise task decomposition from a reasoning perspective an open problem for future studies.

In Section 3.1.1 we have introduced how to generate data with specified properties. However, generating a series of such data that can eventually form a dataset with good diversity and domain coverage requires long-term scheduling. To this end, dataset-wise task decomposition dynamically adjusts the conditions used at each stage of multi-step generation to ensure the overall dataset grows in the right direction:

Specifically, S3 Wang et al. 2023b targets the most frequently mislabeled categories at each iteration, according to the performance of the downstream model trained on previously generated data. Similarly, Honovich et al. 2023b; Shao et al. 2023 utilize a generate-then-expand paradigm, to enhance the diversity of the overall dataset accordingly. Some other methods also leverage specific data structures to model the pathways of data generation. For example, Explore-Instruct Wan et al. 2023 models the domain space as a tree structure and continually refines the generated data along with tree traversal to promote both the specialization and domain coverage of the generated data.

2 Data Curation

After the preceding steps, one may excessively generate overflowing and theoretically unlimited data Dgen\mathcal{D}_{\text{gen}}. However, these datasets often comprise a considerable portion of noisy, worthless, or even toxic samples, which primarily stems from two causes. Firstly, LLMs can inevitably produce corrupted samples with incorrect labels due to the hallucination problem. Secondly, ineffective prompts containing ambiguous descriptions can trick the model into generating irrelevant or redundant samples. Consequently, directly utilizing these low-quality data without proper processing may have a significant negative impact.

To address this, plenty of data curation approaches have been studied, which mainly fall into two dominant groups of high-quality sample filtering and label enhancement as elaborated below.

Sample filtering aims to weed out undesired low-quality samples and obtain a more helpful subset Dcurated ⁣⊂ ⁣Dgen\mathcal{D}_{\text{curated}}\!\subset\!\mathcal{D}_{\text{gen}}. These methods typically design heuristic criteria or re-weighting functions to rerank samples for filtering, as shown in Figure 4.

For methods based on heuristic metrics, the key step is to design appropriate criteria based on the learning dynamics, such as confidence score (Seedat et al. 2023), influence function Ye et al. 2022b, and generation ability Meng et al. 2022. SuperGen Meng et al. 2022 employs the estimated generation probability to identify samples most related to the desired label. Seedat et al. 2023 discard samples with both low confidence and low uncertainty. Some other methods assume that clean samples are prone to hold similar predictions under different conditions and employ cross-condition consistency for filtering. Specifically, such consistency can be between LLM and downstream classifier Yu et al. 2023c, between multiple executions Ye et al. 2023, or between neighboring data points Seedat et al. 2023. Chen et al. 2023b leverage the powerful text understanding capabilities of LLMs to assess the quality of different samples and filter out those with low scores. Results show that Alpagasus Chen et al. 2023b, trained on a much smaller but curated dataset, surpasses the original Alpaca Taori et al. 2023 across several benchmarks, underscoring the importance of data curation.

On the other hand, re-weighting methods believe all data are valuable but with varying importance. Thus, they assign larger weights to correctly annotated or influential samples during downstream utilization Zhang et al. 2023b; Gao et al. 2023a; Meng et al. 2023. For instance, SunGen Gao et al. 2023a proposes an adaptive bi-level re-weighting algorithm without human annotations. FewGen Meng et al. 2023 designs a discriminative meta-learning objective to adjust sample weights and demarcate the nuanced differences between different labels.

2.2 Label Enhancement

Label enhancement methods strive to rectify the potentially erroneous annotations in generated samples. Due to confirmation bias, it is unrealistic for LLMs to identify their own mistakes. To address this, recent works either rely on human intervention or incorporate a student model for human-free knowledge distillation.

A straightforward strategy for label refinery is to include human efforts to re-annotate the corrupted samples Chung et al. 2023a; Wang et al. 2021; Pangakis et al. 2023. Wang et al. 2021 proposed to actively select samples with the lowest confidence for human re-labeling. Pangakis et al. 2023 and Liu et al. 2022a further emphasize the importance of human review and suggest comparing annotations from humans and LLMs guided by the same codebook. Despite the simplicity, these methods can lead to considerable labeling costs and can be unrealistic in practical deployment.

To reduce the labeling cost, a more pragmatic human-free paradigm is developed which involves auxiliary student models for knowledge distillation and label refinery Xiao et al. 2023; Zhao et al. 2023a; Saad-Falcon et al. 2023. These methods rely on the weakly supervised ability of student models and hypothesize that a student distilled from the LLM teacher can produce superior labels. The seminal work FreeAL Xiao et al. 2023 proposes a collaborative framework, where a student model is leveraged to distill the high-quality task-related knowledge from the weak annotations and in return feedback LLMs for label refinery. MCKD Zhao et al. 2023a designs a multistage distillation pipeline with data-split training and cross-partition labeling to avoid overfitting on noisy labels. With the expanding abilities and availability of LLMs, the incorporation of auxiliary student models will play a more crucial role as a cost-effective alternative to human intervention.

3 Data Evaluation

Before the employment of generated data, it is important to evaluate the quality and application effectiveness of the data, to ensure its value to downstream tasks. The current mainstream evaluation methods can be roughly divided into two categories: direct and indirect, which evaluate the quality of Dgen\mathcal{D}_{\text{gen}} individually and through its effectiveness on downstream tasks, respectively.

Ideally, automatic evaluation of the LLMs’ generation results can be easily realized with ground truths from existing datasets, if available Zhu et al. 2023. However, for open-ended data, human-based evaluation is necessitated. A straightforward idea is to provide some generated samples to human experts, who will then determine whether they are correct, according to which we can estimate the overall generation quality Wang et al. 2023e. Theoretically, the larger the sample size, the more accurate the estimation results will be, but the labor it costs will correspondingly get higher. To this end, a reliable auxiliary model can be leveraged for a more comprehensive yet cost-effective evaluation of the generated data in replace of human experts Chung et al. 2023b. Considering that most models can only process contents of limited length, appropriate information extraction can reduce the burden of the auxiliary model and contribute to a more precise prediction of whether a sample contains factual errors Lee et al. 2022.

The quantification of data diversity primarily employs vocabulary statistics and sample relevance calculations. Vocabulary statistics Yu et al. 2023b, such as vocabulary size and N-gram frequency, provide a straightforward and intuitive approach. However, they struggle to capture the semantic information of a dataset. The calculation of sample relevance compensates for this limitation effectively. The most common measures of sample correlation are based on cosine similarity Wang et al. 2023b and sample distance Chung et al. 2023b, which can better capture the contextual and semantic diversity of the dataset. Furthermore, these metrics can also be leveraged to select in-context demonstrations edemoe_{\text{demo}} Wang et al. 2023e that are more dissimilar with the previously generated samples, thereby leading to more diversified generation results.

3.2 Indirect Evaluation

The performance of downstream models trained on the generated data can also reflect the generation quality to some extent Yu et al. 2023b; Chung et al. 2023b. Specifically, the impact of synthetic data can be evaluated from multiple dimensions except for the specialized capabilities of the downstream models. For example, TruthfulQA enables the assessment of a model’s ability to identify true claims Sun et al. 2023; NIV2 is employed to evaluate a model’s language comprehension and reasoning abilities across multiple tasks Wang et al. 2023e.

For open-ended benchmarks, evaluation by humans or auxiliary models is necessitated due to the absence of standardized answers. To fully leverage the preference outputs of the auxiliary models, multiple evaluation strategies have been designed, such as response ranking Xu et al. 2023a, four-level rating system Wang et al. 2023e and Elo scores Bai et al. 2022. To further reduce evaluation costs, Sun et al. 2023; Xu et al. 2023a utilize the automatic evaluation framework based on GPT-4 proposed by Vicuna for evaluation. However, general LLMs may lack enough knowledge for domain-specific tasks, which hinders them to provide effective evaluation Bran et al. 2023. Therefore, collecting human assessment data to fine-tune open-source models for evaluation purposes is an important practice in real-world scenarios He et al. 2023. Other techniques like Peng et al. 2024; Peng et al. 2023 remain to be further explored.

Future Directions

Current multi-step generation algorithms depend on the model’s understanding of task requirements, requiring it to perform complex logical reasoning with limited information. However, in real-world complex scenarios, this limited information may not adequately support effective decision-making. For instance, the generation of mathematical problem-solution pairs entails multiple reasoning steps and may necessitate the utilization of calculator tools for validation. To date, there remains a lack of systematic investigation on how to activate the reasoning and planning capabilities of LLMs for autonomous synthetic data generation. Inspired by prevalent LLMs-based agents like HuggingGPT Shen et al. 2023 and MetaGPT Hong et al. 2023, we believe it would also be quite valuable to develop a data generation agent for industrial applications.

2 Knowledge Enhancement

Recent research has found that LLMs’ knowledge is long-tailed and biased Navigli et al. 2023; Fei et al. 2023. Lacking specific domain knowledge, LLMs tend to generate biased, monotonous, and even unfaithful data. Though we have introduced how to mildly guide the data generation with task specification and conditional prompting in the previous sections, such methods still hold strong limitations and are not conducive to scalable implementation. Instead, we believe that developing automated condition controls directly on mature domain knowledge bases will significantly improve the efficiency of knowledge enhancement. For example, we can establish certain links between the LLMs and external knowledge graphs Ji et al. 2022 or retrieve augmentation from the website Gao et al. 2023b, which is helpful for the definition, decomposition, and reasoning of data features throughout the entire generation process. Additionally, with enhanced domain knowledge, we may also better assess the quality of generated data or even develop automatic evaluation systems. Overall, we believe that knowledge-driven data generation will be a key focus for future studies.

3 Synergy between Large & Small LMs

In Section 3.2, we introduced the use of small domain-specific models for data curation. In particular, FreeAL Xiao et al. 2023 has shown the feasibility of low-cost data curation with integrated collaboration between large and small models. The idea of leveraging real-time feedback provided by automated performance evaluation during the data generation process to guide the corresponding adjustments in the following generation hints at an important research direction. However, the exploitation of small LMs at the current stage is simply based on prediction confidence. In the future, we are looking forward to seeing more diversified collaboration modes between large and small models to improve the quality of generated data, e.g., usage of various output information, new design of collaborative architectures, and so on.

4 Human-Model Collaboration

Data, as the source of model intelligence, theoretically cannot be generated completely without human intervention. Otherwise, wild synthetic data that carries noisy, toxic information can easily “poison” a model, even resulting in mode collapse. Due to the inherent bias of LLMs, they can hardly be self-aware of the bias in their generated data and finally deviate from our intentions. Thus, designing a human-friendly interactive system to involves a few necessary human knowledge for annotation and verification is vital and irreplaceable. To date, there is still a lack of a generic framework to standardize and systematize the human-machine collaboration involved in the data production process.

We believe that an appropriate design of such a system must be based on a thorough understanding of the strengths and limitations of human intervention, and should follow the human-centered principle. To achieve sustainable and efficient human involvement, we need comprehensive consideration of various factors such as feasibility, cost, and even labor psychology. For specific examples: (i)-readability and interpretability of the information provided by the LLMs should be ensured to reduce obstacles to human understanding; (ii)-upstream knowledge enrichment or filtering should be carried out to improve the efficiency of human resource utilization and reduce consumption on tasks with low cost-effectiveness; (iii)-incorporating enjoyable interactive features can not only mitigate the negative impact of mechanical data processing tasks on humans but also attract a broader audience.

Conclusion

In this paper, we present a systematic review of advancements in synthetic data generation propelled by Large Language Models (LLMs). We aim to offer guidance to enterprises and organizations on effectively building their domain-specific datasets using LLMs. In the meantime, we endeavor to provide insights into the challenges and opportunities within this field, while also proposing potential directions for future research. We hope that our work can promote the rapid production of large amounts of data in various fields and push the limits of data-centric AI. We also envision a fantastic future, where an LLMs community, endowed with human-like abilities such as bionics and communication, may be constructed to generate data for its own self-improvement.

Limitations

In this paper, we survey existing studies on LLMs-driven synthetic data generation, curation, and evaluation, proposing a generic workflow for real-world practice. Synthetic data generation is a broad topic that involves data and models of various modals, including vision and speech. Due to the page limit, we mainly focus on the objective of text data and LLMs-driven approaches, while leaving investigations in other fields for future work. We will also keep paying attention to the latest work and add more related approaches with more detailed analysis.

Ethics Statement

We believe that our proposed workflow of LLMs-driven synthetic data generation, curation, and evaluation can benefit both researchers who are interested in data-centric AI and industrial producers who are facing data problems. However, the malicious use of such synthetic data also raises ethical concerns that should arouse our vigilance.

Acknowledgements

This work is supported by the Pioneer R&D Program of Zhejiang (No. 2024C01035), NSFC under Grants (No. 62206247), and the Fundamental Research Funds for the Central Universities (No. 226-2024-00049).

References

Appendix A Data Annotation

In the main text, we introduced a series of techniques for general data synthesis. Though annotation can be considered a special type of synthesis with the input of a particular sample as the synthesis condition, there are also approaches specifically suitable for data annotation. Among them, selective annotation is one of the most important practices. Selective annotation represents an optimal tradeoff between expensive and precise human annotation and economic but relatively rough LLMs-based annotationWang et al. 2021; Kocon et al. 2023.

The key to selective annotation is to define a "cost-effective" sample distribution between humans and LLMs. Zhang et al. 2023b; Bansal and Sharma 2023 covers some common selection strategies for LLMs-based annotation, including random selection, maximum entropy selection, least confidence selection and kkmeans selection for thorough comparisons. Results show that uncertainty-based methods, i.e. maximal entropy and least confidence, perform significantly better than the random baseline, with faster convergence and better performance of the downstream model trained on the annotated data. Li et al. 2023a also utilizes uncertainty to estimate LLMs’ annotation capability to effectively allocate the annotation work among humans and LLMs. Su et al. 2023 instead proposes a novel unsupervised, graph-based selective annotation method named vote-kk, to select diverse and representative examples to annotate.

Appendix B Tuning Techniques

Another large body of research pertains to the tuning techniques, such as model fine-tuning Zhao et al. 2023b; Sun et al. 2023; Meng et al. 2023; Kurakin et al. 2023 and soft prompting Chen et al. 2023a, which have already been heavily studied in other fields and can be detailedly referred in Hu et al. 2023; Lu et al. 2023; Wei et al. 2023a; Xiao and Chen 2023. Despite their effectiveness in improving the generation performance, most of the existing approaches are established on the accessibility of the LLMs, while their application on black-box models remains to be further explored.

Appendix C Applications

LLM-driven synthetic data generation has served as a new alternative to traditional human-dependent data collection and demonstrated great potential in various applications, including general tasks, domain-specific tasks, and multimodal tasks.

With the exploding capabilities of LLMs, this generation pipeline has been adopted in a wide range of basic NLP studies, including text classification Ye et al. 2022b; Yu et al. 2023c; Sahu et al. 2022, named entity recognition Xiao et al. 2023, question answering Li and Callison-Burch 2023, relationship extraction He et al. 2023, and natural language inference Zhang et al. 2023b. These studies further underpin diverse applications, such as sentiment recognition Gao et al. 2023a; Ye et al. 2022b, online translation Oh et al. 2023, stance detection Li et al. 2023a and spam identification Smith et al. 2022.

Some domain-specific tasks also impose significant demands on this pipeline, where human annotation can be extremely expensive and impractical, such as medical diagnosis Tang et al. 2023, drug discovery Xiao et al. 2023, clinical trial extraction Xu et al. 2023b, industrial advertisement Zhang et al. 2022 and tabular data analysis Seedat et al. 2023.

Stemming from the simplicity and low cost, this generation paradigm has also exhibited significant promise in multimodal tasks, including text-image retrieval Kritharoula et al. 2023, chat understanding Han et al. 2023, visual question answering Han and Gardent 2023, and multimodal instruction tuning Liu et al. 2023.

Appendix D Benchmark Datasets

In Table 1, we summarize representative benchmark datasets for evaluating models trained through data generation. Among them, ToolBench Qin et al. 2023 is generated by LLMs and is commonly employed to evaluate the performance of LLMs in tool usage proficiency. In most classification task evaluations Li et al. 2023c; Wang et al. 2023b; Sahu et al. 2022, LLMs are infrequently used as test models; instead, small language models trained on generated data are often used, followed by testing on existing benchmarks.