Harnessing large-language models to generate private synthetic text

Alexey Kurakin, Natalia Ponomareva, Umar Syed, Liam MacDermed, Andreas Terzis

Introduction

Machine learning models can memorize their training data and it is possible to extract the training data from a model . Training a model with differential privacy (DP) provably reduces the risk of memorization , which is critical when ML models are trained on sensitive data. However, DP training only ensures that the model does not release private information, and just releasing the model or its predictions is not adequate for many applications. For example, other researchers might want to use the data for analysis, or to build a different predictive model. It would therefore be ideal to release the dataset itself while protecting the privacy of the users that contributed to it.

Local differential privacy has been proposed as a method of preprocessing low-dimensional datasets before public release . Local DP adds noise to individual data points in the training data. While protecting privacy, local DP generally leads to much lower utility, due to the large amount of noise that must be added compared to central differential privacy, where DP is applied to the model or statistical output . Generally there is an the inherent tension between privacy and utility when releasing private datasets: we want to release a dataset that protects the privacy of the underlying data while at the same time we want the dataset to be as useful as the original data for any possible downstream task. Therefore, we focus on central DP and consider generating private synthetic data. Generating such synthetic data involves creating a generative model that learns the original data distribution. To protect the original data, either the generative model should be made private, via DP training, or privacy should be enforced at inference time (e.g., during the generation of synthetic data items, so-called private prediction). Private inference has been shown to be inferior to DP training when a large number of inferences is required . Since we seek to generate at least as much data as in the original dataset, DP training is the clear choice.

Several works proposed using publicly pre-trained large language models (LLM) for private synthetic data generation . This approach involves privately fine-tuning an LLM using class labels as prompts for the model and subsequently sampling from this model. However these attempts have had mixed success: they either reported poor utility even for non-private synthetic data, or had to augment standard NLP loss metrics to assist the LLM to correctly respond to prompts during the generation process. Additionally, none of the previous work considered privacy leakage from a pre-trained LLM itself. The privacy leakage happens because these papers used academic datasets (like IMDB ) as sensitive dataset and they utilized GPT-2 LLM which was pre-trained on these datasets without any privacy guarantees.

Although we follow a similar recipe conceptually, in that we use a DP-finetuned LLM model to generate private synthetic data, we highlight the following differences in our execution of this idea:

Privacy leakage mitigation. We draw attention to the need to account for the data that went into pre-training of the LLMs used for generation. Our de-duplication of the pre-training data ensures that no privacy leakage, possibly present in previous works, takes place.

Reporting: We use a long sequence of text (512 tokens, representing full reviews like IMDB or Yelp) as our privacy unit. Our privacy guarantees (Appendix A) are tight and transparent, and we tune the hyperparameters of the downstream classifier on private synthetic data only.

Method: We demonstrate that the standard approach to private fine-tuning does not yield the desired quality of generated data. Instead of augmenting the LLM’s objective or architecture for fine-tuning as in , we identify a loss function, well known to the NLP community, that is particularly suitable for private fine-tuning. Additionally, we argue that parameter-efficient fine-tuning, especially LoRA tuning, is beneficial for synthetic data generation

Our contributions can be summarized as follows:

We demonstrate state-of-the-art results in terms of quality of synthetic data. Specifically, we show in multiple experiments that the quality of the model trained on private synthetic data is comparable to or even better than the quality of the downstream model trained on real data with DP.

We demonstrate that parameter efficient fine-tuning like prompt-tuning and LoRA-tuning is superior to full fine-tuning when the tuning is performed privately. In particular, LoRA-tuning results in up to 11 percentage points lift in downstream model performance. To the best of our knowledge, we are are the first to demonstrate that parameter-efficient tuning performs better than full fine-tuning when each is combined with DP, whereas the opposite often holds for non-DP training .

We show that generating more synthetic data than the size of the original dataset is helpful, especially for simpler downstream models.

We show that DP synthetic data can be used to tune the hyperparameters of the downstream classifiers. We achieve ranking correlation with the ordering of trials performed on real data of up to 87%, even for ϵ=1\epsilon=1.

Related work

Privacy-preserving synthetic data generation requires that the generated data is both high-fidelity (i.e., exhibits similar distributional characteristics as the original data) and anonymized to preserve the privacy of the users who contributed their data. For complex data like text, images, audio and video, most existing approaches build a generative model, for example a GAN-based model . However, in most previous work the data is anonymized using heuristic methods, without providing formal privacy guarantees. For example, Melamud and Shivade attempted to de-identify summaries of clinical discharge notes using heuristic rules for an LSTM model and only empirically demonstrated the privacy of the synthetic data.

DP-fine tuning is a standard method for fine tuning LLMs that satisfies differential privacy guarantees and has been shown to perform well with appropriate hyperparameter tuning . DP-fine tuning involves using a pre-trained model and a modification of a training algorithm like DP-SGD to fine tune the model on private data.

For private synthetic text generation, Bommasani et al. suggested using a pre-trained GPT-2 model and then DP-fine tuning it on private data with word-level privacy, but did not implement or evaluate any method. In similar vein, Yue et al. DP-fine tuned pre-trained GPT models of various sizes. While they do obtain good results on some of the benchmarks, they also observe up to 25%25\% drop of downstream model accuracy on synthetic data (even without DP) on other benchmarks. Putta et al. attempted a similar recipe on a pre-trained distilGPT2 model, but also reported a large performance drop of the classifier trained on synthetic data. Additionally, they proposed modifying the fine tuning process to also include a discriminator that attempts to distinguish between the labels to improve the separability of learned representations for two binary classes of the text data. Similarly, Mattern et al. proposed augmenting the training objective with an additional term penalizing the generation of sample with the wrong label.

None of the prior work takes into account problem of data contamination between LLM pre-training dataset and dataset used in downstream task. As we show in Appendix D this problem is real. In particular some of both training and test samples examples from downstream datasets could be found in GPT-2 pre-training data, which is used by all prior work. This may potentially invalidate DP-guarantees and may result in overestimated accuracy on downstream tasks.

Additionally none of the works on DP synthetic data mentioned above explored parameter-efficient fine tuning. To the best of our knowledge, we are the first to demonstrate that parameter-efficient finetuning like LoRA tuning can produce better quality synthetic DP data than full finetuning.

Preliminaries

Differential Privacy (DP) is considered the gold standard for ensuring data anonymization. Throughout this work we employ a notion of DP called (ϵ,δ)(\epsilon,\delta)-DP.

Consider neighbouring datasets to be datasets that differ only in addition or removal of one record only. Given non-negative ϵ\epsilon and δ≤1\delta\leq 1, a mechanism A\mathcal{A} is (ϵ,δ)(\epsilon,\delta)-DP if for any two neighbouring datasets DD and D′D^{\prime} and for any S⊆Range(A)S\subseteq\text{Range}(\mathcal{A}),

The ϵ\epsilon and δ\delta values determine the strength of the privacy guarantees, with smaller values corresponding to stronger guarantees. The post-processing property of a DP mechanism means that applying any data-independent transformation to its output will remain DP with the same guarantees.

DP in context of ML models

In context of ML, DP can be introduced either at the input level, during the training of a model (DP-Training), or during model serving (prediction) . DP synthetic data falls into the first category and in general is a harder task than introducing DP during the training. This is because DP synthetic data ensures that any ML model trained on this data is DP with respect to the original training data. This is in contrast with DP-Training that only ensures that a particular ML model is DP. Therefore, it is expected that any model trained on DP synthetic data should perform at most as well as the downstream DP-Trained ML model on real data. However the idea of using a pre-trained generative LLM to aid generation of synthetic data means that we inject massive amount of public data, making the task of DP synthetic data generation less daunting.

The most practical methods of DP-Training for non convex losses are gradient-noise injection methods like DP-SGD , which work by clipping per example gradients to limit the sensitivity of the loss, and noising aggregated clipped gradients with Gaussian noise to make them private. The noise level is proportional to the clipping norm (the sensitivity) and the strength of ϵ\epsilon guarantees. The same recipe can be adopted to adaptive optimizers like Adafactor , where the noised gradients are passed to the optimizer to figure out the optimal learning rate.

LLMs

Throughout the paper we will use the terms of pre-training and fine-tuning of LLMs: pre-training is the initial training of a LLM with a large public dataset, for example C4 . Fine-tuning is an adaptation of a pre-trained model to perform some concrete task, for example question-answering, which involves running several epochs of an optimizer over the additional task training data.

Methodology

As a motivational example, consider the task of medical data sharing for research purposes: a medical provider has a sensitive dataset with patients records and wants to accomplish some machine learning task. They may want to share the dataset with external researchers and academic institutions to get their help in solving the downstream task, while preserving the privacy of the original data.

We assume that we have a sensitive dataset DD consisting of (Dtrain,Dvalid,Dtest)(D_{train},D_{valid},D_{test}), where the privacy of each record must be protected (see additional details on the unit of privacy in Appendix A). We want to accomplish some task on this dataset, such as training some downstream machine learning model. Additionally, we would like to allow a non-trusted third party to be able to perform the downstream task without violating privacy. To achieve this, we aim to create a synthetic dataset DsynthD^{synth}, which is DP with respect to the dataset DD. Our dataset DsynthD^{synth} will consist of synthetic training and validation splits. Figure 1 illustrates our methodology of data generation and evaluation:

Privately finetune (e.g., using DP-Training) a publicly pre-trained generative LLM GG on DtrainD_{train}, using DvalidD_{valid} for hyperparameter tuning. To tune hyperparameters for DP-Training, we follow an algorithm outlined in (Section 5.4.1).

Independently sample G to generate two new synthetic datasets DtrainsynthD^{synth}_{train} and DvalidsynthD^{synth}_{valid} which will serve as synthetic training and validation data.

Train a downstream model MM on DtrainsynthD^{synth}_{train} and use DvalidsynthD^{synth}_{valid} for hyperparameter tuning.

Evaluate the final performance of the model on real dataset DtestD_{test}.

Below we outline how the Prefix-LM way of splitting of the training example into input and target is advantageous for DP-training. Let’s consider some example from the dataset which is tokenized into input prefix p={z1,…,zk}p=\{z_{1},\ldots,z_{k}\} and target x={zk+1,…,zn}x=\{z_{k+1},\ldots,z_{n}\}. Typically, weighted next token prediction cross entropy loss looks like the following: L(z⃗,w⃗,θ)=−∑i=1nwizilog⁡P(z∣z<i,θ)L(\vec{z},\vec{w},\theta)=-\sum_{i=1}^{n}w_{i}z_{i}\log P(z|z_{<i},\theta) where θ\theta - model parameters, z⃗={z1,…,zn}\vec{z}=\{z_{1},\ldots,z_{n}\} is tokenized training example (including input and target tokens), each ziz_{i} as one hot encodings of token, P(z∣z<i)P(z|z_{<i}) is the probability of ii-th token given values z<iz_{<i} of all previous tokens and w⃗={w1,…,wn}\vec{w}=\{w_{1},\ldots,w_{n}\} is weights vector for each token in the loss.

Standard next-token prediction loss assigns weights wi=1w_{i}=1 to all tokens, including those in the prefix pp. As a result prefix tokens will be included in the gradient of the loss ∂L∂θ\frac{\partial L}{\partial\theta}, thus essentially forcing the model to learn the distribution of tokens in the prefix as well. On the other hand, the Prefix-LM formulation assigns zero weights to the prefixOriginal paper only describes bidirectional attention over prefix and omits the description of loss weights. Nevertheless zero weighting of the prefix is implemented in the T5 code., i.e. ∀i≤k:wi=0\forall i\leq k:w_{i}=0, so the total loss looks like the following: LPrefixLM(z⃗,w⃗,θ)=−∑i=k+1nzilog⁡P(z∣z<i,θ)L_{\text{PrefixLM}}(\vec{z},\vec{w},\theta)=-\sum_{i=k+1}^{n}z_{i}\log P(z|z_{<i},\theta)

As a result the LLM is not forced to learn distribution of input prefix pp which we found to be beneficial for differentially-private training. DP-Training adds the noise to all the gradients, in a standard setup this will result in the gradients from the prefix portion being corrupted with the noise. This in turn means that prompting the DP-Trained LLM to generate synthetic data will not work as well as expected. We believe this is the same phenomenon that was observed in works Putta et al. and Mattern et al. , where authors had to add an adversarial head or augment the loss respectively, to aid the model in differentiating different types of prompts. Prefix-LM in turn is a standard loss well known to the community, and this comes with the benefits of knowing approximate hyperparameter values for its tuning. The aforementioned Prefix-LM setup allows to train one model for all the class labels and can be easily extended beyond the binary classification setup.

2 Parameter-efficient fine tuning

Full finetuning of large models is expensive, and, empirically, tuning very large number of weights with DP-finetuning often results in substantial utility drop. Many techniques exist that update the pretraining model without resorting to full model weights update. In this work, we consider two popular ones - Prompt Tuning and LoRA.

Prompt tuning is a technique which prepends a small prompt tensor in front of the model’s input in the embedding space, freezes the rest of the model’s parameters and then finetunes only the prompt tensor weights. We found that combining prompt tuning with differentially private training allows us to achieve much higher utility of the trained generative model compared to full model fine-tuning. This could be explained by the fact that the prompt tensor is much smaller compared to the size of entire model (we used prompt tensor with 20480 parameters vs 8B weights in the full model) and smaller models tend to have smaller gap between private and non-private utility , probably due to the total amount of noise injected during the training. It should be noted that prompt tuning as described in original paper showed very poor utility when trained with differential privacy. We observed that even in the best runs LLM quality metrics (perplexity, next token prediction accuracy) fluctuated significantly. No amount of typical hyperparameter tuning could improve prompt tuning utility in DP-regime.

Borrowing some ideas from and experimenting with various optimizers and ways to initialize prompt tensor proved to be the key to making prompt-tuning work. Eventually we found out that the main culprit of poor utility was prompt tensor initialization. initializes prompt tensor by using embeddings of some real tokens from vocabulary. Changing prompt tensor initialization to random uniform with small range [−0.01,0.01][-0.01,0.01] significantly improved utility. Additionally we observed that change of optimizer from Adafactor to Adam or Momentum helped to make training more stable, which simplified hyperparameter tuning (Appendix E).

LoRA tuning (Low-rank Adaptation) is a technique that freezes all the pre-trained model weights and introduces trainable low-rank decomposition matrices into each dense layer (MLP and Attention). This results in fewer trainable weights than full fine tuning but the number of trainable weights in LoRA is significantly larger than in Prompt tuning. For example, rank 8 LoRA updates 20M trainable parameters, as opposed to 41K prompt tuning vs 8B full fine tuning. Empirically we find (Section 5) that LoRA results in superior performance, surpassing that of both full finetuning and Prompt finetuning and that tuning both MLP layers and Attention blocks is preferred, see Appendices F and J.5 for more details.

As a conclusion, we advocate for the use of parameter-efficient techniques when performing DP-training, with LoRA being the most promising so far.

3 Data sampling

Experiments

In our experiments we used a model with architecture similar to Lamda 8B , which we pre-trained on The Pile dataset using a standard next-token prediction loss. We stress that for our experimental results to be valid we must ensure that the pre-trained model was not itself trained on data that is considered private for the downstream task. For example, the GPT-2 model used in seemingly contained IMDB data in its pre-training dataset , but this model was subsequently used to generate a synthetic version of IMDB, see also appendix D for details. To prevent privacy leakage we modified the pre-training dataset by de-duplicating it against all sensitive datasets used in downstream tasks, following the recipe and scripts from . The outline of the de-duplication approach is as follows. First we tokenized and constructed a suffix array for each involved dataset (The Pile, IMDB, Yelp, AGNews). Then we used the suffix arrays to find common sequences of 50 or more tokens which appear in The Pile and any other dataset. Finally we cut all those common sequences from The Pile dataset. Note that this de-duplication is “stronger” than simply removing the datasets from the Pile. After cutting the sequences we de-tokenized the dataset back to strings and used it for pre-training. Refer to Appendix C for additional details.

Datasets and classification problems

We conducted our experiments on IMDB , Yelp and AGNews datasets. All these datasets only provide a training and test set, so in each case we use the first 90% of the training set for training and the remaining 10% for validation. For each dataset we formulated a binary classification problem (sentiment classification) as the downstream prediction task.

1 Downstream classifier performance

We investigate the utility of using private synthetic data for a downstream task. For each dataset, we consider two types of models. First one is a (encoder-only) BERT model with classification head. BERT is publicly pretrained and then fine tuned using either real data or our generated synthetic data. This model benefits from public pre-training data. We also consider a word-level CNN model that does not utilize any public data. For each model, we report the performance on real data with no DP guarantees (an entry "Real" with ϵ=∞\epsilon=\infty in Table 1). This serves as a upper bound of downstream classifier performance. We also report the performance of doing DP-Training on the downstream classifier directly (entries "Real" with ϵ∈(1,3,10)\epsilon\in(1,3,10), referred to as "DP-on-real" in the text) and report the results on synthetic data generated from fine-tuned (Fine-tuned-SD), prompt tuned (Prompt-tuned-SD) and LoRA-Tuned (LoRA-tuned-SD) models. We would like to highlight however that in the case of using the real data directly for DP-Training, only the resulting downstream model is DP, and the real data can’t be shared freely or used for hyperparameter tuning (or such tuning should be accounted for in privacy guarantees). DP Synthetic data however can be shared freely and used for feature engineering, hyperparameter tuning, etc.

Firstly, our results in Table 1 indicate that obtaining good fidelity non-private synthetic data is possible, contrary to the results reported in and Putta et al. . Both Fine-tuned-SD and LoRA-tuned SD exhibits better performance than Prompt-tuned-SD, in line with current understanding that for a non-DP setting, tuning more model parameters is beneficial . Interestingly, even for non DP setting, downstream models trained on LoRA synthetic data outperform those trained on fully fine-tuned synthetic data in 2 out of 3 datasets.

Private synthetic data

While there is a clear utility drop when going from non-private SD data to private SD, DP LoRA-tuned-SD is a clearly superior way of obtaining DP synthetic data. Prompt-tuned DP SD is better than fully fine tuned DP SD, however LoRA outperforms the Prompt-tuned DP synthetic data in majority of the cases. We hypothesize that it might be due to less total noise being added in DP LoRA models, due to fewer parameters being updated than with the full fine-tuning. Prompt tuning on the other hand updates the minimal number of parameters, however this minimum update hurts the utility of SD, suggesting that like with everything in ML, there is a “sweet spot” on the number of parameters trained with DP.

The difference between the performance is significant, with LoRA-tuned-SD exhibiting of up to 10-11% lift on downstream BERT classifier tasks, compared to model trained on fine-tuned-SD. For CNN model that is more dependent on the quality of the data than BERT (that essentially reaps some additional benefits from Transfer Learning), the results are even more significant, with a boost from prompt-tuned-SD (vs fine-tuned-SD) reaching up to 22%.

Private synthetic data vs DP-Training on real data

To obtain a DP downstream ML model, we can either use DP synthetic training data or introduce DP directly during downstream model training (DP-on-real). As previously mentioned, the former is a harder setup. When comparing BERT models, we can see that private LoRA-tuned-SD achieves performance similar or even superior (e.g., for IMDB and Yelp datasets) to DP-on-real for all levels of privacy, but an additional benefit of such synthetic data is that it can be additionally shared freely and used for hyperperameter tuning and feature engineering. For CNN model, LoRA-tuned-SD (and even prompt-tuned SD) exhibits better performance than DP-on-real. This is due to the fact that private synthetic data benefits from massive amount of public data that was used for pretraining of the LLM (CNN model itself is trained from scratch, as opposed to BERT that is itself a pre-trained model, albeit with smaller amount of public data than the 8b Lamda model we used for SD generation). This indicates that for simpler models synthetic data can be a preferred way of injecting additional public knowledge. This is an interesting result since it is commonly assumed that for Transfer Learning to work, public data should come from a similar distribution as the target data. However in case of synthetic data, we inject public data from different distributions (crawl of the web) than that of the downstream task (e.g. Yelp reviews).

Amount of synthetic data vs downstream classifier performance

We studied how much synthetic data we should generate relative to amount of real data. Table 2 demonstrates that generating more synthetic data can be beneficial, but has diminishing returns for BERT (0.8% lift going from 1x to 3x times the data), with benefits more pronounced for simple models like WordCNN (1.4% lift from increasing the amount of synthetic data 3x).

One can also potentially combine the synthetic data with training with DP on real data, by pre-training the downstream model with DP synthetic data and then fine-tuning with DP on real data. This will however require spreading the privacy budget between DP synthetic data and DP-Training of the downstream classifier. We leave this for future work.

Comparison with prior work

While works below don’t provide sufficient (or any) information on their privacy unit (as we do in Appendix A), we assume that privacy unit that was used is one example (e.g. 1 full yelp or imdb review etc); we also assume central DP setting, that δ\delta values are the same or comparable etc. Additionally, none of the works below take into account the fact that pre-training data might have contained the data they deem private (as we highlight in Appendix D), potentially invalidating their reported DP guarantees.

Yue et al. used Yelp dataset for multiclass (rating) classification, so our results are not directly comparable. Putta et al. used AGNews dataset. Their approach is a combination of next token prediction (similar to our setup) and additional loss term from a new head that attempts to learn to distinguish between various classes directly (instead of simply relying on the prompts in the next token prediciton head). Putta et al. reports 0.867 accuracy of downstream task for ϵ\epsilon of 3, while we obtain 89.6 (the baseline performance of downstream classifier for our and their work is comparable, 0.938, suggesting that we are using comparable downstream classifiers). Mattern et al. suggested a modification of the loss (prompt-mismatch loss, to discourage the generation of text inconsistent with the prompt, like generating a negative review when positive prompt was given). They performed experiments on IMDB dataset. Their best IMDB experiments reporting worse accuracy on DP synthetic data (89.1%89.1\% theirs vs 90.6%90.6\% ours for ϵ=3\epsilon=3). They also seem to have worse performance on real data despite using the same model (BERT-classifier).

2 Tuning downstream model hyperparameters on synthetic data

With the following experiments on IMDB data, we want to demonstrate that private synthetic data is useful for hyperparameter tuning of the downstream classifier. For all our experiments, when tuning the downstream classifier, we use validation accuracy on set-aside portion of synthetic data for hyperparameter selection. We tune weight decay and learning rate for both CNN and BERT models. For synthetic data, we create vectors of accuracy on validation (synthetic) data and performance on real test data for all possible combinations of hyperparameter values tried. We then report the ranking correlation between performance as indicated by validation accuracy (synthetic data) and test accuracy computed on real data. We also report the ranking correlation of accuracies on real validation and real test data, to provide an upper bound. Additionally, we report rank-biased overlap ranking metric , which is a weighted metric that gives more weight to the top of the ranking (we use parameters that give 85% of the weight to the first top 25% of ranking).

Table 3 demonstrates excellent ranking correlation on synthetic data. Interestingly, prompt-tuned synthetic data metrics, in particular the mean and standard deviation of the top 25% of trials, suggest that BERT classifier performance is less sensitive to hyperparameters on better fidelity (prompt or LoRA tuning) data than on worse fidelity data (fine-tuning).

3 Estimating synthetic data quality

It is useful to have an efficient method of evaluating the quality of a synthetic dataset without relying on specific downstream tasks. For one, a key use case for privacy preserving synthetic data is to enable data sharing without a definitive end use case. For another, training the generator LLM has multiple hyperparameters that can be tuned, and it can be prohibitive to evaluate candidate models using full data synthesis and downstream training (which itself might require tuning hyperparameters). Instead, lighter weight proxy metrics can be used. Commonly used proxy metrics are: perplexity, n-gram statistics, and MAUVE . We investigate the effectiveness of each of these metrics by comparing their correlation to downstream performance (Table 4). These metrics are used to compare datasets, and thus their absolute value is uninformative. For n-gram statistics we determine the frequency of unigrams, bigrams, and sample lengths in characters for both the original and synthetic datasets. We then compute the area under the divergence frontier between these two frequency distributions as is done by MAUVE. MAUVE works by computing the difference between two datasets by first embedding each example, then clustering the datasets, followed by comparing (via divergence frontiers) the histogram of cluster membership across the two datasets. It has recently been shown to be an effective metric for synthetic text datasets , which our results support. We compute the MAUVE score as given in Pillutla et al. using the suggested hyperparameters unless noted. We investigated modifying these hyperparameters and confirm they make little difference to the relative ranking, with the notable exception of the model used to embed examples. Unlike the original paper, we find larger models to be much more effective. In particular, embedding using Sentence-T5 has much higher correlation to downstream performance than BERT or any other model we tried. For more details see appendix K. Our results match many of the results given in Kour et al. . All metrics are at least somewhat noisy with standard test-set perplexity performing very well. Given its ease to compute while finetuning, perplexity is our recommended proxy metric when available.

Conclusion

We have shown that training downstream models on DP synthetic training data is an effective alternative to training such models with DP directly on real data for text classification tasks. We explored two methods for privately generating the synthetic training data, both of which involve modifying the weights of an existing LLM. One method privately fine-tuned all the layers of the LLM, while the other method used parameter efficient fine tuning ( ‘prompt-tuning’ and ‘LoRA-tuning‘). Our experiments demonstrated that LoRA tuning is a superior way of obtaining DP-synthetic data, which provides performance on the downstream task that is comparable or even better than directly DP-Training on real data. We showed that the standard NLP Prefix-LM loss is well suited for DP-finetuning. Private synthetic data can be used freely for all the purproses, such as feature engineering, hyperparameter tuning, debugging and monitoring, and sharing, but without any privacy-related concerns. We also showed that while Mauve is a good proxy metric for evaluating the quality of the synthetic data, simpler metrics like perplexity, when available, perform well.

Ethics statement

We expect that our proposed method of generating DP synthetic data will facilitate safer data sharing and that the societal impact will be positive, since entities who own private data but do not necessarily have the knowledge or resources to train predictive models can share private synthetic data with specialists for model creation, benefiting from their expertise without comprising the privacy of the users who contributed their data. The main limitation of our approach is that we only conducted experiments on English datasets, however we expect that methods should work on multilingual datasets as along as public multi-lingual data are available for LLM pre-training.

Reproducibility Statement

All of our experiments are based on open sourced frameworks and public datasets, refer to Appendices H and M. We further provided necessary details to reproduce our experiments in Appendices C, E, F, G and I.

References

Appendix A Our privacy guarantees

To provide all the information needed to understand our privacy guarantees, we follow the guidelines outlined in .

DP setting. We provide central DP guarantee where the service provider is trusted to correctly implement the mechanism.

Data accesses covered: Our DP guarantees apply only to a single training run. We don’t account for hyperparameter tuning in our guarantees.

Final mechanism output: We use DP-Training methods. Only model’s predictions (e.g. synthetic data generated by the DP-Trained model) is released, however the mechanism’s output is technically the full sequence of privatized gradients, and the guarantee also applies at this level. This also means that all the checkpoints are protected/can be released publicly.

Unit of privacy. We consider example-level DP, where each example is a long sequence of text, e.g. a full Yelp review. Maximum length of such unit in tokens is 512, tokens are extracted by SentencePiece algorithm trained on c4. We consider full data protection (full text of a review).

Adjacency definition for “neighboring” datasets": We use add-or-remove adjacency definition.

Type of accounting used: RDP-based accounting.

Accounting assumptions : Poisson sampling was assumed for privacy amplification but shuffling was used in training)

The formal DP statement: We use various levels of ϵ\epsilon values: 1,3,10. Our δ=1training_data_size\delta=\frac{1}{training\_data\_size}

Transparency and verifiability: We are going to open source our code. We use open-sourced t5x framework.

Let DD be a dataset of examples, where each example is a sequence of tokens (such as a full Yelp or IMDB review) provided by a user. We train LLMs to generate synthetic data using a next-token prediction objective. Therefore, when training an LLM, each example is used to generate several sub-examples, one per token in the example. Concretely, consider an example in the dataset consisting of the sequence of tokens t=(t1,…,tn){\mathbf{t}}=(t_{1},\ldots,t_{n}). The ii-th sub-example generated from this example consists of the feature vector xi=(t1,…,ti−1){\mathbf{x}}_{i}=(t_{1},\ldots,t_{i-1}) and the label yi=tiy_{i}=t_{i}. In other words, for each sub-example, the preceding tokens are used to predict the next token. Let w{\mathbf{w}} be the trainable parameters of the model. The loss of example t{\mathbf{t}} is defined

We consider several weighting schemes. Standard normalization sets αi=1n\alpha_{i}=\frac{1}{n} for each sub-example ii, while Prefix-LM normalization, first proposed by Raffel et al. , sets αi=0\alpha_{i}=0 for all i≤ki\leq k and 1n−k\frac{1}{n-k} for all i>ki>k whenever the first kk tokens of an example correspond to the prefix that specifies the learning task.

for our desired values of ϵ≥0{\epsilon}\geq 0 and δ∈\delta\in. Therefore, our privacy guarantees apply to any change to any single example, and not just changes to a single sub-example.

Appendix B Additional related work

While in our paper we concentrate on generating synthetic text data, the task of generating synthetic tabular data has been explored extensively before.

For tabular data, early works on synthetic data concentrated on estimating the utility of the synthetic data as a quality of the statistical answers over the data (synthetic data for query release). Privacy-preserving aspect for such synthetic data was often achieved via DP methods More recently, works that instead evaluated how useful the tabular synthetic data is for some downstream ML model started gaining popularity . In general, the approaches for generating synthetic tabular data can be categorized into Marginal-based and generative models-based. Marginal-based models calculate private marginal distribution by taking marginal distribution over attributes and appropriately privatizing it with some DP-mechanism like Gaussian or Laplace. Then attributes for synthetic instances are sampled from these distributions. Generative models instead attempt to build a function that approximates the data distribution and sample from this function. GAN-based methods were a common choice for such generative models: e.g., DP-GAN , Convolutional GAN and PATE-GAN . A recent benchmark reported that for achieving best downstream ML model performance, marginal-based methods still outperform GAN-based approaches.

Appendix C Language model pre-training.

Model architecture. For synthetic data generation, one can use either decoder only or encoder-decoder models. We used a decoder-only transformer model , which should be more parameter-efficient, with architecture similar to LaMDA . The quality of the pre-trained model is of paramount importance for generating good fidelity synthetic data. We initially experimented with 1B model and found that a significant boost in downstream performance can be achieved by using a higher capacity model.

Therefore we use 8B model for final fine-tuning and prompt tuning. A smaller version (1B) model is used to tune hyperparameters like learning rate, clipping norm, batch etc. The hyperparameter values found using 1B model are used for the final 8B model. Statistics for 1B and 8B models are provided in the Table 5.

Pre-training data. We used public dataset The Pile as a basis for our pre-training data. However we run an extra post-processing step by deduplicating The Pile against all datasets used in downsteam tasks. This deduplication step is necessary to ensure that private data won’t be accidentally learned during model pre-training, which otherwise would invalidate our DP guarantees.

To do deduplication we followed the recipe outlined in . Specifically, we tokenized and constructed a suffix array for each involved dataset (The Pile, IMDB, Yelp, AGNews). Then we used these suffix arrays to find common sequences of 50 of more tokens which appear in The Pile and any other dataset. Finally we cut all those common sequences from The Pile dataset. After cutting the sequences we de-tokenized dataset back to strings and used it for pre-training. It’s important to note that we only deduplicate The Pile against other datasets, we did not run deduplication of The Pile against itself.

Pre-training procedure. We pre-trained our models using T5X codebase , however we adopted few tricks which were used to train open sourced GPT-NeoX model . Details are below.

We pre-trained a model for 380k steps using batch size of 1024 example with 1024 tokens sequence length, which result in training for approximately 400B tokens. We used same SentencePiece tokenizer which was used in original T5 model . Cross entropy loss on next token prediction (teacher-forcing) was used as a training objective. Additionally we employed weight decay of 0.0010.001 and an auxiliary z-loss 10−4log⁡2(Z)10^{-4}\log^{2}(Z) where log⁡(Z)\log(Z) is softmax normalizer. Training was done with AdaFactor optimizer with learning rate schedule min⁡(0.01,1N)\min(0.01,\frac{1}{\sqrt{N}}) where NN is a step counter.

Appendix D Analysis of pre-training data contamination of GPT-2 training data.

Multiple prior works on private synthetic data generation start with a pre-trained GPT-2 model and then finetune it with differential privacy on some downstream dataset. In this section we demonstrate that pre-training data for GPT-2 model includes examples from downstream tasks which invalidates the privacy guarantees of such prior work.

Let’s assume we pre-train model MM on some public dataset DpD_{p} and finetune with DP-SGD on sensitive dataset DsD_{s}. If Dp∩Ds=∅D_{p}\cap D_{s}=\emptyset then we can conclude that the process of obtaining our model MM is differentially private w.r.t. DsD_{s}. However this is no longer the case if intersection of DpD_{p} and DsD_{s} is non-empty.

GPT-2 was pre-trained on a WebText dataset . While the dataset is not released publicly, authors discuss in detail the process of how dataset was constructed by scraping the content of links posted on Reddit website. Additionally, they released a list of top domains https://github.com/openai/gpt-2/blob/master/domains.txt from which the dataset was formed. Specifically, it could be seen that 183080 web-pages from IMDB and 36188 web-pages from Yelp websites were included GPT-2 pre-training data. This by itself, indicates that parts of IMDB and Yelp dataset were include in GPT-2 pre-training data.

We performed further analysis in the following way. We took a subset of OpenWebText dataset - a public re-implementation of WebText. Specifically we used c4/webtextlike from TFDShttps://www.tensorflow.org/datasets/catalog/c4#c4webtextlike. We computed an intersection of IMDB dataset and c4/webtextlike using the approach from . We found that c4/webtextlike contains 136136 distinct examples from IMDB training and test set.

Finally we manually looked at a few of the examples to verify that they are indeed part of IMDB dataset and that they were likely used to construct WebText dataset. Here is a code snippet to obtain one of these examples:

This indicated that indeed at least some portion of the IMDB reviews would have been obtained during WebText dataset construction process and would have been used for GPT-2 pre-training. While it is impossible to say, without an access to the actual pretraining data, what percentage of the downstream tasks data was included for GPT-2, it highlight the importance of our deduplication process for obtaining rigorous privacy guarantees.

Appendix E Details of prompt-tuning.

Training instability with Adafactor optimizer and default prompt initialization. recommends Adafactor as a default optimizer. They also suggest to initialize prompt using pre-trained embeddings of tokens from vocabulary. While we confirm that it works well for non-private prompt tuning of LLM, it appeared to be inadequate for DP-prompt tuning.

With this setup we observed significant training instabilities even in the best DP-runs after significant amount of clipping norm and learning rate tuning, see figure 2.

Changing optimizer and prompt tuning initialization.

To achieve good performance with prompt tuning we experimented with two things.. First of all, we tried various optimizers (including Adam and Momentum) instead of Adafactor. Additionally we experimented with random initialization of prompt tensor.

While changing optimizer didn’t really improve the downstream performance, we found that Adam and Momentum optimizers lead to more stable training runs. Changing prompt tensor initalization to random uniform with small range was the key for improving prompt tuning performance with DP. Table 6 demonstrates that random initialization results in a lift of up to 30% in downstream CNN performance, and up to 10% for BERT model. We hypothesize that proper initialization is very important to DP-Training and addition of noise makes it hard for the model to "recover" from bad initialization.

We did some limited experiments with the scale of random uniform initialization, however for most of them we didn’t run full synthetic data pipeline and only compared LLM perplexity and next token prediction accuracy. These experiments showed that random uniform from range [−0.01,0.01][-0.01,0.01] was the best, while both decreasing and increasing the range led to worse metrics.

We also compared Adam optimizer with fixed learning rate, Adam optimizer with cosine learning rate decay and Momentum optimizer with cosine learning rate decay. As long as learning rate was tuned, all three performed similarly. Eventually we settled on Adam optimizer with fixed learning rate.

Appendix F Details of LoRA

Let’s say one of the layers in the network is represented as dense matrix multiplication operation WhWh where WW is a trainable weight and hh is an input to the layer (typically embedding or hidden state vector). LoRA proposes to replace weight matrix with a sum W+LRW+LR, freeze WW and tune LL and RR matrices. In this notation LL and RR are low-rank matrices such that the result of their multiplication has the same shape as WW. For example if WW is a matrix with the shape n×mn\times m, then LL would be n×rn\times r matrix and RR would be r×mr\times m matrix, where rr is the rank of low-rank adapter.

Similar to , all layers where LoRA is not applied are considered frozen in our setup. We performed LoRA tuning using Adam optimizer with fixed learning rate. Unlike we did not use weight decay because we observed no benefits of weight decay in differentially private LoRA training.

Typically, in a transformer model LoRA can be applied to attention layers and MLP layers. Hu et al. study LoRA only on attention layers. In this work, we tried to applied LoRA to attention layers only (Attention-LoRA), MLP layers only (MLP-LoRA) and both attention and MLP layers (Full-LoRA). Overall we found out that Full-LoRA is the best choice in most cases, however in some cases MLP-LoRA could be slightly better.

We also studied how rank of LoRA affects downstream performance. Similar to we observed that initially increase of LoRA rank lead to increase of accuracy on downstream task, followed by eventual decrease of downstream accuracy with further increase of rank.

Our main results in table 1 are reported using Full-LoRA rank 8 on AGNews, rank 1 for Yelp datasets, and MLP-LoRA rank 32 on IMDB dataset.

Overall we can recommend to use Full-LoRA rank 8 as a good default value, however we can encourage tuning of LoRA parameters when resources allow it and the best possible performance is needed.

See also detailed ablation of LoRA parameters in Appendix J.5.

Appendix G Details of synthetic data sampling.

For data generation/sampling part, LLMs have the following knobs:

temperature: this is a constant by which the logits are divided prior to softmax and subsequent sampling. Large temperature flatten the tokens distributions, making rarer tokens more likely to be selected; it also increases diversity of the data generated but can negatively affect the quality. Our experiments show that default temperature of 1 (so no modification of the tokens distribution) works the best.

topk: given a token distribution, topk determines what portion of the distribution to keep before sampling. Topk is similar to low temperature - e.g. setting topk to top 1000 tokens means that next token will be chosen from the most likely tokens. We keep this parameter unmodified (set to ∞\infty).

numdecodes: is similar to the number of candidates in standard beam search. We use the value of 1 here, resulting in the sequence where each token is the most likely to be returned.

While we did experiment with various values of these hyperparameters (see appendix J.3), we find that default settings already provide enough of diversity of generated examples for our datasets. If further diversity needs to be enforced (e.g. for very large datasets), we recommend increasing the temperature or numdecode values.

Appendix H Details of datasets for downstream tasks

We conducted experiments on IMDB reviews , Yelp reviews and AGNewshttps://www.tensorflow.org/datasets/catalog/ag_news_subset datasets. On IMDB and Yelp we formulated downstream task as binary sentiment classification. On AGNews the downstream task was classification of titles of news articles into one of 4 topics (World, Sports, Business and Sci/Tech).

All of the considered datasets only provided training and test sets. To obtain validation set we split original training set into two chunks in deterministic way using TFDS split slicing API. We used first 90% of the original training set for training, and last 10% for validation.

Some of the dataset statistics for considered dataset and our train/test/validation splits is provided in table 7.

Appendix I Architecture and training of downstream models.

We used two types of models for all downstream experiments. First one is a BERT encoder with a dense layer on top of it, second one is a shallow CNN.

BERT model. For BERT-based classifier we used BERT-Base model pretrained on English datahttps://tfhub.dev/tensorflow/bert_en_uncased_L-12_H-768_A-12/4 with standard BERT tokenization and preprocessing. We put a dense layer on top of pooled output of the BERT encoder to produce classification score. Output of dense layer was a single floating point number per input sequence which was converted to probability using sigmoid function. We trained entire model (including BERT encoder and dense layer) without freezing any layers.

CNN model. In addition to BERT, we used a shallow CNN model without any extra pre-training. Our model architecture followed the idea from with the main difference that we didn’t pre-train embeddings beforehand. To be more specific, our model works as follows. We convert input sequence to lowercase and tokenize it by splitting it on whitespaces and punctuation signs. Then we embed it into 384 dimensional embedding, using vocabulary from most common 30k words from the dataset. Embedding was randomly initialized and trained as a part of downstream training task. Output of embedding layer was passed through 1D convolution with kernel size 3, 256 output filters and RELU activation. This followed by global max pooling along the sequence length. Then we have a fully connected layer with 256 outputs and RELU activation, followed by final logits layer with sigmoid activation.

Training of downstream model. Both BERT and CNN were trained in a similar way. For non-private training we used Adam optimizer and DP-Adam (as implemented in the Keras DPModel library) for private training. Model was regularized with weight decay and we did early stopping (by non-increasing validation accuracy). We choose best hyperparameters by running a sweep over learning rates {10−7,…,10−1}\{10^{-7},\ldots,10^{-1}\} and weight decays {5×10−1,…,5×10−6,0}\{5\times 10^{-1},\ldots,5\times 10^{-6},0\}. Additionally for DP-training baseline on real data we did a sweep of clipping norm over {10−2,…,102}\{10^{-2},\ldots,10^{2}\}.

Appendix J Ablations

We study the influence of various factors on quality of generated synthetic data. We observed that better pre-trained model (in terms of model size and choice of pre-training dataset) leads to better synthetic data in differentially-private case, see details in Appendix J.1.

All standard techniques which are used to improve utility of DP-training apply to both full fine-tuning and parameter-efficient tuning of LLMs. Specifically, it’s usually better to train longer and with larger batch size . Also it’s important to do a hyperparameter sweep over learning rate and clipping norm, see appendix J.2. We also experimented with various parameters of temperature sampling for data synthesis. We found that sampling with T=1T=1, without truncating sampling distribution (i.e. TopK=∞\text{TopK}=\infty) and without performing any filtering usually well and is a reasonable default setting, refer to Appendix J.3.

We also looked into variations of loss formulation for LLM training (Appendix J.4). As explained in Section 4.1 we found that Prefix-LM formulation usually leads to better performance in DP setting, compared to Full LM. We observed that normalizing LLM loss by number of non-padding tokens, as suggested in , makes it easier to tune LLM hyperparameters, especially clipping norm for DP. It also makes it easier for LLM to learn to produce outputs of various length.

We looked into influence of LLM pre-trainind dataset and model size on performance on downstream task. As shown in table 8 larger size of LLM translated into better downstream performance for both BERT and CNN models.

Additionally, we did some limited study comparing pre-training on C4 and The Pile datasets. Initially we expected that C4 (which is essentially a web crawl) better matches text distribution in Yelp and IMDB datasets compared to The Pile (which was intentionally composed to be more diverse). However our experiments showed similar performance on both datasets, so eventually we settled on The Pile, which is easier to download and use.

J.2 Ablation: hyperparmeter tuning for DP-training of LLM

Batch size and number of training steps. For our prompt tuning run on IMDB with ϵ=1\epsilon=1 we did a thorough sweep of various batch sizes and number of training steps of LLM prompt tuning. Results are summarized in table 9 and confirm the general observation that longer training with larger batch is typically better when trained with DP.

Learning rate. In our experiments we found that downstream performance was quite sensitive to learning rate of LLM, see table 10.

J.3 Ablation: sampling parameters.

After training a generative model, we still need to use that model to create a differentially private dataset. Large language models generate sequences one tokens at a time in an autoregressive manner; given previous tokens the model’s final layer emits a probability distribution over all possible next tokens. Instead of sampling directly from this probability distribution, it is common to modify it in three ways:

Temperature: Shaping the token distribution using temperature tt so that the final softmax over the final logits uiu_{i} gives the next token probably as

Top K: Truncating the token distribution so that only the kk most likely tokens are sampled from. All other tokens are given zero probability.

Num Decodes: Repeating the full sampling process (i.e. decoding) N times and then out of the N candidates return only the sample with the highest likelihood.

To analyze the effect of these parameters on dataset quality we performed a sweep over these parameters and computed the downstream performance of models as discussed in section I. Results are shown in tables 11 and 12.

Results from our ablation study on sampling parameters show that while small gains can be gained from tuning these parameters, such gains are modest compared to using the default parameters of t=1.0t=1.0, top-k =∞=\infty, and num decodes =1=1. We note that slightly higher temperatures appear to help in both cases (1.4 for IMDB and 1.2 for Yelp) and in both cases should be paired with either tighter top-k or an increase in decodes. However, using additional decodes are computationally costly and likely not worth the additional cost. Using a temperature less than 1.0 never seems to help.

J.4 Ablation: loss of LLM.

LLMs are commonly trained/fine-tuned with next-token prediction teacher forcing. The loss for each token in this setup is a cross-entropy loss for each token. For an instance that is a collection of tokens (e.g. a full yelp review), the loss is therefore is a sum over per-token losses. It is common to normalize this loss by the number of non padding tokens, which roughly translates into the normalization by the batch size and the average number of tokens for instances in the batch. We recommend to follow this normalization scheme because it makes it much easier to find the appropriate clipping norm for DP-Training. When the loss is not normalized, the appropriate clipping norm can be in the thousands. With normalized loss, a standard clipping norm of 1 or 3 usually works out of the box. For example, for Yelp dataset, without the loss normalization the clipping norm was found to be approx 2000, with accuracy of the fine-tuning of approx 0.29. With loss normalization, the clipping norm of 1 resulted in performance of 0.43.

Note that if it’s feasible to perform full hyperparameter sweep of clipping norm and learning rate, then benefits of loss normalization diminishes. Nevertherless even in this case loss normalization can provide small advantage, see Table 13.

J.5 Ablation: LoRA parameters.

We conducted detailed sweep of LoRA parameters (rank and in which layers to introduce LoRA) on IMDB and AGNews, see Tables 14 and 15. As could be seen from these tables, Full-LoRA performs better than Attention-LoRA and MLP-LoRA in most cases on IMDB dataset. However MLP-LoRA seem to be better choice overall on AGNews. If practitioner has to pick parameters to introduce LoRA beforehand without tuning, then we would recommend Full-LoRA as a reasonable default.

In terms of rank, best performance is typically achieved for ranks in range $.Fromourexperiments,itdoesnotmakesensetoincreaserankabove. From our experiments, it does not make sense to increase rank above32$ because it result in little to no performance gains, however it is more expensive because requires tuning of more parameters.

Appendix K Evaluating synthetic data quality: Mauve robustness.

Mauve has multiple parameters that control its behavior. The most influential such parameters include: the degree of dimensionality reduction (PCA explained variance), the number of clusters to use, the number of samples to use, and the model used to initially embed the samples. Pillutla et al. came to the conclusion that while these affected performance none of these parameters mattered enough to worry about needing to tune for a specific application. They recommended a default setting of buckets =500=500 and explained variance =0.9=0.9, We performed our own ablation studies on these parameters (see Figure 3). As noted in the paper, while we found MAUVE to be robust to most of these parameters, the model mattered a great deal.

Appendix L Evaluating synthetic data quality: Proxy metric experiments.

Proxy metrics are a less expensive method for estimating synthetic data quality. They are particularly useful for helping tune the large number of hyperparameters needed to train and sample from the generator model. This includes tuning the model type (model architecture, pre training process, fine-tuning, prompt-tuning, etc..), the hyper parameters of model training (e.g. epsilon, clipping norm, learning rate, batch size, epochs, etc..), and finally the hyperparameters of the sampling procedure needed to create the final dataset (e.g. temperature, top-k, and decodes). We examined how well various metrics correlate with downstream performance of a final classifier trained on the synthetic data (Figure 4). Since we are interested in selecting the best performing parameters, we want a metric that is most likely to select a high quality model and thus should care primarily about the rank correlation.

While the distribution of text lengths isn’t the most correlated to downstream performance, it still has strong predictive power, and has the advantage of being easy to visualize and understand. We plot the distribution of lengths across all final datasets (across architecture and epsilon) for the IMDB and YELP datasets (figures 5 and 6). We would point out that all synthetic data is truncated to 512 tokens because this was sequence length of the trained LLM. At the same time, some of the real data is longer (17% of imdb and 7% of yelp) .

L.2 Synthetic data bigram frequency comparison

Comparing the distribution if bigrams across datasets can give additional insight into how the synthetic data differs between finetuning approach and epsilon. As one might expect, less noise leads to a smaller distributional gap (figure 7). Full finetuning has the largest gap, followed by prompt tuning, and LoRA being the most aligned with the original distribution.

Appendix M Implementation details and estimates of the required compute

LLM training. We run all our experiments using T5X codebase and used implementation of transformer layers from Flaxformer library. For differentially private training of transformers we used private_text_transformers repository.

We pre-trained 1B model on TPUv3 with 128 cores and 8B model was pretrained on TPUv3 with 1024 cores, both runs took around 8 days. Both finetuning and prompt-tuning of 8B models was done on TPUv3 with 128 cores. Diferentially private finetuning of 8B model required between 20 hours (for shortest run on IMDB) and up to 80 hours for some of the longer runs. Prompt tuning required 1.5 hour for short run and up to 20 hours for longest runs. LoRA-tuning was about 2x slower compared to prompt tuning with 3 hours training for shortest run and up to 2 days for longest runs.

In most experiments we used a total batch size of 1024 examples per optimizer step, however we did run a few ablations with larger batch as well. In order to fit entire batch into memory we used a technique called gradient accumulation. This technique splits entire batch into smaller chunks, sequentially computes gradients over each chunk and then aggregates them together to obtain final gradient for entire batch. We used chunks of 32 examples for full finetuning and LoRA, and chunks of size up to 256 for prompt tuning. Even with difference in chunk sizes, we observed that LoRA is only 2x slower compared to prompt tuning when using the same total effective batch and same number of training steps.

Downstream model. Downstream classifier was implemented in Tensorflow using Keras library for training and TFDS to load datasets.

Each downstream model was run on TPUv2 with 8 cores. To obtain each downstream accuracy number we run a sweep of around 28 different hyperparameter settings. Entire sweep took around 4 hours for CNN model and up to 80 hours combined for BERT model and synthetic dataset of 0.5M examples. Each sweep was repeated 3 times to compute error bars.

Appendix N Examples of generated and real data

Table LABEL:tab:sample_imdb_text shows examples of generated synthetic data.