Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers

Ruisi Zhang, Seira Hidano, Farinaz Koushanfar

Introduction

Natural language processing with its application in various fields have attracted much attention in recent years. With the recent advance in transformer-based language models (LMs), BERT Devlin et al. (2018) and its variants Liu et al. (2019); Xie et al. (2021); Qin et al. (2022) are used to classify text datasets and achieve state-of-the-art performance. However, LMs tend to memorize data during training, which results in unintentional information leakage Carlini et al. (2021).

Model inversion attacks Fredrikson et al. (2015), which invert training samples from the private dataset, has long been applied in the vision domain Yang et al. (2019); Zhang et al. (2018); Wang et al. (2021); Kahla et al. (2022). For text-based datasets, the model inversion attacks have been applied in the medical domain to infer patients’ privacy information. Recent work like KART Nakamura et al. (2020) and Lehman et al. Lehman et al. (2021) consider inverting tabular data with sensitive attributes from the medical datasets. However, they follow a fill-in-blank scheme and fail to reconstruct sentences with fluency from scratch.

In this paper, we focus on reconstructing private training data from fine-tuned LMs at inference time. It has a more general scenario where unauthorized personal data such as chats, comments, reviews, and search history may be used to train LMs. We perform a systemic study on how much private information is leaked via model inversion attacks. There are several challenges when performing model inversion attacks on NLP models. First, the candidate pixel range is 256 in images, but the candidate token range is more than 30,000. Therefore, it is harder to find the exact token for the sentence. Secondly, image inversion is more error-tolerant, i.e., error in some pixels will not affect the overall results. However, errors in some of the tokens will significantly affect the fluency of the texts. Thirdly, current text reconstruction attacks have different settings from model inversion attacks. Most of them fall into two categories: Gradient attack Deng et al. (2021); Zhu et al. (2019), which utilize gradient during distributed training to do attack. Embedding level reconstruct attack Xie and Hong (2021); Pan et al. (2020), which trains a mapping function to reconstruct texts from pre-trained embeddings.

We propose Text Revealer to perform model inversion attack on text data. In the attack, the adversary knows the domain of the private dataset and has access to the target models. Our attack consists of two stages: in the first stage, we collect texts from the same domain as the public dataset and extract high-frequency phrases from the public dataset as templates. Then, we train a GPT-2 as the text generator on the public dataset. In the second stage, we borrow the idea from PPLM Dathathri et al. (2019) to continuously perturb the hidden state in the GPT-2 based on the feedback from the target model. By minimizing the cross-entropy loss, generated text distribution becomes closer to the private datasets. Experiments on the Emotion and Yelp datasets with two target models, BERT and TinyBERT, demonstrate Text Revealer can reconstruct private information with readable contents. In summary, our approach has the following contributions:

∙\bullet We propose Text Revealer, the first model inversion attack for text reconstruction against text classification with transformers, to reconstruct private training data from the target models.

∙\bullet Results on two transformer-based models and two datasets with different lengths have demonstrated Text Revealer can reconstruct private texts with accuracy.

Approach

Adversary’s Capability and Knowledge

We consider the adversary knows the domain of the dataset on which the language model (LM) is fine-tuned. The adversary also has white-box access to the LM. During the attack, given the input sentences or input embeddings, the adversary can get the prediction score over NN classes with the probabilities P=(p1,p2,...,pn)P=(p_{1},p_{2},...,p_{n}). Throughout this paper, let pa(X)p_{a}(X) denote the prediction with a given input XX for a label aa.

2 Attack Construction

As shown in Figure 1, the general attack construction consists of two stages: (1) public dataset collection and analysis, and (2) word embedding perturbation.

Word embedding perturbation

We borrow the text perturbation idea from Plug and Play Language Model(PPLM) Dathathri et al. (2019) but change the optimization objective to perform the model inversion attack. PPLM is a lightweight text generation algorithm that uses an attribute classifier to help the GPT-2 perturb its hidden state and guide the generation.

Experiments

Emotion Dataset Saravia et al. (2018) is a sentence-level emotion classification dataset with six labels: sadness, joy, love, anger, fear, and surprise. Yelp Dataset Zhang et al. (2015) is a document-level review dataset. The reviews are labeled from 1 to 5 stars indicating the user’s preference. Following the split methods in image model-inversion attacks Zhang et al. (2020), we randomly sample 80% of the samples as public dataset and 20% of the samples as the private dataset. The average token length is 20 in the private Emotion dataset and 134 in the private Yelp dataset.

2 Evaluation Metrics

Recovery Rate (RR.) is the percentage of tokens in private dataset recovered by different attack methods. We filtered punctuations, special tokens and NLTK’s stop words Bird et al. (2009) in the private dataset. Attack Accuracy (Acc.) is the classification accuracy using evaluation classifier on inverted texts. According to Zhang et al. (2020), the higher the classification accuracy is, the more private information the texts is considered to be inverted. We use BERT Large as the evaluation classifier and fine-tune it until the accuracy on the private dataset is over 95%. Fluency is the Pseudo Log-Likelihood (PLL) of fixed-length models on inverted textshttps://huggingface.co/docs/transformers/perplexity.

3 Baselines

We compare our method with vanilla model inversion attack Fredrikson et al. (2015) and vanilla text generation Radford et al. (2019).

In this setting, the adversary exploits the classification loss by adjusting text embeddings and returning the texts minimizing the cross-entropy loss. We set the text length to the average length of the private dataset, adjust the text embeddings for 50 epochs and run the same times as the number of templates.

Vanilla text generation (VTG)

In this setting, the adversary uses a GPT-2 model to generate the texts. The GPT-2 is trained on the public dataset and conditioned on our collected templates. We limit the maximum length to the average length of the private dataset. For the attack accuracy on VTG, we calculate the frequency of collected templates under different labels in the public dataset and use the label with the highest frequency as the target label. During the attack, only VTG requires extra annotations to benchmark its performance. For target model TinyBERT and BERT, we run VTG twice and use the results for two target models.

4 Target Models

We use two representative transformer-based LM as our target model: (1) Tiny-BERT Bhargava et al. (2021) and (2) BERT Devlin et al. (2018). The TinyBERT has four layers, 312 hidden units, a feed-forward filter size of 1200, and 6 attention heads. It has 110M parameters. The BERT has 12 layers, 768 hidden units, a feed-forward filter size of 3072, and 12 attention heads. It has 4M parameters.

5 Results

The result of our model inversion attack is summarized in Table 1. We can make the following observations: (1) VTG and Text Revealer tend to invert more tokens and achieve lower PLL compared with VMI. This is because VMI trains the text embeddings from scratch with random initialization, which results in not meaningful combinations of tokens. Even though many tokens can be recovered using VMI, private information still cannot be inferred from the private dataset. (2) Compared with VTG, our algorithm achieves higher recovery rates and higher attack accuracy. By perturbing the hidden state of trained GPT-2, our algorithm can infer more private information from the target model. (3) For smaller Emotion dataset, all three methods achieve high attack accuracy and can invert private information from the private dataset. However, in the larger Yelp dataset, VMI and VTG’s attack accuracy becomes near 20%. It is close to random classification because only 5 classes are in the dataset. (4) For both VMI and Text Revealer, BERT’s recovery rate and attack accuracy are higher than TinyBERT. It means more private information is memorized as the transformer becomes larger.

6 Ablation Studies

In this setting, we fine-tune the GPT-2 on the public dataset and compare the results with vanilla GPT-2 on the Yelp dataset. The results are summarized in Table 2. The table shows that fine-tuned GPT-2 achieves higher recovery rate and attack accuracy. It means more sensitive information is revealed from fine-tuning.

Effectiveness of model inversion attack

We analysis the effectiveness of word perturbation and loss function by comparing with different methods. For word perturbation, we compare Text Revealer’s word perturbation with Gumbel softmax Jang et al. (2016). For Gumbel softmax, we first use fine-tuned GPT-2 to generate original sentences, and set coefficient for tokens in vocabulary. Then, we update the coefficient based on the cross-entropy loss using Gumbel softmax and update the input sentences. From the table, we can find Gumbel softmax and modified perturbation achieve similar attack accuracy. However, the recovery rate and fluency is lower than modified perturbation.

For loss function, we compare cross entropy with Modified entropy loss Song and Mittal (2021). The modified entropy loss makes the loss monotonically decreasing with the prediction probability of the correct label and increasing with the prediction probability of any incorrect label. In this setting, we use the same pipeline as Text Revealer, but change the loss to modified entropy loss to update GPT-2’s hidden state. We can find cross entropy loss achieves best performance out of other loss functions.

7 Analysis

In this setting, we analyze how the inversion accuracy changes with the length of the template and summarize the results in Table 7. The template length is LL, and we extract the first 0.3 LL, 0.5 LL, and 0.7 LL tokens from the template. If the extracted token number is smaller than 1, we choose the first token from the template. For example, if the original template is "if i could give this place 0 star", then 0.7 LL template would be "if i could give this," 0.5 LL template would be "if i could give," and 0.3 LL template would be "if i." As text length becomes shorter, Text Revealer achieves worse recovery rates and attack accuracy. It means longer template length can give GPT-2 more contexts and help it invert more private information.

Inverted examples & Visualization

Following TAG Deng et al. (2021), We display how our inverted examples is approximate the distribution of private dataset at embedding level and sentence level. We use sentences with longest matching subsequence in the private dataset as ground truth. We first use Principal Component Analysis (PCA) Abdi and Williams (2010) to reduce the dimension for ground truth and inverted examples and display the plot in Figure 2. Then, we display the inverted examples at sentence level in Table 5

Public dataset & Private dataset

We use Kendall tau distance Fagin et al. (2003) to measure the token distribution similarity between the private and public datasets. The Kendall tau distance is a metric to measure the top-kk elements correlation between two lists. The larger the distance is, the stronger correlation the two lists have. When ranking the tokens in each dataset, we filtered the special tokens, punctuations, and stop words in NLTK Bird et al. (2009). From the results in Table 6, we can find the word ranking and frequency is close between the private and public datasets.

We also compare the accuracy of the Emotion and the Yelp Dataset in Table 6. Before fine-tuning, the classification accuracy is similar for both private and public datasets. However, after fine-tuning, the accuracy between the private dataset and the public dataset has a large gap, which means some texts are memorized Carlini et al. (2021) in the trained transformers.

Conclusion

In this paper, we present Text Revealer to invert texts from the private dataset. Experiments on different target models and datasets have demonstrated our algorithm can faithfully reconstruct private texts with accuracy. In the future, we are interested in exploring potential defense methods.

References

Appendix A Appendix

Our work has the following limitations. (1) From the attack side, the model inversion attack achieves lower attack accuracy as the dataset becomes larger. It would be interesting to explore better template extract methods to improve the accuracy. (2) From the ethical standpoint, evaluating how sensitive BERT based model is to classification tasks is essential to future defense methods. Due to page limits, we did not go into the defense methods. We will continue to work on it in future work.

A.2 Ethics Statement

Our work focuses on the privacy problems in natural language processing. Even though model inversion attacks proposed in our paper can cause potential data leakage. Evaluating how sensitive BERT based model is to classification tasks is important to future defense methods.

A.3 Hyperparameters and training details

We perform nn-gram analysis on the public dataset. For Emotion dataset, we extract phrase length from 1 word to 3 words with frequency higher than 20. Then, we filter one-word adjectives and permute them with 2-3 word to form template. The total number is 220. For Yelp dataset, we extract phrase length from 3 words to 8 words with frequency higher than 20. The total number is 320. Some of the templates are summarized in Table 7

BERT fine-tuning

The BERT and TinyBERT are trained as target models on the private dataset. For both BERT and TinyBERT, we set the batch size to 8. We use AdamW as the optimizer with initial learning rate set to 5e-5. We train 5 epochs on BERT and 10 epochs on TinyBERT. Other parameters follow the original implementation in Devlin et al. (2018); Jiao et al. (2019).

GPT-2 fine-tuning

The GPT-2 is trained on the public dataset to help Text Revealer reconstruct texts. For GPT-2, we set the batch size to 8, and use AdamW as the optimizer with initial learning rate set to 5e-5. We train 10 epochs for both datasets. Other parameters follow the original implementation in Radford et al. (2019). It has 124M parameters.

Text Revealer

When using the Text Revealer, we set the iteration number to the average length of the private dataset. We set the window mask to 3 and KL loss coefficient to 0. Other parameters follow the original implementation in Dathathri et al. (2019).

Computing Infrastructure

Our code is implemented with PyTorch. Our attack pipeline are all constructed using TITAN Xp. We fine-tune the target LMs and GPT-2 on NVIDIA RTX A6000.

A.4 More inverted examples

We display more inverted examples in Table 8 and Table 9. We show four examples for each method in Emotion dataset and two examples for each method in Yelp dataset.