Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness

Lingjuan Lyu, Xuanli He, Yitong Li

Introduction

Many language applications have involved deep learning techniques to learn text representation through neural models Bengio et al. (2003); Mikolov et al. (2013); Devlin et al. (2019), performing composition over the learned representation for downstream tasks Collobert et al. (2011); Socher et al. (2013). However, the input text often provides sufficient clues to portray the author, such as gender, age, and other important attributes. For example, sentiment analysis tasks often have privacy implications for authors whose text is used to train models. Many user attributes have been shown to be easily detectable from online review data, as used extensively in sentiment analysis results Hovy et al. (2015); Potthast et al. (2017). Private information can take the form of key phrases explicitly contained in the text. However, it can also be implicit. For example, demographic information about the author of a text can be predicted with above chance accuracy from linguistic cues in the text itself Preoţiuc-Pietro et al. (2015).

On the other hand, even the learned representation, rather than the text itself, may still contain sensitive information and incur significant privacy leakage. One might argue that sensitive information like gender, age, location and password should not be leaked out and should have been removed from representation. However, on the intermediate representation level, which is trained from the input text to contain useful features for the prediction task, it can meanwhile encode personal information which might be exploited for adversarial usages, especially a modern deep learning model has vastly more capacity than they need to perform well on their tasks. And, it has been justified that an attacker can recover private variables with higher-than-chance accuracy, only using hidden representation Li et al. (2018); Coavoux et al. (2018). Therefore, the fact that representations appear to be abstract real-numbered vectors should not be misconstrued as being safe.

The naive solution of removing protected attributes is insufficient: other features may be highly correlated with, and thus predictive of, the protected attributes Pedreshi et al. (2008). To tackle with these privacy issues, Li et al. (2018) proposed to train deep models with adversarial learning, which explicitly obscures individuals’ private information, while improves the robustness and privacy of neural representation in part-of-speech tagging and sentiment analysis tasks. In a parallel study, Coavoux et al. (2018) proposed defence methods based on modifications of the training objective of the main model. However, both works provide only empirical improvements in privacy, without any formal guarantees. Prior works have approached formal differential privacy guarantee by training differentially private deep models Abadi et al. (2016); McMahan et al. (2018); Yu et al. (2019). However, these works generally only considered the training data privacy rather than the test data privacy. While cryptographic methods can be used for privacy protection, it could be resource-hungry or overly complex for the user.

To alleviate the above limitations, we take inspirations from differential privacy Dwork and Roth (2014) to provide formal privacy guarantee of the extracted representation from user-authored text. Meanwhile, we propose a robust training algorithm to derive a robust target model to maintain utility, which also offers fairness as a by-product. To the best of our knowledge, our work is the only work to date that can provide formal differential privacy guarantee of the extracted representation, while ensuring fairness.

For the first time, the privacy of the extracted neural representation from text is formally quantified in the context of differential privacy. A novel approach called Differentially Private Neural Representation (DPNR) is proposed to perturb the extracted representation.Also, we prove that masking words via dropout can further enhance privacy.

To maintain utility, we propose a robust training algorithm that incorporates the noisy training representation in the training process to derive a robust target model, which also reduces model discrimination in most cases.

On benchmark datasets across various domains and multiple tasks, we empirically demonstrate that our approach yields comparable accuracy to the non-private baseline on the main task, while significantly outperforms the non-private baseline and adversarial learning on the privacy taskcode and preprocessed datasets are available at: https://github.com/xlhex/dpnlp.git.

Preliminary: Differential Privacy

Differential privacy Dwork and Roth (2014) provides a mathematically rigorous definition of privacy and has become a de facto standard for privacy analysis. Within DP framework, there are two general settings: central DP (CDP) and local DP (LDP).

In CDP, a trusted data curator answers queries or releases differentially private models by using randomisation mechanisms Dwork and Roth (2014); Abadi et al. (2016); Yu et al. (2019). For scenarios where data are sourced from end users, and end users do not trust any third parties, DP should be enforced in a “local” manner to enable end users to perturb their data before publication, which is termed as LDP Dwork and Roth (2014); Duchi et al. (2013). Compared with CDP, LDP offers a stronger level of protection.

In our system, we aim to protect the test-phase privacy of the extracted neural representations from end users, we therefore adopt LDP. LDP has shown the advantage that the data is randomised before individuals disclose their personal information, so the server and the middle eavesdropper can never see or receive the raw data. In terms of LDP mechanisms, randomised response Warner (1965); Duchi et al. (2013) and its variants have been widely used for aggregating statistics, such as frequency estimation, heavy hitter estimation, etc Erlingsson et al. (2014).

Let A\mathchar58D→O\mathcal{A}\mathrel{\mathop{\mathchar 58\relax}}\mathcal{D}\to\mathcal{O} be a randomised algorithm mapping a data entry in D\mathcal{D} to O\mathcal{O}. The algorithm A\mathcal{A} is (ϵ,δ)(\epsilon,\delta)-local differentially private if for all data entries \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}},\mathchoice{\mbox{\boldmath\displaystyle x^{\prime}}}{\mbox{\boldmath\textstyle x^{\prime}}}{\mbox{\boldmath\scriptstyle x^{\prime}}}{\mbox{\boldmath\scriptscriptstyle x^{\prime}}}\in\mathcal{D} and all outputs o∈Oo\in\mathcal{O}, we have

If δ=0\delta=0, A\mathcal{A} is said to be ϵ\epsilon-local differentially private.

A formal definition of LDP is provided in Definition 2.1, The privacy parameter ϵ\epsilon captures the privacy loss consumed by the output of the algorithm: ϵ=0\epsilon=0 ensures perfect privacy in which the output is independent of its input, while ϵ→∞\epsilon\rightarrow\infty gives no privacy guarantee.

For every pair of adjacent inputs x\textstyle x and x′\textstyle x^{\prime}, differential privacy requires that the distribution of \mathcal{A}(\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}) and \mathcal{A}(\mathchoice{\mbox{\boldmath\displaystyle x^{\prime}}}{\mbox{\boldmath\textstyle x^{\prime}}}{\mbox{\boldmath\scriptstyle x^{\prime}}}{\mbox{\boldmath\scriptscriptstyle x^{\prime}}}) are “close” to each other where closeness are measured by the privacy parameters ϵ\epsilon and δ\delta. Typically, the inputs x\textstyle x and x′\textstyle x^{\prime} are adjacent inputs when all the attributes of one record are modified. In real scenario, the adjacent input is an application specific notion. For example, a sentence is divided into several items for every 5 words, and two sentences are considered to be adjacent if they differ by at most 5 consecutive words Wang et al. (2018). In this work, we consider a word-level DP, i.e., two inputs are considered to be adjacent if they differ by at most 1 word. For brevity, we use (ϵ,δ)(\epsilon,\delta)-DP to represent (ϵ,δ)(\epsilon,\delta)-LDP for the rest of the paper. We remark that all the randomisation mechanisms used for CDP, including Laplace mechanism and Gaussian mechanism Dwork and Roth (2014), can be individually used by each party to inject noise into local data to ensure LDP before releasing Lyu et al. (2020a); Yang et al. (2020); Lyu et al. (2020b); Sun and Lyu (2020). In particular, we adopt Laplace Mechanism which ensures ϵ\epsilon-DP with δ=0\delta=0 throughout the paper.

In a nutshell, data universe can be expressed as D=(X,A,Y)\mathcal{D}=(\mathcal{X},\mathcal{A},\mathcal{Y}), which will be convenient to partition as (X,Y)×A(\mathcal{X},\mathcal{Y})\times\mathcal{A} Jagielski et al. (2018). Given one person’s record x\textstyle x, we can write it as a pair \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}=(\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{I},\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S}) where \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{I}\in(\mathcal{X},\mathcal{Y}) represents the insensitive attributes and \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S}\in\mathcal{A} represents the sensitive attributes. Our main goal is to promise differential privacy only with respect to the sensitive attributes. Write \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S}\sim\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}^{\prime}_{S} to denote that \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S} and \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}^{\prime}_{S} differ in exactly one coordinate (i.e. one word/token in NLP domain). An algorithm is (ϵ,δ)(\epsilon,\delta)-differentially private in the sensitive attributes if for all \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{I}\in(\mathcal{X},\mathcal{Y}) and for all \mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S}\sim\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}_{S}^{\prime}\in\mathcal{A} and for all O⊆OO\subseteq\mathcal{O}, we have:

Post-processing. DP enjoys a well-known post-processing property Dwork and Roth (2014): any computation applied to the output of an (ϵ,δ)(\epsilon,\delta)-DP algorithm remains (ϵ,δ)(\epsilon,\delta)-DP. This nice property allows the attacker to implement any sophisticated post-processing function on the privatised representation from the user, without compromising DP or making it less differentially private.

Main Framework

As indicated in §1, uploading raw input or representations to a server takes the risk of revealing sensitive information to the eavesdropper who eavesdrops on the hidden representation and tries to recover private information of the input text. Hence, similar as Coavoux et al. (2018), we consider an attack scenario during inference phase in Figure 1, which consists of three parts: (i) a feature extractor to extract latent representation of any test input x\textstyle x; (ii) a main classifier to predict the label yy from the extracted latent representation; (iii) and an attacker (eavesdropper) who aims to infer some private information z\mathbf{z} contained in x\textstyle x, from the latent representation of x\textstyle x used by the main classifier. In this scenario, each example consists of a triple (\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}},y,\mathbf{z}), where x\textstyle x is an input text, yy is a single label (e.g. topic or sentiment), and z\mathbf{z} is a vector of private information contained in x\textstyle x. Such attack would occur in scenarios where the computation of a neural network is shared across multiple devices. For example, phone users send their learned representations to the cloud for grammar correction or translation Li et al. (2018), or to obtain the classification result, e.g., the topic of the text or its sentiment Li et al. (2017).

2 Methodology

To defend against the middle eavesdropper, we aim to design an approach that can preserve privacy of the extracted test representation from the user without significantly degrading the main task performance. To achieve this goal, we introduce a DP noise layer after a predefined feature extractor (determined by the server), which results in differentially private representation that can be transferred to the server for classification (the topic of the text or its sentiment), as shown in Figure 2.

In terms of model training on the server, theoretically, one could remove the noise layer and conduct non-private training by following Equation 1:

However, doing so may deteriorate test performance, due to the injected noise in the test representation. To improve model robustness to the noisy representation, we put forward a robust training algorithm by incorporating a noise layer which adds the same level of noise as the test phase in the training process as well. Therefore, the robust training objective can be re-written as:

The detailed robust training process on the server is given in Algorithm 1. After the robust target model is built, server then provides a feature extractor ff to the user, as illustrated in Figure 2.

3 Privacy Guarantee

where the coordinates \mathchoice{\mbox{\boldmath\displaystyle r}}{\mbox{\boldmath\textstyle r}}{\mbox{\boldmath\scriptstyle r}}{\mbox{\boldmath\scriptscriptstyle r}}=\{r_{1},r_{2},\cdots,r_{k}\} are i.i.d. random variables drawn from the Laplace distribution defined by Lap(b)Lap(b), where the noise scale b=Δfϵb=\frac{\Delta f}{\epsilon}, ϵ\epsilon is the privacy budget and Δf\Delta f is the sensitivity of the extracted representation.

Note that to apply additive noise mechanism, the sensitivity Δ\Delta of the output representation \mathchoice{\mbox{\boldmath\displaystyle x_{r}}}{\mbox{\boldmath\textstyle x_{r}}}{\mbox{\boldmath\scriptstyle x_{r}}}{\mbox{\boldmath\scriptscriptstyle x_{r}}}=f(\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}}) needs to be determined. Estimating the true sensitivity of xr\textstyle x_{r} is challenging. Instead, we follow Shokri and Shmatikov (2015) to use input-independent bounds by enforcing a range on the extracted representation, hence bounding the sensitivity of each element of the extracted representation with 1, i.e., Δf=1\Delta f=1. Limiting the range of the extracted representation can also improve the training process by helping to avoid overfitting.

A formal statement for the privacy guarantees of Algorithm 2 is provided in Theorem 1.

Let the entries of the noise vector r\textstyle r be drawn from Lap(b)Lap(b) with b=Δfϵb=\frac{\Delta f}{\epsilon}. Then Algorithm 2 is ϵ\epsilon-differentially private.

3.2 Word Dropout Enhances Privacy

In NLP, each input is a sequence composed of words/tokens {w1,⋯ ,wd}\{w_{1},\cdots,w_{d}\}. Under word-level DP, two sentences are considered to be adjacent inputs if they differ by at most 1 word (i.e., 1 edit distance). In this scenario, to lower privacy budget without significantly degrading the inference performance, we borrow the idea of nullification Wang et al. (2018) and apply it to word dropout.

As stated in Theorem 2, word dropout in combination with any ϵ\epsilon-differentially private mechanism provides a tighter privacy bound in the context of word-level DP. A detailed proof follows.

Combine these two cases, and use the fact that Pr⁡[Ini=0]=μ\Pr[I_{ni}=0]=\mu, we have:

Therefore, after dropout, the privacy budget is lowered to ϵ′=ln[(1−μ)exp⁡(ϵ)+μ]\epsilon^{\prime}=ln[(1-\mu)\exp(\epsilon)+\mu].∎

Since the perturbed representation \mathcal{A}(\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}})=f(\mathchoice{\mbox{\boldmath\displaystyle x}}{\mbox{\boldmath\textstyle x}}{\mbox{\boldmath\scriptstyle x}}{\mbox{\boldmath\scriptscriptstyle x}})+\mathchoice{\mbox{\boldmath\displaystyle r}}{\mbox{\boldmath\textstyle r}}{\mbox{\boldmath\scriptstyle r}}{\mbox{\boldmath\scriptscriptstyle r}} is ϵ\epsilon-differentially private, combining dropout beforehand, the privacy budget is lowered to ϵ′=ln[(1−μ)exp⁡(ϵ)+μ]\epsilon^{\prime}=ln[(1-\mu)\exp(\epsilon)+\mu], hence improving privacy guarantee. Apparently, a high value of μ\mu has a positive impact on the privacy but a potential negative impact on the utility. In particular, when μ=1\mu=1, all dd words will be masked, which gives the highest privacy, i.e., ϵ′=0\epsilon^{\prime}=0, but totally destroys inference performance. Hence, a smaller value of μ\mu is preferred to trade off privacy and accuracy.

Experiments

In this section, we conduct comprehensive studies over different tasks and datasets to examine the efficacy of the proposed algorithm from three facets: 1) main task performance, 2) privacy and 3) target model fairness.

We use two natural language processing tasks: 1) sentiment analysis and 2) topic classification, with a range of benchmark datasets across various domains. Table 1 summarises the statistics of the used datasets.

Trustpilot Sentiment dataset Hovy et al. (2015) contains reviews associated with a sentiment score on a five point scale, and each review is associated with 3 attributes: gender, age and location, which are self-reported by users. The original dataset is comprised of reviews from different locations, however in this paper, we only derive tp-us for our study. Following Coavoux et al. (2018), we extract examples containing information of both gender and age, and treat them as the private information. We categorise “age” into two groups: “under 34” (u34) and “over 45” (o45).

1.2 Topic Classification

For topic classification, we focus on two genres of documents: news articles and blog posts.

We use ag news corpus Del Corso et al. (2005). To ensure a fair comparison, we use the corpus preprocessed by Coavoux et al. (2018)https://github.com/mcoavoux/pnet/tree/master/datasets. We use both “title” and “description” fields as the input document.. And the task is to predict the topic label of the document, with four different topics in total.

Regarding the private information in ag, named entities appearing in text are vulnerable to privacy leakage inferred by attackers. In order to simulate the attack, we firstly adopt the NLTK NER system Bird et al. (2009) to recognise all “Person” entities in the corpus. Then we retain the five most frequent person entities and use them as the private information. Due to the sparsity of name entities, each target entity only appears in very few articles. Hence we select the examples containing at least one of these named entities to mitigate the unbalance and data scarcity. Thus, the attacker aims to identify these five entities as five independent binary classification tasks.

We derive a blog posts dataset (blog) from the blog authorship corpus presented Schler et al. (2006). However, the original dataset only contains a collection of blog posts associated with authors’ age and gender attributes but does not provide topic annotations. Thus we follow Coavoux et al. (2018) to run the LDA algorithm Blei et al. (2003) with the topic number of 10 on the whole collection to identify the topic label of each document. Afterwards, we selected posts with single dominating topic (>80%>80\%) and discarded the rest, which results in a dataset with 10 different topics. Similar to TP-US, the private variables are comprised of the age and gender of the author. And the age attribute is binned into two categories, “under 20” (U20) and “over 30” (O30).

For all three datasets, we randomly split the preprocessed corpus into training, development and test by 8:1:1.

2 Evaluation Metrics

Similar to Coavoux et al. (2018), we define sentiment analysis and topic classification as the main tasks, whereas the inference of private information is considered as the auxiliary tasks of attackers. Each auxiliary task is eavesdropped by one attacker.

We use accuracy to assess the performance for both main tasks. The auxiliary tasks are evaluated via the following metrics:

For demographic variables (i.e., gender and age): 1−X1-X, where XX is the average over the accuracies of the prediction by the attacker on these variables.

For named entities: 1−F1-F, where FF is the F1 score between the ground truths and the prediction by the attacker on the presence of all named entities.

We denote the value of 1-XX or 1-FF as empirical privacy, i.e., the inverse accuracy or F1 score of the attacker, higher means better empirical privacy, i.e., lower attack performance.

3 Model Selection

Model and Parameters. For implementation, owing to its success across multiple NLP tasks, we apply BERT base Devlin et al. (2019) to the classification tasks. Specially, BERT takes a text input, then generates a representation which embeds holistic information. We apply a dropout to this representation before a softmax layer, which is responsible for label classification.We run 4 epochs on the training set, and choose the checkpoint with the best loss on the dev set.

After we obtain a well-trained target model, we partition it into two parts, BERT model acts as the feature extractor ff in Figure 2, which could be deployed on users’ devices, while the remaining layers act as the classifier on the server. In our implementation, privacy is enforced in the hidden representation extracted by the feature extractor as shown by Algorithm 2. For attack classifier, we utilise a 2-layer MLP with 512 hidden units and ReLU activation trained over the target model, which delivers the best attack performance on the dev set in our preliminary experiments.

We report the averaged results over 5 independent runs for all experiments.

4 Performance Analysis of Target Model

Firstly, we would like to study how the privacy parameters (ϵ,μ\epsilon,\mu) in Theorem 1 and 2 affect the accuracy of main tasks. We investigate this using different parameter settings, varying one parameter while fixing the other.

To analyse the impact of different privacy budget ϵ\epsilon on accuracy, we choose ϵ∈{0.05,0.1,0.5,1,5}\epsilon\in\{0.05,0.1,0.5,1,5\} with fixed μ=0\mu=0. Noted that to provide reasonable privacy guarantee, ϵ\epsilon should be set below 10 Hamm et al. (2015); Abadi et al. (2016). Moreover, ϵ≤1\epsilon\leq 1 means a relatively tight privacy guarantee. Surprisingly, there is no obvious relationship between accuracy and ϵ\epsilon. We speculate the denoising training procedure of BERT and layernorm Ba et al. (2016) make BERT resistant to the injected noises, which can maintain the performance of the main tasks. We will conduct an in-depth study on this in the future.

Table 2 shows that in most cases, our method can achieve comparable performance to the non-private baseline, across all ϵ\epsilon even when the noise level is high (ϵ=0.05\epsilon=0.05), which validates the robustness of our method to DP noise. It also implies that the DP-noised representation not only preserves privacy, but also retains general information for the main task.

4.2 Impact of Dropout Rate μ𝜇\mu

Similarly, we study how the word dropout rate μ\mu affects accuracy-privacy trade-off. Table 3 reports the performance of different models under different μ∈{0.1,0.3,0.5,0.8}\mu\in\{0.1,0.3,0.5,0.8\} with fixed ϵ=1\epsilon=1. In most cases, as μ\mu becomes larger, accuracy starts to degrade as expected. However, as indicated in Theorem 2, higher μ\mu results in better privacy as well. Moreover, μ=0.5\mu=0.5 can still provide a relatively high accuracy, while privacy budgets are reduced to ϵ′=ln[(1−μ)exp⁡(ϵ)+μ]=0.62\epsilon^{\prime}=ln[(1-\mu)\exp(\epsilon)+\mu]=0.62.

Overall, both results demonstrate that our dpnr can protect privacy of the extracted representations of user-authored text, without significantly affecting the main task performance.

5 Attack Model

Apart from formal privacy guarantee from DP, we use the performance of the diagnostic classifier of the attackers for empirical privacy. To fairly compare with the standard training and adversarial training in previous work Coavoux et al. (2018), we train an attack model that is trying to predict private variables from the representation. We measure the empirical privacy of a hidden representation by the ability of an attacker to predict accurately specific private information from it. If its empirical privacy (c.f., Section 4.2) is low, then an eavesdropper can easily recover information about the input. In contrast, a higher empirical privacy (close to that of a most-frequent label baseline) suggests that xr\textstyle x_{r} mainly contains useful information for the main task, while other private information is erased.

To study the relationship between DP and empirical privacy, we numerically investigate the impact of the different differential privacy budgets on empirical privacy. Recall that the empirical privacy is measured by 1-X/FX/F, and the higher is better. Figure 3 shows that with the increase of the budget, empirical privacy across all datasets demonstrate a decreasing trend, especially for ag, which well aligns with DP where the higher value of ϵ\epsilon implies lower formal privacy guarantee. Since ϵ=0.05\epsilon=0.05 provides the best privacy guarantee, we fix ϵ=0.05\epsilon=0.05 and μ=0\mu=0 as a default setting in the rest of this section, unless otherwise mentioned.

For empirical privacy, we investigate whether our dpnr can provide better attack resistance compared with the adversarial learning (adv) Coavoux et al. (2018) and non-private training method (non-priv), which indicates a lower bound. We also report the majority class prediction (majority) as an upper bound.

Table 4 shows that the attack model can indeed recover private information with reasonable accuracy when targeting towards the non-private representations, manifesting that representations inadvertently capture sensitive information about users, apart from the useful information for the main task. By contrast, our dpnr significantly reduces the amount of information encoded in the extracted representation, as validated by the substantially higher empirical privacy than non-priv across all datasets. We also observe that our dpnr achieves comparable empirical privacy to the majority class (majority), and consistently outperforms the adversarial learning (adv) from Coavoux et al. (2018), which confirms the argument of Elazar and Goldberg (2018) that adversarial learning can not fully remove sensitive demographic traits from the data representations. Conversely, the post-processing property of DP ensures that the privacy loss of the extracted representation cannot be increased even by the most sophisticated attacker.

This claim can be further confirmed by Table 5, which reports the accuracy of the attacker on classifying whether a named entities is absent or presented in the document over agFor space limitation, we only report 3 of 5 entities and the results of other two are similar.. Generally, both adv and dpnr can reduce attack accuracy, misleading the attacker classifier to predict most of the shared representations as majority (A). While our dpnr significantly outperforms both non-priv and adv, corroborating our analysis above.

6 Target Model Fairness

Recently, fairness concern has gained lots of attention in NLP community Bolukbasi et al. (2016); Zhao et al. (2017); Chang et al. (2019); Lu et al. (2018); Sun et al. (2019). Depending on the literature, fairness can have different interpretation. In this section, we further consider the relation between differential privacy and fairness. We ask the research question whether differential privacy noise can help enhance model fairness? We focus on a particular scenario of fairness, that is given a specific demographic variable (e.g. gender) a fair model should deliver an equal or similar performance over the subgroups (e.g. male vs. female) Rudinger et al. (2018); Zhao et al. (2018).

To empirically evaluate the fairness, we take inspirations of Rudinger et al. (2018); Zhao et al. (2018); Li et al. (2018) and partition the test data into sub-groups by the demographic variables, i.e., age, gender and five person entities. Different from predicting demographic variables in attacker (§4.5), we measure the main task accuracy difference among subgroups of demographic variables.

In fact, we noticed dpnr can also help mitigate the bias in the representations with respect to the specific demographic or identity attributes, such that the decisions made by our robust target model are able to improve the fairness among the concerned demographic groups.

First of all, as the distribution of the demographic groups in tp-us and blog datasets is relatively even, hence there is no significant deviation on the main tasks (see Table 6). However, we still observe an noticeable difference for the age group in blog and the gender attribute in tp-us. To help better understand the phenomenon, we perform further analysis by plotting the non-private and differentially private representations of age on blog in Figure 4. It can be clearly observed that the patterns of two subgroups are much easier to be distinguished in the non-private representations, while the differentially private representations mostly mix the representations of “under 20” and “over 30”. We speculate that this is a consequence of the regularising effect of DP.

Table 7 shows the fairness results on ag, where we observe the entity distributions are skewed and the prediction of the non-priv model on the dominant groups is significantly superior to the minority groups, which causes a severe violation in terms of the fairness. Even under such circumstance, our dpnr method can mitigate this skewed bias, achieving more fair prediction than other baselines.

Discussion

Privacy and fairness are two emerging but important areas in NLP community. Prior efforts predominantly focus on either privacy or fairness Li et al. (2018); Coavoux et al. (2018); Rudinger et al. (2018); Zhao et al. (2018); Lyu et al. (2020a), but there is no systematic study on how privacy and fairness are related. This work fills this gap, and discovers the impact of differential privacy on model fairness. We empirically show that privacy and fairness can be simultaneously achieved through differential privacy.

We hope that this work highlights the need for more research in the development of effective countermeasures to defend against privacy leakage via model representation and mitigate model bias in a general sense, and not only specific to a particular attack. More generally, we hope that our work spurs future interest into developing a better understanding of why differential privacy works.

Meanwhile, differential privacy may incur a reduction in the model’s accuracy. It is worthwhile to explore how to get a better trade-off between privacy, fairness and accuracy.

Conclusion and Future Work

In this paper, we take the first effort to build differential privacy into the extracted neural representation of text during inference phase. In particular, we prove that masking the words in a sentence via dropout can further enhance privacy. To maintain utility, we propose a novel robust training algorithm that incorporates a noisy layer into the training process to produce the noisy training representation. Experimental results on benchmark datasets across various tasks, and parameter settings demonstrate that our approach ensures representation privacy without significantly degrading accuracy. Meanwhile, our DP method helps reduce the effects of model discrimination in most cases, achieving better fairness than the non-private baseline. Our work makes a first step towards understanding the connection between privacy and fairness in NLP – which were previously thought of as distinct classes. Moving forward, we believe that our results justify a larger study on various NLP applications and models, which will be our immediate future work.

Acknowledgments

We would like to thank the anonymous reviewers for their valuable feedback.

References