Tempered Sigmoid Activations for Deep Learning with Differential Privacy
Nicolas Papernot, Abhradeep Thakurta, Shuang Song, Steve Chien, Úlfar Erlingsson
Introduction
Machine learning (ML) can be usefully applied to the analysis of sensitive data, e.g., in the domain of healthcare . However, ML models may unintentionally reveal sensitive aspects of their training data, e.g., due to overfitting . To counter this, ML techniques that offer strong guarantees expressed in the framework of differential privacy have been developed. A seminal example is the differentially private stochastic gradient descent, or DP-SGD, of Abadi et al. . The technique is a generally-applicable modification of stochastic gradient descent. In addition to its rigorous privacy guarantees, it has been empirically shown to stop known attacks against the privacy of training data; a representative example being the leaking of secrets .
Beyond privacy, training using DP-SGD offers advantages such as strong generalization and the promise of reusable holdouts . Yet, its advantages have not been without cost: empirically, the test accuracy of differentially private ML is consistently lower than that of non-private learning (e.g., see ). Such accuracy loss may sometimes be inevitable: for example, the task may involve heavy-tailed distributions and noise added by DP-SGD hinders visibility of examples in the tails . However, this does not explain the accuracy loss of differentially private ML on benchmarks that are known to be relatively simple when learning without privacy: e.g., MNIST , FashionMNIST , and CIFAR10 .
This paper is the first to observe that DP-SGD leads to exploding model activations as a deep neural network’s training progresses. This makes it difficult to control the training algorithm’s sensitivity at a minimal impact to its correctness. Indeed, exploding activations cause unclipped gradient magnitudes to also increase, which in turn induces an information loss once the clipping operation is applied to bound gradient magnitudes. This exacerbates the negative impact of noise calibrated to the clipping bound, thus degrading the utility of each gradient step when learning with privacy. Indeed, the gradient clipping of DP-SGD does not bring the nice properties of gradient clipping commonly used to regularize deep learning because DP-SGD clips gradients at the granularity of individual training examples rather than at the level of a batch.
We thus hypothesize that activation functions need to be bounded when learning with DP-SGD. We propose that neural architectures for private learning employ a general family of bounded activations: tempered sigmoids. We note that prior work has explored tempered losses as a means to provide robustness to noise during training . Because the family of tempered sigmoids can—in the limit—represent an approximation of ReLUs on the subset of their domain that is exercised in training, we expect that our approach will perform no worse than current architectures. These architectures use ReLUs as the de facto choice of activation function.
Through both analysis and experiments, we validate the significantly superior performance of tempered sigmoids when training neural networks with DP-SGD. In our analysis, we relate the role of the temperature parameter in tempered sigmoids to the clipping operation of DP-SGD. Unlike prior work, which attempted to adapt the clipping norm to the gradients of each layer’s parameters post hoc to training , we find that tempered sigmoids preserve more of the signal contained in gradients of each layer because they rescale each layer’s activations and better predispose the corresponding layer’s gradients to clipping. We conclude that using tempered sigmoids is a better default activation function choice for private ML.
In summary, our contributions facilitate DP-SGD learning as follows:
We analytically show in Section 3 how tempered sigmoid activations control the gradient norm explicitly, and in turn support faster convergence in the settings of differentially private ML, by eliminating the negative effects of clipping and noising large gradients.
To demonstrate empirically the superior performance of tempered sigmoids, we show in Section 5.1 how using tempered sigmoids instead of ReLU activations significantly improves a model’s private-learning suitability and achievable privacy/accuracy tradeoffs.
We advance the state-of-the-art of deep learning with differential privacy for MNIST, FashionMNIST, and CIFAR10. On these datasets, we find in Section 5.2 that the parameter setting in which tempered sigmoids perform best happens to correspond to the tanh function. On MNIST, our model achieves 98.1% test accuracy for a privacy guarantee of , whereas the previous state-of-the-art reported in the TensorFlow Privacy library was 96.6%. On FashionMNIST, we obtain test accuracy compared to for . Finally, on CIFAR10, we achieve 66.2% test accuracy at in a setup for which prior work achieved 61.6%.
Training-data Memorization, Differential Privacy, and DP-SGD
Machine learning models easily memorize sensitive, personal, or private data that was used in their training, and models may in practice disclose this data—as demonstrated by membership inference attacks and secret extraction results .
To reason about the privacy guarantees of algorithms such as training by stochastic gradient descent, differential privacy has become the established gold standard . Informally, an algorithm is differentially private if it always produces effectively the same output (in a mathematically precise sense), when applied to two input datasets that differ by only one record. Formally, a learning algorithm that trains models from the set is -differentially-private, if the following holds for all training datasets and that differ by exactly one record:
Here, gives the formal privacy guarantee, by placing a strong upper bound on any privacy loss, even in the worst possible case. A lower indicates a stronger privacy guarantee or a tighter upper bound. The factor allows for some probability that the property may not hold (in practice, this is required to be very small, e.g., in inverse proportion to the dataset size).
A very attractive property of differential-privacy guarantees is that they hold true for all attackers—whatever they are probing and whatever their prior knowledge—and that they remain true under various forms of composition. In particular, the output of a differentially-private algorithm can be arbitrarily post processed, without any weakening of the guarantees. Also, if sensitive training data contains multiple examples from the same person (or, more generally, the same sensitive group), -differentially-private training on this data will result in model with a -differential-privacy guarantee for each person, as long as at most training-data records are present per person.
Approach
When training a model with differential privacy, gradients computed during SGD are computed individually for each example (i.e., the gradient computation is not averaged across all samples contained in a minibatch). The gradient for each model parameter is then clipped such that the total norm of the gradient across all parameters is bounded by :
Because this operation is performed on per-example gradients, this allows DP-SGD to control the sensitivity of learning to individual training examples. However, this clipping operation will lead to information loss when some of the signal contained in gradients is discarded because the magnitude of gradients is too large. One way to reduce the magnitude (or at least control it), is to prevent the model’s activations from exploding. This is one of the reasons why common design choices for the architecture of modern deep neural networks make it difficult to optimize model parameters with DP-SGD: prominent activation functions like the REctified Linear Unit (ReLU) are unbounded.
We hypothesize that replacing ReLUs with a bounded activation function prevents activations from exploding and thus keeps the magnitude of gradients to a more reasonable value. This in turn implies that, given a fixed level of privacy guarantee, the clipping operation applied by DP-SGD will discard less signal from gradient updates—eventually resulting in higher performance at test time.
Based on this intuition, we propose replacing the unbounded activations typically used in deep neural networks with a general family of bounded activations: the tempered sigmoids. We note that an idea that is conceptually close to ours, the use of tempered losses, was recently found to provide robustness to noise during training .We experimented with the tempered loss of but did not find any improvements for DP-SGD training. Tempered sigmoids are the family of functions that take the form of:
where controls the scale of the activation, is the inverse temperature, and is the offset. By decreasing the value of , we reduce the magnitude of a neuron’s activation. Complementary to this, the inverse temperature rescales a neuron’s weighted inputs. We note that setting , , and in particular yields the tanh function exactly, i.e., we have .
Controlling the gradient norm with tempered sigmoids.
One of the main issues in practice with DP-SGD is tuning the value of the clipping parameter:
If is set too low, then clipping introduces bias by changing the underlying objective optimized during learning.
Instead if the clipping parameter is set too high, clipping increases variance by forcing DP-SGD to add too much noise. Indeed, recall that DP-SGD adds Gaussian noise with variance to the average of (clipped) per-example gradients. This noise is scaled to the clipping norm such that where is a hyperparameter called the noise multiplier. Thus, large clipping norms lead to noise with large variance being added to the average gradient before it is applied to update model parameters.
It turns out that the temperature parameter in Equation (3) can be used as a knob to control the norm of the gradient of the loss function, and with an appropriate choice avoids these two issues of clipping. In the following, we formalize the relationship between our tempered sigmoids and the clipping of DP-SGD in the context of the binary logistic loss and its multiclass counterpart.
Experimental Setup
We use three common benchmarks for differentially private ML: MNIST, FashionMNIST, and CIFAR10. While the three datasets are considered as “solved” in the computer vision community, achieving high utility with strong privacy guarantees remains difficult on all three datasets . Concretely, the state-of-the-art for MNIST is a test accuracy of given an differential privacy guarantee. With stronger guarantees, the accuracy continues to degrade. In the same privacy-preserving settings, prior approaches achieve a test accuracy of on FashionMNIST. For CIFAR10, a test accuracy of can be achieved given an differential privacy guarantee.
All of our experiments are performed with the JAX framework in Python, on a machine equipped with a 5th generation Intel Xeon processor and NVIDIA V100 GPU acceleration. For both MNIST and FaashionMNIST, we use a convolutional neural network whose architecture is described in Table 2. For CIFAR10, we use the deeper model in Table 2. The choice of architectures is motivated by prior work which showed that training larger architectures is detrimental to generalization when learning with privacy . This can be explained in two ways. Given a fixed privacy guarantee, increasing the number of parameters increases (a) how much each parameter needs to be clipped relatively and (b) how much noise needs to be added, with the norm of noise increasing as a function of the square root of the number of parameters.
When we train these architectures with ReLU activations for both the convolution and fully-connected layers, we are able to exactly reproduce the previous state-of-the-art results mentioned above for MNIST, FashionMNIST, and CIFAR10. To experiment with the tempered sigmoid proposed in Section 3, we implement it in JAX and use it in lieu of the ReLU in the architecture from Table 2 and Table 2. Our code is staged for an open-source release, and we include the code snippet for the tempered sigmoid activation below—to demonstrate the practicality of implementing the change we propose in neural architectures.
from jax.scipy.special import expitdef tempered_sigmoid(x, scale=2., inverse_temp=2., offset=1., axis=-1): return scale * expit(inverse_temp * x) - offsetdef elementwise(fun, **fun_kwargs): """Layer that applies a scalar function elementwise on its inputs.""" init_fun = lambda rng, input_shape: (input_shape, ()) apply_fun = lambda params, inputs, **kwargs: fun(inputs, **fun_kwargs) return init_fun, apply_funTemperedSigmoid = elementwise(tempered_sigmoid, axis=-1)
Evaluating the family of tempered activation functions
For each of the three datasets considered, we use DP-SGD to train a pair of models. The first model uses ReLU whereas the second model uses a tempered sigmoid as the activation for all of its hidden layers (i.e., both convolutional and fully-connected layers). The models are based off the architecture of Table 2 for MNIST and FashionMNIST, or Table 2 for CIFAR10. All other architectural elements are kept identical. In our experiments, we subsequently fine-tuned architectural aspects (i.e., model capacity) as well as the choice of optimizer and its associated hyperparameters, separately for the activation function in each setting (ReLU and tempered sigmoid), to avoid favoring any one choice.
Recall from Section 3 that tempered sigmoids are bounded activations that are parameterized such that their inputs and output can be rescaled—through the inverse temperature and scale parameters respectively—and their output recentered with the offset . Tempered sigmoids help control the norm of the gradient of the loss function, and in turn mitigate some of the negative effects from clipping. In Figure 2, we visualize the influence of the scale , inverse temperature , and offset on the test performance of models trained with DP-SGD and tempered sigmoids .
Tempered sigmoids significantly outperform models trained with ReLU on all three datasets. On MNIST, the best tempered sigmoid achieves test accuracy whereas the baseline ReLU model trained to provide identical privacy guarantees () achieved accuracy. This contributes to bridging the gap between privacy-preserving learning and non-private learning, which results in a test accuracy of for this architecture with both ReLU and tanh. On FashionMNIST, we achieve a best performing model of with tempered sigmoids in comparison with with ReLUs. A non-private model achieves with tanh and with ReLUs. On CIFAR10, the best tempered sigmoid architectures achieve test accuracy whereas the ReLU variant obtained under the same privacy guarantees and the non-private baseline .
From Figure 2, it appears clearly that a subset of tempered sigmoids performs best when learning with DP-SGD on the three datasets we considered. These form a cluster of points which result in models with significantly higher test accuracy. These points are colored in dark green. For each dataset, we compute the average value of the 10% best-performing triplets . On MNIST, the average triplet obtained is , on FashionMNIST , and on CIFAR10 . While this observation may not hold for other datasets, it is thus interesting to note here how these values happen to be close to the triplet setting for MNIST and FashionMNIST—and to a lesser extent for CIFAR10. Recall that this setting corresponds exactly to the tanh function. For this reason, we explore the particular case of tanh next. We seek to understand whether it is able to sustain the significant improvements of tempered sigmoids over ReLU for the datasets we considered.
2 Improving the state-of-the-art on MNIST, FashionMNIST, and CIFAR10 with tanh
We now turn to the particular case of the tanh function to understand the broader implications of our results from Section 5.1 for the three datasets considered: MNIST, FashionMNIST, and CIFAR10. In the following experiment, we find that for these datasets positive results observed on the general family of tempered sigmoids can be reproduced with tanh alone, which is obtained by setting , and in . One of the reasons we focus on the particular example of tanh is that changing the activation function to a tanh does not introduce new hyperparameters in learning: the values of need not be tuned if we choose to train architectures with a tanh.
The tanh was an improvement on all three datasets and its performance is in line with the best test accuracy observed across tempered sigmoids on Figure 2. The test accuracy of the tanh model is on MNIST, on FashionMNIST, and on CIFAR10. Figure 3 visualizes the privacy-utility Pareto curve of the ReLU and tanh models trained with DP-SGD for all three datasets. Rather than plotting the test accuracy as a function of the number of steps, we plot it as a function of the privacy loss (but the privacy loss is a monotonically increasing function of the number of steps). The tanh models outperform their ReLU counterparts consistently regardless of the privacy loss expended.
Impact of tanh on activation norms.
To explain why a simple change of activation functions has such a large positive impact on the model’s accuracy, we conjectured that the bounded nature of the tanh, and more generally tempered sigmoids, prevents activations from exploding during training.
Fine-tuning the optimizer.
To ensure that the comparison between ReLU and tempered sigmoids is fair, we now turn our attention to the training algorithm itself and verify that the superior behavior of tanh holds after a thorough hyperparameter search: this includes the number of filters , learning rate, optimizer, batch size, and number of epochs. We find that it is important to tailor algorithm and hyperparameter choices to the specificities of private learning: an optimizer or learning rate that yields good results without privacy may not perform well with privacy. Among the hyperparameters mentioned above, it is particularly important to fine-tune the learning rate to maximize performance given a fixed privacy budget. This is because the privacy budget limits the number of steps we can possibly take on the training set (as visualized on Figure 3). For example, Table 3 shows how learning rates obtained through a hyperparameter search based on Batched Gaussian Process Bandits vary across the non-DP and DP settings when training on FashionMNIST, but that the choice of optimizer (and in particular whether it is adaptive or not) does not influence results as much.
Table 4 summarizes the results after performing this hyperparameter search for each of the three datasets considered in our experiments. We compare the non-private baseline and the DP-SGD with ReLU baseline to our DP-SGD approach with tempered sigmoids (instantiated by tanh here on our three datasets) after all hyperparameters have been jointly fined-tuned. Even in their own individually-best setting, tempered sigmoids continue to consistently outperform ReLU with 98.1% test accuracy (instead of 96.6% for ReLU) on MNIST, 86.1% test accuracy (instead of 81.9% for ReLU) on FashionMNIST, 66.2% test accuracy (instead of 61.6% for ReLU) on CIFAR10.
Conclusions
Rather than first train a non-private model and later attempt to make it private, we bypass non-private training altogether and directly incorporate specificities of private learning in the selection of activation functions. Selecting a tempered sigmoid as the activation function renders the architecture more suitable for learning with differential privacy. This improves substantially upon the state-of-the-art privacy/accuracy trade-offs on three benchmarks which remain challenging for deep learning with differential privacy: MNIST, FashionMNIST, and CIFAR10. Future work may continue to explore this avenue: model architectures need to be chosen explicitly for privacy-preserving training We expect that in addition to activation functions studied in our work, other architectural aspects can be modified to further reduce the observed gap in performance of private learning compared to non-private learning. In addition, choosing the parameters to be shared across all layers was not a necessity. We found that layer-wise parameters did not improve results on our datasets, but it may be the case for different tasks: this is related to the idea of setting layer-wise clipping norms .
Broader impact
Our work helps make privacy-preserving training more practical. Our analysis and experimental results help practitioners make better choices when design neural architectures for privacy-preserving deep learning. In particular, the conclusions from our paper can readily be applied in real-world machine learning pipelines. For this reason, we expect the broader impact of this work to be generally positive given the numerous applications of machine learning to sensitive datasets. This includes applications in domains like healthcare or language modeling.