See through Gradients: Image Batch Recovery via GradInversion
Hongxu Yin, Arun Mallya, Arash Vahdat, Jose M. Alvarez, Jan Kautz, Pavlo Molchanov
Introduction
Sharing weight updates or gradients during training is the central idea behind collaborative, distributed, and federated learning of deep networks . In the basic setting of federated stochastic gradient descent, each device learns on local data, and shares gradients to update a global model. Alleviating the need to transmit training data offers several key advantages. This keeps user data private, allaying concerns related to user privacy, security, and other proprietary concerns. Further, this eliminates the need to store, transfer, and manage possibly large datasets. With this framework, one can train a model on medical data without access to any individual’s data , or perception model for autonomous driving without invasive data collection .
While this setting might seem safe at first glance, a few recent works have begun to question the central premise of federated learning - is it possible for gradients to leak private information of the training data? Effectively serving as a “proxy” of the training data, the link between gradients to the data in fact offers potential for retrieving information: from revealing the positional distribution of original data , to even enabling pixel-level detailed image reconstruction from gradients . Despite remarkable progress, inverting an original image through gradient matching remains a very challenging task – successful reconstruction of images of high resolution for complex datasets such as ImageNet has remained elusive for batch sizes larger than one.
Emerging research on network inversion techniques offers insights into this task. Network inversion enables noise-to-image conversion via back-propagating gradients on appropriate loss functions to the learnable inputs. Initial solutions were limited to shallow networks and low-resolution synthesis , or creating an artistic effect . However, the field has rapidly evolved, enabling high-fidelity, high-resolution image synthesis on ImageNet from commonly trained classifiers, making downstream tasks data-free for pruning, quantization, continual learning, knowledge transfer, etc. . Among these, DeepInversion yields state-of-the-art results on image synthesis for ImageNet. It enables the synthesis of realistic data from a vanilla pretrained ResNet-50 classifier by regularizing feature distributions through batch normalization (BN) priors.
Building upon DeepInversion , we delve into the problem of batch recovery via gradient inversion. We formulate the task as the optimization of the input data such that the gradients on that data match the ones provided by the client, while ensuring realism of the input data. However, since the gradient is also a function of the ground-truth label, one of the main challenges is to identify the ground-truth label for each data point in the batch. To tackle this, we propose a one-shot batch label restoration algorithm that uses gradients from the last fully connected layer.
Our goal is to recover the exact images that the client possesses. By starting from noisy inputs generated by different random seeds, multiple optimization processes are likely to converge to different minimas. Due to the inherently spatially-invariant nature of convolutional neural networks (CNNs), these resulting images share spatial information but differ in the exact location and arrangement. To allow for improved convergence towards the ground truth images, we compute a registered mean image from all candidates and introduce a group consistency regularization term on every optimization process to reduce deviation. We find that the proposed approach and group consistency regularization provide superior better image recovery compared to prior optimization approaches .
Our non-learning based image recovery method recovers more specific details of the hidden input data when compared to the state-of-the-art generative adversarial networks (GAN), such as BigGAN . More importantly, we demonstrate that a full recovery of individual images of px resolution with high fidelity and visual details, by inverting gradients of the batch, is now made feasible even up to batch size of images.
We introduce GradInversion to recover hidden original images from random noise via optimization given batch-averaged gradients.
We propose a label restoration method to recover ground truth labels using final fully connected layer gradients.
We introduce a group consistency regularization term, based on multi-seed optimization and image registration, to improve reconstruction quality.
We demonstrate that a full recovery of detailed individual images from batch-averaged gradients is now feasible for deep networks such as ResNet-50.
We introduce a new Image Identifiability Precision metric to measure the ease of inversion over varying batch sizes, and identify samples vulnerable to inversion.
Related Work
Image synthesis. GANs have delivered state-of-the-art results for generative image modeling, e.g., BigGAN-deep on ImageNet . Training a GAN’s generator, however, requires access to original data. Multiple works have also looked into training GANs given only a pretrained model , but result in images that lack details or perceptual similarities to original data.
Prior work in security studies image synthesis from a pretrained single network. The model inversion attack by Fredrikson et al. optimizes inputs to obtain class images using gradients from the target model. Follow-up works scale to new threat scenarios, but remain limited to shallow networks. The Secret Revealer exploits priors from auxiliary datasets and trains GANs to guide inversion, scales the attack to modern architectures, but on the datasets with less diverse samples, e.g., MNIST and face recognition.
Though originally aiming at understanding network properties, visualization techniques offer another viable option to generate images from networks. Mahendran et al. explore inversion, activation maximization, and caricaturization to synthesize “natural pre-images” from a trained network . Nguyen et al. use global generative priors to help invert trained networks for images, followed by Plug & Play that boosts up image diversity and quality via latent priors. These methods still rely on auxiliary dataset information, feature embedding, or altered training.
Recent efforts focus on image generation from a pretrained network without any auxiliary information. DeepDream by Mordvintsev et al. hints on “dreaming” new visual features onto images leveraging gradients on inputs, extendable towards noise-to-image conversion. Saturkar et al. extended the approach to more realistic images. The more recent extensions significantly improved state-of-the-art performance on image synthesis from off-the-shelf classifiers, without auxiliary information nor additional training but relying on BN statistics.
Gradient-based inversion. There have been early attempts to invert gradients in pursuit of proxy information of the original data, e.g., the existence of certain training samples or sample properties of the dataset. These methods primarily focus on very shallow networks.
A more challenging task aims at reconstructing the exact images from gradients. The early attempt by Phong et al. brought theoretical insights on this task by showing provable reconstruction feasibility on single neuron or single layer networks. Wang et al. empirically inverted out single image representations from gradients of a -layer network. Along the same line, Zhu et al. pushed gradient inversion towards deeper architectures by jointly optimizing for “pseudo” labels and inputs to match target gradients. The method leads to accurate reconstruction up to pixel-level, while remains limited to continuous models (e.g., ones with sigmoid instead of ReLU) without any strides, and scales up to low-resolution CIFAR datasets. Zhao et al. extend the approach with a label restoration step, hence improving speed of single image reconstruction. The very recent work by Geiping et al. for the first time pushed the boundary towards ImageNet-level gradient inversion - it reconstructs single images from gradients. Despite remarkable progress, the field struggles on ImageNet for any batch size larger than one, when gradients get averaged.
GradInversion
In this section, we explain GradInversion in detail. We first frame the problem of input reconstruction from gradients as an optimization process. Then, we explain our batch label restoration method, followed by the auxiliary losses used to ensure realism and group consistency regularization.
Given a network with weights and a batch-averaged gradient calculated from a ground truth batch with images and labels , our optimization solves for
where refers to ground truth gradient at layer , and the summation, scaled by , runs over all layers. One key yet missing component here is the that initiates the backpropagation. We next explain an effective algorithm for restoring batch-wise label from the gradients of the fully connected classification layer.
2 Batch Label Restoration
Considering the cross-entropy loss for the classification task, the ground truth gradient of of batch size can be decomposed into:
where denotes an original image/label pair. For each image , the gradient w.r.t. the network final logits at index is , where is the post-softmax probability in range (, ), and is the binary presentation of at index among total classes. Consequentially, this leaves \text{sign}\big{(}\nabla_{z_{n,k}}\mathcal{L}(x_{k},y_{k})\big{)} negative iff at the ground truth index, and positive otherwise. However, we do not have access to as gradients are only given w.r.t. the model parameters.
Note that where is the input of the fully connected layer, and is also the output of previous layer. If the previous layer has commonly used activation functions such as ReLU or sigmoid, is always non-negative. This hints on target label existence via signs of a new informative indicator:
where is an -matrix, constructed by summing the tensor along the feature dimension . Interestingly, contains negative values for the ground truth label of each instance. Thus, the column of can be used to restore the ground truth label for the image by simply identifying the index of the negative entry. Zhao et al. explored this rule for single image label restoration. However, we do not have access to in our multi-sample batch setup as the given gradients are averaged over all images.
Motivated by this, we define the -dimensional batch-level vector by averaging along its columns:
The appealing property of is that it can be computed easily from the given gradient for the fully connected layer by summing it along the feature dimension as shown on the right hand side of Eqn. 7.
As noted above, each column in is a vector, containing a single negative peak at the label index and positive otherwise. Since the vector is a linear super-position of ’s columns, from all individual images ’s in the batch, this information can be lost during summation. However, we empirically observe that the encoded positions often possess larger magnitudes . This leaves a negative sign mostly intact when the summation brings in positive values from other images.
To further enable a more robust propagation of negative signs, we utilize column-wise minimum values, instead of summation along the feature dimension for the calculation: to have a sum along the feature dimension be negative, at least one of its positions has to be negative, but not vice versa. This further boosts up the label restoration accuracy, especially when the batch size is large. Thus, we formulate the final label restoration algorithm for batch size as:
with corresponding to the feature embedding dimension before the fully connected layer. The resulting supports Eqn. 3 in subsequent optimization in pursuit for . One limitation of the proposed method is that it assumes non-repeating labels in the batch, which generally holds for a randomly sampled batch of size that is much smaller than the number of classes at for ImageNet.
Even with correct , finding the global minima for remains challenging. The task is under-constrained, suffers from information loss due to non-linearity and pooling layers, and has only one correct solution . We next introduce based on fidelity and group consistency regularization to assist with this optimization.
3 Fidelity (Realism) Regularization
We use the strong prior proposed in DeepInversion to guide the optimization towards natural images. Specifically, we add to the loss function to steer away from unrealistic images with no discernible visual information:
where and are the batch-wise mean and variance estimates of feature maps corresponding to the convolutional layer. By enforcing valid intermediate distributions at all levels, yields convergence towards realistic-looking solutions.
4 Group Consistency Regularization
An additional challenge of gradient-based inversion lies in the exact localization of the target object, due to translational invariance of CNNs. Unlike an ideal scenario where optimization converges to one ground truth, we observe that when repeating the optimization with different seeds, e.g., as in Fig. 2, each optimization process unveils a local minimum that allocates semantically correct image features at all levels, but differs from others – images shift around the ground truth, focusing on slightly different details. During the forward pass, the existence of pooling layers, strided convolutions, and zero-padding, jointly causes spatial equivariance among the restored images, as also observed by Geiping et al. . A combination of the restored images from varying seeds, however, hints at the potential for a better restoration of the final image closer to the ground truth.
We introduce a group consistency regularization term that exploits multiple seeds simultaneously in a joint optimization manner, as shown in Fig 3. Intuitively, a joint exploration with multiple paths can expand and enlarge the search space during gradient descent. However, we have to regularize them to prevent too much divergence, at least in the final stages, given the search for a single target. We optimize each input using the target Eqn. 1. To facilitate information exchange, we regularize all the inputs simultaneously with a new group consistency regularization term:
This leads to our final group consistency regularization shown in Fig. 3. We (i) first compute the pixel-wise mean within the candidate set of size as a coarse registration target, (ii) register each individual image towards the target via , and (iii) obtain the post-registered mean as the target for regularization. We use RANSAC-flow for . As we will show later, group consistency regularization enables consistent improvements in recovery across various evaluation metrics, further closing the gap between reconstructed and original batches.
5 The Final Update
Using all the above losses, we update the input in an iterative manner. To further encourage exploration and diversity, we add pixel-wise random Gaussian noise in each update, inspired by the Langevin updates in energy-based models . Our final optimization steps are:
where corresponds to an optimizer update, denotes randomly sampled noise to encourage exploration, is the learning rate, and re-scales the finally added noise.
Experiments
We evaluate our method for the classification task on the large-scale 1000-class ImageNet ILSVRC 2012 dataset at pixels. We first perform a number of ablations to evaluate the contribution of each component of our method. Then, we show the success of GradInversion and compare with prior art. Finally, we increase the batch size to explore the limits of gradient inversion.
Evaluation metrics. We present visual comparisons of images obtained under different settings and evaluate three quantitative metrics for image similarity. To account for pixel-wise mismatch, we compute: (i) the cosine similarity in FFT frequency response, (ii) post-registration PSNR, and (iii) LPIPS perceptual similarity score between reconstruction and original images.
We first restore labels from the gradients of the fully connected layer. Table 1 summarizes the averaged label restoration accuracy on ImageNet training and validation sets, given K randomly drawn samples divided into varying batch sizes. In a zero-shot method, GradInversion restores original labels accurately, improving upon prior art .
1.2 Batch reconstruction
We next gradually add each proposed loss to the optimization process. Here we focus on a batch of images for algorithm ablations before expanding towards a larger batch size. We summarize results in Table 5 and discuss insights next:
Adding . Adding fidelity regularization immediately improves image quality. Conditioned on image prior, gradient inversion starts to allocate visual details towards individual images, enabling both visual and quantitative improvements in Table 5.
a) Lazy regularization. We observe “lazy” pixel-wise mean as regularization target already brings in performance improvements. Though not yet accommodating for inter-seed variation, pixel-wise mean hints on correct “perceived” positions of the target objects. Objects start to emerge at correct positions with improved orientations.
b) Registration enhancement. We then add in registration to exploit consensus among candidates. We start registration after K initial optimization iterations to allow for sufficient feature emergence, then iterate every iterations. Ideally, each candidate shall be registered to its original image for the best spatial adjustment. While given no such access, registration to pixel-wise mean turns out to be effective. The final registration-based regularization helps close the remaining gap - it improves all evaluation metrics in Table 5. At this stage, GradInversion accurately allocates detailed original contents to individual images, from averaged gradients.
Inverting different networks. We observe that gradients from a stronger feature extractor leak more information - see a quick comparison in Table 3. Self-supervised pretraining of ResNet-50 leads to the best image reconstruction, when compared to a standard training recipe of the same ResNet-50 architecture, and a weaker ResNet-18. We continue our analysis with ResNet-50 MOCO V2 to study the limits of batch reconstruction under gradient inversion.
2 Comparison with the state-of-the-art
We next compare with prior art on the batch size of 8 images with 224x224px. We summarize both qualitative (Fig. 4) and quantitative results (Table 4). We compare with three viable methods for image synthesis:
(i) Gradient inversion : We first compare with prior model inversion methods for gradient matching: (i) deep gradient leakage method by Zhu et al. and (ii) federated gradient inversion by Geiping et al. . We first extend both techniques towards ImageNet batch restoration following the authors’ public open-sourced repository . For an additional fair comparison, we also compare with both methods at batch size one in Fig. 5, and show notable fidelity and localization improvements.
(ii) DeepInversion : We also analyze performance improvements over the baseline DeepInversion method that synthesizes images conditioned on ground-truth labels.
GradInversion outperforms prior art both visually (Fig. 4) and numerically (Table 4). Without label restoration, a joint optimizing to seek for image-label pairs struggles to converge on ImageNet, as also observed by even at batch size one. Total variation prior and magnitude-invariant loss as in help improve reconstruction, but remain too weak to guide optimization towards ground truth. The DeepInversion baseline improves image fidelity as expected, but inverts images with little observable links to the original batch. Projection onto BigGAN’s latent space offers a balance between image fidelity and restored details, but falls short under a weaker guidance from original gradients rather than original images, as projection to latent space is NP-hard and misses visual details .
3 Effect of scaling up the batch size
We next increase the batch size. Our current analysis scales up to batch size using a GB NVIDIA V100 GPU. As shown in Fig. 6, the amount of recoverable image content gradually decreases as batch size increases. As expected, more averaging of gradient information in a batch better protects privacy of an individual image. Surprisingly, GradInversion still unveils a decent amount of original visual information at batch size , and sometimes a viable complete reconstruction, as shown in Fig. 7.
Image Identifiablity Precision (IIP). We formulate a new score that measures the amount of “image-specific” features revealed by gradient inversion. Intuitively, this measures how easy it is to identify a particular image, given only its reconstruction, among all its similar peers in the original dataset. Quantifiably, we calculate the fraction of exact matches between an original image and the nearest neighbor to its reconstruction. The resulting metric, referred to as Image Identifiability Precision (IIP), evaluates gradient inversion strength across varying batch sizes. Fig. 8 plots the IIP curve for GradInversion. As expected, reconstruction efficacy gradually decreases as batch size increases, as also seen in Fig. 6. We make a surprising observation that many samples () can be correctly identified even after averaging gradient from 48 images.
Conclusions
We introduced GradInversion to reconstruct individual images in a batch, given averaged gradients. We showed that the assumption of privacy when sharing gradients from deep networks on complex datasets even at large batch sizes, does not hold. This offers new insights into the development of privacy-preserving deep learning frameworks.
It can also be fruitful to study the underlying mechanism of information transfer that enables original data recovery from gradients. We hope that future work can study vulnerabilities of aggregration-based federated learning , as well as further strengthen them to prevent inversion.
References
Appendix C - Additional Details & Analysis
We next discuss several of our observations when performing GradInversion for the ResNet-50 network (MOCO V2) on the ImageNet1K dataset. We would like to note that these observations hold for the chosen network and optimization settings used, and may not be general in scope. We hope that sharing our experiences would help provide insights for future work.
Vanishing objects. Images recovered from gradient inversion occasionally omit details of original images, as shown in Fig. 11 (a) where the diver and the bird disappear post inversion. The observation is in line with the missing details phenomena during the latent code projection as in StyleGAN2 by Karras et al. .
Texts & digits. GradInversion unveils existence of texts and digits, while their exact details remain blurry. See several quick examples in Fig. 11 (b).
Human faces. Recovery remains harder for samples and distributions in ImageNet that involve human faces (also deemed challenging by BigGAN ). Even though detailed features and patches can be reversed, e.g., mouths, eyes, and noses, etc., they are not correctly arranged spatially, as shown in Fig. 11 (c). We conjecture that this is a result of (i) the under-representation of such distributions in ImageNet1K and (ii) such features being ignored by the network for the classification task. Different observations may hold for other datasets where tasks enforce networks to focus on spatial alignments of facial features, e.g., towards facial recognition.