Reconstructing Training Data from Trained Neural Networks

Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, Michal Irani

Introduction

It is commonly believed that neural networks memorize the training data, even when they are able to generalize well to unseen test data (e.g., (Zhang et al., 2021; Feldman, 2020)). Exploring this memorization phenomenon is of great importance both practically and theoretically. Indeed, it has implications on our understanding of generalization in deep learning, on the hidden representations learnt by neural networks, and on the extent to which they are vulnerable to privacy attacks.

A fundamental question for understanding memorization is:

Are the specific training samples encoded in the parameters of a trained classifier? Can they be recovered from the network parameters?

In this work, we study this question, and devise a novel scheme which allows us to reconstruct a significant portion of the training data from the parameters of a trained neural network alone, without having any additional information on the data. Thus, we provide a proof-of-concept that the learning process can sometimes be reversed: That is, instead of learning a model given a training dataset, it is possible to find the training data given a trained model. In Figure 1 we show how our approach reconstructs images from the CIFAR10 dataset, given a simple trained binary classifier.

Many works try to “crack” neural networks by analyzing and visualizing either their learnt parameters or representations (Erhan et al., 2009; Mahendran and Vedaldi, 2015; Olah et al., 2017, 2020). This is usually done by “inverting” the model, namely finding inputs that are strongly correlated with the model’s activations (Mordvintsev et al., 2015; Yin et al., 2020; Fredrikson et al., 2015). Unsurprisingly, the results are semantically correlated with the training dataset. However, one rarely sees an exact version of a training sample.

Our results have potential negative implications on privacy in deep learning. Our scheme can be viewed as a training-data reconstruction attack, since an adversary might recover sensitive training data. For example, if a medical device includes a model trained on sensitive medical records, an adversary might reconstruct this data and thus violate the privacy of the patients. Privacy attacks in deep learning have been widely studied in recent years (cf. Liu et al. (2021)), but as far as we are aware, the known attacks cannot reconstruct portions of the training data from a trained model.

Our approach relies on theoretical results about the implicit bias in training neural networks with gradient-based methods. The implicit bias has been studied extensively in recent years with the motivation of explaining generalization in deep learning (see Section 2). We use results by Lyu and Li (2019); Ji and Telgarsky (2020), which establish that, under some technical assumptions, if we train a neural network with the binary cross entropy loss, its parameters will converge to a stationary point of a certain margin-maximization problem. This result implies that the parameters of the trained network satisfy a set of equations w.r.t. the training dataset. In our approach, given a trained network, we find a dataset that solves this set of equations w.r.t. the trained parameters.

We show that large portions of the training samples are encoded in the parameters of a trained classifier. We also provide a practical scheme to decode the training samples, without any assumptions on the data. As far as we know, this is the first work that shows that reconstruction of actual training samples from a trained neural network classifier is possible.

Related Work

The most common approach for analysing what is learnt by a neural network is by searching inputs that maximize the class output or the activations of neurons in intermediate layers (Erhan et al., 2009; Olah et al., 2020). Oftentimes this is done via optimization with respect to the model input. Optimizing without any prior on the input usually results in noise inputs. Therefore, most approaches incorporate priors such as smoothness regularization or the use of pre-trained image generators (Mahendran and Vedaldi, 2015; Yosinski et al., 2015; Mordvintsev et al., 2015; Nguyen et al., 2016a, b, 2017) (see Olah et al. (2017) for a comprehensive summary). Optimization w.r.t. the input may also result in adversarial examples (Szegedy et al., 2013; Goodfellow et al., 2014). Recently, (Tsipras et al., 2018; Engstrom et al., 2019) showed that classifiers trained to be robust to adversarial examples tend to learn representations that are more aligned with human vision. This was later utilized by (Santurkar et al., 2019; Mejia et al., 2019) to generate class-conditional images from a trained classifier. While all those approaches indicate that, unsurprisingly, the learnt representations are strongly correlated with the datasets on which the model was trained, none of them demonstrate the reconstruction of exact training samples from the trained models.

Many methods deal with extracting sensitive information from trained models. Perhaps the closest to our approach is model-inversion that aims to reconstruct class representatives from the training data of a trained model (Fredrikson et al., 2015; He et al., 2019; Yang et al., 2019; Yin et al., 2020). It is important to note that the reconstructed images, albeit semantically similar to some input images, are still not actual samples from the training set. Carlini et al. (2021, 2019) demonstrated reconstruction of training data from generative language models. By completing sentences, they reveal sensitive information from the training data. We note that this approach is specific to generative language models, while our approach considers classifiers and is less data specific. Membership-inference attacks (Shokri et al., 2017) aim to determine whether a given data point was used to train the model or not. For these methods to work, the adversary must be able to guess a specific input, whereas our approach does not assume such ability. Lastly, avoiding leakage of sensitive information on the training dataset is the motivation behind differential privacy in machine learning, which has been extensively studied (Abadi et al., 2016; Dwork et al., 2006; Chaudhuri et al., 2011). For an elaborated discussion on the relation of these approaches to ours see Appendix A.

In overparameterized neural networks one might expect overfitting to occur, but it seems that gradient-based methods are biased towards networks that generalize well (Zhang et al., 2021; Neyshabur et al., 2017). Mathematically characterizing this implicit bias is a major problem in the theory of deep learning. Our approach is based on a characterization of the implicit bias of gradient flow in homogeneous neural networks due to Lyu and Li (2019) and Ji and Telgarsky (2020) (see Section 3 for details). The implicit bias of gradient-based methods in neural networks was extensively studied in recent years both for classification tasks (e.g., Soudry et al. (2018); Gunasekar et al. (2018c); Ji and Telgarsky (2018); Nacson et al. (2019); Vardi et al. (2021); Chizat and Bach (2020); Gunasekar et al. (2018a); Moroshko et al. (2020)) and regression tasks (e.g., Gunasekar et al. (2018b); Arora et al. (2019); Azulay et al. (2021); Yun et al. (2020); Woodworth et al. (2020); Razin and Cohen (2020); Li et al. (2020); Vardi and Shamir (2021); Timor et al. (2022)). See Vardi (2022) for a survey.

Background and Reconstruction Scheme

In this section we present our training data reconstruction scheme, as well as provide a brief overview on the theoretical results about implicit bias, which motivate our approach.

Moreover, L(θ(t))→0{\cal L}({\boldsymbol{\theta}}(t))\to 0 as t→∞t\to\infty.

The above theorem guarantees directional convergence to a first order stationary point (of the optimization problem (1)), which is also called Karush–Kuhn–Tucker point, or KKT point for short. The KKT approach allows inequality constraints, and is a generalization of the method of Lagrange multipliers, which allows only equality constraints.

2 Dataset Reconstruction

Suppose we are given a trained neural network with parameters θ{\boldsymbol{\theta}}, and our goal is to reconstruct the dataset that the network was trained on. Although Theorem 3.1 holds asymptotically as the time tt tends to infinity, it suggests that also after training for a finite number of iterations the parameters of the network might approximately satisfy Eq. (2), and the coefficients λi\lambda_{i} satisfy Eq. (4). Since nn is unknown (and so is the number of samples on the margin) we set m≥2nm\geq 2n which represents the number of samples we want to reconstruct (thus, we only need to upper bound nn), and fix yi=1y_{i}=1 for i=1,…,m/2i=1,\ldots,m/2 and yi=−1y_{i}=-1 for i=m/2+1,…,mi=m/2+1,\ldots,m. We define the following losses:

Note that the unknown parameters are the xi\mathbf{x}_{i}’s and λi\lambda_{i}’s, and that θ{\boldsymbol{\theta}} and the yiy_{i}’s are given. The loss LstationaryL_{\text{stationary}} represents the stationarity condition that the parameters of the network satisfy, and LλL_{\lambda} represents the dual feasibility condition. We additionally define LpriorL_{\text{prior}} which represents some prior knowledge we might have about the dataset. For example, if we know that the dataset contains images, prior knowledge would be that each input coordinate (i.e. each pixel) is between and 11. Given no prior knowledge on the data, we can define Lprior≡0L_{\text{prior}}\equiv 0. Finally, we define the reconstruction loss as:

We note that if there exist {xi}i=1n\{\mathbf{x}_{i}\}_{i=1}^{n} and {λi}i=1n\{\lambda_{i}\}_{i=1}^{n} which satisfy the KKT conditions, then there are {xi}i=1m\{\mathbf{x}_{i}\}_{i=1}^{m} and {λi}i=1m\{\lambda_{i}\}_{i=1}^{m} which achieve zero loss in Eq. (8). Indeed, such a solution can be obtained by adding to {xi}i=1n\{\mathbf{x}_{i}\}_{i=1}^{n} additional points xj\mathbf{x}_{j} with λj=0\lambda_{j}=0, or by duplicating some points in {xi}i=1n\{\mathbf{x}_{i}\}_{i=1}^{n} and modifying the λ\lambda’s accordingly. Also, note that since we choose m≥2nm\geq 2n, then we set at least nn labels yiy_{i} to 11 and at least nn labels to −1-1. Hence, there is a solution to Eq. (8) even though we do not know the real distribution of labels in the actual training data.

Intuitively, a reason to believe that there is enough information in Eq. (2) to reconstruct the data, is the following observation: Eq. (2) represents a set of pp equations with O(nd)O(nd) unknown variables, where pp is the number of parameters in the network. In practice, neural networks are often highly overparameterized (i.e., p>ndp>nd), suggesting more equations than variables.

A Simple Experiment in Two Dimensions

To further improve our reconstruction results, we remove some of the extra points which did not converge to a training sample. In Figure 2e we removed points xi\mathbf{x}_{i} with corresponding λi<5\lambda_{i}<5. According to Eq. (2), points with λi=0\lambda_{i}=0 should not affect the parameters, hence their corresponding xi\mathbf{x}_{i} can take any value. In practice, it is sufficient to remove points with a small enough corresponding λi\lambda_{i}. Finally, to remove duplicates, we greedily remove points which are very close to other points. That is, we randomly order the points, and iteratively remove points that are at distance <0.03<0.03 from another point. The final reconstruction result is depicted in Figure 2f.

Results

We conduct experiments on binary classification tasks where images are taken from the MNIST (LeCun et al., 2010) and CIFAR10 (Krizhevsky et al., 2009) datasets and the labels are set to odd vs. even digits (MNIST), and vehicles vs. animalsAutomobile, Truck, Airplane, Ship vs. Bird, Horse, Cat, Dog, Deer, Frog. (CIFAR10). We make sure that the class distribution in the training and test sets is balanced, and normalize the train and test sets by reducing the mean of the training set from both.

We consider MLP architectures. Unless stated otherwise, our models comprise of three fully-connected layers with dimensions dd-10001000-10001000-11 (where dd is the dimension of the input) with ReLU activations. Biases are set to zero except for the first layer, to line up with the theoretical assumption of homogeneous models in Section 3. The parameters are initialized using standard Kaiming He initialization (He et al., 2015) except for the weights of the first layer that are initialized to a Gaussian distribution with standard deviation 10−410^{-4} (see discussion in Subsection 5.2). We train our models using full batch gradient descent for 10610^{6} epochs with a learning rate of 0.010.01. All models achieve zero training error (i.e., all the train samples are labeled correctly), and a training loss <10−6<10^{-6}. To compute the test accuracy, we use the original test sets of MNIST/CIFAR10 with 1000010000/80008000 images respectively, and labeled accordingly.

2 Training Set Reconstruction

We minimize the loss defined in Eq. (8) with α1=1, α2=5, α3=1\alpha_{1}=1,~{}\alpha_{2}=5,~{}\alpha_{3}=1. We initialize xi∼N(0,σxI)\mathbf{x}_{i}\sim\mathcal{N}(0,\sigma_{x}I), where σx\sigma_{x} is a hyperparameter, and λi∼U\lambda_{i}\sim\mathcal{U}. We set the number of reconstructed samples to m=2nm=2n (where nn is the size of the original training set). Note that our loss contains the derivative of ReLU Eq. (6). This derivative is a step function, containing only flat regions which are hard to optimize. We replace the derivative of the ReLU layer (backward function) with a sigmoid, which is the derivative of softplus (a smooth version of ReLU). We use the fact that our inputs are images to penalize values outside the range $.Tothisendweset. To this end we setL_{\text{prior}}(z)=\max\{z-1,0\}+\max\{-z-1,0\}foreachpixelfor each pixelz,andaverageoveralldimensions(pixels)in, and average over all dimensions (pixels) in\mathbf{x}_{i}.Weoptimizeourlossfor. We optimize our loss for100,000iterationsusinganSGDoptimizerwithmomentumiterations using an SGD optimizer with momentum0.9.Weconductatotalof. We conduct a total of100runsusingarandomgridsearchonthehyperparameters(e.g.learningrate,runs using a random grid search on the hyperparameters (e.g. learning rate,\sigma_{x}.SeeAppendixBforfulldetails).Thisresultsin. See Appendix B for full details). This results in100m$ “reconstructed” inputs.

While some xi\mathbf{x}_{i} end up converging to a training sample, some end as noise (similar phenomenon can be observed in 2D in Figure 2d). To identify the reconstructions that are most similar to a training image we use the SSIM metric (Wang et al., 2004).

In Figure 3 we show the best reconstruction results (in terms of SSIM) for models trained on nn=500500 samples from MNIST/CIFAR10 datasets (with test accuracy 88.0%/77.6%88.0\%/77.6\% resp.). Note that the reconstructed images are very similar to the real input data, although a bit noisy. The source of this noise is not entirely clear. Possible reasons may be the complexity of the optimization problem, or the possibility that the trained model has not fully converged to the KKT point of Problem (1).

We observed that small initializations significantly improve the quality of the reconstructed samples. We conjecture that small initialization causes faster convergence to the direction of the KKT point. This is also theoretically implied in Moroshko et al. (2020) (for certain linear models). Similarly, training for more epochs also improves the quality of the reconstruction. In Appendix C we show results for reconstructions from networks trained with standard initialization or trained for much fewer epochs. During the training phase, we used full batch gradient descent, to remain as much aligned to the theoretical setting. In Appendix C we show that our approach can reconstruct training data also from models trained with mini-batch SGD.

3 Practice vs. Theory

In this section we analyze some relations between our experimental results to the theory laid down in Section 3. Given a trained model and its reconstructed samples, we match each training sample to its best reconstruction (in terms of SSIM score). We then plot this SSIM score against Φ(θ;x)\Phi({\boldsymbol{\theta}};\mathbf{x}) (the value of the model’s output on this training sample) – for all training samples. In Figure 4 each cell shows such plot for a given model. The top row shows models trained on the same architecture with different number of training samples (nn), where in the bottom row we show the results for models trained on n=500n=500 training samples, with different architectures (all results are on CIFAR10).

Recall that we do not expect to reconstruct samples that are far from the margin (Subsection 3.2). It is evident from Figure 4 that good reconstructions (e.g., SSIM>0.4>0.4) are obtained for samples that lie on the margin, as expected from theory. The plots indicate that increasing training size makes reconstruction more difficult. Lastly, as seen from the rightmost plot in the bottom row, we manage to get high-quality reconstructions from a non-homogeneous model (trained with biases in all hidden layers). This indicates that our approach may work beyond the theoretical limitations of Theorem 3.1.

4 Comparison to other Reconstruction Schemes

Given a trained model Φ(θ;⋅)\Phi({\boldsymbol{\theta}};\cdot), we search for x\mathbf{x} which maximizes or minimizes Φ(θ;x)\Phi({\boldsymbol{\theta}};\mathbf{x}), corresponding to positive or negative labels. We initialize x∼N(0,σI)\mathbf{x}\sim\mathcal{N}(0,\sigma I) for several values of σ\sigma and optimize w.r.t. the model output (see Appendix B for the choice of hyperparameters). In Figure 5a (left) it is apparent that in our two-dimensional experiment, model inversion successfully reconstructed 77 training samples, which indeed lie on a local minimum or maximum. However, note that our scheme reconstructs all 2020 samples (Figure 2). In high dimensions, namely, in MNIST and CIFAR, while our scheme can reconstruct a large portion of the training set (Figure 3a), model inversion converges to noisy/blurry class representatives that correspond to high/low output values (Figure 5a, right). Such results are typical with model inversion since not all class members from the training set are visually similar (see discussions in Shokri et al. (2017); Melis et al. (2019)).

The weights of the first fully-connected layer have the same dimension as that of the input. One may wonder whether training samples are directly encoded there. In Figure 5b we show the weights that are most similar (SSIM) to a training sample, or all of them in the 2D case. As seen in the 2D case, most weights are in the general direction of a training sample, however the scale is unknown without prior knowledge on the data. For images (MNIST/CIFAR10), not more than 33 or 44 of the weights have resemblance to training samples, while our scheme manages to reconstruct dozens of samples. See Appendix B and C for details and all 10001000 weights of the models.

Discussion and Conclusion

Even though our results are shown for relatively small-scale models, they are the first to show that the parameters of trained networks may contain enough information to fully reconstruct training samples, and the first to reconstruct a substantial amount of training samples. Moreover, the theoretical basis of the implicit bias in neural networks provides an analytic explanation to this phenomenon.

Solving our optimization problem for convolutional neural networks turned out to be more challenging and is therefore a subject of future research. We note that the theoretical results that we rely on (i.e., Theorem 3.1) also covers convolutional neural networks. We believe that the homogeneity restriction might be relaxed, and showed reconstructions also from a non-homogeneous model (Figure 4, bottom-rightmost). We also believe that our method may be extended to multi-class classifiers using an extension of Theorem 3.1. Finally, showing reconstructions on larger models and datasets, or on tabular or textual data are interesting future directions.

On the theoretical side, it is not entirely clear why our optimization problem in Eq. (8) converges to actual training samples, even though there is no guarantee that the solution is unique, especially when using no prior (other than simple bounding to $$). As a final note, our work brings up the question: are samples on margin the only ones that can be recovered from a trained classifier? or there exist better reconstruction schemes to reconstruct even more training samples from a trained neural network.

This project received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No 788535), and ERC grant 754705, and from the D. Dan and Betty Kahn Foundation, and was supported by the Carolito Stiftung.

References

Checklist

Do the main claims made in the abstract and introduction accurately reflect the paper’s contributions and scope? [Yes]

Did you describe the limitations of your work? [Yes]

Did you discuss any potential negative societal impacts of your work? [Yes]

Have you read the ethics review guidelines and ensured that your paper conforms to them? [Yes]

If you are including theoretical results…

Did you state the full set of assumptions of all theoretical results? [Yes]

Did you include complete proofs of all theoretical results? [N/A] We rely on known theoretical results, so proofs are not required.

Did you include the code, data, and instructions needed to reproduce the main experimental results (either in the supplemental material or as a URL)? [Yes]

Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? [Yes]

Did you report error bars (e.g., with respect to the random seed after running experiments multiple times)? [N/A]

Did you include the total amount of compute and the type of resources used (e.g., type of GPUs, internal cluster, or cloud provider)? [Yes]

If you are using existing assets (e.g., code, data, models) or curating/releasing new assets…

If your work uses existing assets, did you cite the creators? [Yes]

Did you mention the license of the assets? [Yes]

Did you include any new assets either in the supplemental material or as a URL? [N/A] We do not have new assets.

Did you discuss whether and how consent was obtained from people whose data you’re using/curating? [N/A] We used only publicly available assets.

Did you discuss whether the data you are using/curating contains personally identifiable information or offensive content? [N/A] We used only publicly available data.

If you used crowdsourcing or conducted research with human subjects…

Did you include the full text of instructions given to participants and screenshots, if applicable? [N/A]

Did you describe any potential participant risks, with links to Institutional Review Board (IRB) approvals, if applicable? [N/A]

Did you include the estimated hourly wage paid to participants and the total amount spent on participant compensation? [N/A]

Appendix

Below we discuss several privacy attacks that have been extensively studied in recent years (see Liu et al. , Jegorova et al. for surveys).

In membership-inference attacks [Shokri et al., 2017, Long et al., 2018, Salem et al., 2018, Yeom et al., 2018, Song and Mittal, 2021] the adversary determines whether a given data point was used to train the model or not. For example, if the model was trained on records of patients with a certain disease, the adversary might learn that an individual’s record appeared in the training set and thus infer that the owner of the record has the disease with high chance. Note that membership inference attacks are significantly different from our attack, as the adversary must choose a specific data point. E.g., if the inputs are images, then the adversary must be able to guess a specific image.

In model-extraction attacks [Tramèr et al., 2016, Oh et al., 2019, Wang and Gong, 2018, Carlini et al., 2020b, Jagielski et al., 2020, Milli et al., 2019, Rolnick and Kording, 2020, Chen et al., 2021] the adversary aims to steal the trained model functionality. In this attack, the adversary only has black-box access with no prior knowledge of the model parameters or training data, and the outcome of the attack is a model that is approximately the same as the target model. It was shown that in certain cases the adversary can reconstruct the exact parameters of the target model. We note that such attacks might be combined with our attack in order to allow extraction of the training dataset in a black-box setting. Namely, in the first stage the model is extracted using model-extraction attacks, and in the second stage the training dataset is reconstructed using our attack.

Model-inversion attacks [Fredrikson et al., 2015] are perhaps the closest to our attack, as they consider reconstruction of input data given a trained model. These attacks aim to infer class features or construct class representatives, given that the adversary has some access (either black-box or white-box) to a model.

Fredrikson et al. showed that a face-recognition model can be used to reconstruct images of a certain person. This is done by using gradient descent for obtaining an input that maximizes the output probability that the face-recognition model assigns to a specific class. Thus, if a class contains only images of a certain individual, then by maximizing the output probability for this class we obtain an image that might be visually similar to an image of that person. It is important to note that the reconstructed image is not an actual example from the training set. Namely, it is an image that contains features which the classifier identifies with the class, and hence it might be visually similar to any image of the individual (including images from the training set). If the class members are not all visually similar (which is generally the case), then the results of model inversion do not look like the training data (see discussions in Shokri et al. and Melis et al. ). For example, if this approach is applied to the CIFAR-10 dataset, it results in images which are not human-recognizable [Shokri et al., 2017]. In Zhang et al. , the authors leverage partial public information to learn a distributional prior via generative adversarial networks (GANs) and use it to guide the inversion process. That is, they generate images where the target model outputs a high probability for the considered class (as in Fredrikson et al. ), but also encourage realistic images using GAN. We emphasize that from the reasons discussed above, this method does not reconstruct any specific training data point. Another approach for model inversion is training a model that acts as an inverse of the target model [Yang et al., 2019]. Thus, the inverse model takes the predicted confidence vectors of the target model as input, and outputs reconstructed data. A recent paper Balle et al. shows a reconstruction attack where the attacker has information about all the data samples except for one. On the theoretical side, Brown et al. prove that in certain settings, models memorize information about training examples, and show reconstruction attacks on some synthetic datasets.

Model inversion and information leakage in collaborative deep learning was studied in, e.g., He et al. , Melis et al. , Hitaj et al. , Zhu et al. , Yin et al. , Huang et al. . Extraction of training data from language models was studied in Carlini et al. , where they use the ability of language models to complete a given sentence in order to reveal sensitive information from the training data. We note that this attack is specific to language models, which are generative models, while our approach considers classifiers and is less specific.

Avoiding leakage of sensitive information on the training dataset is the motivation behind differential privacy in machine learning, which has been extensively studied in recent years [Abadi et al., 2016, Dwork et al., 2006, Chaudhuri et al., 2011]. This approach allows provable guarantees on privacy, but it typically comes with high cost in accuracy. Other approaches for protecting the privacy of the training set, which do not allow such provable guarantees, have also been suggested (e.g., Huang et al. , Carlini et al. [2020a]).

Appendix B Implementation Details

A typical reconstruction runs for about 3030 minutes on a GPU Tesla V-100 3232GB, for reconstructing m=1000m=1000 samples from a model with architecture dd-10001000-10001000-11, and for 100,000100,000 epochs (running times slightly differ with the number of samples mm, number of epochs and the size of the model, but it still takes about this time to run). Our code is implemented in PyTorch [Paszke et al., 2019]. We will release the code.

B.2 Hyperparameters

The intuition behind is to encourage as many samples to lie on a margin, and thus try and reconstruct some sample from the training set.

To sum it all, the hyperparameters of our reconstruction scheme are:

σx\sigma_{\mathbf{x}}, the initial scale of xix_{i} initialization

α\alpha, of the derivative of the modified ReLU

To find the set of hyperparamerers we used Weights&Biases [Biewald, 2020] using a random grid search where the parameters are sampled from the following distributions:

Learning rate, log-uniform in [10−5,1][10^{-5},1]

σx\sigma_{\mathbf{x}}, log-uniform in [10−6,1][10^{-6},1]

λmin\lambda_{\text{min}}, log-uniform in [10−4,1][10^{-4},1]

When searching for hyperparameters for the model inversion results in Subsection 5.4 we use the following:

Learning rate, log-uniform in [10−6,1][10^{-6},1]

σx\sigma_{\mathbf{x}}, log-uniform in [10−7,1][10^{-7},1]

B.3 Post-Processing of Reconstructed Samples

After the reconstruction run ends we want to match the reconstructed samples to samples from the training set. This is done in the following manner:

Scaling. Each reconstructed sample is stretched to fit into the range $$ (by linear transformation of its minimal/maximal values).

Searching Nearest Neighbours. For each training sample from the training set we compute the distance to all reconstructed outputs using NCC [Lewis, 1995].

Voting. For each training sample we compute the mean of all the closest nearest neighbours (all reconstructed samples with NCC score largest than 0.90.9 of the distance to the closest nearest neighbour). Now we have pairs of trainig-sample and its reconstruction.

Sorting. For each pair we compute its SSIM [Wang et al., 2004], and sort the results by descending order.

Appendix C Supplementary Results

In this subsection we provide more details and experiments on each model presented in Figure 4. In Table 1 we show the train loss, test error and test loss of each model from Figure 4. All the models achieved a train accuracy of 100%100\%. We note that adding more training samples improves the test accuracy, while adding more layers keeps the test accuracy approximately the same. In Figures 6-11 we show the best 4545 extracted images (sorted by SSIM score) for the models presented in Figure 4. The reconstructions for the 5050 and 500500 samples with a dd-10001000-10001000-11 architecture is presented in Figure 1 and Figure 3 (top) respectively.

C.2 All Comparisons for Subsection 5.4

In this subsection we provide more detailed results on the comparison to other methods as presented in Subsection 5.4. In Figure 12 and Figure 13 we provide more results from the model inversion attack on models trained on CIFAR10 and MNIST respectively. These are the same models from Figure 5 (a). In this attack, we either minimize or maximize the model’s output w.r.t. a randomly initialized input. In this experiment, half of the initializations were maximized and the other half is minimized. The images are ordered by output of the model, in an increasing order. The results indicate that the model inversion attack mostly converge to similar reconstructions, even with many different initializations and different hyperparameters. Also, these reconstruction are mostly blurry, and probably represent the averages of each class.

In Figure 14 and Figure 15 we show all the weights, as images, of the first fully-connected layer of models trained on CIFAR10 and MNIST respectively. These are the same models as in Figure 3, i.e., there are 10001000 weights. Some weights are indicative of several input samples, e.g., a plane from CIFAR10 and the digits 88 and 55 from MNIST. We note that our reconstruction scheme is able to reconstruct much more samples, and in better quality than is represented in these weights.

C.3 Stretching the Theoretical Limitations

In this section we show results from several experiments which go beyond the theoretical limitations of Theorem 3.1.

In this subsection we consider networks trained with standard initialization scales. We recall that in the experiments presented in Section 5 the first fully-connected layer is initialized to a Gaussian distribution with mean and standard deviation 10−410^{-4}, while the other layers are initialized by standard Kaiming initialization [He et al., 2015]. In Figure 17 and Figure 18 we show reconstructions of a model trained on CIFAR10 on 10 and 50 samples respectively, where the all the layers of the model are initialized by standard Kaiming initialization. The architecture of the model is dd-10001000-10001000-11. We note that although the quality of the reconstructions is lower than when initializing the first layer with a small scale, there is still a strong signal that some of reconstructions correlate with training samples. It is an interesting future direction to improve the reconstruction quality for models with standard initialization.

In Figure 16 (a,b) we plot the SSIM score of each training sample against the output of the model. Note that indeed in these experiments the best SSIM score is lower than from other experiments presented in Figure 4. This corresponds to the lower quality of reconstructions when using standard initialization.

In the experiments from Section 5 we trained each model for 10610^{6} epochs. The reason for this long training time is that Theorem 3.1 gives guarantees only when converging to KKT point. Such a convergence happens only after training until infinity, and longer training time may converge closer to the KKT point. In this section we provide reconstruction results for models trained for only 10410^{4} epochs. Figure 19 and Figure 20 show reconstructions for models trained on 500500 samples from CIFAR10 and MNIST datasets respectively, with an architecture of dd-10001000-10001000-11. It is clear that the quality of the reconstruction is very similar to when training for more epochs, this may indicate that even after significantly less training epochs the model converge sufficiently close to a KKT point.

In Figure 16 (d,e) we plot the SSIM score of each training sample against the output of the model. We note that we are able to reconstruct samples which appear approximately on the margin for both MNIST and CIFAR. In addition, the model for MNIST did not achieve train error, and the margin is still very small. With that said, we are still able to reconstruct a large portion of the data with high quality. This goes beyond our theoretical limitations which have guarantees only for models which successfully label the entire training set.

In the experiments from Section 5 we trained the models using full-batch gradient descent. This was done to align with the theoretical guarantees of Theorem 3.1, which assume training with gradient flow. In Figure 21 we show reconstructions from a model trained with mini-batch SGD, using a batch size of 5050. The model is trained on 500500 images from CIFAR10, and with an architecture of dd-10001000-10001000-11.

In Figure 16 (c) we plot the SSIM score of each training sample against the output of the model. This plot shows that we indeed reconstruct samples that lie on the margin.