Dataset Meta-Learning from Kernel Ridge-Regression
Timothy Nguyen, Zhourong Chen, Jaehoon Lee
Introduction
Datasets are a pivotal component in any machine learning task. Typically, a machine learning problem regards a dataset as given and uses it to train a model according to some specific objective. In this work, we depart from the traditional paradigm by instead optimizing a dataset with respect to a learning objective, from which the resulting dataset can be used in a range of downstream learning tasks.
Our work is directly motivated by several challenges in existing learning methods. Kernel methods or instance-based learning (Vinyals et al., 2016; Snell et al., 2017; Kaya & Bilge, 2019) in general require a support dataset to be deployed at inference time. Achieving good prediction accuracy typically requires having a large support set, which inevitably increases both memory footprint and latency at inference time—the scalability issue. It can also raise privacy concerns when deploying a support set of original examples, e.g., distributing raw images to user devices. Additional challenges to scalability include, for instance, the desire for rapid hyper-parameter search (Shleifer & Prokop, 2019) and minimizing the resources consumed when replaying data for continual learning (Borsos et al., 2020). A valuable contribution to all these problems would be to find surrogate datasets that can mitigate the challenges which occur for naturally occurring datasets without a significant sacrifice in performance.
Question: What is the space of datasets, possibly with constraints in regards to size or signal preserved, whose trained models are all (approximately) equivalent to some specific model?
In attempting to answer this question, in the setting of supervised learning on image data, we discover a rich variety of datasets, diverse in size and human interpretability while also robust to model architectures, which yield high performance or state of the art (SOTA) results when used as training data. We obtain such datasets through the introduction of a novel meta-learning algorithm called Kernel Inducing Points (KIP ). Figure 1 shows some example images from our learned datasets.
We explore KIP in the context of compressing and corrupting datasets, validating its effectiveness in the setting of kernel-ridge regression (KRR) and neural network training on benchmark datasets MNIST and CIFAR-10. Our contributions can be summarized as follows:
We formulate a novel concept of -approximation of a dataset. This provides a theoretical framework for understanding dataset distillation and compression.
We introduce Kernel Inducing Points (KIP ), a meta-learning algorithm for obtaining -approximation of datasets. We establish convergence in the case of a linear kernel in Theorem 1. We also introduce a variant called Label Solve (LS ), which gives a closed-form solution for obtaining distilled datasets differing only via labels.
We explore the following aspects of -approximation of datasets:
Compression (Distillation) for Kernel Ridge-Regression: For kernel ridge regression, we improve sample efficiency by over one or two orders of magnitude, e.g. using 10 images to outperform hundreds or thousands of images (Tables 1, 2 vs Tables A1, A2). We obtain state of the art results for MNIST and CIFAR-10 classification while using few enough images (10K) to allow for in-memory inference (Tables A3, A4).
Compression (Distillation) for Neural Networks: We obtain state of the art dataset distillation results for the training of neural networks, often times even with only a single hidden layer fully-connected network (Tables 1 and 2).
Privacy: We obtain datasets with a strong trade-off between corruption and test accuracy, which suggests applications to privacy-preserving dataset creation. In particular, we produce images with up to 90% of their pixels corrupted with limited degradation in performance as measured by test accuracy in the appropriate regimes (Figures 3, A3, and Tables A5-A10) and which simultaneously outperform natural images, in a wide variety of settings.
We provide an open source implementation of KIP and LS , available in an interactive Colab notebookhttps://colab.research.google.com/github/google-research/google-research/blob/master/kip/KIP.ipynb.
Setup
In this section we define some key concepts for our methods.
Given a learning algorithm (e.g. gradient descent with respect to the loss function of a neural network), let denote the resulting model obtained after training on . We regard as a mapping from datapoints to prediction labels.
We provide some justification for this definition in the Appendix. In this paper, we will measure -approximation with respect to - loss for multiway classification (i.e. accuracy). We focus on weak -approximation, since in most of our experiments, we consider models in the low-data regime with large classification error rates, in which case, sample-wise agreement of two models is not of central importance. On the other hand, observe that if two models have population classification error rates less than , then (2) is automatically satisfied, in which case, the notions of weak-approximation and strong-approximation converge.
We are interested in learning algorithms given by KRR and neural networks. These can be investigated in unison via neural tangent kernels. Furthermore, we study two settings for the usage of -approximate datasets, though there are bound to be others:
(Privacy guarantee) Can an -approximate dataset be found such that the distribution from which it is drawn and the distribution from which the original training dataset is drawn satisfy a given upper bound in mutual information?
Motivated by these questions, we introduce the following definitions:
In other words, datasets produced by have fraction of its entries contain no information about the dataset (e.g. because they have a fixed value or are filled in randomly). Corrupting information is naturally a way of enhancing privacy, as it makes it more difficult for an attacker to obtain useful information about the data used to train a model. Adding noise to the inputs to neural network or of its gradient updates can be shown to provide differentially private guarantees (Abadi et al. (2016)).
Kernel Inducing Points
This leads to our first-order meta-learning algorithm KIP (Kernel Inducing Points), which uses kernel-ridge regression to learn -approximate datasets. It can be regarded as an adaption of the inducing point method for Gaussian processes (Snelson & Ghahramani, 2006) to the case of KRR. Given a kernel , the KRR loss function trained on a support dataset and evaluated on a target dataset is given by
where if and are sets, is the matrix of kernel elements . Here is a fixed regularization parameter. The KIP algorithm consists of optimizing (7) with respect to the support set (either just the or along with the labels ), see Algorithm 1. Depending on the downstream task, it can be helpful to use families of kernels (Step 3) because then KIP produces datasets that are -approximations for a variety of kernels instead of a single one. This leads to a corresponding robustness for the learned datasets when used for neural network training. We remark on best experimental practices for sampling methods and initializations for KIP in the Appendix. Theoretical analysis for the convergence properties of KIP for the case of a linear kernel is provided by Theorem 1. Sample KIP -learned images can be found in Section F.
KIP variations: i) We can also randomly augment the sampled target batches in KIP . This effectively enhances the target dataset , and we obtain improved results in this way, with no extra computational cost with respect to the support size. ii) We also can choose a corruption fraction and do the following. Initialize a random -percent of the coordinates of each support datapoint via some corruption scheme (zero out all such pixels or initialize with noise). Next, do not update such corrupted coordinates during the KIP training algorithm (i.e. we only perform gradient updates on the complementary set of coordinates). Call this resulting algorithm . In this way, is -corrupted according to Definition 5 and we use it to obtain our highly corrupted datasets.
Label solving: In addition to KIP , where we learn the support dataset via gradient descent, we propose another inducing point method, Label Solve (LS ), in which we directly find the minimum of (7) with respect to the support labels while holding fixed. This is simple because the loss function is quadratic in . We refer to the resulting labels
as solved labels. As is the pseudo-inverse operation, is the minimum-norm solution among minimizers of (7). If is injective, using the fact that for injective and surjective (Greville (1966)), we can rewrite (8) as
Experiments
We perform three sets of experiments to validate the efficacy of KIP and LS for dataset learning. The first set of experiments investigates optimizing KIP and LS for compressing datasets and achieving state of the art performance for individual kernels. The second set of experiments explores transferability of such learned datasets across different kernels. The third set of experiments investigate the transferability of KIP -learned datasets to training neural networks. The overall conclusion is that KIP -learned datasets, even highly corrupted versions, perform well in a wide variety of settings. Experimental details can be found in the Appendix.
We focus on MNIST (LeCun et al., 2010) and CIFAR-10 (Krizhevsky et al., 2009) datasets for comparison to previous methods. For LS , we also use Fashion-MNIST. These classification tasks are recast as regression problems by using mean-centered one-hot labels during training and by making class predictions via assigning the class index with maximal predicted value during testing. All our kernel-based experiments use the Neural Tangents library (Novak et al., 2020), built on top of JAX (Bradbury et al., 2018). In what follows, we use FC and Conv to denote a depth fully-connected or fully-convolutional network. Whether we mean a finite-width neural network or else the corresponding neural tangent kernel (NTK) will be understood from the context. We will sometimes also use the neural network Gaussian process (NNGP) kernel associated to a neural network in various places. By default, a neural kernel refers to NTK unless otherwise stated. RBF denotes the radial-basis function kernel. Myrtle- architecture follows that of Shankar et al. (2020), where an -layer neural network consisting of a simple combination of convolutional layers along with average pooling layers are inter-weaved to reduce internal patch-size.
We would have used deeper and more diverse architectures for KIP , but computational limits, which will be overcome in future work, placed restrictions, see the Experiment Details in Section D.
We apply KIP to learn support datasets of various sizes for MNIST and CIFAR-10. The objective is to distill the entire training dataset down to datasets of various fixed, smaller sizes to achieve high compression ratio. We present these results against various baselines in Tables 1 and 2. These comparisons occur cross-architecturally, but aside from Myrtle LS results, all our results involve the simplest of kernels (RBF or FC1), whereas prior art use deeper architectures (LeNet, AlexNet, ConvNet).
We obtain state of the art results for KRR on MNIST and CIFAR-10, for the RBF and FC1 kernels, both in terms of accuracy and number of images required, see Tables 1 and 2. In particular, our method produces datasets such that RBF and FC1 kernels fit to them rival the performance of deep convolutional neural networks on MNIST (exceeding 99.2%). By comparing Tables 2 and A2, we see that, e.g. 10 or 100 KIP images for RBF and FC1 perform on par with tens or hundreds times more natural images, resulting in a compression ratio of one or two orders of magnitude.
For neural network trainings, for CIFAR-10, the second group of rows in Table 2 shows that FC1 trained on KIP images outperform prior art, all of which have deeper, more expressive architectures. On MNIST, we still outperform some prior baselines with deeper architectures. This, along with the state of the art KRR results, suggests that KIP , when scaled up to deeper architectures, should continue to yield strong neural network performance.
For LS , we use a mix of NNGP kernelsFor FC1, NNGP and NTK perform comparably whereas for Myrtle, NNGP outperforms NTK. and NTK kernels associated to FC1, Myrtle-5, Myrtle-10 to learn labels on various subsets of MNIST, Fashion-MNIST, and CIFAR-10. Our results comprise the bottom third of Tables 1 and 2 and Figure 2. As Figure 2 shows, the more targets are used, the better the performance. When all possible targets are used, we get an optimal compression ratio of roughly one order of magnitude at intermediate support sizes.
2 Kernel to Kernel Results
Here we investigate robustness of KIP and LS learned datasets when there is variation in the kernels used for training and testing. We draw kernels coming from FC and Conv layers of depths 1-3, since such components form the basic building blocks of neural networks. Figure A1 shows that KIP -datasets trained with random sampling of all six kernels do better on average than KIP -datasets trained using individual kernels.
For LS , transferability between FC1 and Myrtle-10 kernels on CIFAR-10 is highly robust, see Figure A2. Namely, one can label solve using FC1 and train Myrtle-10 using those labels and vice versa. There is only a negligible difference in performance in nearly all instances between data with transferred learned labels and with natural labels.
3 Kernel to Neural Networks Results
Significantly, KIP -learned datasets, even with heavy corruption, transfer remarkably well to the training of neural networks. Here, corruption refers to setting a random fraction of the pixels of each image to uniform noise between and (for KIP , this is implemented via )Our images are preprocessed so as to be mean-centered and unit-variance per pixel. This choice of corruption, which occurs post-processing, is therefore meant to (approximately) match the natural pixel distribution.. The deterioriation in test accuracy for KIP -images is limited as a function of the corruption fraction, especially when compared to natural images, and moreover, corrupted KIP -images typically outperform uncorrupted natural images. We verify these conclusions along the following dimensions:
Robustness to dataset size: We perform two sets of experiments.
(i) First, we consider small KIP datasets (10, 100, 200 images) optimized using multiple kernels (FC1-3, Conv1-2), see Tables A5, A6. We find that our in-distribution transfer (the downstream neural network has its neural kernel included among the kernels sampled by KIP ) performs remarkably well, with both uncorrupted and corrupted KIP images beating the uncorrupted natural images of corresponding size. Out of distribution networks (LeNet (LeCun et al., 1998) and Wide Resnet (Zagoruyko & Komodakis, 2016)) have less transferability: the uncorrupted images still outperform natural images, and corrupted KIP images still outperform corrupted natural images, but corrupted KIP images no longer outperform uncorrupted natural images.
(ii) We consider larger KIP datasets (1K, 5K, 10K images) optimized using a single FC1 kernel for training of a corresponding FC1 neural network, where the KIP training uses augmentations (with and without label learning), see Tables A7-A10 and Figure A3. We find, as before, KIP images outperform natural images by an impressive margin: for instance, on CIFAR-10, 10K KIP -learned images with 90% corruption achieves 49.9% test accuracy, exceeding 10K natural images with no corruption (acc: 45.5%) and 90% corruption (acc: 33.8%). Interestingly enough, sometimes higher corruption leads to better test performance (this occurs for CIFAR-10 with cross entropy loss for both natural and KIP -learned images), a phenomenon to be explored in future work. We also find that KIP with label-learning often tends to harm performance, perhaps because the labels are overfitting to KRR.
Robustness to hyperparameters: For CIFAR-10, we took 100 images, both clean and 90% corrupted, and trained networks on a wide variety of hyperparameters for various neural architectures. We considered both neural networks whose corresponding neural kernels were sampled during KIP -training those that were not. We found that in both cases, the KIP -learned images almost always outperform 100 random natural images, with the optimal set of hyperparameters yielding a margin close to that predicted from the KRR setting, see Figure 3. This suggests that KIP -learned images can be useful in accelerating hyperparameter search.
Related Work
Coresets: A classical approach for compressing datasets is via subset selection, or some approximation thereof. One notable work is Borsos et al. (2020), utilizing KRR for dataset subselection. For an overview of notions of coresets based on pointwise approximatation of datasets, see Phillips (2016).
Neural network approaches to dataset distillation: Maclaurin et al. (2015); Lorraine et al. (2020) approach dataset distillation through learning the input images from large-scale gradient-based meta-learning of hyper-parameters. Properties of distilled input data was first analyzed in Wang et al. (2018). The works Sucholutsky & Schonlau (2019); Bohdal et al. (2020) build upon Wang et al. (2018) by distilling labels. More recently, Zhao et al. (2020) proposes condensing training set by gradient matching condition and shows improvement over Wang et al. (2018).
Inducing points: Our approach has as antecedant the inducing point method for Gaussian Processes (Snelson & Ghahramani, 2006; Titsias, 2009). However, whereas the latter requires a probabilistic framework that optimizes for marginal likelihood, in our method we only need to consider minimizing mean-square loss on validation data.
Low-rank kernel approximations: Unlike common low-rank approximation methods (Williams & Seeger, 2001; Drineas & Mahoney, 2005), we obtain not only a low-rank support-support kernel matrix with KIP , but also a low-rank target-support kernel matrix. Note that the resulting matrices obtained from KIP need not approximate the original support-support or target-support matrices since KIP only optimizes for the loss function.
Neural network kernels: Our work is motivated by the exact correspondence between infinitely-wide neural networks and kernel methods (Neal, 1994; Lee et al., 2018; Matthews et al., 2018; Jacot et al., 2018; Novak et al., 2019; Garriga-Alonso et al., 2019; Arora et al., 2019a). These correspondences allow us to view both Bayesian inference and gradient descent training of wide neural networks with squared loss as yielding a Gaussian process or kernel ridge regression with neural kernels.
Instance-Based Encryption: A related approach to corrupting datasets involves encrypting individual images via sign corruption (Huang et al. (2020)).
Conclusion
We introduced novel algorithms KIP and LS for the meta-learning of datasets. We obtained a variety of compressed and corrupted datasets, achieving state of the art results for KRR and neural network dataset distillation methods. This was achieved even using the simplest of kernels and neural networks (shallow fully-connected networks and purely-convolutional networks without pooling), which notwithstanding their limited expressiveness, outperform most baselines that use deeper architectures. Follow-up work will involve scaling up KIP to deeper architectures with pooling (achievable with multi-device training) for which we expect to obtain even more highly performant datasets, both in terms of overall accuracy and architectural flexibility. Finally, we obtained highly corrupt datasets whose performance match or exceed natural images, which when developed at scale, could lead to practical applications for privacy-preserving machine learning.
We would like to thank Dumitru Erhan, Yang Li, Hossein Mobahi, Jeffrey Pennington, Si Si, Jascha Sohl-Dickstein, and Lechao Xiao for helpful discussions and references.
References
Appendix A Remarks on Definition of ϵitalic-ϵ{\epsilon}-approximation
Another key feature of our definition is that datapoints of an -approximating dataset must have the same shape as those of the original dataset. This makes our notion of an -approximate dataset more restrictive than returning a specialized set of extracted features from some initial dataset.
Analogues of our -approximation definition have been formulated in the unsupervised setting, e.g. in the setting of clustering data (Phillips, 2016; Jubran et al., 2019).
Appendix B Tuning KIP
For sampling from the target set, which we always do in a class-balanced way, we found larger batch sizes typically perform better on the test set if the train and test kernels agree. If the train and test kernels differ, then smaller batch sizes lead to less overfitting to the train kernel.
Initialization: We tried two sets of initializations. The first (“image init”) initializes to be a subset of . The second (“noise init”) initializes with uniform noise and with mean-centered, one-hot labels (in a class-balanced way). We found image initialization to perform better.
Number of Training Iterations: Remarkably, KIP converges very quickly in all experimental settings we tried. After only on the order of a hundred iterations, independently of the support size, kernel, and corruption factor, the learned support set has already undergone the majority of its learning (test accuracy is within more than 90% of the final test accuracy). For the platforms available to us, using a single V100 GPU, one hundred training steps for the experiments we ran involving target batch sizes that were a few thousand takes on the order of about 10 minutes. When we add augmentations to our targets, performance continues to improve slowly over time before flattening out after several thousands of iterations.
Appendix C Theoretical Results
For the case of a linear kernel, we prove the below convergence theorem:
We discuss the case where is optimized, with the case where both are optimized proceeding similarly. In this case, by genericity, we can assume , else the learning dynamics is trivial. Furthermore, to simplify notation for the time being, assume the dimensionality of the label space is without loss of generality. First, we establish convergence. For a linear kernel, we can write our loss function as
Next, we claim that given a fixed initial , then for sufficiently small , gradient-flow of (A2) starting from cannot converge to a non-global local minima. We proceed as follows. If is a singular value decomposition of , with a diagional matrix of singular values (and any additional zeros for padding), then where denotes the diagonal matrix with the map
We also have the following result about -approximation using the label solve algorithm:
Then yields a strong -approximation of with respect to algorithms (-RR, -RR) and mean-square loss, where
By definition, is the minimizer of
For general kernels, we make the following simple observation concerning the optimal output of KIP .
Appendix D Experiment Details
In all KIP trainings, we used the Adam optimizer. All our labels are mean-centered -hot labels. We used learning rates and for the MNIST and CIFAR-10 datasets, respectively. When sampling target batches, we always do so in a class-balanced way. When augmenting data, we used the ImageGenerator class from Keras, which enables us to add horizontal flips, height/width shift, rotatations (up to 10 degrees), and channel shift (for CIFAR-10). All datasets are preprocessed using channel-wise standardization (i.e. mean subtraction and division by standard-deviation). For neural (tangent) kernels, we always use weight and bias variance and , respectively. For both neural kernels and neural networks, we always use ReLU activation. Convolutional layers all use a filter with stride 1 and same padding.
Compute Limitations: Our neural kernel computations, implemented using Neural Tangents libraray (Novak et al., 2020) are such that computation scales (i) linearly with depth; (ii) quadratically in the number of pixels for convolutional kernels; (iii) quartically in the number of pixels for pooling layers. Such costs mean that, using a single V100 GPU with 16GB of RAM, we were (i) only able to sample shallow kernels; (ii) for convolutional kernels, limited to small support sets and small target batch sizes; (iii) unable to use pooling if learning more than just a few images. Scaling up KIP to deeper, more expensive architectures, achievable using multi-device training, will be the subject of future exploration.
Kernel Parameterization: Neural tangent kernels, or more precisely each neural network layer of such kernels, as implemented in Novak et al. (2020) can be parameterized in either the “NTK” parameterization or “standard” parameterization Sohl-Dickstein et al. (2020). The latter depends on the width of a corresponding finite-width neural network while the former does not. Our experiments mix both these parameterizations for variety. However, because we use a scale-invariant regularization for KRR (Section B), the choice of parameterization has a limited effect compared to other more significant hyperparameters (e.g. the support dataset size, learning rate, etc.) All our final readout layers use the fixed NTK parameterization and all our statements about which parameterization we are using should be interpreted accordingly. This has no effect on the training of our neural networks while for kernel results, this affects the recursive formula for the NTK at the final layer if using standard parameterization (by the changing the relative scales of the terms involved). Since the train/test kernels are consistently parameterized and KIP can adapt to the scale of the kernel, the difference between our hybrid parameterization and fully standard parameterization has a limited affect..
Single kernel results: (Tables 1 and 2) For FC, we used kernels with NTK parametrization. For RBF, our rbf kernel is given by
where is the dimension of the inputs and . We found that treating as a learnable parameter during KIP had mixed resultsOn MNIST it led to very slight improvement. For CIFAR10, for small support sets, the effect was a small improvement on the test set, whereas for large support sets, we got worse performance. and so keep it fixed for simplicity.
For MNIST, we found target batch size equal to 6K sufficient. For CIFAR-10, it helped to sample the entire training dataset of 50K images per step (hence, along with sampling the full support set, we are doing full gradient descent training). When support dataset size is small or if augmentations are employed, there is no overfitting (i.e. the train and test loss/accuracy stay positively correlated). If the support dataset size is large (5K or larger), sometimes there is overfitting when the target batch size is too large (e.g. for the RBF kernel on CIFAR10, which is why we exclude in Table 2 the entries for 5K and 10K). We could have used a validation dataset for a stopping criterion, but that would have required reducing the target dataset from the entire training dataset.
We train KIP for 10-20k iterations and took 5 random subsets of images for initializations. For each such training, we took 5 checkpoints with lowest train and loss and computed the test accuracy. This gives 25 evaluations, for which we can compute the mean and standard deviation for our test accuracy numbers in Tables 1 and 2.
Kernel transfer results: For transfering of KIP images, both to other kernels and to neural networks, we found it useful to use smaller target batch sizes (either several hundred or several thousand), else the images overfit to their source kernel. For random sampling of kernels used in Figure A1 and producing datasets for training of neural networks, we used FC kernels with width 1024 and Conv kernels with width 128, all with standard parametrization.
Neural network results: Neural network trainings on natural data with mean-square loss use mean-centered one-hot labels for consistency with KIP trainings. For cross entropy loss, we use one-hot labels. For neural network trainings on KIP -learned images with label learning, we transfer over the labels directly (as with the images), whatever they may be.
For neural network transfer experiments occurring in Table 1, Table 2, Figure 3, Table A5, and Table A6, we did the following. First, the images were learned using kernels FC1-3, Conv1-2. Second, we trained for a few hundred iterations, after which optimal test performance was achieved. On MNIST images, we trained the networks with constant learning rate and Adam optimizer with cross entropy loss. Learning rate was tuned over small grid search space. For the FC kernels and networks, we use width of 1024. On CIFAR-10 images, we trained the networks with constant learning rate, momentum optimizer with momentum 0.9. Learning rate, L2 regularization, parameterization (standard vs NTK) and loss type (mean square, softmax-cross-entropy) was tuned over small grid search space. Vanilla networks use constant width at each layer: for FC we use width of 1024, for Conv2 we use 512 channels, and for Conv8 we use 128 channels. No pooling layers are used except for the WideResNet architecture, where we follow the original architecture of Zagoruyko & Komodakis (2016) except that our batch normalization layer is stateless (i.e. no exponential moving average of batch statistics are recorded).
For neural network transfer experiments in Figure A3, Tables A7-A10, we did the following. Our KIP -learned images were trained using only an FC1 kernel. The neural network FC1 has an increased width 4096, which helps with the larger number of images. We used learning rate and the Adam optimizer. The KIP learned images with only augmentations used target batch size equal to half the training dataset size and were trained for 10k iterations, since the use of augmentations allows for continued gains after longer training. The KIP learned images with augmentations and label learning used target batch size equal to a tenth of the training dataset size and were trained for 2k iterations (the learned data were observed to overfit to the kernel and have less transferability if larger batch size were used or if trainings were carried out longer).
All neural network trainings were run with 5 random initializations to compute mean and standard deviation of test accuracies.
In Table 2, regularized ZCA preprocessing was used for a Myrtle-10 kernel (denoted with ZCA) on CIFAR-10 dataset. Shankar et al. (2020) and Lee et al. (2020) noticed that for neural (convolutional) kernels on image classification tasks, regularized ZCA preprocessing can improve performance significantly compared to standard preprocessing. We follow the prepossessing scheme used in Shankar et al. (2020), with regularization strength of without augmentation.
Appendix E Tables and Figures
We report various baselines of KRR trained on natural images. Tables A1 and A2 shows how various kernels vary in performance with respect to random subsets of MNIST and CIFAR-10. Linear denotes a linear kernel, RBF denotes the rbf kernel (A9) with , and FC1 uses standard parametrization and width 1024. Interestingly enough, we observe non-monotonicity for the linear kernel, owing to double descent phenomenon Hastie et al. (2019). We include additional columns for deeper kernel architectures in Table A2, taken from Shankar et al. (2020) for reference.
Comparing Tables 1, 2 with Tables A1, A2, we see that 10 KIP -learned images, for both RBF and FC1, has comparable performance to several thousand natural images, thereby achieving a compression ratio of over 100. This compression ratio narrows as the support size increases towards the size of the training data.
Next, Table A3 compares FC1, RBF, and other kernels trained on all of MNIST to FC1 and RBF trained on KIP -learned images. We see that our KIP approach, even with 10K images (which fits into memory), leads to RBF and FC1 matching the performance of convolutional kernels on the original 60K images. Table A4 shows state of the art of FC kernels on CIFAR-10. The prior state of the art used kernel ensembling on batches of augmented data in Lee et al. (2020) to obtain test accuracy of 61.5% (32 ensembles each of size 45K images). By distilling augmented images using KIP , we are able to obtain 64.7% test accuracy using only 10K images.
E.2 KIP and LS Transfer Across Kernels
Figure A1 plots how KIP (with only images learned) performs across kernels. There are seven training scenarios: training individually on FC1, FC2, FC3, Conv1, Conv2, Conv3 NTK kernels and random sampling from among all six kernels uniformly (Avg All). Datasets of size 10, 100, 200 are thereby trained then evaluated by averaging over all of FC1-3, Conv1-3, both with the NTK and NNGP kernels for good measure. Moreover, the FC and Conv train kernel widths (1024 and 128) were swapped at test time (FC width 128 and Conv width 1024), as an additional test of robustness. The average performance is recorded along the y-axis. AvgAll leads to overall boost in performance across kernels. Another observation is that Conv kernels alone tend to do a bit better, averaged over the kernels considered, than FC kernels alone.
In Figure A2, we plot how LS learned labels using Myrtle-10 kernel transfer to the FC1 kernel and vice versa. We vary the number of targets and support size. We find remarkable stability across all these dimensions in the sense that while the gains from LS may be kernel-specific, LS -labels do not perform meaningfully different from natural labels when switching the train and evaluation kernels.