Predicting Neural Network Accuracy from Weights

Thomas Unterthiner, Daniel Keysers, Sylvain Gelly, Olivier Bousquet, Ilya Tolstikhin

Introduction

Deep neural networks (DNNs) are considered state of the art methods for many machine learning problems today. Yet, a deeper understanding of the mechanisms underlying these successes is still lacking. The deep learning phenomena, i.e. various surprising and insightful empirical findings surrounding the efforts to understand DNN training and generalization have recently gained a lot of attention from researchers and practitioners (Zhang et al. 2017; Frankle & Carbin 2019; Zhang et al. 2019). Research in this direction is actively growing, yet many such phenomena remain to be discovered.

This paper discusses the prediction of the accuracy of trained neural networks, using only their weights as inputs. Specifically, we consider convolutional neural networks (CNNs) trained on standard datasets for the popular task of image classification. We see this study as a step towards gaining a deeper understanding of neural network training and performance. Understanding what can be said by looking at the trained weights can be useful in understanding the training process in general. It can also have practical applications such as early stopping of unsuccessful training runs (Domhan et al. 2015).

As a first step in this direction we study CNNs trained in the under-parameterized regime, in which the observed train and test accuracies do not differ substantially. Then we show that our findings appear to transfer to the over-parameterized regime (Belkin et al. 2018). We demonstrate (Section 5.2) that the predictor trained on a collection of very small CNNs is capable of ranking large ResNet models according to train/test accuracy fairly well by looking only at the ResNet’s weights.

The studies presented in this paper may raise more questions than they answer, but we hope that this will serve as starting point for other researchers to make progress in understanding deep learning phenomena. The main contributions of this paper are:

We propose a new formal setting that captures the approach and relates to previous works.

We introduce a new, large dataset with strong baselines and discuss extensive empirical results. The data is of a new modality, mapping trained weights of neural networks to their accuracy.

The experiments show that, somewhat surprisingly, it is possible to predict the accuracy using trained weights alone. Furthermore, only few statistics of the weights are sufficient for high accuracy in prediction.

Experiments on transfer of prediction across architectures and datasets show that it is possible to rank neural network models trained on an unknown dataset just by observing the trained weights, without ever having access to the dataset itself.

Next, we describe a formal setting that considers this and related tasks (Section 2) and discuss related work (Section 3). We introduce a new dataset for this task and present empirical results on our dataset (Section 4). We also discuss the performance of the resulting predictors under domain shift (Section 5).

Formal setting

We will train CNNs on SNS_{N} using hyperparameters λ\lambda and get a particular weight vector W=A(SN,λ)W=\mathcal{A}(S_{N},\lambda), where A\mathcal{A} denotes the learning procedure and WW may be considered a flattened vector containing all the weights. The hyperparameters λ\lambda include architecture-specific details (e.g. number of layers and activation function), optimizer-specific details (e.g. learning rate and initialization variance), and other parameters (e.g. weight regularization and fraction of the training set to use). Notice that the training method A\mathcal{A} may have internal sources of stochasticity, including order of examples in mini-batches or weight initialization. Also note that depending on λ\lambda, the weight vector WW may be of a variable dimension (e.g. for varying number of layers).

2 Domain shift

Rather than discovering properties of DNNs that are specific to a particular dataset or architecture (which nevertheless could be interesting on its own), we are even more interested in those that hold across various datasets and architectures. In that sense, domain shift provides a setting close to what we actually are interested in: observing any sort of positive transfer between different datasets and architectures would indicate that there are properties of DNNs that transfer. Our goal is to demonstrate the existence of these invariant properties and study them.

Related work

There are only few works that consider the problem setting described in Section 2. The most relevant of these are (Jiang et al. 2019; Yak et al. 2019; Eilertsen et al. 2020; Martin & Mahoney 2020; Martin et al. 2020).

The overall setting and motivations of Eilertsen et al. 2020 are similar to ours. However, the main difference is that instead of predicting the accuracy, the authors focus on predicting the hyperparameters λ\lambda using the weights WW (the opposite direction of the black arrow in Figure 1).

Concurrent works from Martin & Mahoney 2020; Martin et al. 2020 confirm our findings by showing that more complex statistics derived from weight matrices (Martin & Mahoney 2018) correlate well with the performance of state-of-the-art models in vision and language processing.

Jiang et al. 2019 and Yak et al. 2019 both investigate how to predict the generalization gap, i.e. the difference between training and test set performance, of a neural network based on the hidden activations of training set examples. Jiang et al. 2019 train large CNN/ResNet architectures on CIFAR datasets and approximate the minimal distances to the class boundary for each data point in each hidden layer. They use this margin distribution to train a linear regressor that predicts generalization gaps. Yak et al. 2019 expand upon this work by training a large number of small fully-connected networks on different variations of a generated spiral dataset. They replace the linear predictor with a recurrent neural network to handle varying neural network depth, and show that predictions transfer between small fully-connected architectures and varying synthetic datasets. Both works heavily rely on the margins in the intermediate layers of the networks. These margins can not be computed analytically and require a computationally expensive approximation procedure (Elsayed et al. 2018), which is not guaranteed to be accurate. Margin approximation also involves an inference pass over the training set SNS_{N}. Our estimators F^\hat{F} use only weights of the networks (or their simple statistics) to predict the accuracy. As the weights are (one of) the most important characteristic of a trained DNN, it is interesting to study this connection without requiring information about the training set SNS_{N}. We show that these estimators transfer to networks trained on unobserved natural image datasets and with ResNet32 architectures. Finally, experimental design utilized in these previous works may lead to an undesirable leakage, as discussed in Section 4.1.

DeChant et al. 2019 train ResNets and other large architectures on CIFAR and ImageNet datasets. They demonstrate that it is possible to tell whether or not the network will make a mistake on one particular image by looking at the activations of that image in the network’s layers.

The relation between the train and test accuracies is the central question of statistical learning theory (Vapnik 1998; Shalev-Shwartz & Ben-David 2014). Jiang et al. 2020 recently performed a large scale empirical study analyzing correlation between various generalization error bounds and network performance.

A problem somewhat similar to ours has been studied in the context of hyperparameter optimization and neural architecture search (NAS). (Streeter 2019a; Streeter 2019b) propose procedures that select good hyperparameter values based on previous exploration. To apply early stopping to unsuccessful runs, Swersky et al. 2014 and Domhan et al. 2015 predict the final performance of a neural network based on few training iterations. Similar techniques were applied in NAS to select candidate architectures, where the prediction is usually based on hyperparameters, architectures, information about the dataset, and performance measurements of similar architectures, see (Baker et al. 2017; Istrate et al. 2019) and references therein.

Experiments: Small CNN Zoo

Results reported in this section are based on a new dataset which we call the Small CNN Zoo Made publicly available together with the code reproducing the experiments at https://github.com/google-research/google-research/tree/master/dnn_predict_accuracy. It contains weights of a fixed CNN architecture trained on 4 different image datasets using a large number of different hyperparameter configurations. For each network, accuracy and cross-entropy loss on the train and test data are available.

To enable predicting accuracy from the flattened weight vector, we keep the number of weights in the architecture small: 3 convolutional layers with 16 filters each, followed by global average pooling and a fully connected layer, for a total of 4 970 learnable weights. As a result, the best test accuracies we obtain on CIFAR10 and SVHN are 56% and 78%, respectively, which is far below state of the art. However, it is worth pointing out that the smallest CNN architectures achieving above 90% test accuracy on CIFAR10 that we are aware of require on the order of 10610^{6} parameters (Lin et al. 2014; Springenberg et al. 2015), i.e. 200x more, and work on RGB inputs, while we ignore away color information.

We train on 4 natural image classification problems: MNIST (LeCun et al. 2010), Fashion MNIST (Xiao et al. 2017), grayscale CIFAR10 (CIFAR10-GS) (Krizhevsky 2009), and grayscale SVHN (SVHN-GS) (Netzer et al. 2011). Global average pooling and using grayscale allows us to apply the same architecture across all four datasets.

Instead of stopping training when networks converge or reach a certain level of accuracy, we train each CNN for 86 epochs. We do so because we want to study CNNs under general conditions: properties discovered by only looking at converged models may not hold for intermediate steps.

The distribution of the CNN models with respect to their test/train accuracy is reported in Figure 2. MNIST, Fashion MNIST, and CIFAR10-GS all have balanced classes and the histograms peak at around 10%—the accuracy of a random or constant prediction. SVHN-GS is unbalanced with the largest class containing around 19% of the samples. Here many models seem to converge to the constant majority class prediction, which explains the shifted peak.

We do not observe overfitting in the Small CNN Zoo dataset, even though some of the models were trained only on 10% of the training examples. Likely, this is due to the small architecture used. Based on this dataset we may gain insights on why and how neural networks train, but it is less likely that the dataset will directly lead to deeper understanding generalization.

Why not use multiple seeds? We use one random seed per hyperparameter configuration. This avoids having models that are too similar between the train and test splits of the CNN collections, which possibly leads to a leakage. In fact, this may point to a possible shortcut taking place in the studies of Jiang et al. 2019. The authors used 3 random seeds per hyperparameter configuration and did not enforce that the models trained with the same hyperparameters (but different random seeds) were allocated to the same split. Inspecting their dataset closer shows that the variance in generalization gap between the networks that only differ in random seed is orders of magnitude smaller than the average variance between all networks (10−510^{-5} vs 10−310^{-3}). A similar shortcut may take place in the studies of Yak et al. 2019. Here, the authors did not use the same hyperparameters with different seeds for training, but they trained networks with the same hyperparameters on versions of the synthetic datasets that were generated using different seeds.

2 Training the estimators

Types of estimators We explore three different estimators: logit-linear model (L-Linear), gradient boosting machine using regression trees (GBM), and a fully-connected DNN. All three methods were trained to minimize MSE. Each of these 3 methods comes with its own hyperparameters and initial experiments showed that it is important to tune them.

Training protocol and metrics Each of the 4 CNN collections is divided into two splits: 15k CNNs are used for the training split and the remaining ones were held out for the test split. The entire training and hyperparameter selection for the models took place on the training splits. The test splits are used only once to evaluate the single best model that we chose based on the 3-fold cross-validation performed on the training split.

We performed hyperparameter selection by evaluating 1k unique hyperparameter configurations sampled randomly and independently from pre-specified ranges for every combination of estimator type, input features, and CNN collection.

In all experiments we use MSE as the training objective. We also compute the mean absolute deviation and the coefficient of determination or R2R^{2} score. The R2R^{2} score normalizes the MSE of the estimator F^\hat{F} by the MSE of the best constant prediction. Larger R2R^{2} scores correspond to better predictions and the score never exceeds 1. For further details on the Small CNN Zoo dataset and the experimental setup we refer to Supplementary A.

3 Empirical results

In the experiments, GBM and DNN models always produce significantly better results than the logit-linear model. In some cases, the DNN model achieves slightly better results than GBM, but overall it is on par or significantly worse than GBM. These conclusions hold across all 4 datasets and the corresponding results are shown for one of the datasets (CIFAR10-GS) and a selection of input features in Table 1. In the interest of space we therefore only report the results for GBM in the following. All numbers for other models can be found in Supplementary A.5.

Table 2 presents the results of training the GBM models with different input features on the 4 CNN collections.

Using flattened weights First, we notice that a naive baseline of using the entire flattened vector WW already achieves a rather strong performance across all 4 datasets. Interestingly, almost the same performance can be recovered just by using the parameters of the last dense layer W4W^{4}, while using any other (convolutional) layer alone results in a noticeably worse performance. This observation is consistent with feature importance measurements produced by the GBM model (Supplementary B), which indicate that parameters of the last dense layer were among the most informative and frequently used ones.

Notably, Eilertsen et al. 2020 also report a strong performance of the per-layer statistics in their work.

Interpreting the R2R^{2} score and MSE values MSE provides an absolute measure of the model performance and on its own does not tell us much about the model: the value of 10−410^{-4} can correspond to a good and bad performance depending on the problem. The R2R^{2} score is a relative measure: it compares the MSE of the model to the MSE of a constant prediction. Moreover, R2R^{2} score is scale invariant and multiplying the outputs by a constant won’t change the metric. In Table 2 we use the R2R^{2} scores because we find them slightly easier to interpret: a non-positive value indicate that we are not doing better than fitting a constant predictor and values close to 1 point at stronger performance. The MSE values are reported in the Supplementary A.5 and scatter plots with raw predictions and true targets can be found in Supplementary D.

4 Ablation studies

Table 2 (upper block) shows that the parameter vector of a trained CNN alone contains a strong signal regarding the network’s accuracy. To understand more about the nature of this signal, we performed additional studies.

We tried several other input features for the estimators, including (i) the hyperparameter configuration λ\lambda (containing 7 parameters) used while training the CNN, (ii) the concatenation (λ,W)(\lambda,W) of the hyperparameters λ\lambda with the entire weight vector WW, and (iii) the weight statistics similar to W~ ⁣L\widetilde{W}_{\!L} computed only for a subset of the layers: W~ ⁣L4\widetilde{W}_{\!L}^{4} for the final dense layer and W~ ⁣L1,4\widetilde{W}_{\!L}^{1,4} for the combination of the first convolutional and the final dense layers. The results are reported in Table 2 (second block).

Statistics for subsets of layers Motivated by the fact that using the weights of the last dense layer W4W^{4} is as good as using the whole weight vector WW we tested whether statistics for a subset of the layers is enough to recover the performance based on W~ ⁣L\widetilde{W}_{\!L}. Curiously, the statistics of the last dense layer W~ ⁣L4\widetilde{{W}}_{\!L}^{4} perform worse. The results improve if we add the statistics of the first convolutional layer W~ ⁣L1,4\widetilde{{W}}_{\!L}^{1,4}, but they are still slightly worse than with all layers.

Permutation and scale invariance We also examined how the estimator’s predictions change as we modify its inputs. Notice that two ReLU CNNs with parameters WW and c⋅Wc\cdot W have exactly the same test/train accuracy (but not the same cross-entropy loss) for any real value c>0c>0, because their outputs h(X;W)h(X;W) and h(X;c⋅W)h(X;c\cdot W) coincide for all inputs XX. The same is true for any CNN if we permute the order of filters/channels consistently across all layers. We want to emphasize that we did not incorporate these inductive biases in any of the estimators we trained. Nevertheless, it may be interesting to test whether these (or similar) invariances emerge naturally in the trained estimators.

For a given estimator F^\hat{F} trained with the entire weight vectors WW, we tested several ways of modifying its inputs W↦φ(W)W\mapsto\varphi(W), including multiplying it with various positive factors and permuting it in several different ways. Then we looked at the absolute difference ∣F^(φ(W))−F^(W)∣|\hat{F}\bigl(\varphi(W)\bigr)-\hat{F}(W)| across multiple CNNs WW (from the test split of the same CNN collection F^\hat{F} was trained on) and various types of modifications φ\varphi. We report a short summary of this study here. Details can be found in Supplementary C.

The Mean Absolute Deviation (MAD) of modifications φ\varphi that we tried spanned a range between 0.010.01 and 0.130.13. Scaling the weights φ(W)=c⋅W\varphi(W)=c\cdot W with c∈{2,10,100}c\in\{2,10,100\} or permuting parameters within each of the first 3 convolutional layers leads to MADs less than 0.050.05. The estimator is more sensitive to permutations within the final dense layer, which leads to a MAD of 0.06. Global permutation of the entire vector WW (without preserving the layers) or scaling with small constants c∈{10−1,10−3}c\in\{10^{-1},10^{-3}\} all lead to a MAD larger than 0.110.11. Summarizing, the estimator is not too sensitive to the order of parameters in the convolutional layers, and much more sensitive to permutations within the final dense layer. The estimator is invariant to scaling the weights with positive factors larger than 1 and changes its predictions significantly for factors smaller than 1.

5 Understanding observed behaviors: first steps

We notice that networks trained on CIFAR10-GS with SGD (Figure 3, right) perform significantly worse than those trained with Adam/RMSProp (Figure 3, left). It is also interesting to note how the majority of the networks trained with SGD align along a line in this 2D space. Among the networks trained with Adam/RMSProp, we observe two well-separated groups: the strongly performing ones in the upper-right corner (also depicted in the zoomed-in subplot) and the ones with near-chance performance (the blue “tentacles” in the bottom part). Further analysis reveals that these two groups can be perfectly separated from each other by looking at the bias maxima in the final dense layer (not shown in the plots): the bias maxima are below 0.1 for the badly performing models (the “tentacles”) and above 0.1 for all the rest of the networks. In future work we would like to understand better what causes these “symptoms” during training and investigate ways to alleviate them.

Transfer to new architectures and datasets

In the previous section we showed using the Small CNN Zoo dataset that strong predictors of accuracy based on weights exist. Next we want to explore the domain shift setting introduced in Section 2.2 and study whether the predictors can handle networks trained on unobserved datasets or with different architectures. We emphasize that throughout this section the models were not fine-tuned or adjusted to the new collections in any way.

First we look at how the GBM models transfer across the CNN collections. Two examples of such experiments are shown in Figure 4. The figures demonstrate that the MSE of the predictions may not be the best metric to look at. The “drift” of points away from the diagonal line (which corresponds to zero MSE) is likely due to the difference in average accuracy between various datasets. Most of the networks in the MNIST collection achieve an accuracy higher than 60%, while the best accuracy for CIFAR10-GS was 55%. Nevertheless, we see that networks with higher accuracy tend to receive higher prediction values. In other words, the predictors are doing a reasonable job in ranking the networks. We can use Kendall’s τ\tau rank correlation coefficient to measure the quality of ranking. It ranges from -1 (anti-ranking) to 1 (perfect ranking) and takes values around 0 for random ranking.

Table 3 contains the values of Kendall’s τ\tau coefficient for all possible transfer experiments performed on the Small CNN Zoo (and 2D plots similar to Figure 4 are reported in Supplementary D). The smallest coefficient of 0.6 corresponds to the transfer from SVHN-GS to the MNIST collection. It is perhaps surprising that the rank test shows such a large correlation. We want to highlight that when training CNNs on the 4 datasets we only scale the pixel values to the $$ interval and do not perform any other standardization. We would expect the moments of the pixel values for the MNIST dataset to be very different from those of SVHN-GS and this difference in distributions to affect the form of the filters in the convolutional layers.

2 Networks trained with different architecture

In this section we want to test whether predictors trained on the Small CNN Zoo can rank larger, over-parametrized networks, capable of overfitting. For this purpose we will use the DEMOGEN collection (Jiang et al. 2019), which contains 216 Wide-ResNet32 models (He et al. 2016) trained on the original (colored) CIFAR10 dataset with the best models achieving 100% training and 93% test accuracy. The collection contains 72 networks for each of the three different architectures: ResNet32x1, ResNet32x2, and ResNet32x4, which differ in the number of filters.

Table 4 reports the τ\tau coefficients demonstrating how well predictions of the GBM model trained on CIFAR10-GS CNN collection correlate with actual accuracies of the networks from DEMOGEN. We compare to both train and test accuracies, because, as discussed in Section 4.1, for the Small CNN Zoo dataset there is no relevant difference between the two and we do not really know which of them the GBM model predicts. As a reference we also report the τ\tau coefficients when using the train accuracy as a proxy for the test one (or vice versa).

All the numbers are significantly larger than zero, indicating that the predictor’s ranking is far from being random. The predictions seem to correlate slightly better with train accuracy than with test. This hints that the predictors trained on the Small CNN Zoo may be using the train accuracy as a shortcut while predicting the test one. The ranking coefficient between the train and test accuracies decreases with network size, which points to increasing overfitting.

To further verify that our findings hold up with other architectures, we show in Supplement E that our findings also hold for Multi-Layer Perceptrons.

Conclusions and future directions

We demonstrated that it is possible to predict the performance of a DNN using only its weights (or simple statistics thereof) as inputs. Surprisingly, these predictions are able to rank networks trained on unobserved natural image datasets/with different large architectures. Whether these predictions can be reduced to simple human-interpretable rules and whether they can be helpful to improve DNN training remains an important open question. It also remains to be explored whether our findings transfer to domains outside of CNNs, e.g. to architectures commonly used in natural language understanding, reinforcement learning, or unsupervised applications.

Our work only used off-the-shelf regression algorithms (GBM and fully-connected DNNs) to predict the network accuracy using its weights. In future it seems natural to try methods with stronger inductive biases. For instance, using Deep Sets approach (Zaheer et al. 2017) to account for the invariance of CNNs w.r.t. the order of the filters and channels could allow us to get even better performance in practical applications, or yield better insights.

We believe our findings open the door to a number of interesting further questions. The idea that most neural network contain a highly efficient sub-network, the “lottery ticket hypothesis” (Frankle & Carbin 2019), recently gained a lot of attention. Morcos et al. 2019 show that these sub-networks transfer across tasks and datasets. An interesting avenue for future research would be to see if a trained classifier is able to identify these sub-networks (or other related properties) from the weights used to initialize a network.

Finally, we share a large dataset of trained CNNs in hope that this will enable the community to further explore this interesting direction of research.

Acknowledgements

We are thankful to Ibrahim Alabdulmohsin, Iuliya Beloshapka, Samy Bengio, Lucas Beyer, Alexey Dosovitskiy, Pierre Foret, Yiding Jiang, Alexander Kolesnikov, Dilip Krishnan, Hossein Mobahi, Shay Moran, Behnam Neyshabur, Sebastian Nowozin, Paul Rubenstein, Hanie Sedghi, Jakob Uszkoreit, and Scott Yak for valuable discussions.

References

Appendix A Further details on the Small CNN Zoo dataset and experiments

This section contains details on the way the Small CNN Zoo was generated and on the results of training the accuracy predictors reported in Tables 1 and 2 of the main text.

A.2 Base CNNs: hyperparameters for training

For each dataset, we sample 30k different hyperparameter configurations of the CNN training:

Optimizer is chosen uniformly from one of the following: vanilla SGD optimizer, Adam optimizer (Kingma & Lei 2014), and RMSProp optimizer;

Learning rate is sampled log-uniformly from [5×10−4,5×10−2][5\times 10^{-4},5\times 10^{-2}];

Dropout rate is sampled uniformly from [0,0.7][0,0.7];

Variance of weight initializer is sampled log-uniformly from [10−3,0.5][10^{-3},0.5];

Type of weight initializer is chosen uniformly from one of the following: Xavier normal (Glorot & Bengio 2010), He normal (He et al. 2015), orthogonal (Saxe et al. 2014), normal, and truncated normal;

Activation function is chosen uniformly from ReLu and hyperbolic tangent;

Fraction of training examples to use is sampled uniformly from {0.1,0.25,0.5,1.0}\{0.1,0.25,0.5,1.0\};

We never used same hyperparameter configuration with several different random seeds.

A.3 Accuracy predictors: types of the models

We use three types of predictors: logit-linear models, gradient boosted machine using desicion trees (GBM), and fully-connected ReLu networks (DNN).

A.4 Accuracy predictors: hyperparameters for training

For each of the 3 types of predictors and each of the 4 CNN collections we perform hyperparameter selection by evaluating 1k unique configurations:

For the GBM accuracy predictor we use the following protocol. Refer to the Light-GBM documentation for the exact meaning of the parameters:

num_leaves is sampled uniformly from [20,104][20,10^{4}];

learning_rate is sampled log-uniformly from [10−2,10−1][10^{-2},10^{-1}];

max_bin is sampled uniformly from {26−1,27−1,28−1,}\{2^{6}-1,2^{7}-1,2^{8}-1,\};

min_child_weight is sampled uniformly from {1,2,3,4,5}\{1,2,3,4,5\};

reg_lambda is sampled uniformly from [10−3,100][10^{-3},100];

reg_alpha is sampled uniformly from [10−6,5][10^{-6},5];

subsample is sampled uniformly from {0.1,0.2,…,0.9,1}\{0.1,0.2,\dots,0.9,1\};

colsample_bytree for the high dimensional inputs (all weights WW, weights of the second and third convolutional layers W2W^{2} and W3W^{3}, and concatenation of all weights with the hyperparameters (λ,W)(\lambda,W)) is sampled log-uniformly from [10−2,10−1][10^{-2},10^{-1}], for lower dimensional inputs is sampled uniformly from [0.7,1][0.7,1];

For the DNN accuracy predictor we use the following protocol:

Number of layers is sampled uniformly from {3,4,…,9}\{3,4,\dots,9\};

Number of units is sampled uniformly from {256,257,…,511}\{256,257,\dots,511\};

Dropout rate is sampled uniformly from [0,0.2][0,0.2];

Learning rate is sampled log-uniformly from [10−3,0.5][10^{-3},0.5];

Variance of weight initializer is sampled log-uniformly from [10−3,0.1][10^{-3},0.1];

Optimizer is chosen randomly from Adam and SGD;

Batch size is sampled uniformly from {64,128,256,512}\{64,128,256,512\};

Type of weight initializer is chosen uniformly from one of the following: Xavier normal (Glorot & Bengio 2010), He normal (He et al. 2015), orthogonal (Saxe et al. 2014), normal, and truncated normal;

Sigmoid transform is applied to the final layer output.

A.5 Accuracy predictors: detailed empirical results

Tables 5 and 6 contain both R2R^{2} scores and MSE values for all three types of predictors trained on all four CNN collections. Standard deviations capture the variability when training the predictors on three folds of the cross-validation. Every entry in the Tables 5 and 6 is obtained by evaluating 1k hyperparameter configurations of the accuracy predictor (as described in Section A.4) and picking the best one using 3-fold cross validation. Then the best configuration is evaluated on the holdout test split of the CNN collection. The resulting numbers are reported in the tables.

Appendix B GBM importance plots

Figure 5 presents importance values for various entries of the weight vector WW when training the GBM accuracy predictor. Importance values reported in the figure are based on the number of times a single feature (a particular entry of the vector WW in our case) was chosen in the nodes of the trees. Higher numbers correspond to more important (more frequently used) features. We see that all four models make extensive use of parameters of the final dense layer. Among those, biases seem to be slightly more important than weights.

Appendix C Permutation and scale invariance

This section contains the results of a study on how accuracy estimator’s predictions change as we modify its inputs. For a given accuracy predictor F^\hat{F} trained using weight vectors WW as inputs we test several ways of modifying its inputs W↦φ(W)W\mapsto\varphi(W):

Permuting the order of parameters within each layer of WW;

Permuting the order of parameters within all three convolutional layers of WW;

Permuting the order of parameters in the final dense layer;

Multiplying all elements of WW by a constant c>0c>0.

For every type of permutation we try two options: (a) permuting biases and weights jointly, allowing them to mix and (b) permuting biases and weights separately, without mixing them.

We test these modifications with the GBM predictor F^\hat{F} trained using weight vectors WW as inputs on the CIFAR10-GS CNN collection, which has the R2R^{2} score of 0.97. We use uniformly sampled random permutations and scale factors c∈{10−3,10−1,2,10,100}c\in\{10^{-3},10^{-1},2,10,100\}. For every type of modification φ\varphi we take the absolute differences ∣F^(φ(W))−F^(W)∣|\hat{F}\bigl(\varphi(W)\bigr)-\hat{F}(W)| between the predictions on the modified and original CNNs respectively. Then we average them across 1000 CNNs WW from the test split of the CIFAR10-GS CNN collection. Results are reported in Table 9.

Appendix D Detailed results on the transfer experiments

Figure 6 contains the results for all possible transfer experiments performed on Small CNN Zoo as described in Section 5.1. Diagonal plots correspond to the holdout test evaluation of four GBM models.

Appendix E Results on Multi-Layer Perceptrons

To verify that our results do not only apply to CNNs, we performed experiments on fully connected Multi-Layer Perceptrons (MLPs). We trained 10k MLPs each on CIFAR10 and SVHN, using the same hyperparameters as in the Small CNN Zoo (see Section A.2), except that we also sampled the number of hidden units in each layer to be either 8, 16, 32 or 64. This gave us four different neural network sizes for each dataset. To save computation time and verify that our observations also hold with different estimators, we used a Random Forest estimator with 32 trees in all of the experiments. The estimator was trained on networks from one specific dataset and hidden-unit size, and evaluated on all other settings. We used the same weight statistics W~ ⁣L\widetilde{W}_{\!L} as in the main text as input features. Table 10 shows the results of this experiment. We then verified that these results transfer across network architectures and datasets: The resulting Kendall’s τ\tau coefficients are listed in Table 11. Together, these results confirm our CNN findings, namely that it is possible to predict the performance of an MLP based on its weights, and that this prediction transfers across models of different sizes as well as across datasets.