On Power Laws in Deep Ensembles
Ekaterina Lobacheva, Nadezhda Chirkova, Maxim Kodryan, Dmitry Vetrov
Introduction
Neural networks provide state-of-the-art results in a variety of machine learning tasks, however, several neural network’s aspects complicate their usage in practice, including overconfidence , vulnerability to adversarial attacks , and overfitting . One of the ways to compensate these drawbacks is using deep ensembles, i. e. the ensembles of neural networks trained from different random initialization . In addition to increasing the task-specific metric, e. g. accuracy, the deep ensembles are known to improve the quality of uncertainty estimation, compared to a single network. There is yet no consensus on how to measure the quality of uncertainty estimation. Ashukha et al., consider a wide range of possible metrics and show that the calibrated negative log-likelihood (CNLL) is the most reliable one because it avoids the majority of pitfalls revealed in the same work.
Increasing the size of the deep ensemble, i. e. the number of networks in the ensemble, is known to improve the performance . The same effect holds for increasing the size of a neural network, i. e. the number of its parameters. Recent works show that even in an extremely overparameterized regime, increasing leads to a higher quality. These works also mention a curious effect of non-monotonicity of the test error w. r. t. the network size, called double descent behaviour.
In figure 1, left, we may observe the saturation and stabilization of quality with the growth of both the ensemble size and the network size . The goal of this work is to study the asymptotic properties of CNLL of deep ensembles as a function of and . We investigate under which conditions and w. r. t. which dimensions the CNLL follows a power law for deep ensembles in practice. In addition to the horizontal and vertical cuts of the diagram shown in figure 1, left, we also study its diagonal direction, which corresponds to the increase of the total parameter count.
The power-law behaviour of deep ensembles has previously been touched in the literature. Geiger et al., consider simple shallow architectures and reason about the power-law behaviour of the test error of a deep ensemble as a function of when , and of a single network as a function of when . Kaplan et al., , Rosenfeld et al., investigate the behaviour of single networks of modern architectures and empirically show that their NLL and test error follow power laws w. r. t. the network size . In this work, we perform a broad empirical study of power laws in deep ensembles, relying on the practical setting with properly regularized, commonly used deep neural network architectures. Our main contributions are as follows:
for the practically important scenario with NLL calibration, we derive the conditions under which CNLL of an ensemble follows a power law as a function of when ;
we empirically show that, in practice, the following dependencies can be closely approximated with a power law on the whole considered range of their arguments: (a) CNLL of an ensemble as a function of the ensemble size ; (b) CNLL of a single network as a function of the network size ; (c) CNLL of an ensemble as a function of the total parameter count;
based on the discovered power laws, we make several practically important conclusions regarding the use of deep ensembles in practice, e. g. using a large single network may be less beneficial than using a so-called memory split — an ensemble of several medium-size networks of the same total parameter count;
we show that using the discovered power laws for , and having a small number of trained networks, we can predict the CNLL of the large ensembles and the optimal memory split for a given memory budget.
Theoretical view
The primary goal of this work is to perform the empirical study of the conditions under which NLL and CNLL of deep ensembles follow a power law. Before diving into a discussion about our empirical findings, we first provide a theoretical motivation for anticipating power laws in deep ensembles, and discuss the applicability of this theoretical reasoning to the practically important scenario with calibration.
Ashukha et al., emphasize that the comparison of the NLLs of different models with suboptimal softmax temperature may lead to an arbitrary ranking of the models, so the comparison should only be performed after calibration, i. e. with optimally selected temperature . The model-average CNLL of an ensemble of size , measured on the whole dataset , is defined as follows:
To sum up, in this section we derived an asymptotic power law for LE-NLL that may be treated as another definition of CNLL, and that closely approximates the commonly used CNLL in practice.
Experimental setup
We use the logarithmic scale to pay more attention to the small differences between values for large . For a fixed , optimizing the given loss is equivalent to fitting the linear regression model with one factor in the space — (see fig. 1, right as an example).
NLL as a function of ensemble size
When the temperature grows, the general behaviour is that approaches -1 more and more tightly. This behaviour breaks for the ensembles of small networks (blue lines). The reason is that the number of trained small networks is large, and the NLL for large ensembles with high temperature is noisy in log-scale, so the approximation of NLL with power law is slightly worse than for other settings, as confirmed in the rightmost plot of fig. 2. Nevertheless, these approximations are still very close to the data, we present the corresponding plots in Appendix E.2.
In figure 3, we observe that for WideResNet, parameter decreases, as becomes large, and starts growing for large . For VGG, this effect also occurs in a light form but is almost invisible at the plot. This suggests that large networks gain less from the ensembling, and therefore the ensembles of larger networks are less effective than the ensembles of smaller networks. We suppose, the described effect is a consequence of under-regularization (the large networks need more careful hyperparameter tuning and regularization), because we also observed the described effect in a more strong form for the networks with all regularization turned off, see Appendix H. However, the described effect might also be a consequence of the decreased diversity of wider networks , and needs further investigation.
NLL as a function of network size
Single network. Figure 4, left shows the NLL with and the CNLL of a single VGG on the CIFAR-100 dataset as a function of the network size. We observe the double descent behaviour of the non-calibrated NLL, which could not be approximated with a power law for the considered range of . The calibration removes the double-descent behaviour, and allows a close power-law approximation, as confirmed in the middle plot of figure 4. Interestingly, parameter is close to , which coincides with the results of Geiger et al., derived for the test error. The results for other dataset–architecture pairs are given in Appendix I.
Nakkiran et al., observe the double descent behaviour of accuracy as a function of the network size for highly overfitted networks, when training networks without regularization, with label noise, and for much more epochs than is usually needed in practice. In our practical setting, accuracy and CNLL are the monotonic functions of the network size, while for the non-calibrated NLL, the double descent behaviour is observed. Ashukha et al., point out that accuracy and CNLL usually correlate, so we hypothesize that the double descent may be observed for CNLL in the same scenarios when it is observed for accuracy, while the non-calibrated NLL exhibits the double descent at the earlier epochs in these scenarios. To sum up, our results support the conclusions of that the comparison of the NLL of the models of different sizes should only be performed with an optimal temperature.
NLL as a function of the total parameter count
In the previous sections, we analyzed the vertical and horizontal cuts of the -plane shown in figure 1, left. In this section, we analyze the diagonal cuts of this space. One direction of diagonal corresponds to the fixed total parameter count, later referred to as a memory budget, and the orthogonal direction reflects the increasing budget.
Another, practically important, effect is that the lower envelope may be reached at . In other words, for a fixed memory budget, a single network may perform worse than an ensemble of several medium-size networks of the same total parameter count, called a memory split in the subsequent discussion. We refer to the described effect itself as a Memory Split Advantage effect (MSA effect). We further illustrate the MSA effect in figure 5, right, where each line corresponds to a particular memory budget, the x-axis denotes the number of networks in the memory split, and the lowest CNLL, denoted at the y-axis, is achieved at for all lines. We consistently observe the MSA effect for different settings and metrics, i. e. CNLL and accuracy, for a wide range of budgets, see Appendix J. We note that the MSA-effect holds even for budgets smaller than the standard budget. We also show in appendix J that using the memory split with the relatively small number of networks is only moderately slower than using a single wide network, in both training and testing stages. We describe the memory split advantage effect in more details in .
Prediction based on power laws
One of the advantages of a power law is that, given a few starting points satisfying the power law, one can exactly predict values for any . In this section, we check whether the power laws discovered in section 4 are stable enough to allow accurate predictions.
We use the CNLL of the ensembles of sizes as starting points, and predict the CNLL of larger ensembles. We firstly conduct the experiment using the values of starting points, obtained by averaging over a large number of runs. In this case, the CNLL of large ensembles may be predicted with high precision, see Appendix K. Secondly, we conduct the experiment in the practical setting, when the values of starting points were obtained using only 6 trained networks (using 6 networks allows the more stable estimation of CNLL of ensembles of sizes ). The two left plots of figure 6 report the error of the prediction for the different ensemble sizes and network sizes of VGG and WideResNet on the CIFAR-100 dataset. The plots for other settings are given in Appendix K. The experiment was repeated 10 times for VGG and 5 times for WideResNet with the independent sets of networks, and we report the average error. The error is orders smaller than the value of CNLL, and based on this, we conclude that the discovered power laws allow quite accurate predictions.
In section 6, we introduced memory splitting, a simple yet effective method of improving the quality of the network, given the memory budget . Using the obtained predictions for CNLL, we can now predict the optimal memory split (OMS) for a fixed by selecting the optimum at a specific diagonal of the predicted (, )-plane, see Appendix K for more details. We show the results for the practical setting with 6 given networks in figure 6, right. The plots depict the number of networks in the true and predicted OMS; the network size can be uniquely determined by and . In most cases, the discovered power laws predict either the exact or the neighboring split. If we predict the neighboring split, the difference in CNLL between the true and predicted splits is negligible, i. e. of the same order as the errors presented in figure 6, left.
Related Work
Deep ensembles and overparameterization. The two main approaches to improve deep neural networks accuracy are ensembling and increasing network size. While a bunch of works report the quantitative influence of the above-mentioned techniques on model quality , few investigate the qualitative side of the effect. Some recent works consider a simplified or narrowed setup to tackle it. For instance, Geiger et al., similarly discover the power laws in test error w. r. t. model and ensemble size for simple binary classification with hinge loss, and give a heuristic argument supporting their findings. We provide an extensive theoretical and empirical justification of similar claims for the calibrated NLL using modern architectures and datasets. Other layers of works on studying neural network ensembles and overparameterized models include but not limited to the Bayesian perspective , ensembles diversity improvement techniques , neural tangent kernel (NTK) view on overparameterized neural networks , etc.
Power laws for predictions. A few recent works also empirically discover power laws with respect to data and model size and use them to extrapolate the performance on small models/datasets to larger scales . Their findings even allow estimating the optimal compute budget allocation given limited resources. However, these studies do not account for the ensembling of models and the calibration of NLL.
MSA-effect. Concurrently with our work, Kondratyuk et al., investigate a similar effect for budgets measured in FLOPs. Earlier, an MSA-like effect has also been noted in . However, the mentioned works did not consider the proper regularization of networks of different sizes and did not propose the method for predicting the OMS, while both aspects are important in practice.
Conclusion
In this work, we investigated the power-law behaviour of CNLL of deep ensembles as a function of ensemble size and network size and observed the following power laws. Firstly, with a minor modification of the calibration procedure, CNLL as a function of follows a power law on the wide finite range of , starting from , but with the power parameter slightly higher than the one derived theoretically. Secondly, the CNLL of a single network follows a power law as a function of the network size on the whole reasonable range of network sizes, with the power parameter approximately the same as derived. Thirdly, the CNLL also follows a power law as a function of the total parameter count (memory budget). The discovered power laws allow predicting the quality of large ensembles based on the quality of the smaller ensembles consisting of networks with the same architecture. The practically important finding is that for a given memory budget, the number of networks in the optimal memory split is usually much higher than one, and can be predicted using the discovered power laws. Our source code is available at https://github.com/nadiinchi/power_laws_deep_ensembles.
Broader Impact
In this work, we provide an empirical and theoretical study of existing models (namely, deep ensembles); we propose neither new technologies nor architectures, thus we are not aware of its specific ethical or future societal impact. We, however, would like to point out a few benefits gained from our findings, such as optimization of resource consumption when training neural networks and contribution to the overall understanding of neural models. As far as we are concerned, no negative consequences may follow from our research.
Acknowledgments and Disclosure of Funding
We would like to thank Dmitry Molchanov, Arsenii Ashukha, and Kirill Struminsky for the valuable feedback. The theoretical results presented in section 2 were supported by Samsung Research, Samsung Electronics. The empirical results presented in sections 4, 5, 6, 7 were supported by the Russian Science Foundation grant №19-71-30020. This research was supported in part through the computational resources of HPC facilities at NRU HSE. Additional revenues of the authors for the last three years: Stipend by Lomonosov Moscow State University, Travel support by ICML, NeurIPS, Google, NTNU, DESY, UCM.
References
Appendix A Theory
Let us fix some and split the expectation of the remainder into two terms:
Applying the Hoeffding’s inequality to allows us to exponentially bound the probability of deviation from by :
Taking inequality (9) into account, we can bound the first term in (8) as follows:
where .
In the second term of (8), as , the remainder term can be expanded as a Taylor series and thus bounded as follows:
From that we derive the following bound on the second term in (8):
A.2 Power law in NLL with temperature applied after averaging
where . By applying a similar Taylor expansions trick as in Proposition 1 to the function under expectation in (14), we arrive at:
We can see that the first term in (16) is non-negative, while the second one is, on the contrary, always negative. High values of temperature (i. e. ) make the last term negative as well, which, in turn, may lead to negative coefficient in the power law of the right-hand part of (15). We observe this effect in practice: at certain values of temperature applied after averaging, the NLL starts increasing as a function of , see Appendix C.2. We could not apply a similar technique, as was provided in section A.1, to derive the exact asymptote in (15), when the temperature is applied after averaging.
and hence Proposition 1 is applicable. The power law parameter in this case is always positive.
A.3 Lower envelope of power laws with continuous temperature
In this subsection, we show that even when the set of temperatures is uncountable, the lower envelope of power laws asymptotically follows the power law. This argument generalizes our discussion at the end of section 2.
By the definition of , the following inequalities hold:
The first of the right inequalities in (21) can be obtained via multiplication of the left inequalities by and , respectively, and summation. The second inequality is obtained after substituting the first one into the right-hand part of the first left inequality.
As , we come to (20) with , .
Appendix B Experimental details
Data. We conduct experiments on CIFAR-100 and CIFAR-10 datasets, each containing 50000 training and 10000 testing examples. For tuning hyperparameters, we randomly select 5000 training examples as a validation set. After choosing optimal hyperparameters, we retrain the models on the full training dataset. We use a standard data augmentation scheme: zero-padding with 4 pixels on each side, random cropping to produce images, and horizontal mirroring with probability .
In several experiments presented in the Appendix, we train networks without regularization, with all hyperparameters being the same for all network sizes. By this, we mean that we set weight decay and dropout rate to zero, do not use data augmentation, and use an initial learning rate 10 times smaller than in the reference implementation to ensure that the training converges for all considered models.
Computing infrastructure. VGG networks were trained on NVIDIA Teals P100 GPU. Training one network of the standard size / smallest considered size / largest considered size took approximately 1 hour / 20 minutes / 4.5 hours. WideResNet networks were trained on NVIDIA Tesla V100 GPU. Training one network of the standard size / smallest considered size / largest considered size took approximately 5.5 hours / 50 minutes / 32 hours.
Test-time cross-validation. Ashukha et al., utilize a so-called test-time cross-validation to obtain an unbiased, low-variance estimate of the CNLL using the publicly available test set. The test-time cross-validation implies that the test set is randomly split into two equal parts five times. For each split, one part is used to find the optimal temperature, and the other one — to measure the CNLL, and vice versa. Finally, the CNLL is averaged over ten measurements.
Appendix C Calibration of ensembles: applying temperature before or after averaging
In section 2, we introduced two ways of applying temperature to the ensemble, namely before and after averaging. In this subsection, we empirically compare these two calibration procedures. Figure 7 shows the results for VGG on CIFAR-100, for setting with and without regularization. The difference between CNLL values is low in all the cases, hence the two procedures perform similarly. In most of the cases, particularly in practically important cases of ensembling networks of medium and large sizes, calibration with applying temperature before averaging performs slightly better (there are a lot of green pixels in the heatmaps). The results for other dataset-architecture pairs are similar. To sum up, the procedure with applying temperature before averaging can be used in practice instead of the standard one, with temperature applied after averaging, without loss in the quality.
C.2 Dynamics of NLL with fixed temperature and CNLL
Appendix D Comparison of CNLL and LE-NLL
In this section, we empirically show that moving the minimum operation outside the expectation in equation (3) does not change the value of CNLL a lot in practice. In other words, we compare the values of CNLL (3), commonly used in practice, and LE-NLL (4), utilized in section 2, for different dataset–architecture pairs. We consider two scenarios for computing CNLL: (a) when the optimal temperature is chosen using the whole test set, as it is the case for LE-NLL, and (b) when the test-time cross-validation is utilized . In figure 9 we depict the difference between CNLL and LE-NLL for different values of ensemble size and network size , for VGG on CIFAR-100. We observe that the difference is negligible, compared to the values of LE-NLL. The relative difference for all values of and is bounded by / for scenarios (a) and (b) respectively. For WideResNet on CIFAR-100, the relative difference is bounded by / , for VGG on CIFAR-10 — / , for WideResNet on CIFAR-10 — / for scenarios (a) and (b) respectively. In all the the experiments in the paper, we use CNLL computed using test-time cross-validation.
Appendix E NLL with fixed temperature as a function of ensemble size
Parameter decreases when the temperature grows, since lower temperatures result in more contrast predictions, and ensembling smooths them. Parameter is a non-monotonic function of , and its optimum reflects the optimal temperature for the “infinite”-size ensemble. The optimal temperature may be greater or less than one, depending on the dataset–architecture combination.
E.2 Power law approximation for small networks with high temperature
Appendix F Convergence of temperature
To show that the optimal temperature converges when ensemble size increases, we plot optimal vs. for different dataset–architecture pairs in figure 12. We average the optimal temperature over runs, i. e. different trained ensembles, and over folds in test-time cross-validation.
Appendix G Power law approximation of CNLL as a function of ensemble size
Appendix H CNLL of the ensemble of unregularized networks
Appendix I Power law approximation of NLL as a function of network size
Appendix J Power law approximation of NLL as a function of the memory budget. MSA effect
The timing analysis of the memory splitting procedure. One might wonder, how much slower is training and prediction with the ensemble of several medium-size networks compared to a single large network. In table 1, we list the training and prediction time for a single VGG on CIFAR-100, for different network sizes. We conduct this experiment on GPU conducting training / prediction with several networks sequentially, using the training batch size of 64 and the testing batch size of 1024. We observe that using the memory split with the relatively small number of networks (which is the case in practice) is only moderately slower than using a single wide network. For example, for budget , with a single network / memory split of 4 networks, testing takes 111 / 132 sec, while one training epoch takes 42 / 64 seconds. Please note that these numbers are given for the case when the networks in the memory split are trained / validated sequentially, while the memory split allows parallel training / validation.
Appendix K Predictions based on power laws
Finding optimal memory splits in practice. Let’s say we have a budget and want to find an optimal memory split (MS). Algorithm 1 describes the procedure of finding the optimal MS using the discovered power laws. With this algorithm, we need to train networks of size for where is a number of networks in an optimal MS, and after finding , we also need to train lacking networks for the optimal MS. If we do not use power-law predictions we can train MSs one by one (one network of size , then two networks of size , etc.) while the quality of the MS starts to degrade. In this case we need to train networks of size for . As a result, if , power-law predictions allow training fewer networks, and the higher the higher the gain.
Appendix L Additional experiments with ImageNet
Ashukha et al., released the weights of the ensembles of ResNet50 networks (commonly used size) trained on ImageNet. We used their data to empirically confirm the power-law behaviour of NLL and CNLL as functions of the ensemble size. Figure 21 shows that the resulting power law approximations fit the data well, supporting the results presented in section 4.