Diversity with Cooperation: Ensemble Methods for Few-Shot Classification

Nikita Dvornik, Cordelia Schmid, Julien Mairal

Introduction

Convolutional neural networks have become standard tools in computer vision to model images, leading to outstanding results in many visual recognition tasks such as classification , object detection , or semantic segmentation . Massively annotated datasets such as ImageNet or COCO seem to have played a key role in this success. However, annotating a large corpus is expensive and not always feasible, depending on the task at hand. Improving the generalization capabilities of deep neural networks and removing the need for huge sets of annotations is thus of utmost importance.

While such a grand challenge may be addressed from different complementary points of views, e.g., large-scale unsupervised learning , self-supervised learning , or by developing regularization techniques dedicated to deep networks , we choose in this paper to focus on variance-reduction principles based on ensemble methods.

Specifically, we are interested in few-shot classification, where a classifier is first trained from scratch on a medium-sized annotated corpus—that is, without leveraging external data or a pre-trained network, and then we evaluate its ability to adapt to new classes, for which only very few annotated samples are provided (typically 1 or 5). Unfortunately, simply fine-tuning a convolutional neural network on a new classification task with very few samples has been shown to provide poor results , which has motivated the community to develop dedicated approaches.

The dominant paradigm in few-shot learning builds upon meta-learning , which is formulated as a principle to learn how to adapt to new learning problems. These approaches split a large annotated corpus into classification tasks, and the goal is to transfer knowledge across tasks in order to improve generalization. While the meta-learning principle seems appealing for few-shot learning, its empirical benefits have not been clearly established yet. There is indeed strong evidence that training CNNs from scratch using meta-learning performs substantially worse than if CNN features are trained in a standard fashion—that is, by minimizing a classical loss function relying on corpus annotations; on the other hand, learning only the last layer with meta-learning has been found to produce better results . Then, it was recently shown in that simple distance-based classifiers could achieve similar accuracy as meta-learning approaches.

Our paper goes a step further and shows that meta-learning-free approaches can be improved and significantly outperform the current state of the art in few-shot learning. Our angle of attack consists of using ensemble methods to reduce the variance of few-shot learning classifiers, which is inevitably high given the small number of annotations. Given an initial medium-sized dataset (following the standard setting of few-shot learning), the most basic ensemble approach consists of first training several CNNs independently before freezing them and removing the last prediction layer. Then, given a new class (with few annotated samples), we build a mean centroid classifier for each network and estimate class probabilities—according to a basic probabilistic model—of test samples based on the distance to the centroids . The obtained probabilities are then averaged over networks, resulting in higher accuracy.

While we show that the basic ensemble method where networks are trained independently already performs well, we introduce penalty terms that allow the networks to cooperate during training, while encouraging enough diversity of predictions, as illustrated in Figure 1. The motivation for cooperation is that of easier learning and regularization, where individual networks from the ensemble can benefit from each other. The motivation for encouraging diversity is classical for ensemble methods , where a collection of weak learners making diverse predictions often performs better together than a single strong one. Whereas these two principles seem in contradiction with each other at first sight, we show that both principles are in fact useful and lead to significantly better results than the basic ensemble method. Finally, we also show that a single network trained by distillation to mimic the behavior of the ensemble also performs well, which brings a significant speed-up at test time. In summary, our contributions are three-fold:

We introduce mechanisms to encourage cooperation and diversity for learning an ensemble of networks. We study these two principles for few-shot learning and characterize the regimes where they are useful.

We show that it is possible to significantly outperform current state-of-the-art techniques for few-shot classification without using meta-learning.

As a minor contribution, we also show how to distill an ensemble into a single network with minor loss in accuracy, by using additional unlabeled data.

Related Work

In this section, we discuss related work on few-shot learning, meta-learning, and ensemble methods.

Typical few-shot classification problems consist of two parts called meta-training and meta-testing . During the meta-training stage, one is given a large-enough annotated dataset, which is used to train a predictive model. During meta-testing, novel categories are provided along with few annotated examples, and we evaluate the capacity of the predictive model to retrain or adapt, and then generalize on these new classes.

Meta-learning approaches typically sample few-shot learning classification tasks from the meta-training dataset, and train a model such that it should generalize on a new task that has been left aside. For instance, in a “good network initialization” is learned such that a small number of gradient steps on a new problem is sufficient to obtain a good solution. In , the authors learn both the network initialization and an update rule (optimization model) represented by a Long-Term-Short-Memory network (LSTM). Inspired by few-shot learning strategies developed before deep learning approaches became popular , distance-based classifiers based on the distance to a centroid were also proposed, e.g., prototypical networks , or more sophisticated classifiers with attention . All these methods consider a classical backbone network, and train it from scratch using meta-learning.

Recently, these meta-learning were found to be sub-optimal . Specifically, better results were obtained by training the network on the classical classification task using the meta-training data in a first step, and then only fine-tuning with meta-learning in a second step . Others such as simply freeze the network obtained in the first step, and train a simple prediction layer with meta-learning, which results in similar performance. Finally, the paper demonstrates that simple baselines without meta-learning—based on distance-based classifiers—work equally well. Our paper pushes such principles even further and shows that by appropriate variance-reduction techniques, these approaches can significantly outperform the current state of the art.

Ensemble methods.

It is well known that ensemble methods reduce the variance of estimators and subsequently may improve the quality of prediction . To gain accuracy from averaging, various randomization or data augmentation techniques are typically used to encourage a high diversity of predictions . While individual classifiers of the ensemble may perform poorly, the quality of the average prediction turns out to be sometimes surprisingly high.

Even though ensemble methods are costly at training time for neural networks, it was shown that a single network trained to mimic the behavior of the ensemble could perform almost equally well —a procedure called distillation—thus removing the overhead at test time. To improve the scalability of distillation in the context of highly-parallelized implementations, an online distillation procedure is proposed in . There, each network is encouraged to agree with the averaged predictions made by other networks of the ensemble, which results in more stable models. The objective of our work is however significantly different. The form of cooperation they encourage between networks is indeed targeted to scalability and stability (due to industrial constraints), but online distilled networks do not necessarily perform better than the basic ensemble strategy. Our goal, on the other hand, is to improve the quality of prediction and do better than basic ensembles.

To this end, we encourage cooperation in a different manner, by encouraging predictions between networks to match in terms of class probabilities conditioned on the prediction not being the ground truth label. While we show that such a strategy alone is useful in general when the number of networks is small, encouraging diversity becomes crucial when this number grows. Finally, we show that distillation can help to reduce the computational overhead at test time.

Our Approach

In this section, we present our approach for few-shot classification, starting with preliminary components.

We now explain how to perform few-shot classification with a fixed feature extractor and a mean centroid classifier.

During meta-testing, we are given a new dataset Dq={xi,yi}i=1nkD_{q}=\{x_{i},y_{i}\}_{i=1}^{nk}, where nn is a number of new categories and kk is the number of available examples for each class. The (xi,yi)(x_{i},y_{i})’s represent image-label pairs. Then, we build a mean centroid classifier, leading to the class prototypes

Finally, a test sample xx is assigned to the nearest centroid’s class. Simple mean-centroid classifiers have proven to be effective in the context of few-shot classification , which is confirmed in the following experiment.

Motivation for mean-centroid classifier.

We report here an experiment showing that a more complex model than (1) does not necessarily lead to better results for few-shot learning. Consider indeed a parametrized version of (1):

where the weights αij\alpha_{i}^{j} can be learned with gradient descent by maximizing the likelihood of the probabilistic model

where d(⋅,⋅)d(\cdot,\cdot) is a distance function, such as Euclidian distance or negative cosine similarity. Since the coefficients are learned from data and not set arbitrarily to 1/k1/k as in (1), one would potentially expect this method to produce better classifiers if appropriately regularized. When we run the evaluation of the aforementioned classifiers on 1000 5-shot learning tasks sampled from miniImagenet-test (see experimental section for details about this dataset), we get similar results on average: 77.28±0.46%77.28\pm 0.46\% for (1) vs. 77.01±0.50%77.01\pm 0.50\% for (2), confirming that learning meaningful parameters in this very-low-sample regime is difficult.

2 Learning ensembles of deep networks

During meta-training, one needs to minimize the following loss function over a training set {xi,yi}i=1m\{x_{i},y_{i}\}_{i=1}^{m}:

When training an ensemble of KK networks fθkf_{\theta_{k}} independently, one would solve (4) for each network separately. While these terms may look identical, solutions provided by deep neural networks will typically differ when trained with different initializations and random seeds, making ensemble methods appealing in this context.

In this paper, we are interested in ensemble of networks, but we also want to model relationships between its members; this may be achieved by considering a pairwise penalty function ψ\psi, leading to the joint formulation:

where θˉ\bar{\theta} is the vector obtained by concatenating all the parameters θj\theta_{j}. By carefully designing the function ψ\psi and setting up appropriately the parameter γ\gamma, it is possible to achieve desirable properties of the ensemble, such as diversity of predictions or collaboration during training.

3 Encouraging diversity and cooperation

To reduce the high variance of few-shot learning classifiers, we use ensemble methods trained with a particular interaction function ψ\psi, as in (5). Then, once the parameters θj\theta_{j} have been learned during meta-training, classification in meta-testing is performed by considering a collection of KK mean-centroid classifiers associated to the basic probabilistic model presented in Eq. (3). Given a test image, the KK class probabilities are averaged. Such a strategy was found to perform empirically better than a voting scheme.

As we show in the experimental section, the choice of pairwise relationship function ψ\psi significantly influences the quality of the ensemble. Here, we describe three different strategies, which all provide benefits in different regimes, starting by a criterion encouraging diversity of predictions.

One way to encourage diversity consists of introducing randomization in the learning procedure, e.g., by using data augmentation or various initializations. Here, we also evaluate the effect of an interaction function ψ\psi that acts directly on the network predictions. Given an image xx, two models parametrized by θi\theta_{i} and θj\theta_{j} respectively lead to class probabilities pi=σ(fθi(x))p_{i}=\sigma(f_{\theta_{i}}(x)) and pj=σ(fθj(x))p_{j}=\sigma(f_{\theta_{j}}(x)). During training, pip_{i} and pjp_{j} are naturally encouraged to be close to the assignment vector eye_{y} in {0,1}d\{0,1\}^{d} with a single non-zero entry at position yy, where yy is the class label associated to xx and dd is the number of classes.

From , we know that even though only the largest entry of pip_{i} or pjp_{j} is used to make predictions, other entries—typically not corresponding to the ground truth label yy—carry important information about the network. It becomes then natural to consider the probabilities p^i\hat{p}_{i} and p^j\hat{p}_{j} conditioned on not being the ground truth label. Formally, these are obtained by setting to zero the entry yy in pip_{i} and pjp_{j} renormalizing the corresponding vectors such that they sum to one. Then, we consider the following diversity penalty

When combined with the loss function, the resulting formulation encourages the networks to make the right prediction according to the ground-truth label, but then they are also encouraged to make different second-best, third-best, and so on, choice predictions (see Figure 1). This penalty turns out to be particularly effective when the number of networks is large, as shown in the experimental section. It typically worsens the performance of individual classifiers on average, but make the ensemble prediction more accurate.

Cooperation.

Apparently opposite to the previous principle, encouraging the conditional probabilities p^i\hat{p}_{i} to be similar—though with a different metric—may also improve the quality of prediction by allowing the networks to cooperate for better learning. Our experiments show that such a principle alone may be effective, but it appears to be mostly useful when the number of training networks is small, which suggests that there is a trade-off between cooperation and diversity that needs to be found.

Specifically, our experiments show that using the negative cosine—in other words, the opposite of (6)—is ineffective. However, a penalty such as the symmetrized KL-divergence turned out to provide the desired effect:

By using this penalty, we managed to obtain more stable and faster training, resulting in better performing individual networks, but also—perhaps surprisingly—a better ensemble. Unfortunately, we also observed that the gain of ensembling diminishes with the number of networks in the ensemble since the individual members become too similar.

Robustness and cooperation.

Given experiments conducted with the two previous penalties, a trade-off between cooperation and diversity seems to correspond to two regimes (low vs. high number of networks). This motivated us to develop an approach designed to achieve the best trade-off. When considering the cooperation penalty (7), we try to increase diversity of prediction by several additional means. i) We randomly drop some networks from the ensemble at each training iteration, which causes the networks to learn on different data streams and reduces the speed of knowledge propagation. ii) We introduce Dropout within each network to increase randomization. iii) We feed each network with a different (crop, color) transformation of the same image, which makes the ensemble more robust to input image transformations. Overall, this strategy was found to perform best in most scenarios (see Figure 2).

4 Ensemble distillation

As most ensemble methods, our ensemble strategy introduces a significant computional overhead at training time. To remove the overhead at test time, we use a variant of knowledge distillation to compress the ensemble into a single network fwf_{w}. Given the meta-training dataset DbD_{b}, we consider the following cost function on example (x,y)(x,y):

When distillation is performed on the dataset DbD_{b}, the network fwf_{w} mimics the behavior of the ensemble on a specific data distribution. However, new categories are introduced at test time. Therefore, we also tried distillation by using additional unnannotated data, which yields slightly better performance.

Experiments

We now present experiments to study the effect of cooperation and diversity for ensemble methods, and start with experimental and implementation details.

We use mini-ImageNet and tiered-ImageNet which are derived from the original ImageNet dataset and Caltech-UCSD Birds (CUB) 200-2011 . Mini-ImageNet consists of 100 categories—64 for training, 16 for validation and 20 for testing—with 600 images each. Tiered-ImageNet is also a subset of ImageNet that includes 351 class for training, 97 for validation and 160 for testing which is 779,165 images in total. The splits are chosen such that the training classes are sufficiently different from the test ones, unlike in mini-ImageNet. The CUB dataset consists of 11,788 images of birds of more than 200 species. We adopt train, val, and test splits from , which were originally created by randomly splitting all 200 species in 100 for training, 50 for validation, and 50 for testing.

Evaluation.

In few-shot classification, the test set is used to sample NN 5-way classification problems, where only kk examples of each category are provided for training and 15 for evaluation. We follow and test our algorithms for k=k= 1 and 5 and NN is set to 1 0001\,000. Each time, classes and corresponding train/test examples are sampled at random. For all our experiments we report the mean accuracy (in %) over 1 0001\,000 tasks and 95%95\% confidence interval.

Implementation details.

For all experiments, we use the Adam optimizer with an initial learning rate 10−410^{-4}, which is decreased by a factor 10 once during training when no improvement in validation accuracy is observed for pp consecutive epochs. For mini-ImageNet, we use p=10p=10, and 20 for the CUB dataset. When distilling an ensemble into one network, pp is doubled. We use random crops and color augmentation during training as well as weight decay with parameter λ=\lambda= 5⋅10−45\cdot 10^{-4}. All experiments are conducted with the ResNet18 architecture , which allows us to train our ensembles of 2020 networks on a single GPU. Input images are then re-scaled to the size 224×224224\times 224, and organized in mini-batches of size 16. Validation accuracy is computed by running 5-shot evaluation on the validation set. During the meta-testing stage, we take central crops of size 224×224224\times 224 from images and feed them to the feature extractor. No other preprocessing is used at test time. When building a mean centroid classifier, the distance dd in (3) is computed as the negative cosine similarity , which is rescaled by a factor 10. For a fair comparison, we have also evaluated ensembles composed of ResNet18 with input image size 84×8484\times 84 and WideResNet28 with input size 80×8080\times 80. All details are reported in Appendix. For reproducibility purposes, our implementation will be made available at http://thoth.inrialpes.fr/research/fewshot_ensemble/.

2 Ensembles for Few-Shot Classification

In this section, we study the effect of ensemble training with pairwise interaction terms that encourage cooperation or diversity. For that purpose, we analyze the link between the size of ensembles and their 1- and 5-shot classification performance on the mini-ImageNet and CUB datasets.

When models are trained jointly, the data stream is shared across all networks and weight updates happen simultaneously. This is achieved by placing all models on the same GPU and optimizing the loss (5). When training a diverse ensemble, we use the cosine function (6) and selected the parameter γ=1\gamma=1 that performed best on the validation set among the tested values (10i10^{i}, for i=−2,…,2i=-2,\ldots,2) for n=5n=5 and n=10n=10 networks. Then, this value was kept for other values of nn. To enforce cooperation between networks, we use the symmetrized KL function (7) and selected the parameter γ=10\gamma=10 in the same manner. Finally, the robust ensemble strategy is trained with the cooperation relationship penalty and the same parameter γ\gamma, but we use Dropout with probability 0.1 before the last layer; each of the network is dropped from the ensemble with probability 0.2 at every iteration; different networks receive different transformation of the same image, i.e. different random crops and color augmentation.

Results.

Table 1 and Table A1 of Appendix summarize the few-shot classification accuracies of ensembles trained with our strategies and compare with basic ensembles. On the mini-ImageNet dataset, the results for 1- and 5-shot classification are consistent with each other. Training with cooperation allows smaller ensembles (n≤5n\leq 5) to perform better, which leads to higher individual accuracy of the ensemble members, as seen in Figure 2. However, when n≥10n\geq 10, cooperation is less effective, as opposed to the diversity strategy, which benefits from larger nn. As we can see from Figure 2, individual members of the ensemble become worse, but the ensemble accuracy improves substantially. Finally, the robust strategy seems to perform best for all values of nn in almost all settings. The situation for the CUB dataset is similar, although we notice that robust ensembles perform similarly as the diversity strategy for n=20n=20.

3 Distilling an ensemble

We distill robust ensembles of all sizes to study knowledge transferability with growing ensemble size. To do so, we use the meta-training dataset and optimize the loss (8) with parameters T=10T=10 and α=0.8\alpha=0.8. For the strategy using external data, we randomly add at each iteration 8 images (without annotations) from the COCO dataset to the 16 annotated samples from the meta-training data. Those images contribute only to the distillation part of the loss (8). Table 1 and Table A1 of Appendix display model accuracies for mini-ImageNet and CUB datasets respectively. For 5-shot classification on mini-ImageNet, the difference between ensemble and its distilled version is rather low (around 1%), while adding extra non-annotated data helps reducing this gap. Surprisingly, 1-shot classification accuracy is slightly higher for distilled models than for their corresponding full ensembles. On the CUB dataset, distilled models stop improving after n=5n=5, even though the performance of full ensembles keeps growing. This seems to indicate that the capacity of the single network may have been reached, which suggests using a more complex architecture here. Consistently with such hypothesis, adding extra data is not as helpful as for mini-ImageNet, most likely because data distributions of COCO and CUB are more different.

In Tables 2, 3, we also compare the performance of our distilled networks with other baselines from the literature, including current state-of-the-art meta-learning approaches, showing that our approach does significantly better on the mini-ImageNet and tiered-ImageNet datasets.

4 Study of relationship penalties

There are many possible ways to model relationship between the members of an ensemble. In this subsection, we study and discuss such particular choices.

As noted by , class probabilities obtained by the softmax layer of a network seem to carry a lot of information and are useful for distillation. However, after meta-training, such probabilities are often close to binary vectors with a dominant value associated to the ground-truth label. To make small values more noticeable, distillation uses a parameter TT, as in (8). Given such a class probability computed by a network, we experimented such a strategy consisting of introducing new probabilities p^=σ(p/T)\hat{p}=\sigma(p/T), where the contributions of non ground-truth values are emphasized. When used within our diversity (6) or cooperation (7) penalties, we however did not see any improvement over the basic ensemble method. Instead, we found that computing the class probabilities conditioned on not being the ground truth label, as explained in Section 3.3, would perform much better.

This is illustrated on the following experiment with two network ensembles of size n=5n=5. We enforce similarity on the full probability vectors in the first one, computed with softmax at T=10T=10 following , and with conditionally non-ground-truth probabilities for the second one as defined in Section 3.3. When using the cooperation training formulation, the second strategy turns out to perform about 1% better than the first one (79.79 % vs 80.60%), when tested on MiniImageNet. Similar observations have been made using the diversity criterion. In comparison, the basic ensemble method without interactions achieves about 80%80\%.

Choice of relationship function.

In principle, any similarity measure could be used to design a penalty encouraging cooperation. Here, we show that in fact, selecting the right criterion for comparing probability vectors (cosine similarity, L2 distance, symmetrized KL divergence), is crucial depending on the desired effect (cooperation or diversity). In Table 4, we perform such a comparison for an ensemble with n=5n=5 networks on the MiniImageNet dataset for a 5−5-shot classification task, when plugging the above function in the formulation 5, with a specific sign. The parameter γ\gamma for each experiment is chosen such that the performance on the validation set is maximized.

When looking for diversity, the cosine similarity performs slightly better than negative L2 distance, although the accuracies are within error bars. Using negative KLsim\text{KL}_{\text{sim}} with various γ\gamma was either not distinguishable from independent training or was hurting the performance for larger values of γ\gamma (not reported on the table). As for cooperation, positive KLsim\text{KL}_{\text{sim}} gives better results than L2 distance or negative cosine similarity. We believe that this behavior is due to important difference in the way these functions compare small values in probability vectors. While negative cosine or L2 losses would penalize heavily the largest difference, KLsim\text{KL}_{\text{sim}} concentrates on values that are close to 0 in one vector and are greater in the second one.

5 Performance under domain shift

Finally, we evaluate the performance of ensemble methods under domain shift. We proceed by meta-training the models on the mini-ImageNet training set and evaluate the model on the CUB-test set. The following setting was first proposed by and aims at evaluating the performance of algorithms to adapt when the difference between training and testing distributions is large. To compare to the results reported in the original work, we adopt their CUB test split. Table 5 compares our results to the ones listed in . We can see that neither the full robust ensemble nor its distilled version are able to do better than training a linear classifier on top of a frozen network. Yet, it does significantly better than distance-based approaches (denoted by cosine classifier in the table). However, if a diverse ensemble is used, it achieves the best accuracy. This is not surprising and highlights the importance of having diverse models when ensembling weak classifiers.

Conclusions

In this paper, we show that distance-based classifiers for few-shot learning suffer from high variance, which can be significantly reduced by using an ensemble of classifiers. Unlike traditional ensembling paradigms where diversity of predictions is encouraged by various randomization and data augmentation techniques, we show that encouraging the networks to cooperate during training is also important.

The overall performance of a single network obtained by distillation (with no computational overhead at test time) leads to state-of-the-art performance for few shot learning, without relying on the meta-learning paradigm. While such a result may sound negative for meta-learning approaches, it may simply mean that a lot of work remains to be done in this area to truly learn how to learn or to adapt.

Acknowledgment

This work was supported by the ERC grant number 714381 (SOLARIS project), the ERC advanced grant ALLEGRO and grants from Amazon and Intel.

References

Appendix

In this section, we elaborate on training, testing and distillation details of the proposed ensemble methods for different datasets, network architectures and input image resolution.

Training ResNet18 on 84x84 images on mini-ImageNet. For all experiments, we use ResNet18 with input image size 84x84, train with the Adam optimizer with an initial learning rate 3⋅10−43\cdot 10^{-4}, which is decreased by a factor 10 once during training when no improvement in validation accuracy is observed for pp consecutive epochs. We use p=20p=20 for training individual models, p=30p=30 for training ensembles and when distilling the model. When distilling an ensemble into one network, pp is doubled. We use random crops and color augmentation during training as well as weight decay with parameter λ=\lambda= 5⋅10−45\cdot 10^{-4}. At training time we use random crop, color transformation and adding random noise as data augmentation. During the meta-testing stage, we take central crops of size 224×224224\times 224 from images and feed them to the feature extractor. No other preprocessing is used at test time. The parameters used in distillation are the same as in Section 4.3 of the paper.

Training WideResNet28 on 80x80 images on mini-ImageNet. For all experiments, we use WideResNet28 with input image size 80x80, train with the Adam optimizer with an initial learning rate 1⋅10−41\cdot 10^{-4}, which is decreased by a factor 10 once during training when no improvement in validation accuracy is observed for pp consecutive epochs. We use p=20p=20 for training individual models, p=30p=30 for training ensembles and when distilling the model. When distilling an ensemble into one network, pp is doubled. We use random crops and color augmentation during training as well as weight decay with parameter λ=\lambda= 5⋅10−45\cdot 10^{-4}. We also set a dropout rate inside convolutional blocks to be 0.5 as described in. At training time we use random crop and color transformation only as data augmentation. During the meta-testing stage, we take central crops of size 80×8080\times 80 from images and feed them to the feature extractor. No other preprocessing is used at test time. The parameters used in distillation are the same as in Section 4.3 of the paper. Here, the maximal ensemble size we evaluated is 10 and not 20 due to memory limitations on available GPUs. Therefore, to construct an ensemble of size 20 we merge two ensembles of size 10, that were trained independently.

Training ResNet18 on 224x224 images on tiered-ImageNet For all experiments, we use ResNet18 with input image size 224x224, train with the Adam optimizer with an initial learning rate 3⋅10−43\cdot 10^{-4}, which is decreased by a factor 10 once during training when no improvement in validation accuracy is observed for pp consecutive epochs. We use p=20p=20 for training individual models, ensembles and for distilation. We use random crops and color augmentation during training as well as weight decay with parameter λ=\lambda= 1⋅10−41\cdot 10^{-4}. At training time we use random crop and color transformation. During the meta-testing stage, we take central crops of size 224×224224\times 224 from images and feed them to the feature extractor. No other preprocessing is used at test time. The parameters used in distillation are the same as in Section 4.3 of the paper.

B Additional Results

In this section we report and analyze the performance of different ensemble types depending on their size for different network architectures and input image resolutions.

Few-shot Classification with ResNet18 on 224x224 images on CUB. The results for 1- and 5-shot classification on CUB are presented in Table A1. Training details and Figure summary of the results are discussed in Experimental section of the paper.

Few-shot Classification with ResNet18 on 84x84 images on mini-ImageNet. The results for 1- and 5-shot classification on MiniImageNet are presented in Table A3 and summarized in Figure A1. We can see that Cooperation training is the most successful here for all ensemble sizes <20<20 and other training strategies that introduce diversity tend to perform worse. This happens because single networks are far from overfitting the training set (as opposed to the case with 224x224 input size) and forcing diversity acts as harmful regularization. In contrary, cooperation training enforces useful learning signal and helps ensemble members achieve higher accuracy. Only for n=20n=20 where diversity matters more, robust ensembles perform the best.

Few-shot Classification with WideResNet28 on 80x80 images on mini-ImageNet. Results for 1- and 5-shot classification on MiniImageNet are presented in Table A2 and summarized in Figure A1. In this case we can see again that Diverse training does not help since the networks do not memorize the training set. Robust ensembles outperform other training regimes emphasizing the importance of the proposed solution that generalizes across architectures.