Ranking Neural Checkpoints
Yandong Li, Xuhui Jia, Ruoxin Sang, Yukun Zhu, Bradley Green, Liqiang Wang, Boqing Gong
Introduction
There is an increasing number of pre-trained deep neural networks (DNNs), which we call checkpoints. We may produce hundreds of intermediate checkpoints when we sweep through various learning rates, optimizers, and losses to train a DNN. Furthermore, semi-supervised and self-supervised learning make it feasible to harvest DNN checkpoints with scarce or no labels. Fine-tuning has become a de facto standard to adapt the pre-trained checkpoints to target tasks. It leads to faster convergence and better performance on the downstream tasks.
However, not all checkpoints are equally useful for a target task, and some could even under-perform a randomly initialized checkpoint (cf. Section 2.2). This paper is concerned with ranking neural checkpoints, which aims to measure how effectively fine-tuning can transfer knowledge from the pre-trained checkpoints to the target task. The measurement should be generic enough for all the neural checkpoints, meaning that it works without knowing any pre-training details (e.g., pre-training examples, hyper-parameters, losses, early stopping stages, etc.) of the checkpoints. It also should be lightweight, ideally without training on the downstream task, to make it practically useful. We may use the measurement to choose the top few checkpoints before running fine-tuning, which is computationally more expensive than calculating the measurements.
Ranking neural checkpoints is crucial. Some domains or applications lack large-scale human-curated data, like medical images , raising a pressing need for high-quality pre-trained checkpoints as a warm start for fine-tuning. Fortunately, there exist hundreds of thousands of checkpoints of popular neural network architectures. For instance, many computer vision models are built upon ResNet , Inception-ResNet , and VGG . As a result, we can construct a candidate pool by collecting the checkpoints released by different groups, for various tasks, and over distinct datasets.
It is nontrivial to rank the checkpoints for a downstream task. We explain this point by drawing insights from the related, yet arguably easier, task transferability problem , which aims to provide high-level guidance about how well a neural network pre-trained in one task might transfer to another. However, not all checkpoints pre-trained in the same source task transfer equally well to the target task . The pre-training strategy also matters. Zhai et al. find that combining supervision with self-supervision improves a network’s transfer results on downstream tasks. He et al. also show that self-supervised pre-training benefits object detection more than its supervised counterpart under the same fine-tuning setup.
We may also appreciate the challenge in ranking neural checkpoints by comparing it with another related line of work: predicting DNNs’ generalization gaps . Jiang et al. use a linear regressor to predict a DNN’s generalization gap, i.e., the discrepancy between its training and test accuracies, by exploring the training data’s margin distributions. Other signals studied in the literature include network complexity and noise stability. Ranking neural checkpoints is more challenging than predicting a DNN’s generalization gap. Unlike the training and test sets that share the same underlying distribution, the downstream task may be arbitrarily distant from the source task over which a checkpoint is pre-trained. Moreover, we do not have access to the pre-training data at all. Finally, instead of keeping the networks static, fine-tuning dramatically changes all weights of the checkpoints.
We establish a neural checkpoint ranking benchmark (NeuCRaB) to study the problem systematically. NeuCRaB covers various checkpoints pre-trained on widely used, large-scale datasets by different training strategies and architectures at a range of early stopping stages. It also contains diverse downstream tasks, whose training sets are medium-sized, making it practically meaningful to rank and fine-tune existing checkpoints. Pairing up all the checkpoints and downstream tasks, we conduct careful fine-tuning with thorough hyper-parameter sweeping to obtain the best transfer accuracy for each checkpoint-downstream-task pair. Hence, we know the groundtruth ranking of the checkpoints for each downstream task according to the final accuracies (over the test/validation sets).
A functional checkpoint ranking measurement should be highly correlated with the groundtruth ranking and, equally importantly, incurs as low computation cost as possible. We study several intuitive methods for ranking the neural checkpoints. One is to freeze the checkpoints as feature extractors and use a linear classifier to evaluate the features’ separability on the target task. Another is to run fine-tuning for only a few epochs (to avoid heavy computation) and then evaluate the resulting networks on the target task’s validation set. We also estimate the mutual information between labels and the features extracted from a checkpoint.
Finally, we propose a lightweight measure, named Gaussian LEEP (LEEP), to rank checkpoints based on the recently proposed log expected empirical prediction (LEEP) . LEEP was originally designed to measure between-task transferabilities. It cannot handle the checkpoints pre-trained by unsupervised or self-supervised learning since it requires all checkpoints to have a classification head. Its computation cost could blow up when the classification head corresponds to a large output space. Moreover, it depends on the classification head’s probabilistic output, which, unfortunately, is often overly confident .
To tackle the above problems, we replace the checkpoints’ output layer with a Gaussian mixture model (GMM). This simple change kills two birds with one stone. On the one hand, GMM’s soft assignment of input to clusters seamlessly applies to LEEP, resulting in the lightweight, effective LEEP measure that works regardless of the checkpoints’ output types. On the other hand, since we fit GMM to the target task’s data, instead of the pre-training data of a different source task, the cluster assignment probabilities are likely more calibrated than the classification probabilities for the target task, if there exist classification heads.
The Neural Checkpoint Ranking Benchmark (NeuCRaB)
which defines the groundtruth ranking list for task .
Denote by all measures that return a ranking score for any checkpoint-task pair under a computation budget . A measure gives rise to the following ranking scores for a task ,
where we underscore the computation budget in the measure .
Our objective in ranking neural checkpoints is to find the best ranking measure in expectation,
where is a metric evaluating the ranking scores against the test accuracies . Section 2.3 details the evaluation methods used in this work. Equipped with such a ranking measure , we can identify the checkpoints that potentially transfer to a downstream task better than the others without resorting to heavy computation.
Following the design principle of , we study diverse downstream tasks including Caltech101 , Flowers102 , Sun397 , and Patch Camelyon . These tasks are representative of general object recognition, fine-grained object recognition, scenery image classification, and medical image classification, respectively. Table 1 in Appendix A.1 provides more details of these tasks. A common theme is that their training sets are all medium-sized, making it especially beneficial to leverage pre-trained checkpoints to avoid overfitting.
2 Neural Checkpoints 𝒞𝒞\mathcal{C}
Thanks to the broad use of DNNs, one may collect neural checkpoints of various types from multiple sources. To simulate this situation, we construct a rich set of checkpoints and separate them into three groups according to the pre-training strategies and network architectures.
Group I: Checkpoints of mixed supervision. The first group of checkpoints are pre-trained with mixed supervision till convergence, including supervised learning, self-supervised learning, semi-supervised learning, and the discriminators or encoders in deep generative models. It consists of 16 ResNet-50s . We borrow 14 models pre-trained on ImageNet from . Among them, four are pre-trained by self-supervised learning (Jigsaw , Relative Patch Location , Exemplar , and Rotation ), six are the discriminators of generative models (WAE-UKL , WAE-GAN, WAE-MMD , Cond-BigGAN, Uncond-BigGAN , and VAE ), two are based on semi-supervised learning (Semi-Rotation-10% and Semi-Exemplar-10% ), one is by fully supervised learning (Sup-100%-Img ), and one is trained with a hybrid supervised loss (Sup-Exemplar-100% ). We also add two supervised checkpoints pre-trained on iNaturalist (Sup-100%-Inat) and Places365 (Sup-100%-Pla) , respectively. Using the evaluation procedure (cf. equation (1)), we obtain their final accuracies on the downstream tasks described in Section 2.1.
Figure 1 shows the best fine-tuning accuracies offset by their mean for better visualization, and Table 4 (in Appendix) contains the absolute accuracy values. We include the training from scratch (From-Scratch) for comparison. Most of the checkpoints yield significantly better fine-tuning results than From-Scratch. Some of the discriminators in generative models, however, under-perform From-Scratch. The highest-performance checkpoints change from one downstream task to another.
Group II: Checkpoints at different pre-training stages. This group comprises 12 ResNet-50s pre-trained by fully supervised learning on ImageNet, iNaturalist, and Places-365. We save a checkpoint right after each learning rate decay, resulting in four checkpoints per dataset. Figure 2 and Table 5 in Appendix show the best fine-tuning accuracies over the four downstream tasks, where Img-90k refers to the checkpoint trained on ImageNet for 90k iterations. Interestingly, the downstream tasks favor different pre-training sources, indicating the necessity of studying between-task transferabilities . However, the source task information may be not known for all checkpoints. Moreover, the converged model over a source task does not always transfer the best to a downstream task (cf. Img-270k vs. Img-300k on Camelyon, Inat-270k vs. Inet-300k on Flowers102, etc.). We hence construct this NeuCRaB for studying the ranking of neural checkpoints without accessing how one pre-trained the checkpoints over which dataset.
Group III: Checkpoints of heterogeneous architectures. Kornblith et al. show that better network architectures can learn better features that can be transferred across vision-based tasks. Therefore, we construct the third group of checkpoints by using different neural architectures. Four of them belong to the Inception family , one is Inception-ResNet-v2 , six come from the MobileNet family , and two are from the ResNet-v1 family . We train them on ImageNet till convergence. Figure 3 and Table 6 in Appendix visualize their fine-tuning accuracies on the four downstream tasks.
3 Evaluation Metrics ℳℳ\mathcal{M}
We use multiple metrics (cf. in eq. (3)) to evaluate the checkpoint ranking measures.
A practitioner may have resources to test up to checkpoints for their task of interest. We consider it a success if a measure ranks the highest-performance checkpoint into the top . A measure’s Recall@ is the ratio between the number of downstream tasks on which it succeeds and the total number of tasks. We employ and in the experiments.
Given a task, a ranking measure returns an ordered list of the checkpoints. If the measure selects a high-performing checkpoint to the top despite that it misses the highest-performance one, we do not want to overly penalize it. This Rel@ is the ratio between the best fine-tuning accuracy on the downstream task with the top checkpoints and the the best fine-tuning accuracy with all the checkpoints.
We incorporate Pearson’s to compute the linear correlation between a measure’ ranking scores and the evaluation procedure’s final accuracies .
We also include Kendall’s to measure the ordinal association between a ranking measure and the evaluation procedure for each task. After all, what matter is the order of the checkpoints rather than the precise ranking scores.
Checkpoint Ranking Methods
In this section, we describe some intuitive neural checkpoint ranking methods. These methods strive to achieve high correlation with the checkpoint evaluation procedure at low computation cost.
If there is no constraint over computing, the evaluation procedure itself becomes the gold ranking measure. Hence, a natural ranking method is the fine-tuning with early stopping, by which the model is far from convergence. The premature models’ test accuracies are the ranking scores. Experiments reveal that it is hard to forecast from the premature models.
2 Linear Classifiers
We derive the second ranking method also from the evaluation procedure , which replaces a checkpoint’s output layer by a linear classifier tailored for the downstream task. We train the linear classifier while freezing the other layers. The ranking score equals the classifier’s test accuracy. It is worth mentioning that self-supervised learning often adopts this practice as well to evaluate the learned feature representations. We shall see that the linear separability of the features extracted from a checkpoint is a strong indicator of the performance of fine-tuning the full checkpoint.
3 Mutual Information
Suppose the extracted features’ quality well correlates with a checkpoint’s final accuracy on a downstream task. Besides the linear separability above, we can rank the checkpoints by their mutual information between the high-dimensional features and discrete labels of the downstream task. We employ the state-of-the-art mutual information estimator , where controls the trade-off between variance and bias. It is a variational lower bound parameterized by a neural network. Belghazi et al. report that the neural estimators generally outperform prior mutual information estimations, especially when the variables are high-dimensional. We use the code released by the authors to calculate .
4 LEEP for the Checkpoints with Classification Heads
To rank the checkpoints pre-trained over classification source tasks, the recently proposed LEEP measure is directly applicable despite that it was originally designed for between-task transfer. Denote by the classification space of a checkpoint . We can interpret , the -th (softmax) output element, as the probability of classifying the input into the class . Given a downstream task and its test set , the LEEP ranking score for the checkpoint is calculated by
where is the empirical conditional distribution of the downstream task’s label given the source label , and is a “dummy” classifier, which firstly draws a label from the checkpoint and then draws a class from the conditional distribution .
which gives rise to the conditional distribution .
In the experiments, LEEP and the linear classifier are the second best ranking methods for the checkpoints pre-trained for classification. However, LEEP’s computation cost is high when a checkpoint’s classification output is high-dimensional (e.g., iNaturalist contains more than 8000 classes). Besides, its softmax estimation of the classification probability is often poorly calibrated . Finally, it does not apply to the checkpoints with no classification heads.
5 𝒩𝒩\mathcal{N}LEEP
We propose a variation to LEEP that applies to all types of checkpoints including those obtained from unsupervised learning and self-supervised learning. It can also avoids the overly confident softmax.
Feeding the training data of a downstream task into a checkpoint, we obtain their feature representations. The representations are thousands of dimensions, depending on the checkpoint’s neural architecture. We reduce their dimension by using the principal component analysis (PCA). Denote by the resultant low-dimensional representation of the input .
which is arguably more reliable than the class assignment probability output by the softmax classifier because we fit GMM to the downstream task’s training data, whereas the softmax classifier is learned from a different source task.
Hence, we arrive at an improved ranking measure, named LEEP, by replacing , the probability of classifying an input to the class , in equations (4–5) by the posterior distribution .
Experiments on NeuCRaB
There are free parameters in each of the ranking methods. Before presenting the main results, we study how the free parameters in LEEP affect its checkpoint ranking performance. Figure 2 illustrates LEEP’s Kendall’s values over Groups I and II with different PCA feature dimensions and the numbers of Gaussian components. Each Kendall’s is averaged across all the downstream tasks; the higher, the better. Along the vertical axes, we change the feature dimensions by keeping different percentages of the PCA energies; PCA50 means the percentage is 50%. Along the horizontal axes, we adopt different numbers of Gaussian components in GMM; means the number is twice the class number of the downstream task. Notably, the Kendall’s values remain relatively stable. In the remaining experiments with LEEP, we fix the PCA energy to 80% and the number of Gaussian components five times the class number of a downstream task.
Tables 1, 2, and 3 show the checkpoint ranking methods’ performance on Groups I (checkpoints of mixed supervision), II (different pre-training stages), and III (heterogeneous architectures), respectively. We also union the three groups and present the corresponding ranking performance in Table 2 in Appendix. The numbers in the tables are the average over all downstream tasks. In addition to the evaluation metrics detailed in Section 2.3, the GFLOPS column measures the ranking methods’ computing performance; the lower, the better.
We report multiple variations of the ranking methods in the tables. Fine-tuning is computationally expensive, so we stop it after one or five epochs. The linear classifiers are less so as we save the feature representations of downstream tasks’ after one forward pass to the checkpoints. We report the linear classifiers’ ranking results after training them for one epoch, five epochs, and convergence. We test and in the mutual information estimator. Additionally, we experiment with after reducing the feature dimensions by using PCA.
2 Main Findings
In each column of Tables 1, 2, 3, and Table 2 in Appendix, we highlight the best and second best by the bold font and underscore, respectively.
The mutual information fails to rank high-performing checkpoints to the top and even produces negative Pearson and Kendall correlations, probably because of the features’ high dimensions. Reducing the feature dimensions by PCA significantly improves the mutual information’s ranking performance; MI w/ PCA (=0.01) leads to the second best Rel@1, Recall@3 and Rel@3 among the ranking methods in Group III, the checkpoints of heterogeneous neural architectures. Varying in the mutual information estimator can control the trade-off between variance and bias. MI w/ and w/o PCA (=0.01) perform better than MI w/ and w/o PCA (=0.50), respectively. It indicates that neural checkpoint ranking requires low-bias MI estimator since smaller means low-bias but high-variance estimation.
Fine-tuning up to some epochs turns out the worst ranking methods because it leads to low correlation with the groundtruth ranking and yet incurs heavy computation. Similarly, training the linear classifier up to one or five epochs does not perform well except in Group II. These results indicate that it is difficult to forecast the checkpoints’ final performance from premature models. Fine-tuning (5 epochs) and Linear (5 epochs) perform better than Fine-tuning (1 epoch) and Linear (1 epoch) in terms of Person and Kendall correlation, respectively. However, they all fail to select the top checkpoint in Group I and Group III since they produce lower Recall@1 and Recall@3 than others. One possible reason is that the evaluation accuracies of checkpoints in the early stage tend to have large variance.
Feature qualities before fine-tuning the checkpoints. If we train the linear classifiers till convergence, they become the best in Group II, and the second best checkpoint ranking method in Groups I and III in terms of Pearson and Kendall correlations. It can also produce better Recall@1 and Recall@3 than Linear (1 epoch) and Linear (5 epoch) in Groups I, II and III since the evaluation accuracies of converged models are more stable than models in the early training stage. Note that the linear classifiers’ accuracies, i.e., the ranking scores, imply the linear separability of the features extracted by the checkpoints. Recall that the mutual information with PCA feature dimension reduction is among the second best (Rel@1, Recall@3 and Rel@3) in Group III. Since both methods measure the feature representations’ quality by the downstream tasks’ labels, we conjecture that the quality of the features is a strong indicator of the checkpoints’ final fine-tuning performance on the downstream tasks. It would be interesting to study other feature quality measures beyond the linear separability and mutual information in future work.
LEEP performs consistently well in all the groups of checkpoints over all the evaluation metrics with the lowest computation cost . In contrast, the original LEEP measure is not applicable to Group I, the checkpoints of mixed supervision, because it requires that the checkpoints have a classification output layer. Overall, LEEP is the second best over all evaluation metrics among the ranking methods in Groups II and III, whose checkpoints all have a classification output layer. Specifically, LEEP can produce the second best Recall@1, Recall@3 and Rel@3 in Group II, and the best Recall@3, the best Rel@3 and the second best Kendall correlation in Group III. It is a more consistent indicator than fine-tuning, linear classifier, or MI based ranking methods. However, LEEP can not produce better results than LEEP, and it requires slightly larger GFLOPS due to the extra computation cost from the classification head.
We conjecture that LEEP outperforms LEEP mainly because GMMs calibrate the posterior probabilities better than the checkpoints’ softmax classifiers. The checkpoint ranking quality of LEEP score hinges on the performance of the ‘dummy classifier’ – , and is the key element to calculate it. However, can be poorly calibrated and it can not represent a true probability. In contrast, used in LEEP is indeed the probability that the sample belongs to one cluster from a mixture of Gaussian distributions and it can remedy the poor-calibrated problem in LEEP.
Computational costs. Moreover, we highlight the GFLOPS column in the tables. LEEP and LEEP exhibit a clear advantage over the other checkpoint ranking methods in terms of computing. The main reason is that LEEP and LEEP can avoid intensive computation from neural network training, and they only require one forward pass through the training data.
Comparing different groups of the checkpoints. Checkpoint ranking on different groups of checkpoints varies in degrees of difficulty. The most challenging group is Group III, the checkpoints of heterogeneous neural architectures. All the ranking methods produce lower correlations with the groundtruth ranking, and they can barely select the top checkpoints in this group. The main reason is that the neural architectures matter for transfer learning . Besides, heterogeneous neural architectures can demonstrate various performance even if we train them from scratch on downstream tasks. Ranking neural checkpoints by the feature representations of the last layer is not sufficient for those checkpoints. We may explore more advanced ranking methods considering the structures of the deep neural networks in the future.
Checkpoint ranking on Group II is easier than on Group I since all the ranking methods can achieve relatively better results over all evaluation metrics in Group II. The results indicate that checkpoints with various training strategies (Group I) can bring more complex knowledge from source domains, comparing with checkpoints with different early stopping stages (Group II). In addition, fine-tuning the entire models and training linear classifiers up to one or five epochs perform significantly better on Group II since those ranking methods are based on early stopping as well.
Additional experiments in the supplementary materials. To simulate a sufficiently large pool of checkpoints in the real applications, we finally combine the checkpoints in Group I, II, and III into one large group and conduct checkpoint ranking experiments on it. We also add one more group of checkpoints with ResNet-101s to evaluate the checkpoint ranking on deeper models. Please see more details in Appendix A.3 and A.4. We also take object detection and instance segmentation as downstream tasks and conduct preliminary experiments on VOC and Cityscapes . Please refer to Appendix A.6 to see detailed discussions.
Although the benchmark can be easily extended to many downstream tasks in other modalities, e.g., voice, text, and cross-modal modalities, we steer our attention into comparing several intuitive ranking measures on the variants of checkpoints, covering different training strategies, source domains, and architectures at a range of early stopping stages. We formalize the checkpoint ranking idea, demonstrate the existence of an effective yet lightweight measure, LEEP, and hope it can shed light on more efficient ranking methods and practical applications.
Related Work
Our work is broadly related to task transferability and neural networks’ generalization gap.
Task transferability. A task usually refers to a joint distribution over input and label. Task transferability aims to predict how well a deep neural network pre-trained on a source task transfers to the target task. One may estimate the task transferability by data similarities regardless of models being used. Some work in this line includes conditional entropy , data set distance as optimal transport , -relatedness , -distance , and discrepancy distance . Besides, Poole et al. derived information theoretic bounds. These methods are generally hard to compute in practice and rely on the availability of the source data. Some recent task transferability estimators involve both data and the models. Taskonomy is a fully computation method, where task similarity scores are obtained by transfer learning experiments. Dwivedi et al. analyzed the representation similarities to construct a task taxonomy. Besides the models trained on source tasks, all these methods also require a fine-tuned or independently trained model from the target task. In contrast, our work aims to find checkpoint ranking measures that are lightweight in computing and requires no access to the source tasks.
Recent works demonstrated that using pre-trained checkpoints that have similar feature representations as the target task’s representations can improve transfer learning . Song et al. employed attribution maps to compare two models and then quantified transferabilities by the similarity of two models. Those approaches all require a converged model on target datasets, incurring intensive computation. However, we want to design a lightweight method for ranking checkpoints, ideally without any training procedures.
Predicting neural networks’ generation gap. The difference between a model’s performance on the training data versus its performance on test data is known as the generalization gap. It is practically useful and theoretically impactful to predict a neural network’s generalization gap. Most recent work does so by finding a set of features that is predictive of the generalization, e.g., by estimating data margins . Jiang et al. and Yak et al. demonstrate how the margin signatures of a neural network can predict the generalization gap with small errors. Besides, the network complexity and noise stability are also useful cues . Our problem substantially differs from predicting the neural networks’ generalization gap, which is concerned with the training and test data sets that share the same underlying distribution. We instead care about the results after fine-tuning a network’s checkpoint.
Conclusion
Deep learning has triumphed over many fields in both research and real-world applications. There must exist hundreds of thousands of DNNs trained and released by various groups. To this end, it is natural to select an existing, promising DNN checkpoint as a warm start to a training procedure when solving a new task. How to identify useful checkpoints from a large pool for the target task? Towards answering this question, we present NeuCRaB, a thorough benchmark covering diverse downstream tasks and pre-trained DNN checkpoints, along with LEEP, a lightweight, effective checkpoint ranking measure.
The experiments with linear classifiers and mutual information (after PCA) reveal that the features extracted from the checkpoints are good indicators of the checkpoints’ potential in transfer learning. It is worth exploring other ways of evaluating the features’ quality in future work. It is also interesting to investigate the checkpoints’ inherent signatures, such as topology and stability to noise, which might be informative of their transferabilities. Finally, some learning-based methods in predicting networks’ generalization gaps are also promising for the checkpoint ranking problem.
References
Appendix A Appendix
In this appendix, we provide the following details to support the main text:
Training details of pre-training and fine-tuning.
Comparison results on the combined group of checkpoints in Groups I, II and III.
Another group of checkpoints with ResNet101s at different pre-training stages.
Neural checkpoints ranking on object detection and instance segmentation.
In this section, we describe the datasets used for the downstream tasks as shown in Table 4. More specifically, Caltech101 contains 101 classes, including animals, airplanes, chairs and etc, the image size varies from 200 to 300 pixels per edge. Flowers102 have 102 classes, with 40 to 248 training images per class, each image has at least 500 pixels. Patch Camelyon contains 327,680 images of histopathologic scans of lymph node sections with image size of 96x96, which is collected to predict the presence of metastatic tissue. Sun397 is a scenery benchmark with 397 classes, including cathedral, staircase, shelter, river, or archipelago. There are at least 100 images per class. The images are in 200x200 or higher resolutions. We believe the dataset portfolio well represents a broad set of vision tasks.
A.2 Hyper-parameter Sweep
We adopt the similar experiment setting as in to fine-tune the neural networks on the downstream tasks. Specifically, we set the batch size to 512 and use SGD with momentum of 0.9. We do not use weight decay for fine-tuning, and we set it to be 0.01 times the learning rate when training from scratch. We perform per-task hyper-parameter search. For each task, we sweep the learning rate in 0.0001, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.5 and the training step in 2500, 5000, 10000, 15000, 20000, 400000. We incorporate inception data augmentation for pre-training checkpoints and we do not use data-augmentation when we fine-tune the neural networks on the downstream tasks to emphasize the effect of transfer learning.
A.3 Comparison results on all checkpoints in Groups I, II, III
To obtain a comprehensive analysis, we also consolidate the checkpoints from Group I, II and III into one group (including 41 checkpoints in total) and then apply the ranking methods on it. Table 5 shows the comparison results. The results further evaluate our observations in Section 4 of the main text. LEEP performs consistently well on the big group of checkpoints with lowest computation cost. Linear separability of the feature representation is also a good indicator for ranking a large group of neural checkpoints. Fine-tuning with early stopping and mutual information estimator produce poor correlations. The ranking qualities of different ranking methods on the large group of checkpoints are in sharper contrast than on small groups. For instance, the Pearson’s of LEEP vs. Finetune (5 epochs) on the large group is 83.71 vs. 27.84 but they perform 72.84 vs. 68.47 on Group II (Table 2 in the main text). It indicates that LEEP is a low-variance and low-bias checkpoint ranking estimator, while early stopping may produce high-variance ranking results.
A.4 Group IV: Supervised ResNet101s
We incorporate another group of checkpoints, including 12 ResNet101 models pre-trained by fully supervised learning on ImageNet , iNaturalist , and Places-365 . We obtain the checkpoints in the same way as we have done for Group II, but with ResNet101 architecture. We want to study how different model architecture and model size affect the ranking quality.
Figure 3 and Table 10 show the fine-tuning accuracy on 4 downstream tasks. The relative fine-tuning accuracies are similar to the accuracies on Group II. We also observe that a converged checkpoint does not necessarily demonstrates the best performance on the downstream tasks (cf. Img-270k is better than Img-300k on Flowers102 ). Table 6 shows the comparison results of ranking methods on those checkpoints. The relative performance among the ranking methods is similar to what they do in Group II (Table 2 in the main text). Except that they perform better on ResNet101s, e.g., Linear (converged) can achieve 68.60 in terms of Kendall’s on ResNet50s versus 73.48 on ResNet101s, LEEP can get 72.84 in terms of Pearson’s on ResNet50s versus 83.22 on ResNet101s. The observation reveals that the ranking of deeper checkpoints may be more predictable than shallow ones.
A.5 More experimental results on Groups I-IV
We show more comparison results on NeuCRaB in this section. Figures 4 and 5 show the best fine-tuning accuracies offset by their mean (for better visualization) on Groups II and III, respectively. Table 7, 8, 9, 10 demonstrate the absolute best fine-tuning accuracies on Groups I-IV, respectively.
A.6 Neural checkpoint ranking for object detection and instance segmentation
We also evaluate on object detection and segmentation tasks, and show the results in Tables 11. Specifically, we incorporate the recent self-supervised MoCo models (MoCov1, MoCov2, MoCov2-800epoch) and a ResNet50 model (supervised pretrained on ImageNet) into a new group of checkpoints. We evaluate checkpoint ranking on Pascal VOC (object detection) and Cityscapes (instance segmentation). In order to adapt LEEP to detection and segmentation tasks, we assign multiple ground truth labels for one image if it includes multiple object categories and extract the image-level features to perform GMM. We adapt LEEP to detection and segmentation tasks by assigning multi-labels to images with multiple object categories. The experiment results demonstrate that LEEP consistently outperforms the fine-tune and linear evaluation based approachs. We plan to include more diverse downstream tasks in NeuCRaB to facilitate future research.