The Lottery Tickets Hypothesis for Supervised and Self-supervised Pre-training in Computer Vision Models

Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, Zhangyang Wang

Introduction

Deep neural networks pre-trained on large-scale datasets prevail as general-purpose feature extractors . Moving beyond the most traditional greedy unsupervised pre-training , the most popular pre-training in computer vision (CV) is arguably to train the model for supervised classification on ImageNet . Such supervised pre-training enables the network to learn a hierarchy of generalizable features ; it is widely acknowledged to not only benefit the subsequent fine-tuning on other visual classification datasets (especially in small datasets and few-shot learning ), but also to accelerate/improve the training for different, more complicated types of downstream vision tasks, such as object detection and semantic segmentation .

Several state-of-the-art self-supervised pre-training, such as simCLR and MoCo , have demonstrated that it is instead possible to use unlabeled data in pre-training. Their methods refer to no actual labels in pre-training, but instead leverage self-generated pseudo labels or contrasting augmented views . Impressively, self-supervised pre-training yields pre-trained weights with comparable or even better transferability and generalization, for various downstream tasks, compared to their supervised pre-training counterparts.

A few recent efforts have shown to successfully scale up pre-training in CV. That is perhaps most natural for self-supervised pre-training, since unlabeled images are cheap and easily accessible. Chen et al. investigated to boost simCLR with massive unlabeled data in a task-agnostic way, and pointed out the key ingredient to be the use of big (deep and wide) networks during pretraining and fine-tuning. The authors found that, the fewer the labels, the more this approach (task-agnostic use of unlabeled data) benefits from a bigger network. After fine-tuning, the big network is reduced into a much smaller one with little performance loss by using task-specific distillation. We additionally note the latest works suggesting that supervised fine-tuning can also scale up to larger models and datasets beyond ImageNet .

The extraordinary cost of pre-training can be amortized by transferring to many downstream tasks. However, such explosive sizes of pre-trained models can even make fine-tuning computationally demanding, urging us to ask: can we aggressively trim down the complexity of pre-trained models, without damaging their downstream transferability? Note that, the question asked is drastically different from the conventional scope of model compression in CV, where a model is trained, compressed and/or tuned on the same dataset and specific task. In comparison, any simplification for a pre-trained model has to ensure its intact transferability to a variety of possible downstream tasks.

To address this research gap, we turn our attention to lottery ticket hypothesis (LTH) , a fast-rising field that investigates the sparse trainable subnetworks within full dense networks. The original LTH demonstrated small-scale networks contain sparse matching subnetworks capable of training in isolation from initialization to full accuracy. In other words, we could have trained smaller networks from the start if only we had known which subnetworks to choose. Recent investigations showed those matching subnetworks to transfer between related classification tasks. However, no study has closely examined the tantalizing possibility of universal transferability in LTH for CV models, i.e., if we treat the pre-trained weights as our initialization, whether matching subnetworks still exist in the pre-training models, that also enjoy the same downstream transfer performance? Are there universal subnetworks that can transfer to many tasks with no degradation in performance?

The paper carries out the first comprehensive experimental study to seek these desired universal matching subnetworks, from both supervised and self-supervised pre-trained CV models. Our principled methodology bridges pre-training and LTH from two perspectives: i) Initialization via pre-training. In the previous larger-scale settings of LTH for CV , the matching subnetworks are found at an early point in training. Instead, we aim to identify these matching subnetworks from dense pre-trained models (self-supervised or supervised), which creates an initialization directly amenable to sparsification. ii) Transfer learning. Finding the matching subnetwork is an expensive investment, usually costing multiple rounds of pruning and re-training. To justify this extra investment, the found subnetwork must be able to be reused by various downstream tasks, as illustrated in Figure 1.

The course of this study presents the following findings:

Using iterative unstructured magnitude pruning , we identify matching sub-networks up to 67.23%67.23\%, 59.04%59.04\%, 95.60%95.60\% sparsity, at pre-trained weights from ImageNet-equipped supervised pre-training, simCLR and MoCo, respectively. We also find matching subnetworks at pre-trained initialization with sparsity from 73.79%73.79\% to 98.20%98.20\% in a variety of classification, detection and segmentation downstream tasks.

Subnetworks at 67.23%67.23\%, 59.04%59.04\% and 59.04%59.04\% sparsity, found respectively using supervised ImageNet, simCLR and MoCo pre-training, are universally transferable to diverse downstream classification tasks with nearly the same accuracies.

Subnetworks at 73.79%73.79\%/48.80%48.80\%, 48.80%48.80\%/36.00%36.00\% and 73.79%73.79\%/83.22%83.22\% sparsity, found respectively by supervised ImageNet, simCLR and MoCo, can transfer to downstream detection/segmentation tasks without sacrificing performance.

Unlike previous matching subnetworks found at random initialization or early in training, we show that those identified at pre-trained initialization are more sensitive to structure perturbations. Also, different pre-training ways tend to yield diverse mask structures and perturbation sensitivities.

Lastly, pruning from larger pre-trained models can also produce better transferable matching subnetworks.

Practically speaking, this work sets the first step toward replacing large pre-trained models with smaller subnetworks, enabling much more efficient downstream tuning without inhibiting transfer performance. As pre-training becomes increasingly central in the CV field, our results shed light on the relevance of LTH in this new paradigm.

Related Works

A trained deep network could be pruned of excess capacity . Pruning algorithms can be grouped into unstructured and structured : the former sparsifies based on weight magnitudes; while the latter considers hardware-friendliness by removing channels and so on.

The discovery of LTH deviates from the convention of after-training pruning, and points to the existence of independently trainable sparse subnetworks from scratch that can match the performance of dense networks. Follow-up investigations scale up LTH by rewinding approaches , that re-initializes the subnetwork from the early training stage checkpoint rather than from scratch. LTH has been widely explored in image classification , natural language processing , generative adversarial networks , graph neural networks , and reinforcement learning . Most of them adopt (iterative) unstructured weight magnitude pruning . pioneer to study the transferability of the subnetworks identified on one image classification task to another. However, studying the universal transferability of LTH at pre-trained initializations among diverse CV tasks remains untouched.

One most relevant work to ours is from the natural language processing (NLP) field: the authors found universally transferable sparse matching subnetworks (at 40% to 90% sparsity), from the pre-trained initialization of BERT models . Finding their work inspiring, we stress that transplanting their NLP findings to our CV models is highly nontrivial due to multiple barriers: (1) pre-training BERT in uses only a self-supervised objective called “masked language model” (MLM) , while pre-training CV models has a significant variety of popular options, ranging from the supervised fashion , to self-supervision yet with numerous objectives ; (2) BERT models consist of self-attention and fully-connected sub-layers, differing much from the standard convolutional architectures in CV; (3) further complicating the issue is that different CV downstream tasks are known to rely on different priors and invariances; for example, while classification often calls on shift invariance, detection assumes location shift equivariance . That questions the feasibility of asking for one mask to transfer among them all. Such complicacy is well manifested by our delicate observations.

Pre-training in Computer Vision.

Supervised ImageNet pre-training has been a main CV workhorse . The recent surge of self-supervised pre-training suggest the potential of unlabeled data; examples include recovering the artificially corrupted inputs , predicting pseudo-labels , or contrasting augmented views . The state-of-the-art simCLR and MoCo pre-training can reduce the amount of labels needed for tuning downstream image classifiers, by two magnitudes.

Pre-trained networks are usually subsequently fine-tuned, with the architectures unchanged. One exception is which is the first to adapt the backbone architecture to fit different target datasets. It pre-trains a large super-net that contains many weight-shared sub-nets that can individually operate.

Preliminaries and Setups

In this section, we provide the detailed experimental settings and our approaches to find matching subnetworks.

Pre-training.

For the supervised pre-training, we use the official pre-trained ResNet-50The official Pytorch model zoo at https://pytorch.org/docs/stable/torchvision/models.html on the ImageNet dataset . For the self-supervised, we adopt the pre-trained ResNet-50 models with simCLRThe official simCLR model zoo at https://github.com/google-research/simclr and MoCov2The official MoCov2 model zoo at https://github.com/facebookresearch/moco on ImageNet.

Datasets, Training and Evaluation.

All pre-training experiments are conducted on ImageNet. For downstream tasks, we consider classification, object detection and semantic segmentation on multiple datasets. We use four natural image and one synthetic datasets to verify the transferability on classification: Fashion-MNIST , SVHN , CIFAR-10 , CIFAR-100 , and VisDA2017 . These datasets vary remarkably in terms of sample size, color space, resolution, image source, and classes. Following , we train object detection models on the combined training and validation set of Pascal VOC 2012 and Pascal VOC 2007 , then evaluate them on the Pascal VOC 2007 test set. We train and evaluate semantic segmentation models on Pascal VOC 2012 training and validation sets. We follow the standard hyperparameters and evaluation metricsFor detection experiments, we report the other evaluation metrics, AP50 and AP75 in the supplement. The technical details of calculating the retrieval accuracy for simCLR and MoCo pre-training tasks are also included in the supplement. for all pre-training and downstream tasks, as in Table 1.

Subnetworks.

1. Matching subnetworks. Following the definition in , a subnetwork f(x;m⊙θ,γ)f(x;m\odot\theta,\gamma) is matching if it satisfies the following condition:

That is, matching subnetworks perform no worse than the full dense models under the same training algorithm AtT\mathcal{A}_{t}^{\mathcal{T}} and evaluation metric ET\mathcal{E}^{\mathcal{T}}.

2. Winning ticket. If f(x;m⊙θ,γ)f(x;m\odot\theta,\gamma) is a matching subnetwork with θ=θp\theta=\theta_{p} for AtT\mathcal{A}_{t}^{\mathcal{T}}, it is a winning ticket for AtT\mathcal{A}_{t}^{\mathcal{T}}.

3. Universal subnetwork. A subnetwork f(x;m⊙θ,γTi)f(x;m\odot\theta,\gamma_{\mathcal{T}_{i}}) with task-specific configurations of γTi\gamma_{\mathcal{T}_{i}}, is universal for tasks {Ti}i=1N\{\mathcal{T}_{i}\}_{i=1}^{N} if and only if it is matching for each AtiTi\mathcal{A}^{\mathcal{T}_{i}}_{t_{i}}. The task set {Ti}i=1N\{\mathcal{T}_{i}\}_{i=1}^{N} could be a group of (diverse) downstream tasks, such as classification, detection and segmentation.

Pruning Methods.

To find the subnetworks f(x;m⊙θ,γ)f(x;m\odot\theta,\gamma), we adopt the classical iterative magnitude pruning (IMP) approach that is commonly used by the LTH literature . We prune the network by first training the unpruned dense network to completion on a task T\mathcal{T} (i.e., applying AtT\mathcal{A}^{\mathcal{T}}_{t}) and then removing a portion of weights with the globally smallest magnitudes . As revealed by previous works, in order to identify the most competitive matching subnetworks, the process needs to be iteratively repeated for several rounds. Algorithm 1 outlines the full IMP procedure in the supplement.

Although beyond the current scope, our future work plans to examine the practical speedup results on a hardware platform for our training and/or inference phases. For example, in the range of 70%-90% unstructured sparsity, XNNPACK has already shown significant speedups over dense baselines on smartphone processors. Integrating structured pruning will be another future direction of our interest .

Transfer of Pre-training Winning Tickets

In this section, we first show that there exist winning tickets using the pre-trained initialization on both self-supervised and supervised pre-training tasks. As shown in Figure 2, we find winning tickets with 67.23%67.23\%, 59.04%59.04\% and 95.60%95.60\% sparsity for supervised ImageNet, self-supervised simCLR and MoCo pre-training tasks.

Then, we investigate to what extent IMP subnetworks found for pre-training tasks can (universally) transfer to different downstream tasks. We ask the following questions:

Q1: Are winning tickets f(x;mP⊙θp,⋅)f(x;m_{\mathcal{P}}\odot\theta_{p},\cdot), found on the pre-training task P\mathcal{P}, also winning tickets for other downstream tasks T\mathcal{T}?

Q2: Are there common patterns in the transferability of winning tickets from different pre-trainings (e.g., supervised versus self-supervised)?

Q3: Can the transferred subnetworks f(x;mP⊙θp,⋅)f(x;m_{\mathcal{P}}\odot\theta_{p},\cdot) outperform the subnetworks f(x;mT⊙θi,⋅)f(x;m_{\mathcal{T}}\odot\theta_{i},\cdot) (θi∈{θ0,θ5%,θp}\theta_{i}\in\{\theta_{0},\theta_{5\%},\theta_{p}\}Early weight rewinding improves the quality of found matching subnetworks. As indicated by , the best rewinding points usually lie in the first 1%∼5%1\%\sim 5\% training epochs. We take 5%5\% for default comparison.), found on a specific task T\mathcal{T}?

Subnetworks f(x;mT⊙θp,⋅)f(x;m_{\mathcal{T}}\odot\theta_{p},\cdot), found on a specific downstream task with pre-trained weights, can be considered as “performance upbound” for all our IMP subnetworks. f(x;mT⊙θp,⋅)f(x;m_{\mathcal{T}}\odot\theta_{p},\cdot) is identified as matching subnetworks with the sparsity (98.20%98.20\%, 91.41%91.41\%, 73.79%73.79\%) for CIFAR-10, (91.41%91.41\%, 91.41%91.41\%, 20.00%20.00\%) for CIFAR-100, (91.41%91.41\%, 95.60%95.60\%, 91.41%91.41\%) for SVHN, (89.26%89.26\%, 96.48%96.48\%, 73.79%73.79\%) for Fashion-MNIST, and (73.79%73.79\%, 59.04%59.04\%, 67.23%67.23\%) for VisDA2007.

2 Transfer to Detection and Segmentation

Training detection and segmentation models commonly starts from pre-trained initializations . We compare the transferred subnetworks with (mP,θpm_{\mathcal{P}},\theta_{p}) versus the downstream task subnetworks with (mT,θpm_{\mathcal{T}},\theta_{p}), as shown in Figure 5. Observations are organized as follows:

Figure 5 demonstrates it is manageable to find transferable winning tickets on the detection and segmentation with the sparsity (73.79%73.79\%, 48.80%48.80\%, 73.79%73.79\%) and (48.80%48.80\%, 36.00%36.00\%, 83.22%83.22\%) for supervised ImageNet pre-training, self-supervised simCLR and MoCo pre-training tasks respectively.

Analyzing Properties of Pre-training Tickets

We also calculate the number of completely pruned (zero) kernels of subnetworks in Figure 6, which roughly reveals the weight clustering status in the sparse models. We observe that the remaining weights of subnetworks identified on the MoCo pre-training task are more clustered (i.e. more zero kernels) than the ones from ImageNet and simCLR, until reaching an extreme sparsity like 95.60%95.60\%.

Specifically, we provide kernel-wise heatmap visualizations of subnetworks with 79.03%79.03\% sparsity in Figure 7. We find that the completely pruned (zero) kernels are mainly clustered in the early layers of subnetworks, and appear rarely in the later layers. Among three kinds of subnetworks, the one from MoCo has the most dispersed distribution of completely pruned kernels. In general, more structured sparse subnetworks (i.e., more all-zero kernels) may have a stronger potential for hardware speedup .

2 Pre-training versus Random Initialization

Comparing the randomly pruned subnetworks in Figure 8, we observe that pre-trained initialization consistently benefits the accuracy until subnetworks reaching some high sparsity (e.g., 67.23%67.23\%). After that, the performance of random pruned subnetworks is no longer affected by different initializations.

3 More Ablation Studies for Pre-training

reveals that heavily compressed, large transformer models achieve higher performance than lightly compressed, small transformer models in natural language processing. We re-confirm this claim for self-supervised simCLR pre-training, in terms of the transferabilityIn the supplement, we also report the pre-training task performance of subnetworks generated from small- and large-scale pre-trained simCLR. of found matching subnetworks.

In Figure 9, with the same number of remaining weights, subnetworks pruned from simCLRFor a fair comparison, here we adopt the simCLRv2 pre-trained ResNet-152 and ResNet-50 models, since only simCLRv2 released the official pre-trained ResNet-152 model. pre-trained ResNet-152, achieve consistently superior accuracy on the downstream CIFAR-100 task than the ones from simCLR pre-trained ResNet-50 (around one-third size of ResNet-152). At least for simCLR, pruning from larger pre-trained models produces better transferable matching subnetworks.

Our observation is also aligned with the advocates of , to first pretrain a big model and then compress it. The key difference is that, uses standard model compression (knowledge distillation) after downstream fine-tuning is done; in contrast, our results can be seen as a possible second pre-training stage: after the initial pre-training (and before any fine-tuning), performing IMP to find equally-capable matching subnetwork with far fewer parameters.

Temperature Hyperparameter.

The temperature scaling hyperparameter is known to play a significant role in the quality of the simCLR pre-training . It motivates us to investigate the impact of the temperature scaling factor on the transferability of pre-training winning tickets found in Section 4. Without loss of the generality, we consider the subnetworks with the sparsity from 67.23%67.23\% to 73.79%73.79\%. Specifically, we start from training subnetworks at the sparsity level 67.23%67.23\% for 1010 epochs, on the simCLR task with different temperature scaling factors. Then, they are pruned to the level of 73.79%73.79\% sparsity by IMP. Finally, subnetworks are fine-tuned and evaluated on the downstream CIFAR-100 task. Results in Table 2 show that found subnetworks have close transfer performance if the temperature scaling factor lies in a moderate range (i.e., [0.1,0.50.1,0.5]), and the performance will degrade at extreme temperatures (e.g., 20.0).

Conclusion

We study the lottery ticket hypothesis in the context of CV pre-training, via both supervised (e.g., ImageNet classification) and self-supervised (e.g., simCLR and MoCo) ways. Despite the complicacy of our goal, by performing IMP from the pre-trained initializations, we are consistently able to find matching subnetworks at non-trivial sparsity levels, that can be independently trained to full model performance, on both pre-training and downstream tasks. We also present a detailed discussion of cross-task universal transferability.

Acknowledgments

We are grateful fpr the MIT-IBM Watson AI Lab, in particular John Cohn for generously providing the computing resources necessary to conduct this research. Wang’s work is in part supported by the NSF Energy, Power, Control, and Networks (EPCN) program (Award number: 1934755), and by an IBM faculty research award.

References

Appendix A More Technical Details

Following the routines in previous LTH works, the algorithm 1 outlines the full iterative magnitude pruning (IMP) procedure.

A.2 Top-1 Retrieval Accuracy

Here we presents the detailed calculation of top-1 retrieval accuracy for self-supervised pretraining tasks, including simCLR and MoCo . Given a batch of data with nn samples, {z1,⋯ ,zn}\{z_{1},\cdots,z_{n}\} and {z1′,⋯ ,zn′}\{z^{\prime}_{1},\cdots,z^{\prime}_{n}\} donates the feature representations from the two branches of simCLR or MoCo models. ziz_{i} and zi′z^{\prime}_{i} are computed from the same input sample with different data augmentations.

Appendix B More Experimental Results

B.2 Faster RCNN and SSD Detection Results

We notice that detection results of Faster RCNN and SSD with the simCLR pre-training show inferior and unsatisfactory performances, compared with the reported number in BYOL . Although BYOL is implemented with Tensorflow (we use Pytorch) and also has an extra residual block for backbone network, the performance gap is not neglectable. To address this, multiple authors have worked to carefully tune all hyperparameters (learning rate, batch size, training iterations), and thoroughly compared implementation details side-to-side (batch norm, input resolution, etc.). However, we still cannot close the gap. Hence while our results on MoCo and ImageNet are very consistent, we cannot exclude the marginal possibility that simCLR implementation is specifically sensitive to Pytorch versus Tensorflow frameworks (unfortunately, not uncommon) for some reason. Therefore, we put Faster RCNN and SSD detection results with the simCLR pre-training in the appendix as failure cases, and note that it hardly affects any of our main observations/conclusions.

B.3 Ablation about Larger Pre-training Models

Figure 12 collects the pre-training task performance of subnetworks generated from small- and large-scale pre-trained simCLR models. We observe that heavily compressed, large simCLR models (e.g., ResNet-50) obtain superior performance to lightly compressed, small simCLR models (e.g., ResNet-152), which is consistent with . However, subnetworks found on the small-scale pre-trained simCLR model show a slightly better top-1 retrieval accuracy after the sparsity approaches an extreme level.