Adversarial Robustness: From Self-Supervised Pre-Training to Fine-Tuning

Tianlong Chen, Sijia Liu, Shiyu Chang, Yu Cheng, Lisa Amini, Zhangyang Wang

Introduction

Supervised training of deep neural networks requires massive, labeled datasets, which may be unavailable and costly to assemble . Self-supervised and unsupervised training techniques attempt to address this challenge by eliminating the need for manually labeled data. Representations pretrained through self-supervised techniques enable fast fine-tuning to multiple downstream tasks, and lead to better generalization and calibration . Examples of tasks proven to attain high accuracy through self-supervised pretraining include position predicting tasks (Selfie , Jigsaw ), rotation predicting tasks (Rotation ), and a variety of other perception tasks .

The labeling and sample efficiency challenges of deep learning are further exacerbated by vulnerability to adversarial attacks. For example, Convolutional Neural Networks (CNNs), are widely leveraged for perception tasks, due to high predictive accuracy. However, even a well-trained CNN suffers from high misclassification rates when imperceivable perturbations are applied the input . As suggested by , the sample complexity of learning an adversarially robust model with current methods is significantly higher than that of standard learning. Adversarial training (AT) , the state-of-the-art model defense approach, is also known to be computationally more expensive than standard training (ST). The above facts make it especially meaningful to explore:

Can appropriately pretrained models play a similar role for adversarial training as they have for ST? That is, can they lead to more efficient fine-tuning and better, adversarially-robust generalization?

Self-supervision has only recently been linked to the study of robustness. An approach is offered in , by incorporating the self-supervised task as a complementary objective, which is co-optimized with the conventional classification loss through the method of AT . Their co-optimization approach presents scalability challenges, and does not enjoy the benefits of pretrained embeddings. Further, it leaves many unanswered questions, especially with respect to efficient tuning, which we tackle in this paper.

This paper introduces a framework for self-supervised pretraining and fine-tuning into the adversarial robustness field. We motivate our study with the following three scientific questions:

Is an adversarially pretrained model effective in boosting the robustness of subsequent fine-tuning?

Which provides the better accuracy and efficiency: adversarial pretraining or adversarial fine-tuning?

How does the type of self-supervised pretraining task affect the final model’s robustness?

Our contributions address the above questions and can be summarized as follows:

We demonstrate for the first time that robust pretrained models leveraged for adversarial fine-tuning result in a large performance gain. As illustrated by Figure 1, the best pretrained model from a single self-supervised task (Selfie) leads to 3.83% on robust accuracyThroughout this paper, we follow to adopt their defined standard accuracy and robust accuracy, as two metrics to evaluate our method’s effectiveness: a desired model shall be high in both. and 1.3% on standard accuracy on CIFAR-10 when being adversarially fine-tuned, compared with the strong AT baseline. Even performing standard fine-tuning (which consumes fewer resources) with the robust pretrained models improves the resulting model’s robustness.

We systematically study all possible combinations between pretraining and fine-tuning. Our extensive results reveal that adversarial fine-tuning contributes to the dominant portion of robustness improvement, while robust pretraining mainly speeds up adversarial fine-tuning. That can also be read from Figure 1 (smaller marker sizes denote less training epochs needed).

We experimentally show that the pretrained models resulting from different self-supervised tasks have diverse adversarial vulnerabilities. In view of that, we propose to pretrain with an ensemble of self-supervised tasks, in order to leverage their complementary strengths. On CIFAR-10, our ensemble strategy further contributes to an improvement of 3.59% on robust accuracy, while maintaining a slightly higher standard accuracy. Our approach establishes a new benchmark result on standard accuracy (86.04%86.04\%) and robust accuracy (54.64%54.64\%) in the setting of AT.

Related Work

Numerous self-supervised learning methods have been developed in recent years, including: region/component filling (e.g. inpainting and colorization ); rotation prediction ; category prediction ; and patch-base spatial composition prediction (e.g., Jigsaw and Selfie ). All perform standard training, and do not tackle adversarial robustness. For example, Selfie , generalizes BERT to image domains. It masks out a few patches in an image, and then attempts to classify a right patch to reconstruct the original image. Selfie is first pretrained on unlabeled data and fine-tuned towards the downstream classification task.

Adversarial robustness.

Many defense methods have been proposed to improve model robustness against adversarial attacks. Approaches range from adding stochasticity , to label smoothening and feature squeezing , to denoising and training on adversarial examples . A handful of recent works point out that those empirical defenses could still be easily compromised . Adversarial training (AT) provides one of the strongest current defenses, by training the model over the adversarially perturbed training data, and has not yet been fully compromised by new attacks. showed AT is also effective in compressing or accelerating models while preserving learned robustness.

Several works have demonstrated model ensembles to boost adversarial robustness, as the ensemble diversity can challenge the transferability of adversarial examples. Recent proposals formulate the diversity as a training regularizer for improved ensemble defense. Their success inspires our ensembled self-supervised pretraining.

Unlabeled data for adversarial robustness.

Self-supervised training learns effective representations for improving performance on downstream tasks, without requiring labels. Because robust training methods have higher sample complexity, there has been significant recent attention on how to effectively utilize unlabeled data to train robust models.

Results show that unlabeled data can become a competitive alternative to labeled data for training adversarially robust models. These results are concurred by , who also finds that learning with more unlabeled data can result in better adversarially robust generalization. Both works use unlabeled data to form an unsupervised auxiliary loss (e.g., a label-independent robust regularizer or a pseudo-label loss).

To the best of our knowledge, is the only work so far that utilizes unlabeled data via self-supervision to train a robust model given a target supervised classification task. It improves AT by leveraging the rotation prediction self-supervision as an auxiliary task, which is co-optimized with the conventional AT loss. Our self-supervised pretraining and fine-tuning differ from all above settings.

Our Proposal

In this section, we introduce self-supervised pretraining to learn feature representations from unlabeled data, followed by fine-tuning on a target supervised task. We then generalize adversarial training (AT) to different self-supervised pretraining and fine-tuning schemes.

Selfie : By masking out select patches in an image, Selfie constructs a classification problem to determine the correct patch to be filled in the masked location.

Rotation : By rotating an image by a random multiple of 90 degrees, Rotation constructs a classification problem to determine the degree of rotation applied to an input image.

Jigsaw : By dividing an image into different patches, Jigsaw trains a classifier to predict the correct permutation of these patches.

Supervised Fine-tuning

AT versus standard training (ST)

2 AT meets self-supervised pretraining and fine-tuning

It is worth noting that our study on the integration of AT with a pretraining+fine-tuning scheme (Pi,Fj)(\mathcal{P}_{i},\mathcal{F}_{j}) provided by Tables 1-2 is different from , which conducted one-shot AT over a supervised classification task integrated with a rotation self-supervision task.

In order to explore the network robustness against different configurations {(Pi,Fj)}\{(\mathcal{P}_{i},\mathcal{F}_{j})\}, we ask: is AT for robust pretraining sufficient to boost the adversarial robustness of fine-tuning? What is the influence of fine-tuning strategies (partial or full) on the adversarial robustness of image classification? How does the type of self-supervised pretraining task affect the classifier’s robustness?

We provide detailed answers to the above questions in Sec. 4.3, Sec. 4.4 and Sec. 4.5. In a nutshell, we find that robust representation learnt from adversarial pretraining is transferable to down-stream fine-tuning tasks to some extent. However, a more significant robustness improvement is obtained by adversarial fine-tuning. Moreover, AT for full fine-tuning outperforms that for partial fine-tuning in terms of both robust accuracy and standard accuracy (except the Jigsaw-specified self-supervision task). Furthermore, different self-supervised tasks demonstrate diverse adversarial vulnerability. As will be evident later, such diversified tasks provide complementary benefits to model robustness and therefore can be combined.

3 AT by leveraging ensemble of multiple self-supervised learning tasks

Spurred by , we quantify the diversity-promoting regularizer gg through the orthogonality of input gradients of different self-supervised pretraining losses,

Experiments and Results

Dataset Details We consider four different datasets in our experiments: CIFAR-10, CIFAR-10-C , CIFAR-100 and R-ImageNet-224 (a specifically constructed “restricted” version of ImageNet, with resolution 224×224224\times 224). For the last one, we indeed to demonstrate our approach on high-resolution data despite the computational challenge. We follow to choose 10 super classes which contain a total of 190 ImageNet classes. The detailed classes distribution of each super class can be found in our supplement.

For the ablation study of different pretraining dataset sizes, we sample more training images from the 80 Million Tiny Images dataset where CIFAR-10 was selected from. Using the same 10 super classes, we form CIFAR-30K (i.e., 30,000 for images), CIFAR-50K, CIFAR-150K for training, and keep another 10,000 images for hold-out testing.

Dataset Usage In Sec. 4.3, Sec. 4.4 and Sec. 4.5, for all results, we use CIFAR-10 training set for both pretraining and fine-tuning. We evaluate our models on the CIFAR-10 testing set and CIFAR-10-C. In Sec. 4.6, we use CIFAR-10, CIFAR-30K, CIFAR-50K, CIFAR-150K and R-ImageNet-224 for pretraining, and CIFAR-10 training set for fine-tuning, while evaluating on CIFAR-10 testing set. We also validate our approaches on CIFAR-100 in the supplement. In all of our experiments, we randomly split the original training set into a training set and a validation set (the ratio is 9:1).

2 Implementation Details

Model Architecture: For pretraining with the Selfie task, we identically follow the setting in . For Rotation and Jigsaw pretraining tasks, we use ResNet-50v2 . For the fine-tuning, we use ResNet-50v2 for all. Each fine-tuning network will inherit the corresponding robust pretrained weights to initialize the first three blocks of ResNet-50v2, while leaving the remaining blocks randomly initialized.

Training & Evaluation Details: All pretraining and fine-tuning tasks are trained using SGD with 0.9 momentum. We use batch sizes of 256 for CIFAR-10, ImageNet-32 and 64 for R-ImageNet-224. All pretraining tasks adopt cosine learning rates. The maximum and minimum learning rates are 0.1 and 10−610^{-6} for Rotation and Jigsaw pretraining; 0.025 and 10−610^{-6} for Selfie pretraining; and 0.001 and 10−810^{-8} for ensemble pretraining. All fine-tuning phases follow a multi-step learning rate schedule, starting from 0.1 and decayed by 10 times at epochs 30 and 50 for a 100 epochs training.

Evaluation Metrics & Model Picking Criteria: We follow to use: i) Standard Testing Accuracy (TA): the classification accuracy on the clean test dataset; II) Robust Testing Accuracy (RA): the classification accuracy on the attacked test dataset. In our experiments, we use TA to pick models for a better trade-off with RA. Results of models picked using RA criterion are included in the supplement.

3 Adversarial self-supervised pertraining & fine-tuning helps classification robustness

We systematically study all possible configurations of pretraining and fine-tuning considered in Table 1 and Table 2, where recall that the expression (Pi,Fj)(\mathcal{P}_{i},\mathcal{F}_{j}) denotes a specified pretraining+fine-tuning scheme. The baseline schemes are given by the end-to-end standard training (ST), namely, (P1,F3)(\mathcal{P}_{1},\mathcal{F}_{3}) and the end-to-end adversarial training (AT), namely, (P1,F4)(\mathcal{P}_{1},\mathcal{F}_{4}). Table 3 shows TA, RA, and iteration complexity of fine-tuning (in terms of number of epochs) under different pretraining+fine-tuning strategies involving different self-supervised pretraining tasks, Selfie, Rotation and Jigsaw. In what follows, we analyze the results of Table 3 and provide additional insights.

We begin by focusing on the scenario of integrating the standard pretraining strategy P2\mathcal{P}_{2} with fine-tuning schemes F3\mathcal{F}_{3} and F4\mathcal{F}_{4} used in baseline methods. Several observations can be made from the comparison (P2,F3)(\mathcal{P}_{2},\mathcal{F}_{3}) vs. (P1,F3)(\mathcal{P}_{1},\mathcal{F}_{3}) and (P2,F4)(\mathcal{P}_{2},\mathcal{F}_{4}) vs. (P1,F4)(\mathcal{P}_{1},\mathcal{F}_{4}) in Table 3. 1) The use of self-supervised pretraining consistently improves TA and/or RA even if only standard pretraining is conducted; 2) The use of adversarial fine-tuning F4\mathcal{F}_{4} (against standard fine-tuning F3\mathcal{F}_{3}) is crucial, leading to significantly improved RA under both P1\mathcal{P}_{1} and P2\mathcal{P}_{2}; 3) Compared (P1,F4)(\mathcal{P}_{1},\mathcal{F}_{4}) with (P2,F4)(\mathcal{P}_{2},\mathcal{F}_{4}), the use of self-supervised pretraining offers better eventual model robustness (around 3%3\% improvement) and faster fine-tuning speed (almost saving the half number of epochs).

Next, we investigate how the adversarial pretraining (namely, P3\mathcal{P}_{3}) affects the eventual model robustness. It is shown by (P3,F1)(\mathcal{P}_{3},\mathcal{F}_{1}) and (P3,F2)(\mathcal{P}_{3},\mathcal{F}_{2}) in Table 3 that the robust feature representation learnt from P3\mathcal{P}_{3} benefits adversarial robustness even in the case of partial fine-tuning, but the use of adversarial partial fine-tuning, namely, (P3,F2)(\mathcal{P}_{3},\mathcal{F}_{2}), yields a 30%30\% more improvement. We also observe from the case of (P3,F3)(\mathcal{P}_{3},\mathcal{F}_{3}) that the standard full fine-tuning harms the robust feature representation learnt from P3\mathcal{P}_{3}, leading to 0%0\% RA. Furthermore, when the adversairal full fine-tuning is adopted, namely, (P3,F4)(\mathcal{P}_{3},\mathcal{F}_{4}), the most significant robustness improvement is acquired. This observation is consistent with (P2,F4)(\mathcal{P}_{2},\mathcal{F}_{4}) against (P2,F3)(\mathcal{P}_{2},\mathcal{F}_{3}).

Third, at the first glance, adversarial full fine-tuning (namely, F4\mathcal{F}_{4}) is the most important step to improve the final mode robustness. However, adversarial pretraining is also a key, particularly for reducing the computation cost of fine-tuning; for example, less than 5050 epochs in (P3,F4)(\mathcal{P}_{3},\mathcal{F}_{4}) vs. 9999 epochs in the end-to-end AT (P1,F4)(\mathcal{P}_{1},\mathcal{F}_{4}).

Last but not the least, we note that the aforementioned results are consistent against different self-supervised prediction tasks. However, Selfie and Rotation are more favored than Jigsaw to improve the final model robustness. For example, in the cases of adversarial pretraining followed by standard and adversarial partial fine-tuning, namely, (P3,F1)(\mathcal{P}_{3},\mathcal{F}_{1}) and (P3,F2)(\mathcal{P}_{3},\mathcal{F}_{2}), Selfie and Rotation yields at least 3.5%3.5\% improvement in RA. As the adversarial full fine-tuning is used, namely, (P3,F4)(\mathcal{P}_{3},\mathcal{F}_{4}), Selfie and Rotation outperform Jigsaw in both TA and RA, where Selfie yields the largest improvement, around 2.5%2.5\% in both TA and RA.

4 Comparison with one-shot AT regularized by self-supervised prediction task

Figure 3 presents the multi-dimensional performance comparison of our approach vs. the baseline method in . As we can see, our approach yields 1.97%1.97\% improvement on TA while 0.74%0.74\% degradation on RA. However, our approach yields consistent robustness improvement in defending all 1212 unforeseen attacks, where the improvement ranges from 1.03%1.03\% to 6.53%6.53\%. Moreover, our approach separates pretraining and fine-tuning such that the target image classifier can be learnt from a warm start, namely, the adversarial pretrained representation network. This mitigates the computation drawback of one-shot AT in , recalling that our advantage in saving computation cost was shown in Table 3. Next, Figure 4 presents the performance of our approach under different types of self-supervised prediction task. As we can see, Selfie provides consistently better performance than others, where Jigsaw performs the worst.

5 Diversity vs. Task Ensemble

In what follows, we show that different self-supervised prediction tasks demonstrate a diverse adversarial vulnerability even if their corresponding RAs remain similar. We evaluate such a diversity through the transferability of adversarial examples generated from robust classifiers fine-tuned from the adversarially pretrained models using different self-supervised prediction tasks. We then demonstrate the performance of our proposed adversarial pretraining method (4) by leveraging an ensemble of Selfie, Rotation, and Jigsaw.

In Figure 2, we demonstrate the effectiveness of our proposed adversarial pretraining via diversity-promoted ensemble (AP + DPE) given in (4). Here we consider 44 baseline methods: 33 single task based adversarial pretraining, and adversarial pretraining via standard ensemble (AP + SE), corresponding to λ=0\lambda=0 in (4). As we can see in Table 5, AP + DPE yields at least 1.17%1.17\% improvement on RA while at most 3.02%3.02\% degradation on TA, comparing with the best single fine-tuned model. In addition to the ensemble at the pretraining stage, we consider a simple but the most computationally intensive ensemble strategy, an averaged predictions over three final robust models learnt using adversarial pretraining P3\mathcal{P}_{3} followed by adversarial fine-tuning F4\mathcal{F}_{4} over Selfie, rotation, and Jigsaw. As we can see in Table 6, the best combination, ensemble of three fine-tuned models, yields at least 3.59%3.59\% on RA while maintains a slight higher TA. More results of other ensemble configurations can be found in the supplement.

6 Ablation Study and Analysis

For comparison fairness, we fine-tune all models in the same CIFAR-10 dataset. In each ablation, we show results under scenarios (P3,F2\mathcal{P}_{3},\mathcal{F}_{2}) and (P3,F4\mathcal{P}_{3},\mathcal{F}_{4}), where P3\mathcal{P}_{3} represents adversarial pretraining, F2\mathcal{F}_{2} represents partial adversarial fine-tuning and F4\mathcal{F}_{4} represents full adversarial fine-tuning. More ablation results can be found in the supplement.

As shown in Table 7, as the pretraining dataset grows larger, the standard and robust accuracies both demonstrate steady growth. Under the (P3,F4\mathcal{P}_{3},\mathcal{F}_{4}) scenario, when the pretraining data size increases from 30K to 150K, we observe a 0.97% gain on robust accuracy with nearly the same standard accuracy. That aligns with the existing theory . Since self-supervised pretraining requires no label, we could in future grow the unlabeled data size almost for free to continuously boost the pretraining performance.

Ablation of defense approaches in pretraining

In Table 8, we use random smoothing in place of AT to robustify pretraining, while other protocols remain all unchanged. We obtain consistent results to using adversarial pretraining: robust pretraining speed up adversarial fine-tuning and helps final model robustness, while the full adversarial fine-tuning contributes the most to the robustness boost.

Conclusions

In this paper, we combine adversarial training with self-supervision to gain robust pretrained models, that can be readily applied towards downstream tasks through fine-tuning. We find that adversarial pretraining can not only boost final model robustness but also speed up the subsequent adversarial fine-tuning. We also find adversarial fine-tuning to contribute the most to the final robustness improvement. Further motivated by our observed diversity among different self-supervised tasks in pretraining, we propose an ensemble pretraining strategy that boosts robustness further. Our results observe consistent gains over state-of-the-art AT in terms of both standard and robust accuracy, leading to new benchmark numbers on CIFAR-10. In the future, we are interested to explore several promising directions revealed by our experiments and ablation studies, including incorporating more self-supervised tasks, extending the pretraining dataset size, and scaling up to high-resolution data.

References