Robust Pre-Training by Adversarial Contrastive Learning

Ziyu Jiang, Tianlong Chen, Ting Chen, Zhangyang Wang

Introduction

Label efficiency and model robustness are two desirable characteristics when it comes to building machine learning models, including the training of deep neural networks. Traditional supervised learning of deep networks requires a lot of labeled data, whose annotation costs can be prohibitively high. Utilizing unlabeled data, e.g., by various unsupervised or semi-supervised learning techniques, is thus of booming interests. One particular branch of unsupervised learning based on contrastive learning has shown great promise recently . The latest contrastive learning framework can largely improves label efficiency of the deep networks, e.g., surpassing the fully-supervised AlexNet when fine-tuned on only 1% of the labels.

The labeling scarcity is even amplified when we come to adversarially robust deep learning , i.e., to training deep models that are not fooled by maliciously crafted, although imperceivable perturbations. As suggested in , the sample complexity of adversarially robust generalization is significantly higher than that of standard learning. Recent results advocated unlabeled data to be a powerful backup for training adversarially robust models as well, by using unlabeled data to form an auxiliary loss (e.g., a label-independent robust regularizer or a pseudo-label loss).

Only lately, researchers start linking self-supervised learning to adversarial robustness. proposed a multi-task learning framework that incorporates a self-supervised objective to be co-optimized with the conventional classification loss. The latest work introduced adversarial training into self-supervision, to provide general-purpose robust pre-trained models for the first time. However, existing work is based on ad-hoc “pretext” self-supervised tasks, and as shown in , the learned unsupervised representations can be largely improved with contrastive learning, a new family of approaches for self-supervised learning.

In order to learn data-efficient robust models, we propose to integrate contrastive learning with adversarially robust deep learning. Our work is inspired by recent success of contrastive learning, where the model is learned by maximizing consistency of data representations under differently augmented views . This fits particularly well with adversarial training, as one cause of adversarial fragility could be attributed to the non-smooth feature space near samples, i.e., small input perturbations can result in large feature variations and even label change.

By properly combining adversarial learning and contrastive pre-training (i.e. SimCLR ), we could achieve the desirable feature consistency. The resultant unsupervised pre-training framework, called Adversarial Contrastive Learning (ACL), is thoroughly discussed in Section 2. As the pre-training plays an increasingly important role for adversarial training , by leveraging more powerful contrastive pre-training of unsupervised representations, we further contribute to pushing forward the state-of-the-art adversarial robustness and accuracy.

On CIFAR-10/CIFAR-100, we extensively benchmark our pre-trained models on two adversarial fine-tuning settings: fully-supervised (using all training data and their labels) and semi-supervised (using all training data, but only partial/few shot labels). ACL pre-training leads to new state-of-the-art in both robust and standard accuracies. For example, in the fully-supervised setting, we obtain 2.99% robust and 2.14% standard accuracy improvements, over the previous best adversarial pre-training method on CIFAR-10. For the semi-supervised setting, our approach surpasses the state-of-the-art method , by 0.66% robust and 5.87% standard accuracy, when 10% labels are available; that gains can become even more remarkable when we go further to extremely low label rates, e.g., 1%.

Our approach

where D\mathcal{D} denotes the training set. The standard training (ST) could be viewed as a special case ϵ=0\epsilon=0.

Among the few existing works utilizing unlabeled data, we introduce the Unsupervised Adversarial Training (UAT) algorithm which yields the state-of-the-art performance so far. The authors introduce three variants: (i) UAT with Online Targets, which exploits a smoothness regularizer on the unlabeled data to minimize the divergence between a normal sample with its adversarial counter-part; (ii) UAT with Fixed Targets, which first trains a CNN using standard training (no AT applied) and then uses the trained model to supply pseudo labels for unlabeled data, finally combining them with labeled data towards AT altogether; and (iii) UAT++, which combines both types of unsupervised losses and was shown to outperform either individually.

2 Adversarial Contrastive Learning as pre-training: three options

SimCLR can be considered as Standard-to-Standard (S2S) contrasting, as both branches are natural images, and adversarial training nor attack is involved. In our proposed method, Adversarial Contrast Learning (ACL), we inject robustness components into the SimCLR framework, and discuss three candidate algorithms (Fig. 1 (b) - (d)).

Intuitive as it might look, there is a pitfall for A2S implementation. In both SimCLR (S2S) and A2A, their two backbones completely share all parameters. In A2S, while we still hope the adversarial and standard encoder branches to share all convolutional weights, it is important to allow both encoders to maintain their own independent batch normalization (BN) parameters.

During fine-tuning and testing, we by default employ the BN for adversarial branch, since it leads to higher robustness which is our main goal. We also briefly study the effect of using either BN on the fine-tuning performance as well as the pre-trained representations in Sec. 3.2.

Our final unsupervised loss consists of a contrastive loss term on the former pair (through two standard branches) and another contrastive loss term on the latter pair (through two adversarial branches). The two terms are by default equally weighted (α=1\alpha=1 in Algorithm. 1). The four branches share all convolutional weights except BN: the two standard branches use the same set of BN parameters, while the two adversarial branches have their one set of BN parameters too.

Compared to A2A which aggressively demands minimizing the “worst-case” consistencies, the final dual-stream loss is aligned with the classical AT’s objective , i.e., balancing between standard and worst-case (robust) loss terms. The algorithm of DS can be found in Algorithm. 1. We experimentally verify that Dual-Stream is the best among the three ACL variants. A comprehensive comparison of the three variants is presented in Section 3.2.

2.1 Towards supervised fine-tuning and semi-supervised training

For each of the three variants, we will use their pre-trained network weights as the initialization for subsequent fine-tuning. For the fully supervised setting, we identically follow the adversarial fine-tuning scheme described in .

For semi-supervised training, we employ a three-step routine: i) we first perform ACL to obtain a pre-trained model (denoted as PM) on all unlabeled data ; ii) we next adopt PM as the initialization to perform standard training using only labeled data, and use the trained model to generate pseudo-labels for all unlabeled data; iii) lastly, we leverage PM as initialization again, and perform AT over all data (including unlabeled data and their pseudo labels). The routine in principle follows except the new pre-training step (i). The specific loss that we minimize in step iii is:

The first two terms in Eqn. (3) represent a convex combination of CE and distillation losses for labeled data (α\alpha is a coefficient ∈\in (0,1)). The third term is the distillation loss for unlabeled data. The fourth term is feature-consistency robustness loss for all data (with weight 1λ\frac{1}{\lambda}) following , respectively. ∣Nl∣+∣Nu∣|N_{l}|+|N_{u}| represent the batch size (including both labeled and unlabeled data).

Experiments and analysis

Experimental settings: We evaluate three datasets: CIFAR-10, CIFAR-10-C , CIFAR-100. We follow to employ two metrics: 1) Standard Testing Accuracy (TA): The classification accuracy on clean testing images; and 2) Robust Testing Accuracy (RA): the classification accuracy on adversarial perturbed testing images. Evaluation on unforeseen attacks is also considered .

Experiments are organized as following: Section 3.1 first compares our best proposed option: ACL with Dual-Stream (DS), with existing methods, first as pre-training for supervised adversarial fine-tuning (Section 3.1.1); then as part of semi-supervised adversarial training (Section 3.1.2). Section 3.2 then delves into an ablation study comparing the three ACL options: A2A, A2S and DS.

We compare our ACL Dual Stream (DS), with the state-of-the-art robust pre-training , on CIFAR-10 and CIFAR-100. For the latter, we choose its single best pretext task, Selfie, as the comparison subject (and call it so hereinafter). By default, we use ResNet-18 as the backbone all experiments hereinafter, as was also adopted by , unless otherwise specified. ACL (DS) is employed for training the encoder of ResNet-18 (excluding the fully connected (FC) Layer), and then the pre-trained encoder is added with the zero-initialized FC layer to start fine-tuning.

We compare ACL (DS) and Selfie by conducting adversarial fine-tuning over their pre-trained models. For the fully-supervised tuning, we employ the loss in TRADE for adversarial fine-tuning, with the regularization weight 1/λ1/\lambda as 6. The fine-tuned models are selected based on the held-out validation RA. We use SGD with 0.9 momentum and batch size 128. By default, we fine-tune 25 epochs, with initial learning rate set as 0.1 and then decaying by 10 times at epoch 15 and 20. As mentioned in Section 2.2, we report ACL (DS) results using the adversarial branch BN.

As demonstrated in Table 1, while both Selfie and ACL (DS) lead to higher precision and robustness compared to the model trained from random initialization, the proposed ACL (DS) yields a significant improvement on [TA, RA] by [2.14%\mathbf{2.14\%}, 2.99%\mathbf{2.99\%}] on CIFAR-10, and [2.14%\mathbf{2.14\%}, 3.58%\mathbf{3.58\%}] on CIFAR-100, respectively. Those gains are significant, especially considering that the pre-training approach in already outperformed all previous state-of-the-arts robust training approaches. By surpassing , ACL (DS) establishes the new benchmark TA/RA numbers on CIFAR-10 and CIFAR-100.

We then apply ACL (DS) on Wide-Resnet-34-10 , which is also a standard model employed for testing adversarial robustness , to verify if the proposed pre-training can also help for big models. In the supervised fine-tuning experiment with CIFAR10, the random initialization (which equals TRADE) yields [84.33%, 55.46%]We run the official code of . The RA drops by 1% compared to what reported, since we test with the PGD attack starting from random initializations, which leads to stronger adversarial attacks.. By simply using the ACL (DS) pretrained model as the initialization, we obtain [TA, RA] of [85.12%, 56.72%], where our pre-training also contributes to a nontrivial [0.79%, 1.26%] margin for TRADE on the big model.

We further demonstrate that our robustness gain can consistently hold against unforeseen attacks, by leveraging the 19 unforeseen attacks used by on CIFAR-10 with ResNet-18. As shown in Fig. 2, our approach achieves improvement on most of the unforeseen attacks (except gaussian noise, where the performance drop marginally by 0.23%0.23\%). Remarkably, the proposed approach leads to an improvement of 2.42%2.42\% when averaged over all unforeseen attacks.

1.2 ACL for semi-supervised adversarial training

Following the procedure in Section 2.2.1, we compare our ACL (DS)-based semi-supervised training, with (1) Selfie , following the identical routine as our method; and (2) the state-of-the-art UAT++, described in . We emphasize that both Selfie and our method are plugged in to improve over the semi-supervised framework like UAT++ (described Sec. 2.2.1), by supplying pre-training, rather than re-inventing a new different semi-supervised algorithm. The results are summarized in Table 2.

We first follow the setting in to use 10% training labels together with the entire (unlabeled) training set. We observe that, adopting Selfie as pre-training can improve TA by 1.37% with RA almost unchanged, compared to the vanilla UAT++ without any pre-training. That endorses the notable value of pre-training of data-efficient adversarial learning. Moreover, along the same routine, ACL (DS) outperforms Selfie by large extra margins of [4.50%, 0.72%] in [TA ,RA], manifesting the superior quality of our proposed pre-training over . Our result is also the new state-of-the-art TA-RA trade-off achieved in this same setting, as we are aware of.

Encouraged by the promising data efficiency indicated above, we explore a challenging extreme setting that no previous method seems to have tried before: only using 1% labels together with the unlabeled training set. All methods’ hyperparameters are tuned to our best efforts using grid search. As shown in Table 2 bottom row, the advantage of using ACL (DS) pre-training becomes even more remarkable, with [TA, RA] only minorly decreased [1.00%, 0.42%] in this extreme case, compared to using 10% labels. In comparison, both vanilla UAT++ and Selfie suffer from catastrophic performance drops (∼\sim13% - 30%). This result points to the tremendous potential of boosting the label efficiency of adversarial training much higher than the current.

To further understand why our proposed method leads to such impressive margins, we check the accuracy of pseudo labels generated in step ii. We find that ACL (DS) leads to much higher standard accuracy of 86.73% (computed over the entire training set) even with 1% labels, while vanilla UAT++ and Selfie can only achieve 37.67% and 46.75%, respectively. This echos the previous finding in that the quality of pseudo labels constitutes the main bottleneck for semi-supervised adversarial training; and it indicates that our proposed robust pre-training inherits the strong label-efficient learning capability from contrastive learning .

2 Ablation study for ACL

We now report the ablation study among our proposed three ACL options, with two extra natural baselines: SimCLR that embeds no adversarial learning (a.k.a, Standard-to-Standard, or S2S); and random initialization (RI) without any pre-training.

As shown in Tab. 3, while SimCLR (S2S) improves TA over random initialization (RI) by 1.67%, it has little impact on RA. That is as expected, since this pre-training considers no adversarial perturbations. A2A boosts RA marginally more than S2S thanks to introducing adversarial perturbations, but pays a dramatic price of TA (even 1.22% lower than no pre-training). Similar to our previous conjecture, it shows that A2A’s overly aggressive, worst-case only consistency enforcement degrades the feature quality. A2S starts to achieve more favorable RA, but its TA is still slightly below S2S. Eventually, DS jointly optimize both standard and adversarial feature consistencies, ensuring the feature quality and robustness simultaneously. Not only DS surpassed SimCLR by 2.85% RA, but more interestingly, it also outperforms SimCLR by 0.62% TA. That suggests DS to be a potentially more favored pre-training, even for standard training as well. We consider this to align with another recent observation that adversarial examples as data augmentation can improve standard image recognition too : we leave it to be more thoroughly examined for future work.

For A2S and S2S, we compare three scenarios: 1) using one single BN for both standard and adversarial branches; 2) using dual BN for training, and adopt the standard branch BN (denoted as θbn\boldsymbol{\theta}_{\textit{bn}}) for testing; 3) using dual BN for training, and adopt the adversarial branch BN (θbnadv\boldsymbol{\theta}_{\textit{bn}^{\textit{adv}}}) for testing. As shown in Tab. 4, both methods see notable performance drops when using single BNs that mix standard and adversarial feature statistics, which align with the findings in . For dual BN, our default choice of θbnadv\boldsymbol{\theta}_{\textit{bn}^{\textit{adv}}} leads to more favored RA meanwhile still strong TA. Otherwise if we switch to θbn\boldsymbol{\theta}_{\textit{bn}}, TA performance will be emphasized to become higher, while RAs drop slightly yet remain competitive.

To evaluate the learned representations, adopted a linear separability evaluation protocol, where a linear classifier is trained on top of the frozen pre-trained encoder. The test accuracy is used as a proxy for representation quality: higher accuracy indicates features to be more discriminative and desirable. Here we adapt the protocol by also adversarially training the linear classifier, and evaluate the features learned by ACL (DS) pre-training in terms of both TA and RA. We call the RA obtained by the adversarially trained linear classifier as adversarial linear separability of the features.

Tab. 5 compares on two different ways training the linear classifier, and also on employing either set of BNs (θbn\boldsymbol{\theta}_{\textit{bn}} or θbnadv\boldsymbol{\theta}_{\textit{bn}^{\textit{adv}}}) to output features. A good level of adversarial linear separability, i.e., 44.22%44.22\% RA, is observed under θbnadv\boldsymbol{\theta}_{\textit{bn}^{\textit{adv}}}. Interestingly, if we switch from θbnadv\boldsymbol{\theta}_{\textit{bn}^{\textit{adv}}} to θbn\boldsymbol{\theta}_{\textit{bn}}, we immediately lose the adversarial linear separability (close to 0% RA). That again confirms the distinct statistics captured by two BNs .

Lastly, we dissect how the robustness grows during the supervised tuning process, when starting from random initialization and ACL (DS) pre-training, respectively. As indicated in Fig. 3, the robust accuracy of ACL (DS) pre-trained models jumps to 47.38% after one epoch fine-tuning, while it costs randomly initialized models 74 epochs more to achieve the same. Additionally, if we fine-tune the ACL (DS) pre-trained model longer and do not anneal the learning rate earlier, we will also see the robust accuracy decrease before the decay point and end up with 1.0% robustness drop, which was reported in and termed as the adversarial over-fitting phenomenon.

Related work and discussions

ACL demonstrates the value of contrastive-style unsupervised pre-training, for adversarially robust deep learning. Our methodology is deeply rooted in the recent surge of ideas from self-supervised learning (esp. as pre-training ), contrastive learning (esp. ), and adversarial training (esp. the unlabeled approaches ). We review those relevant fields to gradually unfold our rationale.

Training a supervised deep network from scratch requires tons of labeled data, whose annotation costs can be prohibitively high. One promising option is to first pre-train a feature representation using unlabeled data, which is later fine-tuned for down-stream tasks using (few-shot) labeled data. In general, pre-training is discovered to stabilize the training of very deep networks , and usually improves the generalization performance, especially when the labeled data volume is not high. Such pre-training & fine-tuning scheme further facilitates general task transferability and few-shot learnability of each down-steam task. Therefore, pre-trained models have gained popularity in computer vision , natural language processing and so on.

Unsupervised pre-training initially adopted the reconstruction loss or its variants . Such simple, generic loss learns to preserves input data information, but not to enforce more structural priors on the data. Recently, self-supervised learning emerges to train unsupervised dataset in a supervised manner, by framing a supervised learning task in a special form to predict only a subset of input information using the rest. The self-supervised task, also known as pretext task, guides us to a supervised loss function, leading to learned representations conveying structural or semantic priors that can be beneficial to downstream tasks. Classical pretext tasks include position predicting , order prediction , rotation prediction , and a variety of other perception tasks .

Since adversarially robust deep learning is more data-demanding than standard learning , it is natural to leverage unlabeled data using self-supervised learning too. Among other options , the latest work adopted the pre-training & fine-tuning scheme, which is plug-and-play for different down-stream tasks: we hence choose to follow this setting.

2 Contrastive Learning brings feature consistency that fits adversarial robustness

Handcrafted pretext tasks largely rely on heuristics, which can limit the generality of the learned representations. An emerging family of self-supervised methods leverages a contrastive loss for learning representations, and have shown promising results. The contrastive loss is typically defined for minimizing some distance between positive pairs while maximizing it between negative pairs, which allows the model to learn representations without reconstructing the raw input.

The use of contrastive loss for representation learning dates back to , and has been widely studied since then . As shown in , the setup of contrastive tasks via non-trivial data augmentation seems to be one major success factor behind contrastive learning. Other aspects, such as CNN architectures, non-linear transformation network in between CNN and contrastive loss, loss functions, batch size and/or memory bank, can all meaningfully impact the performance of learned representations. By combining these factors, a recently proposed self-supervised learning method, SimCLR , show that it can learn representations on par with supervised learning while requiring no human annotation. Essentially, SimCLR empowers the unsupervised feature by encouraging feature invariance to specified data transformations. Such feature-level invariance is well-known to be desirable for standard generalization of CNNs, yet is often not met by state-of-the-art CNNs . Therefore, enforcing consistency w.r.t. data augmentations has been shown effective for unsupervised learning .

Although existing contrastive learning literature discussed their boosts on the standard generalization, we realize that the feature consistency (as a result of contrasting) is valuable for robustness too. One cause of adversarial fragility could be attributed to the non-smooth feature space near samples, i.e., small input perturbations can result in large feature variations and even label change. Enforcing consistency during training w.r.t. perturbations, has been thus shown to immediately help adversarial robustness . A closely relevant work showed that a combination of stochasticity and diverse augmentations, plus a feature divergence consistency loss, improved the robustness and uncertainty estimates of fully-supervised image classifiers.

The above observations constitute our conjecture that contrastive learning could be a superior option for adversarial pre-training to other classical self-supervision tasks. Compared to those pretexts adopted by , the feature consistency appears to be more directly targeted to the robustness issue.

Conclusion

In this work, we leverage the contrastive learning to enhance the adversarial robustness via self-supervised pre-training. Our method is motivated by the cause of adversarial fragility, and we discuss several options to inject adversarial perturbations to reduce it. Through experiments in both supervised fine-tuning and semi-supervised learning settings, we demonstrate the proposed adversarial contrastive learning framework can lead to models that are both label-efficient and robust. Potential future work includes investigating the defense of larger models and datasets , and the incorporation of more diverse adversarial perturbations.

Acknowledgments

This work is supported by a US Army Research Office Young Investigator Award W911NF2010240.

Broader Impact

Defending machine learning models against adversarial attacks is a crucial component towards the goal of making AI systems more secure and trustworthy. Our proposed framework of Adversarial Contrastive Learning can be a powerful tool to improve model robustness in a data-efficient fashion. It advances the latest achievement in robustness-aware pre-training, and the results further raise the state-the-art bars for both supervised and semi-supervised adversarial training. We expect our techniques to contribute to the grand goal of building more secured and trustworthy AI.

References

Supplementary Material

This supplementary material contains a comparison with a closely related work called Augmix that we could not include in the main paper due to space restrictions. Augmix also considers utilizing feature divergence consistency loss for improving robustness. However, consistency is enforced as a regularization term in the loss while the proposed approach is incorporated as a pre-training method.

When comparing with Augmix , we employ the proposed Adversarial Contrastive Learning (ACL) with Dual Stream (DS) following the same settings in the main paper. To ensure the fair comparison, we keep the identical settings with Augmix except for the parameters that are not used for ACL (DS), where we grid search for the best configuration. Note that we employ the adversarial training loss of UAT++ for adversarial training with Augmix in all experiments since we can not get competitive robustness with the adversarial training loss used in .

We compare the proposed approach with Augmix under supervised training settings in CIFAR10. As shown in Table 6, ACL (DS) obtains an improvement of [0.46%0.46\%, 0.44%0.44\%] on [TA, RA] compared to Augmix, which proves the superiority of enforcing consistency in pre-training.