RobustART: Benchmarking Robustness on Architecture Design and Training Techniques

Shiyu Tang, Ruihao Gong, Yan Wang, Aishan Liu, Jiakai Wang, Xinyun Chen, Fengwei Yu, Xianglong Liu, Dawn Song, Alan Yuille, Philip H. S. Torr, Dacheng Tao

Introduction

Deep neural networks (DNNs) have achieved remarkable performance across a wide range of applications . However, DNNs are susceptible to adversarial examples ; i.e., adding deliberately crafted perturbations imperceptible to humans could easily lead DNNs to wrong predictions, threatening both digital and physical deep learning applications . Besides, some prior works show that DNNs are also vulnerable to natural noises . These phenomena demonstrate that deep learning systems are not inherently secure and robust.

Since DNNs have been integrated into various safety-critical scenarios (e.g., autonomous driving, face recognition) , improving model robustness and further constructing robust deep learning systems in practice is now growing in importance. Benchmarking the robustness of deep learning models paves a critical path to better understanding and further improving model robustness . Existing benchmarks focus on evaluating the performance of commonly used adversarial defense methods , however, there are no comprehensive studies of how architecture design and general training techniques affect robustness. These factors reflect the inherent nature of model robustness, and slight differences may override the gains from the defenses . Thus, a comprehensive benchmark and study of the influence of architecture design and training techniques on robustness are highly important for understanding and improving the robustness of deep learning models.

In this work, we propose RobustART, the first comprehensive robustness benchmark on ImageNet regarding architecture design and training techniques towards diverse noise types. We systematically study over a thousand model architectures (49 prevalent human-designed off-the-shelf model architectures including CNNs , Transformers , and MLP-Mixers ; 1200+ architectures generated by Neural Architecture Search ) and 10+ mainstream training techniques (e.g., data augmentation, adversarial training, weight averaging, label smoothing, optimizer choices). To thoroughly study robustness against noises from different sources, we evaluate diverse noise types including adversarial noises , natural noises (e.g., ImageNet-C , -P , -A , and -O ), and system noises (see Table I for the top-5 robust architectures against different noises). Our large-scale experiments revealed several insights with respect to general model design for the first time: (1) for Transformers and MLP-Mixers, adversarial training is universally effective for improving the robustness against all types of noises (adversarial, natural, and system noises); (2) given comparable model sizes and aligned training settings, CNNs (e.g., ResNet) perform stronger on natural and system noises, while Transformers (e.g., ViT) are more robust on adversarial noises; and (3) for almost all model families, increasing model capacity improves model robustness; however, increasing model sizes or adding extra data cannot improve robustness for some lightweight architectures.

Comprehensive benchmark. We provide the first comprehensive robustness benchmark on ImageNet, regarding architecture design and general training techniques. We investigate more than one thousand model architectures and 10+ training techniques on multiple noise types (adversarial, natural, and system noises).

Open-source framework. We open-source our whole framework, which consists of a model zoo (100+ pre-trained models), an open-source toolkit, and all datasets (i.e., ImageNet with different noise types). The framework provides an open platform for the community, enabling comprehensive robustness evaluation.

In-depth analyses. Based on extensive experiments, we revealed and substantiated several insights that reflect the inherent relationship between robustness and architecture design with different training techniques.

Our benchmark http://robust.art/ provides an open-source platform and framework for comprehensive evaluations, better understanding of DNNs, and robust architectures design. Together with existing benchmarks on defenses, we could build a more comprehensive robustness benchmark and ecosystem involving more perspectives.

Background and Related Work

In this section, we provide a brief overview of existing work on adversarial attacks and defenses, as well as robustness benchmark and evaluation.

In particular, adversarial attacks can be divided into two types: (1) white-box attacks, in which adversaries have the complete knowledge of the target model and can fully access it; and (2) black-box attacks, in which adversaries do not have full access to the target model and only have limited knowledge of it, e.g., can only obtain its prediction without knowing the architecture and weights. A plethora of work has been devoted to performing adversarial attacks . Goodfellow et al. first proposed the Fast Gradient Sign Method (FGSM) which leverages the gradient information of input image to craft adversarial examples efficiently. Madry et al. proposed the Project Gradient Descent (PGD) attack that is similar with FGSM but iterates more steps to generate adversarial examples, after each step it projects noises to the ϵ\epsilon-ball around input images. Carlini & Wagner (C&W) attack is an optimization-based attack that aims to find adversarial perturbations to minimize the specific object function. It transforms a general constrained optimization problem into minimizing an empirically-chosen object function. Universal Adversarial Perturbation (UAP) aims to find the input-agnostic noises that could cause model misclassification when added on input images from various classes. AutoAttack is one of the state-of-the-art adversarial attacks that combine four different attacks to form a perameter-free and computationally affordable attack to test adversarial robustness. It includes two extensions of PGD attack (APGD-CE and APGD-DLR) and two existing adversarial attacks (FAB attack and Square Attack ) to further boost the strength of attack.

On the other hand, various defense approaches have been proposed to improve model robustness against adversarial examples . Specifically, adversarial training minimizes the worst case loss within some perturbation region for classifiers, by augmenting the training set {x(i),y(i)}i=1...n\{x^{(i)},y^{(i)}\}_{i=1...n} with adversarial examples. Defensive Distillation uses the change of softmax temperature to obsfucate the gradient information of model outputs and thus prevents the white-box attack from using gradient to generate adversarial examples. Zhang et al. stated the trade-off between robustness and accuracy and proposed TRADES to defend adversarial attacks. TRADES introduces a regularization term to induce a shift of the decision boundaries away from the training data points. Kannan et al. proposed Adversarial Logit Pairing (ALP) which not only minimizes the model prediction error on adversarial examples but also tries to minimize the distance of model logits between clean examples and the corresponding adversarial examples.

2 Robustness benchmark and evaluation

A number of works have been proposed to evaluate the robustness of deep neural networks . Su et al. first investigated the adversarial robustness of 18 models on ImageNet. DEEPSEC is a platform for adversarial robustness analysis, which incorporates 16 adversarial attacks, 13 adversarial defenses, and several utility metrics. Carlini et al. discussed the methodological foundations, reviewed commonly accepted best practices, and the suggested checklist for evaluating adversarial defenses. RealSafe is another benchmark for evaluating adversarial robustness on image classification tasks (including 15 attacks and 16 defenses). Pang et al. conducted a thorough empirical study of the training tricks of representative adversarial training methods to improve the robustness on CIFAR-10. More recently, RobustBench is developed to track and evaluate the state-of-the-art adversarial defenses on CIFAR-10 and CIFAR-100.

Compared to existing benchmarks, our benchmark has the following characteristics: (1) comprehensive evaluation of different model architectures under aligned training settings and general training schemes over different architectures, while prior works focus on training schemes specialized for improving the performance against limited noise types, e.g., defenses against adversarial examples or common corruptions; (2) all noise types are based on ImageNet, while prior benchmarks mainly focus on image datasets of much smaller sizes, e.g., CIFAR-10 and CIFAR-100; and (3) systematic study of diverse noise types, while prior benchmarks mainly focus on evaluating adversarial robustness.

Robustness Benchmark on Architecture Design and Training Techniques

Existing robustness benchmarks mainly evaluate adversarial defenses, and lack a comprehensive study of the effect of architecture design and general training techniques on robustness. These factors reflect the inherent nature of the model, and comprehensively benchmarking their effects on robustness lays the foundation for building robust DNNs. Therefore, we propose the first comprehensive robustness benchmark on ImageNet considering architecture design and training techniques.

Our main goal is to investigate the effects of two orthogonal factors on robustness, which are architecture design and training techniques. Therefore, we build a comprehensive repository containing: (1) models with different architectures, but trained with the same (aligned) techniques; and (2) models with the same architecture, but trained with different techniques. These fine-grained ablation studies enable a better understanding of how the core model design choices contribute to the robustness.

To conduct a thorough evaluation and explore the robustness trends of different architecture families, our repository tends to cover as many architectures as possible. As for the CNNs, we choose the classical large architectures including ResNet-series (ResNet, ResNeXt, WideResNet) and DenseNet, the lightweight ones including ShuffleNetV2 and MobileNetV2, the reparameterized architecture RepVGG, the NAS models including RegNet, EfficientNet and MobileNetV3, and sub-networks sampled from the BigNAS supernet. As for the recently prevalent Vision Transformers, we reproduce ViT, DeiT, ViTAE, and Swin Transformer. Besides, we also include the MLP based architecture MLP-Mixer. All the models available are listed in Table II. In total, we collect 49 human-designed off-the-shelf model architectures and sampled 1200 subnet architectures from the supernets. For fair comparisons of robustness, as for all human-designed off-the-shelf models we keep the aligned training technique setting (e.g., the same data augmentation techniques).

1.2 Training techniques

An increasing number of techniques have been proposed in recent years to train deep learning models with higher accuracy. Some of them have been proved to be influential for model robustness by previous literature , e.g., data augmentation. However, with investigations only on limited architectures and noise types, these conclusions might not be able to unveil the intrinsic nature behind. Besides, there are broader training techniques but their relationships with robustness are still ambiguous. To give more exact answers about the influence on robustness, we summarize 10+ mainstream training techniques into five different categories (see Table IV) and systematically benchmark their effects on robustness. The relevant pre-trained models using different training techniques are open-sourced.

2 Evaluation strategies

For all of our experiments, we use the large-scale ImageNet-1K dataset , which contains 1,000 classes of colored images of size 224 * 224 with 1,281,167 training examples and 50,000 test instances.

2.2 Noise types

There exist various noise types in the real-world scenarios but most papers only consider a fraction of them (e.g., adversarial noises), which may fail to fully benchmark model robustness. This inspires us to collect diverse noise sources for a more comprehensive robustness profiling. As shown in Table IV, we categorize the noise format into three types: adversarial noise, natural noise, and system noise. The robustness of a model should take all these noises into consideration.

Natural noises. There are many formats of natural noises in the real world, and we choose four typical natural noise datasets (i.e., ImageNet-C, ImageNet-P, ImageNet-A, and ImageNet-O). See Section E.3.2 in the supplementary materials for more details.

System noises. Besides, there always exists system noises in the training or inference stages for models deployed in the industry. Slight differences in pre-processing operations may cause different model performances. For example, different decoders (i.e., RGB and YUV) or resize modes (i.e., bilinear, nearest, and cubic scheme) would introduce system noises and in turn influence model robustness. Thus, we use the ImageNet-S dataset to evaluate model robustness for system noises, which consists of 10 different industrial operations. See Section E.3.3 in the supplementary materials for more details.

2.3 Metrics

Adversarial Noise Robustness. To evaluate the robustness against specific adversarial attacks, we use Adversarial Robustness (AR), which is based on commonly used Attack Success Rate (ASR) and calculated by 1 - ASR (higher indicates stronger model). To evaluate the adversarial robustness under the union of different attacks, we use the Worst-Case Attack Robustness (WCAR), which indicates the lower bound of adversarial robustness against multiple adversarial attacks (higher indicates stronger model). The formal definitions are shown in Section E.4.1 of the supplementary materials.

Natural Noise Robustness. For ImageNet-C, prior works usually adopt mean Corruption Error (mCE) to measure model errors, thus we use 1 - mCE to measure the robustness for ImageNet-C (we call it ImageNet-C Robustness). For ImageNet-P, we follow the mean Flip Probability defined in and take the negative of it (Negative mean Flip Probability, NmFP) as our metric. For ImageNet-A, we simply use the classification accuracy of models. For ImageNet-O, we follow and use Area Under the Precision-Recall (AUPR). For the above 4 metrics, higher indicates a stronger model. Formal definitions are shown in Section E.4.2 of the supplementary materials.

System Noise Robustness. For system noises, we use the accuracy to measure the robustness to different decode and resize modes. To further evaluate the model stability with different decode and resize methods, we use the standard deviation of accuracy across all combinations of these modes and take the negative of it (Negative Standard Deviation, NSD). The higher NSD of a model indicates the better tolerance of facing various combinations of system noises. See Section E.4.3 of the supplementary materials for more details.

3 Framework

Our benchmark is built as a modular framework, which consists of 4 core modules (Models, Training, Noises, and Evaluation) and provides them in an easy-to-use way (see Figure 1). All the modules contained are highly extendable for users through API docs on our website. The benchmark is built on Pytorch including necessary libraries such as Foolbox and ART .

Models. We have collected 49 human-designed off-the-shelf models and their corresponding checkpoints (100+ pre-trained models in total), including ResNets, ViTs, MLP-Mixers, etc. Users could flexibly add their model architecture files, load weight configurations, and register their models in the framework to further study the robustness of the customized model architectures.

Training. We provide the implementations and interfaces of all training techniques mentioned in this paper. Users are also able to add customized training techniques to evaluate the robustness.

Evaluation. We provide various evaluation metrics, including WCAR, NmFP, mCE, AUPR, etc. Users are also recommended to add their own customized metrics through our APIs.

To sum up, based on our open-sourced benchmark, users could (1) conveniently use the source files to reproduce all the proposed results and conduct deeper analyses; (2) add new models, training techniques, noises, and evaluation metrics into the benchmark to conduct additional experiments through our APIs; (3) use our pre-trained checkpoints and research results for other downstream applications or as a baseline for comparisons.

Experiments and Analyses

In this section, we first study the influence of architecture design on robustness. In particular, we divide this part into the analysis of human-designed off-the-shelf model architectures, and networks sampled by the neural architecture search.

Human-designed off-the-shelf architectures To conduct fair and rigorous comparisons among different model architectures, we keep the aligned training techniques for each human-designed off-the-shelf architecture. For optimizers, we use SGD for all model families except for ViTs, DeiTs, ViTAEs, Swin Transformers, and MLP-Mixers; we use AdamW for the rest of model families since Transformers and MLP-Mixers are highly sensitive to optimizers (using SGD would cause the failure of training ). For scheduler, we use cosine scheduler with maximum training epoch=100 for all models. For data pre-processing, we use standard ImageNet training augmentation, which consists of the random resized crop, random horizontal flip, color jitter, and normalization. See Section F.1.1 of the supplementary materials for details.

Architectures sampled from NAS supernets We choose MobileNetV3, ResNet (basic block architecture), and ResNet (bottleneck block architecture) as the three typical NAS architectures to train the supernets using BigNAS. For optimizer, we use SGD for all supernets; for scheduler, we use the cosine scheduler with maximum training epoch=100; for data pre-processing, we follow the settings for human-designed off-the-shelf architectures and use standard ImageNet training augmentation. For other hyper-parameters, we use label smooth and set batch size=512. More details are shown in Section F.1.2 of the supplementary materials.

1.2 Human-designed off-the-shelf architectures

To study the effect of human-designed off-the-shelf model architectures on robustness, we choose 49 models from 15 most commonly used architecture families and keep the aligned training settings for each individual model. We report the model robustness against different noise types and standard performance (i.e., clean accuracy) w.r.t. model architectures. We use classical model floating-point operations (FLOPs) and model parameters (Params) to measure model sizes. For more results about the robustness and transferability maps with more attack magnitudes, please refer to Figure 6, 16 and 17 in the supplementary material. As shown in Figure 2 and 3, we can draw observations as follows.

▶\blacktriangleright Model size ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}{\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow} indicates that this factor has a positive correlation with model robustness, while {\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow} indicates the opposite.). For most of the model families (e.g., ResNets, RegNetX), with the increase of the model size (e.g., FLOPs and Params), robustness for adversarial, natural, and system noises improves, and clean accuracy also increases. This indicates the positive influence of model capacity on robustness and task performance within the same model family. However, for some lightweight models such as MobileNetV2, MobileNetV3, and ShuffleNetV2, this conclusion does not entirely hold (i.e., larger models are not necessarily more robust than smaller models). It is even surprising that for EfficientNets, with the increase of the model size, the adversarial robustness decreases. We conjecture it may be due to the increase of input size (for EfficientNets family, the input size monotonically increases from EfficientNet-B0 to B7, e.g., the input size of EfficientNet-B0 is 224, B1 is 240, and B2 is 260).

▶\blacktriangleright Model architecture ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). Based on our large-scale experiments, we can find that the model architecture is a highly significant factor for robustness.

Specifically, we observe that when compared under the aligned training settings, ViT, ViTAE, and DeiT rank top-3 for worst-case attack robustness; for natural noise robustness, ResNeXt, WideResNet, ResNet rank top-3; for system noises, EfficientNet, DenseNet, ResNeXt rank top-3. Therefore, we could empirically reach the conclusion that with comparable sizes and aligned training settings, CNNs (e.g., ResNet) perform stronger on natural and system noises, while Transformers and MLP-Mixers are more robust against adversarial noises.

Moreover, it is interesting to find that although Swin Transformers are mainly based on the self-attention mechanism like ViTs and DeiTs, their robustness performance is more like CNNs. In other words, in contrast to ViTs, Swin Transformers are comparatively more robust to natural and system robustness, and less robust towards adversarial noises. Some mechanisms in Swin Transformers draw on properties of CNNs, such as window attention (introduce the locality like convolution kernel) and hierarchical architecture, thus we conjecture this might be the reason for the special robustness performance of Swin Transformers. We hope all these observations could inspire more in-depth future studies.

▶\blacktriangleright Adversarial transferability. According to the experimental results, we mainly divide the model architectures into three categories: common CNNs (ResNet, ResNeXt, WideResNet, DenseNet, RegNeXt, and RepVGG), lightweight CNNs (MobileNetV2, MobileNetV3, and ShuffleNetV2), and non-CNNs (ViTs, DeiTs, and MLP-Mixers). In most cases, models in each category are more robust against attacks transferred from other categories while less robust against attacks transferred from themselves. Interestingly, Swin Transformers and ViTAEs show “moderate” robustness (i.e., in contrast to other ViTs, Swin Transformers and ViTAEs are more robust against attacks transferred from all other model families, and attacks generated from them also have stronger transferability to all other model families. We speculate that the design of Swin Transformers and ViTAEs has drawn the architecture properties from both CNNs and classical Transformers (e.g., multi-head attention, hierarchical architecture, feature pyramid).

1.3 Architectures sampled from NAS supernets

We then study the robustness of models sampled from BigNAS supernet. Due to the flexibility of sampled subnets, we first study the effect of model size on robustness (we totally sampled 600 subnets from 3 supernets) and then dig into the detailed factors that affect robustness in a more fine-grained manner. Specifically, for each detailed factor (input size, convolution kernel size, model depth, and expand ratio), we fix other model factors and sample 50 subnets from each of the 3 supernets to evaluate robustness (600 subnets in total). As shown in Figure 4, we highlight the key findings below.

▶\blacktriangleright Model size ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}). Within a specific supernet, increasing model size improves model robustness on adversarial (except for the lightweight supernet MobileNetV3, which is similar to conclusions in the human-designed model study) and natural noises, but there is no apparent effect on system noises. We hypothesize that for lightweight models, simply increasing their capacities cannot improve robustness.

▶\blacktriangleright Model depth ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}). In all supernets, increasing the depth of the last stage (i.e., the number of layers in the subnets’ deepest block) could improve the adversarial robustness of the sampled subnets, which shows the deepest stage is of great importance for model robustness in NAS-sampled network. However, for the depth of other stages or the whole network, the relationship is still ambiguous.

▶\blacktriangleright Input size ({\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). In all supernets, when increasing the input size of sampled subnets, their adversarial robustness decreases. This observation is similar to the study on EfficientNets in Sec 4.1.2, indicating that networks with larger input sizes might be less robust against adversarial attacks. However, the effects of input size on robustness for natural and system noises are ambiguous.

▶\blacktriangleright Convolution kernel size ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}). For sampled subnets, increasing the total number of convolution kernel sizes in all stages could improve the adversarial robustness of subnets, while the effects on natural and system noises are unclear. For each stages’ convolution kernel size, there seems no correlation with model robustness.

To sum up, we found that the architecture design has a huge impact on robustness (especially the sizes and families). Meanwhile, the noise diversity is highly important to robustness evaluation (slight differences would cause opposite observations), we therefore recommend researchers to use more comprehensive and diverse noises when evaluating model robustness.

2 Training techniques towards robustness

2.2 Results and analysis

We list some of the representative results in Figure 5, and more results about all training techniques are shown in Figure 18 to 28 in the supplementary materials. From the experimental results, we draw several observations as follows.

(1) Adversarial training ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). For CNNs, adversarial training largely boosts adversarial robustness and robustness against ImageNet-P, but reduces the clean accuracy as well as the robustness against ImageNet-C and system noises. For ViTs and MLP-Mixer, adversarial training also largely improves adversarial robustness and robustness against ImageNet-P, slightly improves the robustness against ImageNet-C and system noises, whereas slightly reduces the clean accuracy (much smaller than the drop on CNNs, e.g., 76.3 →\rightarrow 57.8 for ResNet-50, while 66.6 →\rightarrow 60.1 for ViT-B/16).

(2) Data augmentation ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). For adversarial noises, data augmentation reduces worst-case attack robustness in most cases; for natural noises (ImageNet-C, -P, -A, and -O), data augmentation yields a stronger model in most cases, and the improvement on ViTs and MLP-Mixers is larger than CNNs (e.g., Augmix improves the ImageNet-C robustness of ResNet-50 from 41.4 to 44.3, while improves ImageNet-C robustness of Mixer-B/16 from 26.6 to 40.3); for system noises, Mixup improves robustness while Augmix reduces robustness, showing that different data augmentation techniques might have different impacts on model robustness.

(3) ImageNet-21K pre-training ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). For both CNNs and ViTs, this technique improves robustness on all of the natural noises and reduces robustness on system noises and adversarial noises (under WCAR metric). And the robustness improvement on ViTs is much bigger than CNNs. (e.g., ImageNet-A robustness of ResNet-50 improves from 2.1 to 8.1, while ViT-B/16 improves from 3.8 to 23.1; ImageNet-C robustness of ResNet-50 improves from 41.4 to 42.7, while ViT-B/16 improves from 29.9 to 50.2)

(2) Dropout ({\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}\uparrow}{\color[rgb]{0,1,0}\definecolor[named]{pgfstrokecolor}{rgb}{0,1,0}\downarrow}). For CNNs, Dropout slightly reduces the adversarial robustness in most cases, and the impact on natural and system noises is ambiguous. For ViTs and MLP-mixers, although Dropout still causes a slight drop in adversarial robustness in most cases, it improves model robustness against all of the natural noises and system noises, and the improvement is large when it comes to ViTs (e.g., Dropout boosts ImageNet-C robustness of ViT-B/16 from 29.9 to 40.5).

For other factors that are not mentioned in the main body, they show no evident effects on robustness. To sum up, there still does not exist an “once-for-all” training technique that yields stronger robustness for all architectures and all noises. Some techniques may even pose opposite effects on the same noise type for different architectures (e.g., AdamW improves adversarial robustness of MobileNetV3, while reducing that for ResNet). Thus, for fair comparisons of model robustness, we should align the training techniques.

3 Discussions

In contrast to existing studies that evaluate adversarial defenses, this paper aims to conduct a comprehensive and fair robustness investigation of different architectures and training techniques. For architectures, we align different model architectures (including CNNs, Transformers, and MLP-Mixers) with the same training settings (without extra training techniques) for a fair comparison, which is often ignored by most existing literature. For training techniques, instead of studying one specific training technique with all its implementation tricks, we summarize and evaluate 10+ training techniques over several architectures to help better understand how training techniques will influence model robustness. Through extensive ablation studies (including eliminating the extra variables with an aligned training setting for comparing different model architectures), we comprehensively demonstrate how different model architectures and training techniques impact the robustness. Based on large-scale experiments and fair comparisons, we have obtained a wide variety of observations. In this section, we offer some of the most important and consistent ones and provide further discussions and analyses on them.

Note that, for all the above experiments for adversarial training, ViTs are adversarially trained based on our aligned training settings (i.e., without using extra training techniques, e.g., AutoAugment or Dropout). Therefore, we further evaluate the effect of adversarial training for ViTs under the optimal ViT training settings (e.g., with Mixup, AutoAugment, etc.). In particular, we conduct Half Adversarial Training on a ViT-B/16 using the optimal setting and compare its robustness with another vanilla trained ViT-B/16 model under the same optimal setting. From the results in Table VI, we could reach similar observations, namely this adversarial training scheme could largely boost the robustness of ViT on all types of noises. In fact, we also conduct the standard PGD adversarial training (only feed adversarial examples into models in each mini-batch) for ViT-B/16 under the optimal settings, however, we observe that the model cannot converge. We speculate that it is too difficult for the model to fit the augmented adversarial examples since there have been various data augmentations in the optimal ViT training procedure .

Overall, to the best of our knowledge, we are the first to propose this conclusion through extensive experiments on ImageNet that adversarial training is universally effective for the robustness of Transformers and MLP-Mixers. We highly advocate researchers to use adversarial training to robustify their Transformers and MLP-Mixers in practice and further investigate the nature behind.

Some of the recent works have already studied the adversarial or natural robustness of vision transformers, and reached various conclusions. For adversarial robustness, shows that individual vision transformers are just as vulnerable as their CNN counterparts to white-box adversaries, while and both demonstrates that ViTs possess better adversarial robustness when compared with CNNs. For robustness against natural noises, shows that a ViT with comparable parameters is more robust to image corruptions than the CNN trained with augmentations, and shows that ViTs are more robust than CNNs under corruptions and OOD distributions.

However, as we can see from the above studies, some of their conclusions and observations are contradictory to others. We attribute this phenomenon to the unfair experimental settings and incomplete noises evaluation. For instance, most of the ViTs used in these works are trained with the standard settings for ViTs which consists of various specially-designed data augmentation techniques (e.g., AutoAugment and RandAugment), while they are not used for CNNs. Studies in 4.2 have revealed that data augmentation techniques could reduce adversarial robustness and improve natural robustness in most cases, therefore it will interfere with the correctness of model robustness evaluation when data augmentation techniques are introduced.

Therefore, in Sec 4.1.2, we fairly evaluate the robustness of ViTs and other architectures (e.g., CNNs) under the aligned training settings with diverse noises. Through our comprehensive empirical studies, we finally reach the conclusion that under aligned training settings, vision transformers and MLP-Mixers are more robust against adversarial noises while less robust against natural noises compared to CNNs.

3.3 Architecture is highly important for robustness

Previous works have demonstrated that model architecture is a more critical factor to robustness than model size. However, there exist two drawbacks of these works: on the one hand, they do not take the effect of training techniques into account, which may overestimate the robustness of model architectures; on the other hand, the numbers of models and noises used in the evaluation are comparatively small, which make the evaluation incomprehensive. Therefore, this paper keeps the aligned training settings when comparing the robustness of different model architectures, without introducing the influence of training techniques. Through our large-scale studies, we empirically demonstrate the significance of model architecture to robustness, and we hereby provide several interesting phenomena:

(1) Different architectures show different robustness. Generally, CNNs are more robust to natural and system noises while Transformers and MLP-Mixers are more robust to adversarial noises.

(2) As one of the Transformers architectures, Swin Transformer shows better robustness against natural noises while less robustness against adversarial noises, which is more like CNNs rather than other Transformers.

(3) Slight modifications in the structure (input size, depth, etc) of subnets sampled from supernets would affect their adversarial robustness.

To improve the robustness, besides designing new defense methods, robust architectures could also be taken into consideration. For example, Transformers are more robust for adversarial noises, and CNNs are more robust for natural and system noises.

3.4 There does not exist a “once-for-all” training technique for improving robustness

Previous works have studied the influence of some specific training techniques (e.g., label smoothing) on model robustness. In this work, we summarize and evaluate 10+ training techniques over several architectures to help better understand how training techniques will influence the robustness. We suggest that for fair evaluation, the training settings for different models should be aligned. Although according to our evaluation results, there does not exist a “once-for-all” training technique for improving robustness, we still draw several representative conclusions for researchers to improve their models’ robustness which is listed as follows:

(1) When training lightweight networks like ShuffleNetV2, MobileNetV2, and MobileNetV3, we highly recommend researchers use AdamW optimizer instead of SGD.

(2) When training Transformers or MLP-Mixers, adversarial training is universally effective for model robustness against all types of noises.

(3) When training MLP-Mixer, label smoothing could largely improve the adversarial robustness.

(4) When training Transformers, we recommend researchers use Dropout, which could largely boost the robustness against natural and system noises.

3.5 Noise diversity has huge influences on robustness evaluation

Conclusions

We propose RobustART, the first comprehensive Robustness benchmark on ImageNet regarding ARchitecture design (49 human-designed off-the-shelf architectures and 1200 networks constructed by the neural architecture search) and Training techniques (10+ general techniques including training data augmentation) towards diverse noises (adversarial, natural, and system noises). In contrast to existing studies that focus on defense methods, RobustART aims to conduct a comprehensive and fair comparison of different architectures and training techniques towards robustness. Besides discovering several new findings, our benchmark is highlighted for fertilizing the community by providing: (a) an open-source platform for researchers to use and contribute; (b) 100+ pre-trained models publicly available to facilitate robustness evaluation; and (c) new observations to better understand the mechanism towards robust DNN architectures. We welcome community researchers to join us together to continuously contribute to building this ecosystem to better understand deep learning. We will credit contributions from researchers who are not the authors of this paper on the project website.

References

Appendix A Limitations and Broader Impacts

RobustART has several limitations, and we list them as follows. (1) Although we have included a large number of architectures and training techniques, due to the rapid emergence of new approaches, there might still exist relevant ones that are not evaluated. (2) We focus on image classification tasks in the first version of our benchmark, and we will continuously develop the benchmark to include more challenging tasks, such as object detection. (3) We presented many intriguing phenomena in our large-scale experiments, but have not dived into some new findings to analyze their causes. We will conduct more thorough studies based on these observations in future work, and we believe that they will also attract broad interests from the community. We will keep the benchmark up-to-date, and hope the community can contribute together to make the ecosystem grows.

Our benchmark will facilitate the studies of deep learning robustness and a better understanding of model vulnerabilities to adversarial examples and other types of noises. We hope that this work will help in building robust model architectures for real-world applications.

Appendix B Licenses

The code of RobustART is released under Apache License 2.0. Most model architectures are added to the code with the license chosen by the original author. All pre-trained model checkpoints we provided are produced using our RobustART code and are under Apache License 2.0 as well. The ImageNet-1K, ImageNet-21K, ImageNet-A, ImageNet-O, ImageNet-P, ImageNet-S datasets we use are downloaded from the official release. The ImageNet-C datasets we use are generated according to the official code release on GitHub.

Appendix C Maintenance Plan

To make our RobustART benchmark energetic and sustainable, we will keep maintaining our benchmark in these following aspects.

Maintain our website and leaderboard. We host our website and leaderboard on http://robust.art. We will maintain and update our website and leaderboard to make it a sustainable resource center for robustness. At the same time, we will allow other researchers to upload their results on the leaderboard after review.

Maintain our code base. Our code base is an easy-to-use framework that includes the model, training, noises, and evaluation processes. The code of our framework is hosted on GitHub, we will keep maintaining and updating it.

Maintain our pre-trained models. We provide all pre-trained models in our cloud disk, which have taken more than 44GB of disk space. Every pre-trained model is able to download freely. Moreover, we will keep updating pre-trained models with different architectures and training techniques.

What’s more, we also plan to expand our benchmark to other tasks (i.e., object detection and semantic segmentation) and other datasets, making our RobustART a robust benchmark across major computer vision tasks and datasets.

Appendix D Reproducibility and Run Time

We provide the code to run this benchmark on GitHub where everyone can download from freely. As for the setup steps and instructions about our code, we provided a detailed document which can be found on our website http://robust.art. Following the setup page of this document, users can easily install the required run time environment of this codebase.

Besides, there is a ’Get Started’ page of this document. Users can experience most of the basic functions including model training, add noise and model evaluation of our code by following the instruction on page. For pro developers, we also provided a detailed API document. This API document explains the most important Python class and Python method of our code. Other developers can use this document to modify this code for their needs.

Since our benchmark experiments need us to train multiple models and evaluate them on different kinds of datasets, it needs a large amount of GPU resources. The total cost of our GPU resources to build this benchmark is about 20 GPU years. Most of our experiments are run on Nvidia GTX 1080Ti GPU, some experiments such as ImageNet-21K pre-training which needs more computing resources are run on Nvidia Tesla V100 GPU. For one training experiment, we run it on 16 GPUs parallel. For some training experiments such as ViT, we run it on 32 GPUs parallel to avoid out-of-memory problems.

Appendix E Detailed Information

For all basic experiments we use ImageNet-1K dataset , which contains 1,000 classes of colored images of size 224 * 224 with 1,281,167 training examples and 50,000 test instances. We believe that different from small data sets such as MNIST and CIFAR-10, ImageNet is more like the application data in the real scene, and thus the robustness evaluation experiments under ImageNet are more valuable and meaningful.

E.2 Training Techniques

Knowledge Distillation. knowledge distillation usually refers to the process of transferring knowledge from a large model (teacher) to a smaller one (student). While large models (such as very deep neural networks or ensembles of many models) have higher knowledge capacity than small models. In our experiments, we train a ResNet-18 model with a pre-trained ResNet-50 as a teacher for knowledge distillation.

Self-Supervised Training. Self-Supervised Training is proposed for utilizing unlabeled data with the success of supervised learning. In our experiments, we use MoCo v2 which is a classic momentum contrast self-supervised learning algorithm.

Weight Averaging. Stochastic Weight Averaging is an optimization procedure that averages multiple points along the trajectory of SGD, with a cyclical or constant learning rate. Weight averaging will lead to better results than standard training. It was proved that weight averaging notably improves training of many state-of-the-art deep neural networks over a range of consequential benchmarks, with essentially no overhead.

Weight Re-parameterization. RepVGG realizes a decoupling of the training-time and inference-time architecture by structural re-parameterization technique so that the inference can be faster than training . We set experiment of RepVGG with and without using re-parameterization to find out its influence on robustness.

Label smoothing. Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them. Assume for a small constant ϵ\epsilon the training set label yy is correct with probability 1−ϵ1-\epsilon and incorrect otherwise. Label Smoothing regularizes a model based on a softmax with kk output values by replacing the hard 0 and 1 classification targets with targets of ϵk−1\frac{\epsilon}{k-1} and 1−ϵ1-\epsilon respectively.

Dropout. Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability. To find out the influence of dropout on robustness, we trained and tested the models with and without dropout.

Data Augmentation. The performance of neural networks often improves with the amount of training data. Data augmentation is a technique to artificially create new training data from existing training data. We use Mixup and Augmix as two kinds of data augmentation methods to find out their influences.

Large-Scale Pre-training. It has been proved that pre-training using a large-scale dataset can improve the performance of neural networks. To find out the influence of large-scale dataset, we set up experiments using pre-training with ImageNet-21K dataset, which consists of 14,197,122 images (about 14 times larger than ImageNet-1K), each tagged in a single-label fashion by one of 21,841 possible classes.

Adversarial Training. Adversarial training is one of the most effective defense methods towards adversarial noises. During training, it generates adversarial examples using some attacks and then feed these adversarial examples into the model to calculate the loss function and update model weights. The most commonly used adversarial training method is PGD adversarial training, which uses PGD attack to generate adversarial examples during training procedure.

SGD and AdamW. Optimizers are algorithms or methods used to minimize the loss function or to maximize the efficiency of production. To Find out the influence of different optimizers, we use two kinds of commonly used optimizers in our experiments. Stochastic Gradient Descent (SGD) is the most basic form of gradient descent. SGD subtracts the gradient multiplied by the learning rate from the weights. Despite its simplicity, SGD has strong theoretical foundations. AdamW is a stochastic optimization method that modifies the typical implementation of weight decay in Adam, by decoupling weight decay from the gradient update.

E.3 Noises

AutoAttack. The AutoAttack proposes two extensions of the PGD-attack: APGD-CE and APGD-DLR, which overcome failures due to suboptimal step size and problems of the objective function. It then combines the two novel attacks with two complementary existing attacks (white-box FAB attack and black-box square attack ) to form a parameter-free and computationally affordable attack. AutoAttack is powerful on many classifiers with various defense methods, thus becoming a good choice for evaluating model robustness in adversarial literature.

Carlini&Wagner attack (C&W). The C&W attack is a family of optimization-based attacks for finding adversarial perturbations that minimize the given loss function. It proposes to transform a general constrained optimization problem into an unconstrained optimization formulation using an empirically chosen loss function. The C&W attack enables strong attack ability, while the whole optimization procedure is time-consuming.

E.3.2 Natural noises

ImageNet-A. ImageNet-A is a dataset of real-world adversarially filtered images that fool current ImageNet classifiers. To build this dataset, they first download numerous images related to an ImageNet class. Thereafter they delete the images that fixed ResNet-50 classifiers correctly predict. With the remaining incorrectly classified images, they manually select visually clear images and make this dataset .

ImageNet-O. ImageNet-O is a dataset of adversarially filtered examples for ImageNet out-of-distribution detectors. To create this dataset, they download ImageNet-22K and delete examples from ImageNet-1K. With the remaining ImageNet-22K examples that do not belong to ImageNet-1K classes, they keep examples that are classified by a ResNet-50 as an ImageNet-1K class with high confidence .

ImageNet-C. The ImageNet-C dataset consists of 15 diverse corruption types applied to validation images of ImageNet. The corruptions are drawn from four main categories: noise, blur, weather, and digital. Each corruption type has five levels of severity since corruptions can manifest themselves at varying intensities. We use all four categories and five levels of severity of this dataset in our benchmark .

ImageNet-P. Like ImageNet-C, ImageNet-P consists of noise, blur, weather, and digital distortions. The dataset also has validation perturbations and difficulty levels. ImageNet-P departs from ImageNet-C by having perturbation sequences generated from each ImageNet validation image. Each sequence contains more than 30 frames .

E.3.3 System noises

The ImageNet-S dataset consists of 3 commonly used decoder types and 7 commonly used resize types. For the decoder, it includes the implementation from Pillow, OpenCV, and FFmpeg. For resize operation, it includes nearest, cubic, hamming, lanczos, area, box, and bilinear interpolation modes from OpenCV and Pillow tools.

During the process of evaluation, Pillow bilinear mode is set as the default resize method when testing the decoding robustness. This setting is also the default training setting in PyTorch official code example. Similarly, Pillow is set as the default image decoding tool (same as PyTorch).

This dataset provides a validation set of ImageNet with different decoding and resize methods, and saves each image file after decoding and resizing it as a 3×width×height3\times width\times height matrix in a .npy file instead of JPEG. According to the commonly used transform on ImageNet test set, it implements pre-processing for images. This dataset provides a matrix of an image after the process of resizing to 3×256×2563\times 256\times 256 then applied a center crop to 3×224×2243\times 224\times 224.

E.4 Metrics

Adversarial Robustness (AR). As widely used in the literature , we choose Attack Success Rate (ASR) as the base evaluation metric. Considering that ASR mainly measures the attack effectiveness while we are about to measure model robustness, we simply modify it to get our adversarial robustness metric:

Worst-Case Attack Robustness (WCAR). To aggregate the model adversarial robustness results under different attacks, we follow instructions in and choose to use the per-example worst-case attack robustness:

E.4.2 Metrics for natural noises

Top-1 Accuracy. We use top-1 accuracy as the metric for ImageNet-A dataset. It can reflect how good performance of this model when facing the images which fool commonly used ResNet-50.

AUPR. We use the area under the precision-recall curve (AUPR) as the metrics for ImageNet-O dataset . This metric requires anomaly scores. Our anomaly score is the negative of the maximum softmax probabilities from a model that can classify the 200 ImageNet-O classes.

NmFP. Denote mm perturbation sequences with S={(x1(i),x2(i),…,xn(i))}i=1m\mathcal{S}=\left\{\left(x_{1}^{(i)},x_{2}^{(i)},\ldots,x_{n}^{(i)}\right)\right\}_{i=1}^{m} where each sequence is made with perturbation pp The “Flip Probability” of network f:X→{1,2,…,1000}f:\mathcal{X}\rightarrow\{1,2,\ldots,1000\} on perturbation sequences S\mathcal{S} is

E.4.3 Metrics for system noises

NSD. We use Negative Standard Deviation (NSD) as the metrics for ImageNet-S system noise, which is the negative value of standard deviation across all accuracy on different decoders and resize methods. The formula of it can be written as NSD=−σ(A)NSD=-\sigma(A), where A={adecoderresize method}A=\{a_{decoder}^{resize~{}method}\}. We use standard deviation because we want to know the stability of a model facing different decoders and resize methods, and we take the negative value of it since we want this value to increase with this model’s performance just like other metrics of this benchmark.

Appendix F Experimental Setup

Here, we provide the details of the experimental settings of our robustness evaluation benchmark.

To conduct fair and rigorous comparisons among different models, we keep the aligned training techniques as much as possible for each human-designed off-the-shelf architecture. For optimizer, we use SGD optimizer with nesterov momentum=0.9 and weight decay=0.0001 for all model families except for ViTs, DeiTs, Swin Transformers, ViTAEs and MLP-Mixers; instead for these five model families, we use AdamW with weight decay=0.05 since Transformers and MLP-Mixers are highly sensitive to optimizers, using SGD would cause the failure of training . For scheduler, we use cosine scheduler with maximum training epoch=100 for all models, for models except for ViTs, DeiTs, Swin Transformers, ViTAEs and MLP-Mixers we use base learning rate (lr)=0.1, warm up lr=0.4, and minimal lr=0.0, and for these five model families we use a much smaller learning rate. For data pre-processing, we use standard ImageNet training augmentation, which consists of random resized crop, random horizontal flip, color jitter, and normalization, for all models including Transformers and CNNs. For all models, we also follow the common settings in network training and use label smooth with ϵ\epsilon=0.1. For other settings, we set batch size=512, the number of loading workers=4, and enable the pin memory.

F.1.2 Architectures sampled from NAS supernets

As for architectures sampled from NAS supernets, we choose MobileNetV3, ResNet (basic block architecture), ResNet (bottleneck block architecture) as three typical NAS architectures to train supernets using BigNAS. During supernet training, each batch we sample 4 subnets from supernet to calculate loss and update network weight. For optimizer, we use SGD for all supernets with nesterov momentum=0.9 and weight decay. For scheduler, we use cosine scheduler with maximum training epoch=100. For data pre-processing, we follow the settings for human-designed off-the-shelf architectures and use standard ImageNet training augmentation consisting of random resized crop, random horizontal flip, color jitter, and normalization. For other hyper-parameters, we use label smooth with ϵ\epsilon=0.1 and batch size=512.

When we study the model size towards robustness, we randomly sample 200 subnets without any factor fixing from each supernet and evaluate subnets’ robustness towards adversarial, natural, and system noises. When we study the influence of a certain factor (i.e., input size, convolution kernel size, model depth, or expand ratio), we first fix all other factors, ensuring they are the same in all sampled subnets, then we randomly sample 50 subnets and evaluate their robustness.

F.2 Settings for Training Techniques

F.3 Settings for Noises

Natural noises. For natural noises, we simply follow the standard settings in ImageNet-C, ImageNet-P, ImageNet-A, and ImageNet-O datasets.

System noises. For system noises, we use the standard setting of ImageNet-S dataset, which can be found in Section E.3.3.

Appendix G Additional Results

We first report the model robustness and standard performance under the view of FLOPs in Figure 6 . For almost all model architectures, the robustness results are the same as those using Params to measure model size. Then we report the detailed results of model robustness under various magnitudes of different adversarial noises in Figure 7, 8, 9, 10, 11, 12, 13, 14, 15. There are 2 model size measurements (FLOPs and Params) and 3 perturbation magnitudes (small, middle and large), resulting in total 6 figure for each attack. In addition, we also report the heatmap of different model architectures under transfer-based adversarial attacks under more perturbation magnitudes in Figure 16, 17. We can see the model robustness results under transfer-based attack with ϵ\epsilon=0.5/255 and ϵ\epsilon=2/255 are almost the same as those under attack with ϵ\epsilon=8/255.

G.2 Training Techniques Towards Robustness

We report the results for the influence of all training techniques studied towards model robustness under various noises in Figure 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28. For each training technique we choose 24 metrics to show its influence on model accuracy and robustness. More results can be found on our website.