Towards Robust Vision Transformer

Xiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li, Ranjie Duan, Shaokai Ye, Yuan He, Hui Xue

Introduction

Following the popularity of transformers in Natural Language Processing (NLP) applications, e.g., BERT and GPT , there has sparked particular interest in investigating whether transformer can be a primary backbone for computer vision applications previously dominated by Convolutional Neural Networks (CNNs). Recently, Vision Transformer (ViT) successfully applies a pure transformer for classification which achieves an impressive speed-accuracy trade-off by capturing long-range dependencies via self-attention. Base on this seminal work, numerous variants have been proposed to improve ViTs from different perspectives containing training data efficiency , self-attention mechanism , introducing convolution or pooling layers , etc. However, these works only focus on the standard accuracy and computation cost, lacking the investigation of the intrinsic influence on model robustness and generalization.

In this work, we take initiatives to explore a ViT model with strong robustness. To this end, we first give an empirical assessment of existing ViT models in Figure 1. Surprisingly, although all ViT variants reproduce the standard accuracy claimed in the paper, some of their modifications may bring devastating damages on the model robustness. A vivid example is PVT , which achieves a high standard accuracy but suffered with large drop of robust accuracy. We show that PVT-Small obtains only 26.6% robust accuracy, which is 14.1% lower than original DeiT-S in Figure 1.

To demystify the trade-offs between accuracy and robustness, we analyze ViT models with different patch embedding, position embedding, transformer blocks and classification head whose impact on the robustness that has never been thoroughly studied. Based on the valuable findings revealed by exploratory experiments, we propose a Robust Vision Transformer (RVT), which has significant improvement on robustness, but also exceeds most other transformers in accuracy. In addition, we propose two new plug-and-play techniques to further boost the RVT. The first is Position-Aware Attention Scaling (PAAS), which plays the role of position encoding in RVT. PAAS improves the self-attention mechanism by filtering out redundant and noisy position correlation and activating only major attention with strong correlation, which leads to the enhancement of model robustness. The second is a simple and general patch-wise augmentation method for patch sequences which adds rich affinity and diversity to training data. Patch-wise augmentation also contributes to the model generalization by reducing the risk of over-fitting. With the above proposed methods, we can build an augmented Robust Vision Transformer∗ (RVT∗). Contributions of this paper are three-fold:

We give a systematic robustness analysis of ViTs and reveal harmful components. Inspired by it, we reform robust components as building blocks as a new transformer, named Robust Vision Transformer (RVT).

To further improve the RVT, we propose two new plug-and-play techniques called position-aware attention scaling and patch-wise augmentation. Both of them can be applied to other ViT models and yield significant enhancement on robustness and standard accuracy.

Experimental results on ImageNet and six robustness benchmarks show that RVT exhibits best trade-offs between standard accuracy and robustness compared with previous ViTs and CNNs. Specifically, RVT-S∗ achieves Top-1 rank on ImageNet-C, ImageNet-Sketch and ImageNet-R.

Related Work

Robustness Benchmarks. The rigorous benchmarks are important for evaluating and understanding the robustness of deep models. Early works focus on the model safety under the adversarial examples with constrained perturbations . In real-world applications, the phenomenon of image corruption or out-of-distribution is more commonly appeared. Driven by this, ImageNet-C benchmarks the model against image corruption which simulates distortions from real-world sources. ImageNet-R and ImageNet-Sketch collect the online images consisting of naturally occurring distribution changes such as image style, to measure the generalization ability to new distributions at test time. In this paper, we adopt all the above benchmarks as the fair-minded evaluation metrics.

Robustness Study for CNNs. The robustness research of CNNs has experienced explosive development in recent years. Numerous works conduct thorough study on the robustness of CNNs and aim to strengthen it in different ways, e.g., stronger data augmentation , carefully designed or searched network architecture, improved training strategy , quantization and pruning of the weights, better pooling or activation functions , etc. Although the methods mentioned above perform well on CNNs, there is no evidence that they also keep the effectiveness on ViTs. A targeted research for improving the robustness of ViTs is still blank.

Robustness Study for ViTs. Until now, there are several works attempting at studying the robustness of ViTs. Early works focus on the adversarial robustness of ViTs. They find that ViTs are more adversarially robust than CNNs and the transferability of adversarial examples between CNNs and ViTs is remarkably low. Follow up works extend the robustness study on ViTs to much common image corruption and distribution shift, and indicate ViTs are more robust learners. Although some findings are consistent with above works, in this paper, we do not make simple comparison of robustness between ViTs and CNNs, but take a step further by analyzing the detailed robust components in ViT and its variants. Based on the analysis, we design a robust vision transformer and introduce two novel techniques to further reduce the fragility of ViT models.

Robustness Analysis of Designed Components

We give the robustness analysis of four main components in ViTs: patch embedding, position embedding, transformer blocks and classification head. DeiT-Ti is used as the base model. All the robustness benchmarks mentioned in section 2 are considered comprehensively. There is a positive correlation between these benchmarks in most cases. Due to the limitation of space, we show the robust accuracy under FGSM adversary in the main body and other results in Appendix A.

F1: Low-level feature of patches helps for the robustness. ViTs tokenize an image by splitting it into patches with size of 16×\times16 or 32×\times32. Such simple tokenization makes the models hard to capture low-level structures such as edges and corners. To extract low-level features of patches, CeiT , LeViT and TNT use a convolutional stem instead of the original linear layer, T2T-ViT leverages self-attention to model dependencies among neighboring pixels. However, these methods merely focus on the standard accuracy. To answer how is the robustness affected by leveraging low-level features of patches, we compare the original linear projection with two new convolution and tokens-to-tokens embedders, proposed by CeiT and T2T-ViT respectively. As shown in Table 2, low-level patch embedding has a positive effect on the model robustness and standard accuracy as more detailed visual features are exploited. Among them tokens-to-tokens embedder is the best, but it has quadratic complexity with the expansion of image size. We adopt the convolutional embedder with less computation cost.

2 Position Embedding

F2: Position encoding is critical for learning shape-bias based semantic features which are robust to texture changes. Besides, existing position encoding methods have no big impact on the robustness. We first explore the necessity of position embeddings. Previous work shows ViT trained without position embeddings has 4% drop of standard accuracy. In this work, we find this gap even can be larger on robustness. In Appendix A, we find with no position encoding, ViT fails to recognize shape-bias objects, which leads to 8% accuracy drop on ImageNet-Sketch. Concerning the ways of positional encoding, learned absolute, sin-cos absolute, learned relative , input-conditioned position representations are compared. In Table 1, the result suggests that most position encoding methods have no big impact on the robustness, and a minority even have a negative effect. Especially, CPE encodes position embeddings conditioned on inputs. Such a conditional position representation makes it changed easily with the input, and causes the poor robustness. The fragility of position embeddings also motivates us to design a more robust position encoding method.

3 Transformer Blocks

F3: An elaborate multi-stage design is required for constructing robust vision transformers. Modern CNNs always start with a feature of large spatial sizes and a small channel size and gradually increase the channel size while decreasing the spatial size. The different sizes of feature maps constitute the multi-stage convolution blocks. As shown by previous works , such a design contributes to the expressiveness and generalization performance of the network. PVT , PiT and Swin employ this design principle into ViTs. To measure the robustness variance with changing of stage distribution, we slightly modify the DeiT-Ti architecture to get five variants (V2-V6) in Table 3. We keep the overall number of transformer blocks consistent to 12 and replace some of them with smaller or larger spatial resolution. Detailed architecture is shown in Appendix A. By comparing with DeiT-Ti, we find all five variants improve the standard accuracy, benefit from the extraction of hierarchical image features. In terms of robustness, transformer blocks with different spatial sizes show different effects. An experimental conclusion is that the model will get worse on robustness when it contains more transformer blocks with large spatial resolution. On the contrary, reducing the spatial resolution gradually at later transformer blocks contributes to the modest enhancement of robustness. Besides, we also observe that having more blocks with larger input spatial size will increase the number of FLOPs and memory consumption. To achieve the best trade-off on speed and performance, we think V2 is the most compromising choice in this paper.

F4: Robustness can be benefited from the completeness and compactness among attention heads, by choosing an appropriate head number. ConViT , Swin and LeViT both use more self-attention heads and smaller dimensions of keys and queries to achieve better performance at a controllable FLOPs. To study how does the number of heads affect the robustness, we train DeiT-Ti with different head numbers. Once the number of heads increases, we meanwhile reduce the head dimensions to ensure the overall feature dimensions are unchanged. Similar with generally understanding in NLP , we find the completeness and compactness among attention heads are important for ViTs. As shown in the Table 4, the robustness and standard accuracy still gain great improvement with the head increasing till to 8. We think that an appropriate number of heads supplies various aspects of attentive information on the input. Such complete and non-redundant attentive information also introduces more fine-grained representations which are prone to be neglect by model with less heads, thus increases the robustness.

F5: The locality constraints of self-attention layer may do harm for the robustness. Vanilla self-attention calculates the pair-wise attention of all sequence elements. But for image classification, local region needs to be paid more attention than remoter regions. Swin limits the self-attention computation to non-overlapping local windows on the input. This hard coded locality of self-attention enjoys great computational efficiency and has linear complexity with respect to image size. Although Swin can also get competitive accuracy, in this work we find such local window self-attention is harmful to the model robustness. The result in Table 2 shows after modifying self-attention to the local version, the robust accuracy is getting worse. We think this phenomenon may be partly caused by the destruction of long-range dependencies modeling in ViTs.

F6: Feed-forward networks (FFN) can be extended to convolutional FFN by encoding multiple tokens in local regions. Such information exchange of local tokens in FFN makes ViTs more robust. LocalViT and CeiT introduce connectivity of local regions into ViTs by adding a depth-wise convolution in feed-forward networks (FFN). Our experiment in Table 2 verifies that the convolutional FFN greatly improves both the standard accuracy and robustness. We think the reason lies in two aspects. First, compared with locally self-attention, convolutional FFN will not damage the long-term dependencies modeling ability of ViTs. The merit of ViTs can be inherited. Second, original FFN only encodes single token representation, while convolutional FFN encodes both the current token and its neighbors. Such information exchange within a local region makes ViTs more robust.

4 Classification Head

F7: Is the classification token (CLS\mathtt{CLS}) important for ViTs? The answer is not, and replacing CLS\mathtt{CLS} with global average pooling on output tokens even improves the robustness. CNNs adopt a global average pooling layer before the classifier to integrate visual features at different spatial locations. This practice also inherently takes advantage of the translation invariance of the image. However, ViTs use an additional classification token (CLS\mathtt{CLS}) to perform classification, are not translation-invariant. To get over this shortcoming, CPVT and LeViT remove the CLS\mathtt{CLS} token and replace it by average pooling along with the last layer sequential output of the Transformer. We compare models trained with and without CLS\mathtt{CLS} token in Table 2. The result shows the adversarial robustness can be greatly improved by removing CLS\mathtt{CLS} token. Also we find removing CLS\mathtt{CLS} token has slight help for the standard accuracy, which can be benefited from the desired translation-invariance.

5 Combination of Robust Components

In the above, we separately analyze the effect of each designed component in the ViTs. To make use of these findings, we combine the selected useful components, listed in follows: 1) Extract low-level feature of patches using a convolutional stem; 2) Adopt the multi-stage design of ViTs and avoid blocks with larger spatial resolution; 3) Choose a suitable number of heads; 4) Use convolution in FFN; 5) Replace CLS\mathtt{CLS} token with token feature pooling. As we find the effects of the above modifications are superimposed, we adopt all of these robust components into ViTs, the resultant model is called Robust Vision Transformer (RVT). RVT has achieved the new state-of-the-art robustness compared to other ViT variants. To further improve the performance, we propose two novel techniques, position-aware attention scaling and patch-wise data augmentation, to train our RVT. Both of them are also applicable to other ViT models.

Position-Aware Attention Scaling

In this section, we introduce our proposed position encoding mechanism called Position-Aware Attention Scaling (PAAS), which modifies the rescaling operation in the dot product attention to a more generalized version. To start with, we illustrate the scaled dot-product attention in transformer firstly. And then the modification of PAAS will be explained.

For preventing extremly small gradients and stabilizing the training process, each element in QKTQK^{T} multiplies by a constant 1d\frac{1}{\sqrt{d}} to be rescaled into a standard range.

Position-Aware Attention Scaling.

where ⊙\odot is the element-wise product. As WpW_{p} is input independent and only determined by the position of each qq, kk in the sequence, our position-aware attention scaling can also serve as a position representation. Thus, we replace the traditional position embedding with our PAAS in RVT. After that the overall self-attention can be decoupled into two parts: the QKTQK^{T} term presents the content-based attention, and Wp/dW_{p}/\sqrt{d} term acts as the position-based attention. This untied design offers more expressiveness by removing the mixed and noisy correlations .

Robustness of PAAS.

As mentioned in section 3.2, most existing position embeddings have no contribution to the model robustness, and some of them even do a negative effect. Differently, our proposed PAAS can improve the model robustness effectively. This superior property relies on the position importance matrix WpW_{p}, which acts as a soft attention mask on each position pair of qq-kk. As shown in Figure 3, we visualize the attention map of 3th query patch in 3th transformer block. Without PAAS, an adversarial input can make some unrelated regions activated and produce a noisy self-attention map. To filter out these noises, PAAS suppresses the redundant positions irrelevant for classification in self-attention map, by a learned small multiplier in WpW_{p}. Finally only the regions important for classification are activated. We experimentally validate that PAAS can provide certain defense power against some white-box adversaries, e.g., FGSM . Not limited to adversarial attack, it also helps to the corruption and out-of-distribution generalization. Details can be referred to section 6.3.

Patch-Wise Augmentation

Image augmentation is a strategy especially important for ViTs since a biggest shortcoming of ViTs is the worse generalization ability when trained on relatively small-size datasets, while this shortcoming can be remedied by sufficient data augmentation . On the other hand, a rich data augmentation also helps with robustness and generalization, which has been verified in previous works . For improving the diversity of the augmented training data, we propose the patch-wise data augmentation strategy for ViTs, which imposes diverse augmentation on each input image patches at training time. Our motivation comes from the difference of ViTs and CNNs that ViTs not only extract intra-patch features but also concern the inter-patch relations. We think the traditional augmentation which randomly transforms the whole image could provide enough intra-patch augmentation. However, it lacks the diversity on inter-patch augmentation, as all of patches have the same transformation at one time. To impose more inter-patch diversity, we retain the original image-level augmentation, and then add the following patch-level augmentation on each image patch. For simplicity, only three basic image transformations are considered for patch-level augmentation: random resized crop, random horizontal flip and random gaussian noise.

Same with the augmentations like MixUp , AugMix , RandAugment , patch-wise augmentation also benefit the model robustness. It effects on the phases after conventional image-level augmentations, and provides the meaningful augmentation on patch sequence input. Different from RandAugment, which adopts augmentations conflicting with ImageNet-C, we only use simple image transform for patch-wise augmentation. It confirms that the most part of robustness improvement is derived from the strategy itself but not the used augmentation. A significant advantage of patch-wise augmentation is that it can be in common use across different ViT models and bring more than 1% and 5% improvement on standard and robust accuracy. Details can be referred to section 6.3.

Experiments

Implementation Details. All of our experiments are performed on the NVIDIA 2080Ti GPUs. We implement RVT in three sizes named by RVT-Ti, RVT-S, RVT-B respectively. All of them adopt the best settings investigated in section 2. For RVT∗, we add PAAS on multiple transformer blocks. The patch-wise augmentation uses the combination of base augmentation introduced in section 6.4. Other training hyperparameters are same with DeiT .

Evaluation Benchmarks. We adopt the ImageNet-1K dataset for training and standard performance evaluation. No other large-scale dataset is needed for pre-training. For robustness evaluation, we test our RVT in three aspect: 1) for adversarial robustness, we test the adversarial examples generated by white-box attack algorithms FGSM and PGD on ImageNet-1K validation set. ImageNet-A is used for evaluating the model under natural adversarial example. 2) for common corruption robustness, we adopt ImageNet-C which consists of 15 types of algorithmically generated corruptions with five levels of severity. 3) for out-of-distribution robustness, we evaluate on ImageNet-R and ImageNet-Sketch . They contain images with naturally occurring distribution changes. The difference is that ImageNet-Sketch only contains sketch images, which can be used for testing the classification ability when texture or color information is missing.

2 Standard Performance Evaluation

For standard performance evaluation, we compare our method with state-of-the-art classification methods including Transformer-based models and representative CNN-based models in Table 5. Compared to CNNs-based models, RVT has surpassed most of CNN architectures with fewer parameters and FLOPs. RVT-Ti∗ achieves 79.2% Top-1 accuracy on ImageNet-1K validation set, which is competitive with currently popular ResNet and RegNet series, but only has 1.3G FLOPs and 10.9M parameters (around 60% smaller than CNNs). With the same computation cost, RVT-S∗ obtains 81.9% test accuracy, 2.9% higher than ResNet-50. This result is closed to EfficientNet-B4, however EfficientNet-B4 requires larger 380×\times380 input size and has much lower throughput.

Compared to Transformer-based models, our RVT also achieves the comparable standard accuracy. We find just combining the robust components can make RVT-Ti get 78.4% Top-1 accuracy and surpass the existing state-of-the-art on ViTs with tiny version. By adopting our newly proposed position-aware attention scaling and patch-wise data augmentation, RVT-Ti∗ can further improve 0.8% on RVT-Ti with little additional computation cost. For other scales of the model, RVT-S∗ and RVT-B∗ also achieve a good promotion compared with DeiT-S and DeiT-B. Although the improvement becomes smaller with the increase of model capacity, we think the advance of our model is still obvious as it strengthen the model ability in various views such as robustness and out-of-domain generalization.

3 Robustness Evaluation

We employ a series of benchmarks to evaluate the model robustness on different aspects. Among them, ImageNet-C (IN-C) calculates the mean corruption error (mCE) as metric. The smaller mCE means the more robust of the model under corruptions. All other benchmarks use Top-1 accuracy on test data if no special illustration. The results are reported in Table 5.

Adversarial Robustness. For evaluating the adversarial robustness, we adopt single-step attack algorithm FGSM and multi-step attack algorithm PGD with steps t=5t=5, step size α=0.5\alpha=0.5. Both attackers perturb the input image with max magnitude ϵ=1\epsilon=1. Table 5 suggests that the adversarial robustness has a strong correlation with the design of model architecture. With similar model scale and FLOPs, most Transformer-based models have higher robust accuracy than CNNs under adversarial attacks. This conclusion is also consistent with . Some modifications on ViTs or CNNs will also weaken or strengthen the adversarial robustness. For example, Swin-T introduces window self-attention for reducing the computation cost but damaging the adversarial robustness, and EfficientNet-B4 uses smooth activation functions which is helpful with adversarial robustness.

We summarize the robust design experiences of ViTs in this work. The resultant RVT model achieves superior performance on both FGSM and PGD attackers. In detail, RVT-Ti and RVT-S get over 10% improvement on FGSM, compared with the previous ViT variants. This advance is further expanded by our PAAS and patch-wise augmentation. Adversarial robustness seems unrelated with standard performance. Although models like Swin-T, TNT-S get higher standard accuracy than DeiT-S, their adversarially robust accuracy is well below the baseline. However, our RVT model can achieve the best trade-off between standard performance and adversarial robustness.

Common Corruption Robustness. To metric the model degradation on common image corruptions, we present the mCE on ImageNet-C (IN-C) in Table 5. We also list some methods from ImageNet-C Leaderboard, which are built based on ResNet-50. Our RVT-S∗ gets 49.4 mCE, which has 4.2 improvement on top-1 method DeepAugment in the leaderboard, and bulids the new state-of-the-art. The result also indicates that Transformer-based models have a natural advantage in dealing with image corruptions. Attributed to its ability of long-range dependencies modeling, ViTs are easier to learn the shape-bias features. Note that in this work we are not considering RandAugment. As a training augmentation of ViTs, RandAugment adopts conflicted augmentation with ImageNet-C and may cause the unfairness of the comparison proposed by .

Out-of-distribution Robustness. We test the generalization ability of RVT on out-of-distribution data by reporting the top@1 accuracy on ImageNet-R (IN-R) and ImageNet-Sketch (IN-SK) in Table 5. Our RVT and RVT∗ also beat other ViT models on out-of-distribution generalization. As the superiority of Transformer-based models on capturing shape-bias features mentioned above, our RVT-S also surpasses most CNN and ViT models and get 35.0% and 46.9% test accuracy on ImageNet-Sketch and ImageNet-R, buliding the new state-of-the-art. Layers Pos. Acc Rob. Emb. Acc 0-1 Ori. 78.2 34.1 Ours 78.4 34.3 0-5 Ori. 78.4 34.6 Ours 78.6 35.2 0-10 Ori. 78.4 34.8 Ours 78.6 35.3 Table 6: Comparison of single and multiple block PAAS. Ori. stands for the learned absolute position embedding in original ViTs. Augmentations Acc Rob. Acc RC GN HF ✓ 78.9 41.5 ✓ 79.0 42.0 ✓ 79.1 41.3 ✓ ✓ 78.8 41.3 ✓ ✓ 79.0 41.9 ✓ ✓ ✓ 79.2 41.7 Table 7: Ablation experiments on patch-wise augmentation. RC, GN, HF represent random resized crop, random gaussian noise and random horizontal flip respectively.

4 Ablation Studies

we conduct ablation studies on the proposed components of PAAS and patch-wise augmentation in this section. Other modifications of RVT are not involved since they have been analyzed in section 2. All of our ablation experiments are based on the RVT-Ti model on ImageNet.

Single layer PAAS vs. Multiple layer PAAS. We evaluate whether using PAAS on multiple transformer blocks can benefit the performance or robustness. The result is suggested in Table 6.3. Learned absolute position embedding in original ViT model is adopted for comparison. With more transformer blocks using PAAS, the standard and robust accuracy gain greater enhancement. After applying PAAS on 5 blocks, the benefit of PAAS gets saturated. There will be the same trend if we replace PAAS with the original position embedding. But the original position embedding is not performed as good as our PAAS on both standard and robust accuracy.

Different types of basic augmentation. Due to the limited training resources, we only test three basic image augmentations: random resized crop, random horizontal flip and random gaussian noise. For random resized crop, we crop the patch according to the scale sampled from [0.85, 1.0], then resize it to original size with aspect ratio unchanged. We set the mean and standard deviation as 0 and 0.01 for random gaussian noise. For each transformation, we set the applying probability p=0.1p=0.1. Other hyper-parameters are consistent with the implementation in Kornia . As shown in Table 6.3, we can see both three augmentations are beneficial of standard and robust accuracy. Among them, random gaussian noise is the better choice as it helps for more robustness improvement.

Combination of basic augmentations. We further evaluate the combination of basic patch-wise augmentations. For traditional image augmentation, combining multiple basic transformation can largely improve the standard accuracy. Differently, as shown in Table 6.3, the benefit is marginal for combining basic patch-wise augmentations, but combination of three is still better than using only single augmentation. In this paper, we adopt the combination of all basic augmentations.

Effect on other ViT architectures. For showing the effectiveness of our proposed position-aware attention scaling and patch-wise augmentation, we apply them to train other ViT models. DeiT-Ti, ConViT-Ti and PiT-Ti are adopted as the base model. The experimental results are shown in Table 8, with combining the proposed techniques into these base models, all the augmented models achieve significant improvement. Specifically, all the improved models yield more than 1% and 5% promotion on standard and robust accuracy on average.

Conclusion

We systematically study the robustness of key components in ViTs, and propose Robust Vision Transformer (RVT) by alternating the modifications which would damage the robustness. Furthermore, we have devised a novel patch-wise augmentation which adds rich affinity and diversity to training data. Considering the lack of spatial information correlation in scaled dot-product attention, we present position-aware attention scaling (PAAS) method to further boost the RVT. Experiments show that our RVT achieves outstanding performance consistently on ImageNet and six robustness benchmarks. Under the exhaustive trade-offs between FLOPs, standard and robust accuracy, extensive experiment results validate the significance of our RVT-Ti and RVT-S.

References

Appendix

Appendix A Additional Results of Robustness Analysis on Designed Components

Here we will show the remaining results of robustness analysis in section 3. As the each component has been discussed detailly, we only give a summary of the results in appendix. We report the additional results of robustness analysis in Table 10, 9, 12, 11, 13 and 14 respectively, where each table presents the results of one or some components. The detailed architecture of models used in robustness analysis on stage distribution is shown in Table 15. Although each robustness benchmark is consistent on the overall trend, we still find some special cases. For example, in Table 12, the V6 version of stage distribution poorly performs on adversarial robustness, but achieves best results on IN-A and IN-R datasets, showing the superior generalization power. Another case is the token-to-token embedder in Table 11. Compared with original linear embedder, token-to-token embedder obtains better results on IN-C, IN-A, IN-R and IN-SK datasets. However, under PGD attacker, it only gets the robust accuracy of 4.7%. The above phenomenon also indicates that using only several robustness benchmarks is biased and cannot get a comprehensive assessment result. Therefore, we advocate that the works about model robustness in future should consider multiple benchmarks. For validating the generality of the proposed techniques, we show the robustness evaluation results when trained on other ViT architectures and larger datasets (ImageNet-22k) in Table 13 and 14.

Appendix B Feature Visualization

In general understanding, intra-class compactness and inter-class separability are crucial indicators to measure the effectiveness of a model to produce discriminative and robust features. We use t-Distributed Stochastic Neighbor Embedding (t-SNE) to visualize the feature sets extracted by ResNet50, DeiT-Ti, Swin-T, PVT-Ti and our RVT respectively. The features are produced on validation set of ImageNet and ImageNet-C. We randomly selected 10 classes for better visualization. As shown in Figure 4, features extracted by our RVT is the closest to the intra-class compactness and inter-class separability. It’s confirmed from the side that our RVT does have the stronger robustness and classification performance.

We also visualize the feature maps of ResNet50, DeiT-S and our proposed RVT-S in Figure 5. Visualized features are extracted on the 5th layer of the models. The result shows ResNet50 and DeiT-S contain a large part of redundant features, highlighted by red boxes. While our RVT-S reduces the redundancy and ensures the diversity of features, reflecting the stronger generalization ability.

Appendix C Loss Landscape Visualization

Loss landscape geometry has a dramatic effect on generalization and trainability of the model. We visualize the loss surfaces of ResNet50 and our RVT-S in Figure 6. RVT-S has a flatter loss surfaces, which means the stability under input changes.