High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks

Haohan Wang, Xindi Wu, Zeyi Huang, Eric P. Xing

Introduction

Deep learning has achieved many recent advances in predictive modeling in various tasks, but the community has nonetheless become alarmed by the unintuitive generalization behaviors of neural networks, such as the capacity in memorizing label shuffled data and the vulnerability towards adversarial examples

To explain the generalization behaviors of neural networks, many theoretical breakthroughs have been made progressively, including studying the properties of stochastic gradient descent , different complexity measures , generalization gaps , and many more from different model or algorithm perspectives .

In this paper, inspired by previous understandings that convolutional neural networks (CNN) can learn from confounding signals and superficial signals , we investigate the generalization behaviors of CNN from a data perspective. Together with , we suggest that the unintuitive generalization behaviors of CNN as a direct outcome of the perceptional disparity between human and models (as argued by Figure 1): CNN can view the data at a much higher granularity than the human can.

However, different from , we provide an interpretation of this high granularity of the model’s perception: CNN can exploit the high-frequency image components that are not perceivable to human.

For example, Figure 2 shows the prediction results of eight testing samples from CIFAR10 data set, together with the prediction results of the high and low-frequency component counterparts. For these examples, the prediction outcomes are almost entirely determined by the high-frequency components of the image, which are barely perceivable to human. On the other hand, the low-frequency components, which almost look identical to the original image to human, are predicted to something distinctly different by the model.

Motivated by the above empirical observations, we further investigate the generalization behaviors of CNN and attempt to explain such behaviors via differential responses to the image frequency spectrum of the inputs (Remark 1). Our main contributions are summarized as follows:

We reveal the existing trade-off between CNN’s accuracy and robustness by offering examples of how CNN exploits the high-frequency components of images to trade robustness for accuracy (Corollary 1).

With image frequency spectrum as a tool, we offer hypothesis to explain several generalization behaviors of CNN, especially the capacity in memorizing label-shuffled data.

We propose defense methods that can help improving the adversarial robustness of CNN towards simple attacks without training or fine-tuning the model.

The remainder of the paper is organized as follows. In Section 2, we first introduce related discussions. In Section 3, we will present our main contributions, including a formal discussion on that CNN can exploit high-frequency components, which naturally leads to the trade-off between adversarial robustness and accuracy. Further, in Section 4-6, we set forth to investigate multiple generalization behaviors of CNN, including the paradox related to capacity of memorizing label-shuffled data (section 4), the performance boost introduced by heuristics such as Mixup and BatchNorm (section 5), and the adversarial vulnerability (section 6). We also attempt to investigate tasks beyond image classification in Section 7. Finally, we will briefly discuss some related topics in Section 8 before we conclude the paper in Section 9.

Related Work

The remarkable success of deep learning has attracted a torrent of theoretical work devoted to explaining the generalization mystery of CNN.

For example, ever since Zhang et al. demonstrated the effective capacity of several successful neural network architectures is large enough to memorize random labels, the community sees a prosperity of many discussions about this apparent ”paradox” . Arpit et al. demonstrated that effective capacity are unlikely to explain the generalization performance of gradient-based-methods trained deep networks due to the training data largely determine memorization. Kruger et al. empirically argues by showing largest Hessian eigenvalue increased when training on random labels in deep networks.

The concept of adversarial example has become another intriguing direction relating to the behavior of neural networks. Along this line, researchers invented powerful methods such as FGSM , PGD , and many others to deceive the models. This is known as attack methods. In order to defend the model against the deception, another group of researchers proposed a wide range of methods (known as defense methods) . These are but a few highlights among a long history of proposed attack and defense methods. One can refer to comprehensive reviews for detailed discussions

However, while improving robustness, these methods may see a slight drop of prediction accuracy, which leads to another thread of discussion in the trade-off between robustness and accuracy. The empirical results in demonstrated that more accurate model tend to be more robust over generated adversarial examples. While argued that the seemingly increased robustness are mostly due to the increased accuracy, and more accurate models (e.g., VGG, ResNet) are actually less robust than AlexNet. Theoretical discussions have also been offered , which also inspires new defense methods .

High-frequency Components & CNN’s Generalization

We first set up the basic notations used in this paper: ⟨x,y⟩\langle\mathbf{x},\mathbf{y}\rangle denotes a data sample (the image and the corresponding label). f(⋅;θ)f(\cdot;\theta) denotes a convolutional neural network whose parameters are denoted as θ\theta. We use H\mathcal{H} to denote a human model, and as a result, f(⋅;H)f(\cdot;\mathcal{H}) denotes how human will classify the data ⋅\cdot. l(⋅,⋅)l(\cdot,\cdot) denotes a generic loss function (e.g., cross entropy loss). α(⋅,⋅)\alpha(\cdot,\cdot) denotes a function evaluating prediction accuracy (for every sample, this function yields 1.01.0 if the sample is correctly classified, 0.00.0 otherwise). d(⋅,⋅)d(\cdot,\cdot) denotes a function evaluating the distance between two vectors. F(⋅)\mathcal{F}(\cdot) denotes the Fourier transform; thus, F−1(⋅)\mathcal{F}^{-1}(\cdot) denotes the inverse Fourier transform. We use z\mathbf{z} to denote the frequency component of a sample. Therefore, we have z=F(x)\mathbf{z}=\mathcal{F}(\mathbf{x}) and x=F−1(z)\mathbf{x}=\mathcal{F}^{-1}(\mathbf{z}).

Notice that Fourier transform or its inverse may introduce complex numbers. In this paper, we simply discard the imaginary part of the results of F−1(⋅)\mathcal{F}^{-1}(\cdot) to make sure the resulting image can be fed into CNN as usual.

We decompose the raw data x={xl,xh}\mathbf{x}=\{\mathbf{x}_{l},\mathbf{x}_{h}\}, where xl\mathbf{x}_{l} and xh\mathbf{x}_{h} denote the low-frequency component (shortened as LFC) and high-frequency component (shortened as HFC) of x\mathbf{x}. We have the following four equations:

where t(⋅;r)t(\cdot;r) denotes a thresholding function that separates the low and high frequency components from z\mathbf{z} according to a hyperparameter, radius rr.

To define t(⋅;r)t(\cdot;r) formally, we first consider a grayscale (one channel) image of size n×nn\times n with N\mathcal{N} possible pixel values (in other words, x∈Nn×n\mathbf{x}\in\mathcal{N}^{n\times n}), then we have z∈Cn×n\mathbf{z}\in\mathcal{C}^{n\times n}, where C\mathcal{C} denotes the complex number. We use z(i,j)\mathbf{z}(i,j) to index the value of z\mathbf{z} at position (i,j)(i,j), and we use ci,cjc_{i},c_{j} to denote the centroid. We have the equation zl,zh=t(z;r)\mathbf{z}_{l},\mathbf{z}_{h}=t(\mathbf{z};r) formally defined as:

We consider d(⋅,⋅)d(\cdot,\cdot) in t(⋅;r)t(\cdot;r) as the Euclidean distance in this paper. If x\mathbf{x} has more than one channel, then the procedure operates on every channel of pixels independently.

With an assumption (referred to as A1) that presumes “only xl\mathbf{x}_{l} is perceivable to human, but both xl\mathbf{x}_{l} and xh\mathbf{x}_{h} are perceivable to a CNN,” we have:

CNN may learn to exploit xh\mathbf{x}_{h} to minimize the loss. As a result, CNN’s generalization behavior appears unintuitive to a human. ∎

Notice that “CNN may learn to exploit xh\mathbf{x}_{h}” differs from “CNN overfit” because xh\mathbf{x}_{h} can contain more information than sample-specific idiosyncrasy, and these more information can be generalizable across training, validation, and testing sets, but are just imperceptible to a human.

As Assumption A1 has been demonstrated to hold in some cases (e.g., in Figure 2), we believe Remark 1 can serve as one of the explanations to CNN’s generalization behavior. For example, the adversarial examples can be generated by perturbing xh\mathbf{x}_{h}; the capacity of CNN in reducing training error to zero over label shuffled data can be seen as a result of exploiting xh\mathbf{x}_{h} and overfitting sample-specific idiosyncrasy. We will discuss more in the following sections.

2 Trade-off between Robustness and Accuracy

We continue with Remark 1 and discuss CNN’s trade-off between robustness and accuracy given θ\theta from the image frequency perspective. We first formally state the accuracy of θ\theta as:

and the adversarial robustness of θ\theta as in e.g., :

where ϵ\epsilon is the upper bound of the perturbation allowed.

With another assumption (referred to as A2): “for model θ\theta, there exists a sample ⟨x,y⟩\langle\mathbf{x},\mathbf{y}\rangle such that:

we can extend our main argument (Remark 1) to a formal statement:

With assumptions A1 and A2, there exists a sample ⟨x,y⟩\langle\mathbf{x},\mathbf{y}\rangle that the model θ\theta cannot predict both accurately (evaluated to be 1.0 by Equation 1) and robustly (evaluated to be 1.0 by Equation 2) under any distance metric d(⋅,⋅)d(\cdot,\cdot) and bound ϵ\epsilon as long as ϵ≥d(x,xl)\epsilon\geq d(\mathbf{x},\mathbf{x}_{l}).

The proof is a direct outcome of the previous discussion and thus omitted. The Assumption A2 can also be verified empirically (e.g., in Figure 2), therefore we can safely state that Corollary 1 can serve as one of the explanations to the trade-off between CNN’s robustness and accuracy.

Rethinking Data before Rethinking Generalization

Our first aim is to offer some intuitive explanations to the empirical results observed in : neural networks can easily fit label-shuffled data. While we have no doubts that neural networks are capable of memorizing the data due to its capacity, the interesting question arises: “if a neural network can easily memorize the data, why it cares to learn the generalizable patterns out of the data, in contrast to directly memorizing everything to reduce the training loss?”

Within the perspective introduced in Remark 1, our hypothesis is as follows: Despite the same outcome as a minimization of the training loss, the model considers different level of features in the two situations:

In the original label case, the model will first pick up LFC, then gradually pick up the HFC to achieve higher training accuracy.

In the shuffled label case, as the association between LFC and the label is erased due to shuffling, the model has to memorize the images when the LFC and HFC are treated equally.

2 Experiments

We set up the experiment to test our hypothesis. We use ResNet-18 for CIFAR10 dataset as the base experiment. The vanilla set-up, which we will use for the rest of this paper, is to run the experiment with 100 epoches with the ADAM optimizer with learning rate set to be 10−410^{-4} and batch size set to be 100100, when weights are initialized with Xavier initialization . Pixels are all normalized to be $.AlltheseexperimentsarerepeatedinMNIST,FashionMNIST,andasubsetofImageNet.TheseeffortsarereportedintheAppendix.Wetraintwomodels,withthenaturallabelsetupandtheshuffledlabelsetup,denoteasMnaturalandMshuffle,respectively;theMshuffleneeds300epochestoreachacomparativetrainingaccuracy.Totestwhichpartoftheinformationthemodelpicksup,forany. All these experiments are repeated in MNIST , FashionMNIST , and a subset of ImageNet . These efforts are reported in the Appendix. We train two models, with the natural label setup and the shuffled label setup, denote as Mnatural and Mshuffle, respectively; the Mshuffle needs 300 epoches to reach a comparative training accuracy. To test which part of the information the model picks up, for any\mathbf{x}inthetrainingset,wegeneratethelow−frequencycounterpartsin the training set, we generate the low-frequency counterparts\mathbf{x}_{l}withwithrsettoset to4,,8,,12,,16$ respectively. We test the how the training accuracy changes for these low-frequency data collections along the training process.

The results are plotted in Figure 3. The first message is the Mshuffle takes a longer training time than Mnatural to reach the same training accuracy (300 epoches vs. 100 epoches), which suggests that memorizing the samples as an “unnatural” behavior in contrast to learning the generalizable patterns. By comparing the curves of the low-frequent training samples, we notice that Mnatural learns more of the low-frequent patterns (i.e., when rr is 44 or 88) than Mshuffle. Also, Mshuffle barely learns any LFC when r=4r=4, while on the other hand, even at the first epoch, Mnatural already learns around 40% of the correct LFC when r=4r=4. This disparity suggests that when Mnatural prefers to pick up the LFC, Mshuffle does not have a preference between LFC vs. HFC.

If a model can exploit multiple different sets of signals, then why Mnatural prefers to learn LFC that happens to align well with the human perceptual preference? While there are explanations suggesting neural networks’ tendency towards simpler functions , we conjecture that this is simply because, since the data sets are organized and annotated by human, the LFC-label association is more “generalizable” than the one of HFC: picking up LFC-label association will lead to the steepest descent of the loss surface, especially at the early stage of the training.

To test this conjecture, we repeat the experiment of Mnatural, but instead of the original train set, we use the xl\mathbf{x}_{l} or xh\mathbf{x}_{h} (normalized to have the standard pixel scale) and test how well the model can perform on original test set. Table 1 suggests that LFC is much more “generalizable” than HFC. Thus, it is not surprising if a model first picks up LFC as it leads to the steepest descent of the loss surface.

3 A Remaining Question

Finally, we want to raise a question: The coincidental alignment between networks’ preference in LFC and human perceptual preference might be a simple result of the “survival bias” of the many technologies invented one of the other along the process of climbing the ladder of the state-of-the-art. In other words, the almost-100-year development process of neural networks functions like a “natural selection” of technologies . The survived ideas may happen to match the human preferences, otherwise, the ideas may not even be published due to the incompetence in climbing the ladder.

However, an interesting question will be how well these ladder climbing techniques align with the human visual preference. We offer to evaluate these techniques with our frequency tools.

Training Heuristics

We continue to reevaluate the heuristics that helped in climbing the ladder of state-of-the-art accuracy. We evaluate these heuristics to test the generalization performances towards LFC and HFC. Many renowned techniques in the ladder of accuracy seem to exploit HFC more or less.

We test multiple heuristics by inspecting the prediction accuracy over LFC and HFC with multiple choices of rr along the training process and plot the training curves.

Batch Size: We then investigate how the choices of batch size affect the generalization behaviors. We plot the results in Figure 4. As the figure shows, smaller batch size appears to excel in improving training and testing accuracy, while bigger batch size seems to stand out in closing the generalization gap. Also, it seems the generalization gap is closely related to the model’s tendency in capturing HFC: models trained with bigger epoch sizes are more invariant to HFC and introduce smaller differences in training accuracy and testing accuracy. The observed relation is intuitive because the smallest generalization gap will be achieved once the model behaves like a human (because it is the human who annotate the data).

The observation in Figure 4 also chips in the discussion in the previous section about “generalizable” features. Intuitively, with bigger epoch size, the features that can lead to steepest descent of the loss surface are more likely to be the “generalizable” patterns of the data, which are LFC.

Heuristics: We also test how different training methods react to LFC and HFC, including

Dropout : A heuristic that drops weights randomly during training. We apply dropout on fully-connected layers with p=0.5p=0.5.

Mix-up : A heuristic that linearly integrate samples and their labels during training. We apply it with standard hyperparameter α=0.5\alpha=0.5.

BatchNorm : A method that perform the normalization for each training mini-batch to accelerate Deep Network training process. It allows us to use a much higher learning rate and reduce overfitting, similar with Dropout. We apply it with setting scale γ\gamma to 1 and offset β\beta to 0.

Adversarial Training : A method that augments the data through adversarial examples generated by a threat model during training. It is widely considered as one of the most successful adversarial robustness (defense) method. Following the popular choice, we use PGD with ϵ=8/255\epsilon=8/255 (ϵ=0.03\epsilon=0.03 ) as the threat model.

We illustrate the results in Figure 5, where the first panel is the vanilla set-up, and then each one of the four heuristics are tested in the following four panels.

Dropout roughly behaves similarly to the vanilla set-up in our experiments. Mix-up delivers a similar prediction accuracy, however, it catches much more HFC, which is probably not surprising because the mix-up augmentation does not encourage anything about LFC explicitly, and the performance gain is likely due to attention towards HFC.

Adversarial training mostly behaves as expected: it reports a lower prediction accuracy, which is likely due to the trade-off between robustness and accuracy. It also reports a smaller generalization gap, which is likely as a result of picking up “generalizable” patterns, as verified by its invariance towards HFC (e.g., r=12r=12 or r=16r=16). However, adversarial training seems to be sensitive to the HFC when r=4r=4, which is ignored even by the vanilla set-up.

The performance of BatchNorm is notable: compared to the vanilla set-up, BatchNorm picks more information in both LFC and HFC, especially when r=4r=4 and r=8r=8. This BatchNorm’s tendency in capturing HFC is also related to observations that BatchNorm encourages adversarial vulnerability .

Other Tests: We have also tested other heuristics or methods by only changing along one dimension while the rest is fixed the same as the vanilla set-up in Section 4.

Model architecture: We tested LeNet , AlexNet , VGG , and ResNet . The ResNet architecture seems advantageous toward previous inventions at different levels: it reports better vanilla test accuracy, smaller generalization gap (difference between training and testing accuracy), and a weaker tendency in capturing HFC.

Optimizer: We tested SGD, ADAM , AdaGrad , AdaDelta , and RMSprop. We notice that SGD seems to be the only one suffering from the tendency towards significantly capturing HFC, while the rest are on par within our experiments.

2 A hypothesis on Batch Normalization

Based on the observation, we hypothesized that one of BatchNorm’s advantage is, through normalization, to align the distributional disparities of different predictive signals. For example, HFC usually shows smaller magnitude than LFC, so a model trained without BatchNorm may not easily pick up these HFC. Therefore, the higher convergence speed may also be considered as a direct result of capturing different predictive signals simultaneously.

To verify this hypothesis, we compare the performance of models trained with vs. without BatchNorm over LFC data and plot the results in Figure 6.

As Figure 6 shows, when the model is trained with only LFC, BatchNorm does not always help improve the predictive performance, either tested by original data or by corresponding LFC data. Also, the smaller the radius is, the less the BatchNorm helps. Also, in our setting, BatchNorm does not generalize as well as the vanilla setting, which may raise a question about the benefit of BatchNorm.

However, BatchNorm still seems to at least boost the convergence of training accuracy. Interestingly, the acceleration is the smallest when r=4r=4. This observation further aligns with our hypothesis: if one of BatchNorm’s advantage is to encourage the model to capture different predictive signals, the performance gain of BatchNorm is the most limited when the model is trained with LFC when r=4r=4.

Adversarial Attack & Defense

As one may notice, our observation of HFC can be directly linked to the phenomenon of “adversarial example”: if the prediction relies on HFC, then perturbation of HFC will significantly alter the model’s response, but such perturbation may not be observed to human at all, creating the unintuitive behavior of neural networks.

This section is devoted to study the relationship between adversarial robustness and model’s tendency in exploiting HFC. We first discuss the linkage between the “smoothness” of convolutional kernels and model’s sensitivity towards HFC (section 6.1), which serves the tool for our follow-up analysis. With such tool, we first show that adversarially robust models tend to have “smooth” kernels (section 6.2), and then demonstrate that directly smoothing the kernels (without training) can help improve the adversarial robustness towards some attacks (section 6.3).

As convolutional theorem states, the convolution operation of images is equivalent to the element-wise multiplication of image frequency domain. Therefore, roughly, if a convolutional kernel has negligible weight at the high-end of the frequency domain, it will weigh HFC accordingly. This may only apply to the convolutional kernel at the first layer because the kernels at higher layer do not directly with the data, thus the relationship is not clear.

Therefore, we argue that, to push the model to ignore the HFC, one can consider to force the model to learn the convolutional kernels that have only negligible weights at the high-end of the frequency domain.

Intuitively (from signal processing knowledge), if the convolutional kernel is “smooth”, which means that there is no dramatics fluctuations between adjacent weights, the corresponding frequency domain will see a negligible amount of high-frequency signals. The connections have been mathematically proved , but these proved exact relationships are out of the scope of this paper.

2 Robust Models Have Smooth Kernels

To understand the connection between “smoothness” and adversarial robustness, we visualize the convolutional kernels at the first layer of the models trained in the vanilla manner (Mnatural) and trained with adversarial training (Madversarial) in Figure 7 (a) and (b).

Comparing Figure 7(a) and Figure 7(b), we can see that the kernels of Madversarial tend to show a more smooth pattern, which can be observed by noticing that the adjacent weights of kernels of Madversarial tend to share the same color. The visualization may not be very clear because the convolutional kernel is only [3 ×\times 3] in ResNet, the message is delivered more clearly in Appendix with other architecture when the first layer has kernel of the size [5 ×\times 5].

3 Smoothing Kernels Improves Adversarial Robustness

The intuitive argument in section 6.1 and empirical findings in section 6.2 directly lead to a question of whether we can improve the adversarial robustness of models by smoothing the convolutional kernels at the first layer.

Following the discussion, we introduce an extremely simple method that appears to improve the adversarial robustness against FGSM and PGD . For a convolutional kernel w\mathbf{w}, we use ii and jj to denote its column and row indices, thus wi,j\mathbf{w}_{i,j} denotes the value at iith row and jjth column. If we use N(i,j)\mathcal{N}(i,j) to denote the set of the spatial neighbors of (i,j)(i,j), our method is simply:

where ρ\rho is a hyperparameter of our method. We fix N(i,j)\mathcal{N}(i,j) to have eight neighbors. If (i,j)(i,j) is at the edge, then we simply generate the out-of-boundary values by duplicating the values on the boundary.

In other words, we try to smooth the kernel through simply reducing the adjacent differences by mixing the adjacent values. The method barely has any computational load, but appears to improve the adversarial robustness of Mnatural and Madversarial towards FGSM and PGD, even when Madversarial is trained with PGD as the threat model.

In Figure 7, we visualize the convolutional kernels with our method applied to Mnatural and Madversarial with ρ=1.0\rho=1.0, denoted as Mnatural(ρ=1.0\rho=1.0) and Madversarial(ρ=1.0\rho=1.0), respectively. As the visualization shows, the resulting kernels tend to show a significantly smoother pattern.

We test the robustness of the models smoothed by our method against FGSM and PGD with different choices of ϵ\epsilon, where the maximum of perturbation is 1.0. As Table 2 shows, when our smoothing method is applied, the performance of clean accuracy directly plunges, but the performance of adversarial robustness improves. In particular, our method helps when the perturbation is allowed to be relatively large. For example, when ϵ=0.09\epsilon=0.09 (roughly 23/25523/255), Mnatural(ρ=1.0\rho=1.0) even outperforms Madversarial. In general, our method can easily improve the adversarial robustness of Mnatural, but can only improve upon Madversarial in the case where ϵ\epsilon is larger, which is probably because the Madversarial is trained with PGD(ϵ=0.03\epsilon=0.03) as the threat model.

Beyond Image Classification

We aim to explore more than image classification tasks. We investigate in the object detection task. We use RetinaNet with ResNet50 + FPN as the backbone. We train the model with COCO detection train set and perform inference in its validation set, which includes 5000 images, and achieve an MAP of 35.6%35.6\%.

Then we choose r=128r=128 and maps the images into xl\mathbf{x}_{l} and xh\mathbf{x}_{h} and test with the same model and get 27.5%27.5\% MAP with LFC and 10.7%10.7\% MAP with HFC. The performance drop from 35.6%35.6\% to 27.5%27.5\% intrigues us so we further study whether the same drop should be expected from human.

The performance drop from the x\mathbf{x} to xl\mathbf{x}_{l} may be expected because xl\mathbf{x}_{l} may not have the rich information from the original images when HFC are dropped. In particular, different from image classification, HFC may play a significant role in depicting some objects, especially the smaller ones.

Figure 8 illustrates a few examples, where some objects are recognized worse in terms of MAP scores when the input images are replaced by the low-frequent counterparts. This disparity may be expected because the low-frequent images tend to be blurry and some objects may not be clear to a human either (as the left image represents).

2 Performance Gain on LFC

However, the disparity gets interesting when we inspect the performance gap in the opposite direction. We identified 1684 images that for each of these images, the some objects are recognized better (high MAP scores) in comparison to the original images.

The results are shown in Figure 9. There seems no apparent reasons why these objects are recognized better in low-frequent images, when inspected by human. These observations strengthen our argument in the perceptual disparity between CNN and human also exist in more advanced computer vision tasks other than image classification.

Discussion: Are HFC just Noises?

To answer this question, we experiment with another frequently used image denoising method: truncated singular value decomposition (SVD). We decompose the image and separate the image into one reconstructed with dominant singular values and one with trailing singular values. With this set-up, we find much fewer images supporting the story in Figure 2. Our observations suggest the signal CNN exploit is more than just random “noises”.

Conclusion & Outlook

We investigated how image frequency spectrum affects the generalization behavior of CNN, leading to multiple interesting explanations of the generalization behaviors of neural networks from a new perspective: there are multiple signals in the data, and not all of them align with human’s visual preference. As the paper comprehensively covers many topics, we briefly reiterate the main lessons learned:

CNN may capture HFC that are misaligned with human visual preference (section 3), resulting in generalization mysteries such as the paradox of learning label-shuffled data (section 4) and adversarial vulnerability (section 6).

Heuristics that improve accuracy (e.g., Mix-up and BatchNorm) may encourage capturing HFC (section 5). Due to the trade-off between accuracy and robustness (section 3), we may have to rethink the value of them.

Adversarially robust models tend to have smooth convolutional kernels, the reverse is not always true (section 6).

Similar phenomena are noticed in the context of object detection (section 7), with more conclusions yet to be drawn.

Looking forward, we hope our work serves as a call towards future era of computer vision research, where the state-of-the-art is not as important as we thought.

A single numeric on the leaderboard, while can significantly boost the research towards a direction, does not reliably reflect the alignment between models and human, while such an alignment is arguably paramount.

We hope our work will set forth towards a new testing scenario where the performance of low-frequent counterparts needs to be reported together with the performance of the original images.

Explicit inductive bias considering how a human views the data (e.g., ) may play a significant role in the future. In particular, neuroscience literature have shown that human tend to rely on low-frequent signals in recognizing objects , which may inspire development of future methods.

Acknowledgements

This material is based upon work supported by NIH R01GM114311, NIH P30DA035778, and NSF IIS1617583. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Institutes of Health or the National Science Foundation.

References