Natural Adversarial Examples

Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, Dawn Song

Introduction

Research on the ImageNet imagenet benchmark has led to numerous advances in classification AlexNet, object detection Huang2017SpeedAccuracyTF, and segmentation He2018MaskR. ImageNet classification improvements are broadly applicable and highly predictive of improvements on many tasks Kornblith2018DoBI. Improvements on ImageNet classification have been so great that some call ImageNet classifiers “superhuman” He2015DelvingDI. However, performance is decidedly subhuman when the test distribution does not match the training distribution hendrycks2019robustness. The distribution seen at test-time can include inclement weather conditions and obscured objects, and it can also include objects that are anomalous.

Recht et al., 2019 Recht2019DoIC remind us that ImageNet test examples tend to be simple, clear, close-up images, so that the current test set may be too easy and may not represent harder images encountered in the real world. Geirhos et al., 2020 argue that image classification datasets contain “spurious cues” or “shortcuts” Geirhos2020ShortcutLI; Arjovsky2019InvariantRM. For instance, models may use an image’s background to predict the foreground object’s class; a cow tends to co-occur with a green pasture, and even though the background is inessential to the object’s identity, models may predict “cow” primarily using the green pasture background cue. When datasets contain spurious cues, they can lead to performance estimates that are optimistic and inaccurate.

To counteract this, we curate two hard ImageNet test sets of natural adversarial examples with adversarial filtration. By using adversarial filtration, we can test how well models perform when simple-to-classify examples are removed, which includes examples that are solved with simple spurious cues. Some examples are depicted in Figure 1, which are simple for humans but hard for models. Our examples demonstrate that it is possible to reliably fool many models with clean natural images, while previous attempts at exposing and measuring model fragility rely on synthetic distribution corruptions geirhos; hendrycks2019robustness, artistic renditions Hendrycks2020TheMF, and adversarial distortions.

We demonstrate that clean examples can reliably degrade and transfer to other unseen classifiers using our first dataset. We call this dataset ImageNet-A, which contains images from a distribution unlike the ImageNet training distribution. ImageNet-A examples belong to ImageNet classes, but the examples are harder and can cause mistakes across various models. They cause consistent classification mistakes due to scene complications encountered in the long tail of scene configurations and by exploiting classifier blind spots (see Section 3.2). Since examples transfer reliably, this dataset shows models have unappreciated shared weaknesses.

The second dataset allows us to test model uncertainty estimates when semantic factors of the data distribution shift. Our second dataset is ImageNet-O, which contains image concepts from outside ImageNet-1K. These out-of-distribution images reliably cause models to mistake the examples as high-confidence in-distribution examples. To our knowledge this is the first dataset of anomalies or out-of-distribution examples developed to test ImageNet models. While ImageNet-A enables us to test image classification performance when the input data distribution shifts, ImageNet-O enables us to test out-of-distribution detection performance when the label distribution shifts.

We examine methods to improve performance on adversarially filtered examples. However, this is difficult because Figure 2 shows that examples successfully transfer to unseen or black-box models. To improve robustness, numerous techniques have been proposed. We find data augmentation techniques such as adversarial training decrease performance, while others can help by a few percent. We also find that a 10×10\times increase in training data corresponds to a less than a 10%10\% increase in accuracy. Finally, we show that improving model architectures is a promising avenue toward increasing robustness. Even so, current models have substantial room for improvement. Code and our two datasets are available at github.com/hendrycks/natural-adv-examples.

Related Work

Out-of-Distribution Detection. For out-of-distribution (OOD) detection hendrycks17baseline; kimin; hendrycks2019oe; hendrycks2019selfsupervised models learn a distribution, such as the ImageNet-1K distribution, and are tasked with producing quality anomaly scores that distinguish between usual test set examples and examples from held-out anomalous distributions. For instance, Hendrycks et al., 2017 hendrycks17baseline treat CIFAR-10 as the in-distribution and treat Gaussian noise and the SUN scene dataset Xiao2010SUNDL as out-of-distribution data. They show that the negative of the maximum softmax probability, or the the negative of the classifier prediction probability, is a high-performing anomaly score that can separate in- and out-of-distribution examples, so much so that it remains competitive to this day. Since that time, other work on out-of-distribution detection has continued to use datasets from other research benchmarks as anomaly stand-ins, producing far-from-distribution anomalies. Using visually dissimilar research datasets as anomaly stand-ins is critiqued in Ahmed et al., 2019 Ahmed2019DetectingSA. Some previous OOD detection datasets are depicted in the bottom row of Figure 3 hendrycks2019oe. Many of these anomaly sources are unnatural and deviate in numerous ways from the distribution of usual examples. In fact, some of the distributions can be deemed anomalous from local image statistics alone. Next, Meinke et al., 2019 Meinke2019TowardsNN propose studying adversarial out-of-distribution detection by detecting adversarially optimized uniform noise. In contrast, we propose a dataset for more realistic adversarial anomaly detection; our dataset contains hard anomalies generated by shifting the distribution’s labels and keeping non-semantic factors similar to the original training distribution.

Spurious Cues and Unintended Shortcuts. Models may learn spurious cues and obtain high accuracy, but for the wrong reasons Lapuschkin2019UnmaskingCH; Geirhos2020ShortcutLI. Spurious cues are a studied problem in natural language processing Cai2017PayAT; Gururangan2018AnnotationAI. Many recently introduced NLP datasets use adversarial filtration to create “adversarial datasets” by sieving examples solved with simple spurious cues Sakaguchi2019WINOGRANDEAA; Bhagavatula2019AbductiveCR; Zellers2019HellaSwagCA; Dua2019DROPAR; Bisk2020PIQARA; Hendrycks2020AligningAW. Like this recent concurrent research, we also use adversarial filtration Sung1995LearningAE, but the technique of adversarial filtration has not been applied to collecting image datasets until this paper. Additionally, adversarial filtration in NLP removes only the easiest examples, while we use filtration to select only the hardest examples and ignore examples of intermediate difficulty. Adversarially filtered examples for NLP also do not reliably transfer even to weaker models. In Bisk et al., 2019 Bisk2019PIQARA BERT errors do not reliably transfer to weaker GPT-1 models. This is one reason why it is not obvious a priori whether adversarially filtered images should transfer. In this work, we show that adversarial filtration algorithms can find examples that reliably transfer to both weaker and stronger models. Since adversarial filtration can remove examples that are solved by simple spurious cues, models must learn more robust features for our datasets.

Robustness to Shifted Input Distributions. Recht et al., 2019 Recht2019DoIC create a new ImageNet test set resembling the original test set as closely as possible. They found evidence that matching the difficulty of the original test set required selecting images deemed the easiest and most obvious by Mechanical Turkers. However, Engstrom et al., 2020 Engstrom2020IdentifyingSB estimate that the accuracy drop from ImageNet to ImageNetV2 is less than 3.6%3.6\%. In contrast, model accuracy can decrease by over 50%50\% with ImageNet-A. Brendel et al., 2018 Brendel2018ApproximatingCW show that classifiers that do not know the spatial ordering of image regions can be competitive on the ImageNet test set, possibly due to the dataset’s lack of difficulty. Judging classifiers by their performance on easier examples has potentially masked many of their shortcomings. For example, Geirhos et al., 2019 geirhos2019 artificially overwrite each ImageNet image’s textures and conclude that classifiers learn to rely on textural cues and under-utilize information about object shape. Recent work shows that classifiers are highly susceptible to non-adversarial stochastic corruptions hendrycks2019robustness. While they distort images with 7575 different algorithmically generated corruptions, our sources of distribution shift tend to be more heterogeneous and varied, and our examples are naturally occurring.

ImageNet-A and ImageNet-O

ImageNet-A is a dataset of real-world adversarially filtered images that fool current ImageNet classifiers. To find adversarially filtered examples, we first download numerous images related to an ImageNet class. Thereafter we delete the images that fixed ResNet-50 resnet classifiers correctly predict. We chose ResNet-50 due to its widespread use. Later we show that examples which fool ResNet-50 reliably transfer to other unseen models. With the remaining incorrectly classified images, we manually select visually clear images.

Next, ImageNet-O is a dataset of adversarially filtered examples for ImageNet out-of-distribution detectors. To create this dataset, we download ImageNet-22K and delete examples from ImageNet-1K. With the remaining ImageNet-22K examples that do not belong to ImageNet-1K classes, we keep examples that are classified by a ResNet-50 as an ImageNet-1K class with high confidence. Then we manually select visually clear images.

Both datasets were manually constructed by graduate students over several months. This is because a large share of images contain multiple classes per image Stock2018ConvNetsAI. Therefore, producing a dataset without multilabel images can be challenging with usual annotation techniques. To ensure images do not fall into more than one of the several hundred classes, we had graduate students memorize the classes in order to build a high-quality test set.

ImageNet-A Class Restrictions. We select a 200200-class subset of ImageNet-1K’s 1,0001,000 classes so that errors among these 200200 classes would be considered egregious imagenet. For instance, wrongly classifying Norwich terriers as Norfolk terriers does less to demonstrate faults in current classifiers than mistaking a Persian cat for a candle. We additionally avoid rare classes such as “snow leopard,” classes that have changed much since 2012 such as “iPod,” coarse classes such as “spiral,” classes that are often image backdrops such as “valley,” and finally classes that tend to overlap such as “honeycomb,” “bee,” “bee house,” and “bee eater”; “eraser,” “pencil sharpener” and “pencil case”; “sink,” “medicine cabinet,” “pill bottle” and “band-aid”; and so on. The 200200 ImageNet-A classes cover most broad categories spanned by ImageNet-1K; see the Supplementary Materials for the full class list.

ImageNet-A Data Aggregation. The first step is to download many weakly labeled images. Fortunately, the website iNaturalist has millions of user-labeled images of animals, and Flickr has even more user-tagged images of objects. We download images related to each of the 200200 ImageNet classes by leveraging user-provided labels and tags. After exporting or scraping data from sites including iNaturalist, Flickr, and DuckDuckGo, we adversarially select images by removing examples that fail to fool our ResNet-50 models. Of the remaining images, we select low-confidence images and then ensure each image is valid through human review. If we only used the original ImageNet test set as a source rather than iNaturalist, Flickr, and DuckDuckGo, some classes would have zero images after the first round of filtration, as the original ImageNet test set is too small to contain hard adversarially filtered images.

We now describe this process in more detail. We use a small ensemble of ResNet-50s for filtering, one pre-trained on ImageNet-1K then fine-tuned on the 200200 class subset, and one pre-trained on ImageNet-1K where 200200 of its 1,0001,000 logits are used in classification. Both classifiers have similar accuracy on the 200200 clean test set classes from ImageNet-1K. The ResNet-50s perform 10-crop classification for each image, and should any crop be classified correctly by the ResNet-50s, the image is removed. If either ResNet-50 assigns greater than 15%15\% confidence to the correct class, the image is also removed; this is done so that adversarially filtered examples yield misclassifications with low confidence in the correct class, like in untargeted adversarial attacks. Now, some classification confusions are greatly over-represented, such as Persian cat and lynx. We would like ImageNet-A to have great variability in its types of errors and cause classifiers to have a dense confusion matrix. Consequently, we perform a second round of filtering to create a shortlist where each confusion only appears at most 1515 times. Finally, we manually select images from this shortlist in order to ensure ImageNet-A images are simultaneously valid, single-class, and high-quality. In all, the ImageNet-A dataset has 7,5007,500 adversarially filtered images.

As a specific example, we download 81,41381,413 dragonfly images from iNaturalist, and after running the ResNet-50 filter we have 8,9258,925 dragonfly images. In the algorithmically diversified shortlist, 1,4521,452 images remain. From this shortlist, 8080 dragonfly images are manually selected, but hundreds more could be selected if time allows.

The resulting images represent a substantial distribution shift, but images are still possible for humans to classify. The Fréchet Inception Distance (FID) Heusel2017GANsTB enables us to determine whether ImageNet-A and ImageNet are not identically distributed. The FID between ImageNet’s validation and test set is approximately 0.990.99, indicating that the distributions are highly similar. The FID between ImageNet-A and ImageNet’s validation set is 50.4050.40, and the FID between ImageNet-A and ImageNet’s test set is approximately 50.2550.25, indicating that the distribution shift is large. Despite the shift, we estimate that our graduate students’ ImageNet-A human accuracy rate is approximately 90%90\%.

ImageNet-O Class Restrictions. We again select a 200-class subset of ImageNet-1K’s 1,0001,000 classes. These 200200 classes determine the in-distribution or the distribution that is considered usual. As before, the 200200 classes cover most broad categories spanned by ImageNet-1K; see the Supplementary Materials for the full class list.

ImageNet-O Data Aggregation. Our dataset for adversarial out-of-distribution detection is created by fooling ResNet-50 out-of-distribution detectors. The negative of the prediction confidence of a ResNet-50 ImageNet classifier serves as our anomaly score hendrycks17baseline. Usually in-distribution examples produce higher confidence predictions than OOD examples, but we curate OOD examples that have high confidence predictions. To gather candidate adversarially filtered examples, we use the ImageNet-22K dataset with ImageNet-1K classes deleted. We choose the ImageNet-22K dataset since it was collected in the same way as ImageNet-1K. ImageNet-22K allows us to have coverage of numerous visual concepts and vary the distribution’s semantics without unnatural or unwanted non-semantic data shift. After excluding ImageNet-1K images, we process the remaining ImageNet-22K images and keep the images which cause the ResNet-50 to have high confidence, or a low anomaly score. We then manually select a high-quality subset of the remaining images to create ImageNet-O. We suggest only training models with data from the 1,0001,000 ImageNet-1K classes, since the dataset becomes trivial if models train on ImageNet-22K. To our knowledge, this dataset is the first anomalous dataset curated for ImageNet models and enables researchers to study adversarial out-of-distribution detection. The ImageNet-O dataset has 2,0002,000 adversarially filtered examples since anomalies are rarer; this has the same number of examples per class as ImageNetV2 Recht2019DoIC. While we use adversarial filtration to select images that are difficult for a fixed ResNet-50, we will show these examples straightforwardly transfer to unseen models.

2 Illustrative Failure Modes

Examples in ImageNet-A uncover numerous failure modes of modern convolutional neural networks. We describe our findings after having viewed tens of thousands of candidate adversarially filtered examples. Some of these failure modes may also explain poor ImageNet-O performance, but for simplicity we describe our observations with ImageNet-A examples.

Consider Figure 6. The first two images suggest models may overgeneralize visual concepts. It may confuse metal with sundials, or thin radiating lines with harvestman bugs. We also observed that networks overgeneralize tricycles to bicycles and circles, digital clocks to keyboards and calculators, and more. We also observe that models may rely too heavily on color and texture, as shown with the dragonfly images. Since classifiers are taught to associate entire images with an object class, frequently appearing background elements may also become associated with a class, such as wood being associated with nails. Other examples include classifiers heavily associating hummingbird feeders with hummingbirds, leaf-covered tree branches being associated with the white-headed capuchin monkey class, snow being associated with shovels, and dumpsters with garbage trucks. Additionally Figure 6 shows an American alligator swimming. With different frames, the classifier prediction varies erratically between classes that are semantically loose and separate. For other images of the swimming alligator, classifiers predict that the alligator is a cliff, lynx, and a fox squirrel. Assessing convolutional networks on ImageNet-A reveals that even state-of-the-art models have diverse and systematic failure modes.

Experiments

We show that adversarially filtered examples collected to fool fixed ResNet-50 models reliably transfer to other models, indicating that current convolutional neural networks have shared weaknesses and failure modes. In the following sections, we analyze whether robustness can be improved by using data augmentation, using more real labeled data, and using different architectures. For the first two sections, we analyze performance with a fixed architecture for comparability, and in the final section we observe performance with different architectures. First we define our metrics.

Metrics. Our metric for assessing robustness to adversarially filtered examples for classifiers is the top-1 accuracy on ImageNet-A. For reference, the top-1 accuracy on the 200 ImageNet-A classes using usual ImageNet images is usually greater than or equal to 90%90\% for ordinary classifiers.

Our metric for assessing out-of-distribution detection performance of ImageNet-O examples is the area under the precision-recall curve (AUPR). This metric requires anomaly scores. Our anomaly score is the negative of the maximum softmax probabilities hendrycks17baseline from a model that can classify the 200200 ImageNet-O classes. The maximum softmax probability detector is a long-standing baseline in OOD detection. We collect anomaly scores with the ImageNet validation examples for the said 200200 classes. Then, we collect anomaly scores for the ImageNet-O examples. Higher performing OOD detectors would assign ImageNet-O examples lower confidences, or higher anomaly scores. With these anomaly scores, we can compute the area under the precision-recall curve auprbaseline. Random chance levels for the AUPR is approximately 16.67%16.67\% with ImageNet-O, and the maximum AUPR is 100%100\%.

More Labeled Data.

One possible explanation for consistently low ImageNet-A accuracy is that all models are trained only with ImageNet-1K, and using additional data may resolve the problem. Bau et al., 2017 Bau2017NetworkDQ argue that Places365 classifiers learn qualitatively distinct filters (e.g., they have more object detectors, fewer texture detectors in conv3) compared to ImageNet classifiers, so one may expect an error distribution less correlated with errors on ImageNet-A. To test this hypothesis we pre-train a ResNet-50 on Places365 zhou2017places, a large-scale scene recognition dataset. After fine-tuning the Places365 model on ImageNet-1K, we find that accuracy is 1.56%1.56\%. Consequently, even though scene recognition models are purported to have qualitatively distinct features, this is not enough to improve ImageNet-A performance. Likewise, Places365 pre-training does not improve ImageNet-O detection, as its AUPR is 14.88%14.88\%. Next, we see whether labeled data from ImageNet-A itself can help. We take baseline ResNet-50 with 2.17%2.17\% ImageNet-A accuracy and fine-tune it on 80%80\% of ImageNet-A. This leads to no clear improvement on the remaining 20%20\% of ImageNet-A since the top-1 and top-5 accuracies are below 2%2\% and 5%5\%, respectively.

Last, we pre-train using an order of magnitude more training data with ImageNet-21K. This dataset contains approximately 21,00021,000 classes and approximately 1414 million images. To our knowledge this is the largest publicly available database of labeled natural images. Using a ResNet-50 pretrained on ImageNet-21K, we fine-tune the model on ImageNet-1K and attain 11.41%11.41\% accuracy on ImageNet-A, a 9.24%9.24\% increase. Likewise, the AUPR for ImageNet-O improves from 16.20%16.20\% to 21.86%21.86\%, although this improvement is less significant since ImageNet-O images overlap with ImageNet-21K images. Academic researchers rarely use datasets larger than ImageNet due to computational costs, using more data has limitations. An order of magnitude increase in labeled training data can provide some improvements in accuracy, though we now show that architecture changes provide greater improvements.

Architectural Changes.

Another useful architecture change is self-attention. Convolutional neural networks with self-attention Hu2018GatherExciteE are designed to better capture long-range dependencies and interactions across an image. We consider the self-attention technique called Squeeze-and-Excitation (SE) Hu2018SqueezeandExcitationN, which won the final ImageNet competition in 2017. A ResNet-50 with Squeeze-and-Excitation attains 6.17%6.17\% accuracy. However, for larger ResNets, self-attention does little to improve ImageNet-O detection.

We consider the ResNet-50 architecture with its residual blocks exchanged with recently introduced Res2Net v1b blocks Gao2019Res2NetAN. This change increases accuracy to 14.59%14.59\% and the AUPR to 19.5%19.5\%. A ResNet-152 with Res2Net v1b blocks attains 22.4%22.4\% accuracy and 23.9%23.9\% AUPR. Compared to data augmentation or an order of magnitude more labeled training data, some architectural changes can provide far more robustness gains. Consequently future improvements to model architectures is a promising path towards greater robustness.

We now assess performance on a completely different architecture which does not use convolutions, vision Transformers Dosovitskiy2020AnII. We evaluate with DeiT touvron2020deit, a vision Transformer trained on ImageNet-1K with aggressive data augmentation such as Mixup. Even for vision Transformers, we find that ImageNet-A and ImageNet-O examples successfully transfer. In particular, a DeiT-small vision Transformer gets 19.0% on ImageNet-A and has a similar number of parameters to a Res2Net-50, which has 14.6% accuracy. This might be explained by DeiT’s use of Mixup, however, which provided a 4% ImageNet-A accuracy boost for ResNets. The ImageNet-O AUPR for the Transformer is 20.9%, while the Res2Net gets 19.5%. Larger DeiT models do better, as a DeiT-base gets 28.2% accuracy on ImageNet-A and 24.8% AUPR on ImageNet. Consequently, our datasets transfer to vision Transformers and performance for both tasks remains far from the ceiling.

Conclusion

We found it is possible to improve performance on our datasets with data augmentation, pretraining data, and architectural changes. We found that our examples transferred to all tested models, including vision Transformers which do not use convolution operations. Results indicate that improving performance on ImageNet-A and ImageNet-O is possible but difficult. Our challenging ImageNet test sets serve as measures of performance under distribution shift—an important research aim as models are deployed in increasingly precarious real-world environments.

References

Appendix

Expanded Results

Full results with various architectures are in Table 1.

2 More OOD Detection Results and Background

Works in out-of-distribution detection frequently use the maximum softmax baseline to detect out-of-distribution examples hendrycks17baseline. Before neural networks, using the reject option or a k+1k+1st class was somewhat common Bartlett2008ClassificationWA, but with neural networks it requires auxiliary anomalous training data. New neural methods that utilize auxiliary anomalous training data, such as Outlier Exposure hendrycks2019oe, do not use the reject option and still utilize the maximum softmax probability. We do not use Outlier Exposure since that paper’s authors were unable to get their technique to work on ImageNet-1K with 224×224224\times 224 images, though they were able to get it work on Tiny ImageNet which has 64×6464\times 64 images. We do not use ODIN since it requires tuning hyperparameters directly using out-of-distribution data, a criticized practice hendrycks2019oe.

We evaluate three additional out-of-distribution detection methods, though none substantially improve performance. We evaluate method of Devries2018LearningCF, which trains an auxiliary branch to represent the model confidence. Using a ResNet trained from scratch, we find this gets a 14.3%14.3\% AUPR, around 2% less than the MSP baseline. Next we use the recent Maximum Logit detector Hendrycks2020ScalingOD. With DenseNet-121 the AUPR decreases from 16.1%16.1\% (MSP) to 15.8%15.8\% (Max Logit), while with ResNeXt-101 (32×832\times 8d) the AUPR of 20.5%20.5\% increases to 20.6%20.6\%. Across over 10 models we found the MaxLogit technique to be slightly worse. Finally, we evaluate the utility of self-supervised auxiliary objectives for OOD detection. The rotation prediction anomaly detector Hendrycks2019UsingSL was shown to help improve detection performance for near-distribution yet still out-of-class examples, and with this auxiliary objective the AUPR for ResNet-50 does not change; it is 16.2%16.2\% with the rotation prediction and 16.2%16.2\% with the MSP. Note this method requires training the network and does not work out-of-the-box.

3 Calibration

In this section we show ImageNet-A calibration results.

Our second uncertainty estimation metric is the Area Under the Response Rate Accuracy Curve (AURRA). Responding only when confident is often preferable to predicting falsely. In these experiments, we allow classifiers to respond to a subset of the test set and abstain from predicting the rest. Classifiers with quality uncertainty estimates should be capable identifying examples it is likely to predict falsely and abstain. If a classifier is required to abstain from predicting on 90% of the test set, or equivalently respond to the remaining 10% of the test set, then we should like the classifier’s uncertainty estimates to separate correctly and falsely classified examples and have high accuracy on the selected 10%. At a fixed response rate, we should like the accuracy to be as high as possible. At a 100% response rate, the classifier accuracy is the usual test set accuracy. We vary the response rates and compute the corresponding accuracies to obtain the Response Rate Accuracy (RRA) curve. The area under the Response Rate Accuracy curve is the AURRA. To compute the AURRA in this paper, we use the maximum softmax probability. For response rate pp, we take the pp fraction of examples with highest maximum softmax probability. If the response rate is 10%, we select the top 10% of examples with the highest confidence and compute the accuracy on these examples. An example RRA curve is in Figure 10 .

ImageNet-A Classes

The 200 ImageNet classes that we selected for ImageNet-A are as follows. goldfish, great white shark, hammerhead, stingray, hen, ostrich, goldfinch, junco, bald eagle, vulture, newt, axolotl, tree frog, iguana, African chameleon, cobra, scorpion, tarantula, centipede, peacock, lorikeet, hummingbird, toucan, duck, goose, black swan, koala, jellyfish, snail, lobster, hermit crab, flamingo, american egret, pelican, king penguin, grey whale, killer whale, sea lion, chihuahua, shih tzu, afghan hound, basset hound, beagle, bloodhound, italian greyhound, whippet, weimaraner, yorkshire terrier, boston terrier, scottish terrier, west highland white terrier, golden retriever, labrador retriever, cocker spaniels, collie, border collie, rottweiler, german shepherd dog, boxer, french bulldog, saint bernard, husky, dalmatian, pug, pomeranian, chow chow, pembroke welsh corgi, toy poodle, standard poodle, timber wolf, hyena, red fox, tabby cat, leopard, snow leopard, lion, tiger, cheetah, polar bear, meerkat, ladybug, fly, bee, ant, grasshopper, cockroach, mantis, dragonfly, monarch butterfly, starfish, wood rabbit, porcupine, fox squirrel, beaver, guinea pig, zebra, pig, hippopotamus, bison, gazelle, llama, skunk, badger, orangutan, gorilla, chimpanzee, gibbon, baboon, panda, eel, clown fish, puffer fish, accordion, ambulance, assault rifle, backpack, barn, wheelbarrow, basketball, bathtub, lighthouse, beer glass, binoculars, birdhouse, bow tie, broom, bucket, cauldron, candle, cannon, canoe, carousel, castle, mobile phone, cowboy hat, electric guitar, fire engine, flute, gasmask, grand piano, guillotine, hammer, harmonica, harp, hatchet, jeep, joystick, lab coat, lawn mower, lipstick, mailbox, missile, mitten, parachute, pickup truck, pirate ship, revolver, rugby ball, sandal, saxophone, school bus, schooner, shield, soccer ball, space shuttle, spider web, steam locomotive, scarf, submarine, tank, tennis ball, tractor, trombone, vase, violin, military aircraft, wine bottle, ice cream, bagel, pretzel, cheeseburger, hotdog, cabbage, broccoli, cucumber, bell pepper, mushroom, Granny Smith, strawberry, lemon, pineapple, banana, pomegranate, pizza, burrito, espresso, volcano, baseball player, scuba diver, acorn,

n01443537, n01484850, n01494475, n01498041, n01514859, n01518878, n01531178, n01534433, n01614925, n01616318, n01630670, n01632777, n01644373, n01677366, n01694178, n01748264, n01770393, n01774750, n01784675, n01806143, n01820546, n01833805, n01843383, n01847000, n01855672, n01860187, n01882714, n01910747, n01944390, n01983481, n01986214, n02007558, n02009912, n02051845, n02056570, n02066245, n02071294, n02077923, n02085620, n02086240, n02088094, n02088238, n02088364, n02088466, n02091032, n02091134, n02092339, n02094433, n02096585, n02097298, n02098286, n02099601, n02099712, n02102318, n02106030, n02106166, n02106550, n02106662, n02108089, n02108915, n02109525, n02110185, n02110341, n02110958, n02112018, n02112137, n02113023, n02113624, n02113799, n02114367, n02117135, n02119022, n02123045, n02128385, n02128757, n02129165, n02129604, n02130308, n02134084, n02138441, n02165456, n02190166, n02206856, n02219486, n02226429, n02233338, n02236044, n02268443, n02279972, n02317335, n02325366, n02346627, n02356798, n02363005, n02364673, n02391049, n02395406, n02398521, n02410509, n02423022, n02437616, n02445715, n02447366, n02480495, n02480855, n02481823, n02483362, n02486410, n02510455, n02526121, n02607072, n02655020, n02672831, n02701002, n02749479, n02769748, n02793495, n02797295, n02802426, n02808440, n02814860, n02823750, n02841315, n02843684, n02883205, n02906734, n02909870, n02939185, n02948072, n02950826, n02951358, n02966193, n02980441, n02992529, n03124170, n03272010, n03345487, n03372029, n03424325, n03452741, n03467068, n03481172, n03494278, n03495258, n03498962, n03594945, n03602883, n03630383, n03649909, n03676483, n03710193, n03773504, n03775071, n03888257, n03930630, n03947888, n04086273, n04118538, n04133789, n04141076, n04146614, n04147183, n04192698, n04254680, n04266014, n04275548, n04310018, n04325704, n04347754, n04389033, n04409515, n04465501, n04487394, n04522168, n04536866, n04552348, n04591713, n07614500, n07693725, n07695742, n07697313, n07697537, n07714571, n07714990, n07718472, n07720875, n07734744, n07742313, n07745940, n07749582, n07753275, n07753592, n07768694, n07873807, n07880968, n07920052, n09472597, n09835506, n10565667, n12267677,

‘Stingray;’ ‘goldfinch, Carduelis carduelis;’ ‘junco, snowbird;’ ‘robin, American robin, Turdus migratorius;’ ‘jay;’ ‘bald eagle, American eagle, Haliaeetus leucocephalus;’ ‘vulture;’ ‘eft;’ ‘bullfrog, Rana catesbeiana;’ ‘box turtle, box tortoise;’ ‘common iguana, iguana, Iguana iguana;’ ‘agama;’ ‘African chameleon, Chamaeleo chamaeleon;’ ‘American alligator, Alligator mississipiensis;’ ‘garter snake, grass snake;’ ‘harvestman, daddy longlegs, Phalangium opilio;’ ‘scorpion;’ ‘tarantula;’ ‘centipede;’ ‘sulphur-crested cockatoo, Kakatoe galerita, Cacatua galerita;’ ‘lorikeet;’ ‘hummingbird;’ ‘toucan;’ ‘drake;’ ‘goose;’ ‘koala, koala bear, kangaroo bear, native bear, Phascolarctos cinereus;’ ‘jellyfish;’ ‘sea anemone, anemone;’ ‘flatworm, platyhelminth;’ ‘snail;’ ‘crayfish, crawfish, crawdad, crawdaddy;’ ‘hermit crab;’ ‘flamingo;’ ‘American egret, great white heron, Egretta albus;’ ‘oystercatcher, oyster catcher;’ ‘pelican;’ ‘sea lion;’ ‘Chihuahua;’ ‘golden retriever;’ ‘Rottweiler;’ ‘German shepherd, German shepherd dog, German police dog, alsatian;’ ‘pug, pug-dog;’ ‘red fox, Vulpes vulpes;’ ‘Persian cat;’ ‘lynx, catamount;’ ‘lion, king of beasts, Panthera leo;’ ‘American black bear, black bear, Ursus americanus, Euarctos americanus;’ ‘mongoose;’ ‘ladybug, ladybeetle, lady beetle, ladybird, ladybird beetle;’ ‘rhinoceros beetle;’ ‘weevil;’ ‘fly;’ ‘bee;’ ‘ant, emmet, pismire;’ ‘grasshopper, hopper;’ ‘walking stick, walkingstick, stick insect;’ ‘cockroach, roach;’ ‘mantis, mantid;’ ‘leafhopper;’ ‘dragonfly, darning needle, devil’s darning needle, sewing needle, snake feeder, snake doctor, mosquito hawk, skeeter hawk;’ ‘monarch, monarch butterfly, milkweed butterfly, Danaus plexippus;’ ‘cabbage butterfly;’ ‘lycaenid, lycaenid butterfly;’ ‘starfish, sea star;’ ‘wood rabbit, cottontail, cottontail rabbit;’ ‘porcupine, hedgehog;’ ‘fox squirrel, eastern fox squirrel, Sciurus niger;’ ‘marmot;’ ‘bison;’ ‘skunk, polecat, wood pussy;’ ‘armadillo;’ ‘baboon;’ ‘capuchin, ringtail, Cebus capucinus;’ ‘African elephant, Loxodonta africana;’ ‘puffer, pufferfish, blowfish, globefish;’ ‘academic gown, academic robe, judge’s robe;’ ‘accordion, piano accordion, squeeze box;’ ‘acoustic guitar;’ ‘airliner;’ ‘ambulance;’ ‘apron;’ ‘balance beam, beam;’ ‘balloon;’ ‘banjo;’ ‘barn;’ ‘barrow, garden cart, lawn cart, wheelbarrow;’ ‘basketball;’ ‘beacon, lighthouse, beacon light, pharos;’ ‘beaker;’ ‘bikini, two-piece;’ ‘bow;’ ‘bow tie, bow-tie, bowtie;’ ‘breastplate, aegis, egis;’ ‘broom;’ ‘candle, taper, wax light;’ ‘canoe;’ ‘castle;’ ‘cello, violoncello;’ ‘chain;’ ‘chest;’ ‘Christmas stocking;’ ‘cowboy boot;’ ‘cradle;’ ‘dial telephone, dial phone;’ ‘digital clock;’ ‘doormat, welcome mat;’ ‘drumstick;’ ‘dumbbell;’ ‘envelope;’ ‘feather boa, boa;’ ‘flagpole, flagstaff;’ ‘forklift;’ ‘fountain;’ ‘garbage truck, dustcart;’ ‘goblet;’ ‘go-kart;’ ‘golfcart, golf cart;’ ‘grand piano, grand;’ ‘hand blower, blow dryer, blow drier, hair dryer, hair drier;’ ‘iron, smoothing iron;’ ‘jack-o’-lantern;’ ‘jeep, landrover;’ ‘kimono;’ ‘lighter, light, igniter, ignitor;’ ‘limousine, limo;’ ‘manhole cover;’ ‘maraca;’ ‘marimba, xylophone;’ ‘mask;’ ‘mitten;’ ‘mosque;’ ‘nail;’ ‘obelisk;’ ‘ocarina, sweet potato;’ ‘organ, pipe organ;’ ‘parachute, chute;’ ‘parking meter;’ ‘piggy bank, penny bank;’ ‘pool table, billiard table, snooker table;’ ‘puck, hockey puck;’ ‘quill, quill pen;’ ‘racket, racquet;’ ‘reel;’ ‘revolver, six-gun, six-shooter;’ ‘rocking chair, rocker;’ ‘rugby ball;’ ‘saltshaker, salt shaker;’ ‘sandal;’ ‘sax, saxophone;’ ‘school bus;’ ‘schooner;’ ‘sewing machine;’ ‘shovel;’ ‘sleeping bag;’ ‘snowmobile;’ ‘snowplow, snowplough;’ ‘soap dispenser;’ ‘spatula;’ ‘spider web, spider’s web;’ ‘steam locomotive;’ ‘stethoscope;’ ‘studio couch, day bed;’ ‘submarine, pigboat, sub, U-boat;’ ‘sundial;’ ‘suspension bridge;’ ‘syringe;’ ‘tank, army tank, armored combat vehicle, armoured combat vehicle;’ ‘teddy, teddy bear;’ ‘toaster;’ ‘torch;’ ‘tricycle, trike, velocipede;’ ‘umbrella;’ ‘unicycle, monocycle;’ ‘viaduct;’ ‘volleyball;’ ‘washer, automatic washer, washing machine;’ ‘water tower;’ ‘wine bottle;’ ‘wreck;’ ‘guacamole;’ ‘pretzel;’ ‘cheeseburger;’ ‘hotdog, hot dog, red hot;’ ‘broccoli;’ ‘cucumber, cuke;’ ‘bell pepper;’ ‘mushroom;’ ‘lemon;’ ‘banana;’ ‘custard apple;’ ‘pomegranate;’ ‘carbonara;’ ‘bubble;’ ‘cliff, drop, drop-off;’ ‘volcano;’ ‘ballplayer, baseball player;’ ‘rapeseed;’ ‘yellow lady’s slipper, yellow lady-slipper, Cypripedium calceolus, Cypripedium parviflorum;’ ‘corn;’ ‘acorn.’

n01498041, n01531178, n01534433, n01558993, n01580077, n01614925, n01616318, n01631663, n01641577, n01669191, n01677366, n01687978, n01694178, n01698640, n01735189, n01770081, n01770393, n01774750, n01784675, n01819313, n01820546, n01833805, n01843383, n01847000, n01855672, n01882714, n01910747, n01914609, n01924916, n01944390, n01985128, n01986214, n02007558, n02009912, n02037110, n02051845, n02077923, n02085620, n02099601, n02106550, n02106662, n02110958, n02119022, n02123394, n02127052, n02129165, n02133161, n02137549, n02165456, n02174001, n02177972, n02190166, n02206856, n02219486, n02226429, n02231487, n02233338, n02236044, n02259212, n02268443, n02279972, n02280649, n02281787, n02317335, n02325366, n02346627, n02356798, n02361337, n02410509, n02445715, n02454379, n02486410, n02492035, n02504458, n02655020, n02669723, n02672831, n02676566, n02690373, n02701002, n02730930, n02777292, n02782093, n02787622, n02793495, n02797295, n02802426, n02814860, n02815834, n02837789, n02879718, n02883205, n02895154, n02906734, n02948072, n02951358, n02980441, n02992211, n02999410, n03014705, n03026506, n03124043, n03125729, n03187595, n03196217, n03223299, n03250847, n03255030, n03291819, n03325584, n03355925, n03384352, n03388043, n03417042, n03443371, n03444034, n03445924, n03452741, n03483316, n03584829, n03590841, n03594945, n03617480, n03666591, n03670208, n03717622, n03720891, n03721384, n03724870, n03775071, n03788195, n03804744, n03837869, n03840681, n03854065, n03888257, n03891332, n03935335, n03982430, n04019541, n04033901, n04039381, n04067472, n04086273, n04099969, n04118538, n04131690, n04133789, n04141076, n04146614, n04147183, n04179913, n04208210, n04235860, n04252077, n04252225, n04254120, n04270147, n04275548, n04310018, n04317175, n04344873, n04347754, n04355338, n04366367, n04376876, n04389033, n04399382, n04442312, n04456115, n04482393, n04507155, n04509417, n04532670, n04540053, n04554684, n04562935, n04591713, n04606251, n07583066, n07695742, n07697313, n07697537, n07714990, n07718472, n07720875, n07734744, n07749582, n07753592, n07760859, n07768694, n07831146, n09229709, n09246464, n09472597, n09835506, n11879895, n12057211, n12144580, n12267677.

ImageNet-O Classes

The 200 ImageNet classes that we selected for ImageNet-O are as follows.

‘goldfish, Carassius auratus;’ ‘triceratops;’ ‘harvestman, daddy longlegs, Phalangium opilio;’ ‘centipede;’ ‘sulphur-crested cockatoo, Kakatoe galerita, Cacatua galerita;’ ‘lorikeet;’ ‘jellyfish;’ ‘brain coral;’ ‘chambered nautilus, pearly nautilus, nautilus;’ ‘dugong, Dugong dugon;’ ‘starfish, sea star;’ ‘sea urchin;’ ‘hog, pig, grunter, squealer, Sus scrofa;’ ‘armadillo;’ ‘rock beauty, Holocanthus tricolor;’ ‘puffer, pufferfish, blowfish, globefish;’ ‘abacus;’ ‘accordion, piano accordion, squeeze box;’ ‘apron;’ ‘balance beam, beam;’ ‘ballpoint, ballpoint pen, ballpen, Biro;’ ‘Band Aid;’ ‘banjo;’ ‘barbershop;’ ‘bath towel;’ ‘bearskin, busby, shako;’ ‘binoculars, field glasses, opera glasses;’ ‘bolo tie, bolo, bola tie, bola;’ ‘bottlecap;’ ‘brassiere, bra, bandeau;’ ‘broom;’ ‘buckle;’ ‘bulletproof vest;’ ‘candle, taper, wax light;’ ‘car mirror;’ ‘chainlink fence;’ ‘chain saw, chainsaw;’ ‘chime, bell, gong;’ ‘Christmas stocking;’ ‘cinema, movie theater, movie theatre, movie house, picture palace;’ ‘combination lock;’ ‘corkscrew, bottle screw;’ ‘crane;’ ‘croquet ball;’ ‘dam, dike, dyke;’ ‘digital clock;’ ‘dishrag, dishcloth;’ ‘dogsled, dog sled, dog sleigh;’ ‘doormat, welcome mat;’ ‘drilling platform, offshore rig;’ ‘electric fan, blower;’ ‘envelope;’ ‘espresso maker;’ ‘face powder;’ ‘feather boa, boa;’ ‘fireboat;’ ‘fire screen, fireguard;’ ‘flute, transverse flute;’ ‘folding chair;’ ‘fountain;’ ‘fountain pen;’ ‘frying pan, frypan, skillet;’ ‘golf ball;’ ‘greenhouse, nursery, glasshouse;’ ‘guillotine;’ ‘hamper;’ ‘hand blower, blow dryer, blow drier, hair dryer, hair drier;’ ‘harmonica, mouth organ, harp, mouth harp;’ ‘honeycomb;’ ‘hourglass;’ ‘iron, smoothing iron;’ ‘jack-o’-lantern;’ ‘jigsaw puzzle;’ ‘joystick;’ ‘lawn mower, mower;’ ‘library;’ ‘lighter, light, igniter, ignitor;’ ‘lipstick, lip rouge;’ ‘loupe, jeweler’s loupe;’ ‘magnetic compass;’ ‘manhole cover;’ ‘maraca;’ ‘marimba, xylophone;’ ‘mask;’ ‘matchstick;’ ‘maypole;’ ‘maze, labyrinth;’ ‘medicine chest, medicine cabinet;’ ‘mortar;’ ‘mosquito net;’ ‘mousetrap;’ ‘nail;’ ‘neck brace;’ ‘necklace;’ ‘nipple;’ ‘ocarina, sweet potato;’ ‘oil filter;’ ‘organ, pipe organ;’ ‘oscilloscope, scope, cathode-ray oscilloscope, CRO;’ ‘oxygen mask;’ ‘paddlewheel, paddle wheel;’ ‘panpipe, pandean pipe, syrinx;’ ‘park bench;’ ‘pencil sharpener;’ ‘Petri dish;’ ‘pick, plectrum, plectron;’ ‘picket fence, paling;’ ‘pill bottle;’ ‘ping-pong ball;’ ‘pinwheel;’ ‘plate rack;’ ‘plunger, plumber’s helper;’ ‘pool table, billiard table, snooker table;’ ‘pot, flowerpot;’ ‘power drill;’ ‘prayer rug, prayer mat;’ ‘prison, prison house;’ ‘punching bag, punch bag, punching ball, punchball;’ ‘quill, quill pen;’ ‘radiator;’ ‘reel;’ ‘remote control, remote;’ ‘rubber eraser, rubber, pencil eraser;’ ‘rule, ruler;’ ‘safe;’ ‘safety pin;’ ‘saltshaker, salt shaker;’ ‘scale, weighing machine;’ ‘screw;’ ‘screwdriver;’ ‘shoji;’ ‘shopping cart;’ ‘shower cap;’ ‘shower curtain;’ ‘ski;’ ‘sleeping bag;’ ‘slot, one-armed bandit;’ ‘snowmobile;’ ‘soap dispenser;’ ‘solar dish, solar collector, solar furnace;’ ‘space heater;’ ‘spatula;’ ‘spider web, spider’s web;’ ‘stove;’ ‘strainer;’ ‘stretcher;’ ‘submarine, pigboat, sub, U-boat;’ ‘swimming trunks, bathing trunks;’ ‘swing;’ ‘switch, electric switch, electrical switch;’ ‘syringe;’ ‘tennis ball;’ ‘thatch, thatched roof;’ ‘theater curtain, theatre curtain;’ ‘thimble;’ ‘throne;’ ‘tile roof;’ ‘toaster;’ ‘tricycle, trike, velocipede;’ ‘turnstile;’ ‘umbrella;’ ‘vending machine;’ ‘waffle iron;’ ‘washer, automatic washer, washing machine;’ ‘water bottle;’ ‘water tower;’ ‘whistle;’ ‘Windsor tie;’ ‘wooden spoon;’ ‘wool, woolen, woollen;’ ‘crossword puzzle, crossword;’ ‘traffic light, traffic signal, stoplight;’ ‘ice lolly, lolly, lollipop, popsicle;’ ‘bagel, beigel;’ ‘pretzel;’ ‘hotdog, hot dog, red hot;’ ‘mashed potato;’ ‘broccoli;’ ‘cauliflower;’ ‘zucchini, courgette;’ ‘acorn squash;’ ‘cucumber, cuke;’ ‘bell pepper;’ ‘Granny Smith;’ ‘strawberry;’ ‘orange;’ ‘lemon;’ ‘pineapple, ananas;’ ‘banana;’ ‘jackfruit, jak, jack;’ ‘pomegranate;’ ‘chocolate sauce, chocolate syrup;’ ‘meat loaf, meatloaf;’ ‘pizza, pizza pie;’ ‘burrito;’ ‘bubble;’ ‘volcano;’ ‘corn;’ ‘acorn;’ ‘hen-of-the-woods, hen of the woods, Polyporus frondosus, Grifola frondosa.’

n01443537, n01704323, n01770081, n01784675, n01819313, n01820546, n01910747, n01917289, n01968897, n02074367, n02317335, n02319095, n02395406, n02454379, n02606052, n02655020, n02666196, n02672831, n02730930, n02777292, n02783161, n02786058, n02787622, n02791270, n02808304, n02817516, n02841315, n02865351, n02877765, n02892767, n02906734, n02910353, n02916936, n02948072, n02965783, n03000134, n03000684, n03017168, n03026506, n03032252, n03075370, n03109150, n03126707, n03134739, n03160309, n03196217, n03207743, n03218198, n03223299, n03240683, n03271574, n03291819, n03297495, n03314780, n03325584, n03344393, n03347037, n03372029, n03376595, n03388043, n03388183, n03400231, n03445777, n03457902, n03467068, n03482405, n03483316, n03494278, n03530642, n03544143, n03584829, n03590841, n03598930, n03602883, n03649909, n03661043, n03666591, n03676483, n03692522, n03706229, n03717622, n03720891, n03721384, n03724870, n03729826, n03733131, n03733281, n03742115, n03786901, n03788365, n03794056, n03804744, n03814639, n03814906, n03825788, n03840681, n03843555, n03854065, n03857828, n03868863, n03874293, n03884397, n03891251, n03908714, n03920288, n03929660, n03930313, n03937543, n03942813, n03944341, n03961711, n03970156, n03982430, n03991062, n03995372, n03998194, n04005630, n04023962, n04033901, n04040759, n04067472, n04074963, n04116512, n04118776, n04125021, n04127249, n04131690, n04141975, n04153751, n04154565, n04201297, n04204347, n04209133, n04209239, n04228054, n04235860, n04243546, n04252077, n04254120, n04258138, n04265275, n04270147, n04275548, n04330267, n04332243, n04336792, n04347754, n04371430, n04371774, n04372370, n04376876, n04409515, n04417672, n04418357, n04423845, n04429376, n04435653, n04442312, n04482393, n04501370, n04507155, n04525305, n04542943, n04554684, n04557648, n04562935, n04579432, n04591157, n04597913, n04599235, n06785654, n06874185, n07615774, n07693725, n07695742, n07697537, n07711569, n07714990, n07715103, n07716358, n07717410, n07718472, n07720875, n07742313, n07745940, n07747607, n07749582, n07753275, n07753592, n07754684, n07768694, n07836838, n07871810, n07873807, n07880968, n09229709, n09472597, n12144580, n12267677, n13052670.