Adversarial Examples that Fool Detectors

Jiajun Lu, Hussein Sibai, Evan Fabry

Introduction

An adversarial example is an example that has been adjusted to produce a wrong label when presented to a system at test time. In the literature, adversarial examples with imperceivable perturbations and unexpected properties (e.g. transferability) are one of the biggest mysteries of neural networks. There is a range of constructions that yield adversarial examples for image classifiers, and there is good evidence that small imperceivable adjustments suffice. Furthermore, Athalye et al. show that it is possible to build a physical object with visible perturbation patterns that is persistently misclassified by standard image classifiers from different view angles at a roughly fixed distance. There is good evidence that adversarial examples built for one classifier will fool others, too . The success of these attacks can be seen as a warning not to use highly non-linear feature constructions without having strong mathematical constraints on what these constructions do; but taking that position means one cannot use methods that are largely accurate and effective.

Detectors are not classifiers. A classifier accepts an image and produces a label. In contrast, a detector, like Faster RCNN , identifies bounding boxes that are “worth labelling”, and then generates labels (which might include background) for each box. The final label generation step employs a classifier. However, the statistics of how bounding boxes cover objects in a detector are complex and not well understood. Some modern detectors like YOLO 9000 predict boxes and labels using features on a fixed grid, resulting in fairly complex sampling patterns in the space of boxes, and this means that pixels outside a box may participate in labelling that box. Another important difference is that detectors usually have RoI pooling or feature map resizing, which might be effective at disrupting adversarial patterns. To date, no successful adversarial attack on a detector has been demonstrated. In this paper, we demonstrate successful adversarial attacks on Faster RCNN, which generalize to YOLO 9000.

We also discuss the generalization ability of adversarial examples. We say that an adversarial perturbation generalizes if, when circumstances (digital or physical) change, the corresponding images remain adversarial. For example, a perturbation of a stop sign generalizes over different distances if it remains adversarial when the camera approaches the stop sign. An example generalizes better if it remains adversarial for more cases (e.g. changes of detector, background and lighting). If an adversarial example cannot generalize, it is not a threat in the majority of real world systems.

Our contributions in this paper are as follows:

We demonstrate a method to construct adversarial examples that fool Faster RCNN digitally; examples produced by our method are reliably either missed or mislabeled by the detector. These examples without modification also fool YOLO 9000, indicating that the construction produces examples that can transfer across models.

Our adversarial examples can be physically created successfully, and they still can fool detectors in suitable circumstances. They can also slip through recent strong image processing defenses against adversarial examples.

In practice, we find that adversarial examples require quite large disruptions to the pattern on the object in order to fool detectors. Physical adversarial examples require bigger disruptions than digital examples to succeed.

Background

Adversarial examples are of interest mainly because the adjustments required seem to be very small and are easy to obtain . Numerous search procedures generate adversarial examples ; all searches look for an example that is (a) “near” a correctly labelled example (typically in L1L_{1} or L2L_{2} norm), and (b) mislabelled. Printing adversarial images then photographing them can retain their adversarial property , which suggests that adversarial examples might exist in the physical world. Their existence could cause a great deal of mischief. There is some evidence that it is difficult to build physical examples that fool a stop sign detector . In particular, if one actually takes a video of an existing adversarial stop sign, the adversarial pattern does not appear to affect the performance of the detector by much. Lu et al. speculated that this might be because adversarial patterns were disrupted by being viewed at different scales, rotations, and orientations. This created some discussion. OpenAI demonstrated a search procedure that could produce an image of a cat that was misclassified when viewed at multiple scales . There is some blurring of the fur texture on the cat they generate, but this would likely be imperceptible to most observers. OpenAI also demonstrated an adversarial image of a cat that was misclassified when viewed at multiple scales and orientations . However, there are significant visible artifacts on that image; few would think that it had not been obviously tampered with.

Recent work has demonstrated physical objects that are persistently misclassified from different angles at a roughly fixed distance . The search procedure manipulates the texture map T{\cal T} of the object. The procedure samples a set of viewing conditions Vi{\cal V}_{i} for the object, then renders to obtain images I(Vi,T)I({\cal V}_{i},{\cal T}). Finally, the procedure adjusts the texture map to obtain images that are (a) close to the original images and (b) have high probability of misclassification. The adversarial properties of the resulting objects are robust to the inevitable errors in color, etc., in producing physical objects from digital representations.

Defences: There is fair evidence that it is hard to tell whether an example is adversarial (and so (a) evidence of an attack and (b) likely to be misclassified) or not . Current procedures to build adversarial examples for deep networks appear to subvert the feature construction implemented by the network to produce odd patterns of activation in late stage ReLU’s; this can be exploited to build one form of defence . There is some evidence that other feature constructions admit adversarial attacks, too . However, adversarial attacks typically introduce unnatural (if small) patterns into images, and image processing methods that remove such patterns yield successful defenses. Guo et al. showed that cropping and rescaling, bit depth reduction, JPEG compression and decompression, resampling and reconstructing using total variation criteria, and image quilting all provide quite effective ways of removing adversarial patterns .

Detectors and classifiers: It is usual to attack classifiers, and all the attacks we are aware of attack on classifiers. However, for many applications, classifiers are not useful by themselves. Road sign is a good example. A road sign classifier would be applied to images that consist largely of a road sign (e.g. those of ). But there is little application need for a road sign classifier except as a component of a road sign detector, because it is unusual in practice to deal with images that consist largely of road sign. Instead, one usually deals with images that contain many things, and must find and label the road sign. It is natural to study road sign classifiers (e.g. ) because image classification remains difficult and academic studies of feature constructions are important. But there is no particular threat posed by an attack on a road sign classifier. An attack on a road sign detector is an entirely different matter. For example, imagine the danger if one could get a template that, with a can of spray paint, could ensure that a detector reads a stop sign as a yield sign (or worse!). As a result, it is important to know whether (a) such examples could exist and (b) how robust their adversarial property is in practice.

Recently, Evtimov et al. have shown several physical stop signs that are misclassified . They cropped the stop signs from the frames before presenting them to the classifier. By cropping, they have proxied the box-prediction process in a detector; however, their attack is not intended as an attack on a detector (the paper does not use the word “detector”, for example). Lu et al. showed that their construction does not fool a standard detector , likely because the cropping process does not proxy a detector’s box selection well, and suggested that constructing an adversarial example that fools a detector might be hard. Figure 1 shows their stop signs presented in are reliably detected by Faster RCNN.

Method

We propose a method to generate digital and physical adversarial examples that are robust to changes of viewing conditions. Our registration and reconstruction based approach generates adversarial perturbations from video sequences of an object with moving cameras. We require the objects in the videos to be accurately aligned in 3D space. We can easily register stop signs as they are 2D polygons. Moreover, we can accurately register face images to a virtual 3D face model. Hence, we perform our experiments on these two types of data.

We use the stop sign example to demonstrate our attack, which extends to other objects that are registered from image domain to some root coordinate system (e.g. faces in section 3.2). We search for an adversarial pattern that (a) looks like a stop sign and (b) fools Faster RCNN. We select a set of NN diverse frames Ii{\cal I}_{i} as the training examples to generate the pattern. A stop sign is represented as a texture map T{\cal T} in some root coordinate system. We construct (currently by hand) correspondences between eight vertices on the stop sign instances in training frames and the vertices of T{\cal T}. We use these correspondences to estimate a viewing map Mi{\cal M}_{i}, which maps the texture T{\cal T} in the root coordinate system to the appropriate pattern in training frame Ii{\cal I}_{i}. We also incorporate in Mi{\cal M}_{i} the illumination intensity, which is estimated by computing the average intensity over the stop sign in the image. Relative illumination intensities are used to scale the adversarial perturbations. We write I(Mi,T){\cal I}({\cal M}_{i},{\cal T}) for the image frame obtained by superimposing T{\cal T} on the frame Ii{\cal I}_{i} using the mapping Mi{\cal M}_{i}; Bs(I){\cal B}_{s}({\cal I}) for the set of stop sign bounding boxes obtained by applying Faster RCNN to the image I{\cal I}; and ϕs(b)\phi_{s}(b) for the score produced by Faster RCNN’s classifier for a stop sign in box bb. To produce an adversarial example, we minimize the mean score for a stop sign produced by Faster RCNN in all training images as a function of T{\cal T}, that is

possibly subject to some constraint on T{\cal T}, such as being close to a normal stop sign in L2L_{2} distance. We have also investigated minimizing the maximum score for all the stop sign proposals, and found that minimizing the mean score gives slightly better results.

Minimization procedure: First, we compute ∇TΦ(T)\nabla_{\cal T}\Phi({\cal T}) by computing \nabla_{\cal I}{\begin{array}[]{c}\mbox{mean}\\ {{b\in{\cal B}_{s}({\cal I}({\cal M}_{i},{\cal T})}}\end{array}}\phi_{s}(b). The gradients in the frame coordinate system are mapped to the root coordinate system with inverse view mapping Mi−1{\cal M}^{-1}_{i}, and then are cropped to the extent of T{\cal T} in that coordinate system. We average gradients mapped from all NN training frames. However, directly using the gradients to take large steps frequently stalls the optimization process. Instead, we find that computing the descent direction with the sign of the gradients for given pattern T(n){\cal T}^{(n)} (nn stands for iteration number) facilitates the optimization process.

We choose a very small step length ϵ\epsilon such that ϵd(n)\epsilon{\bf{d}}^{(n)} represents an update of one least significant bit, which leads to an optimization step of form

The optimization process usually takes hundreds or even thousands of steps. One termination criterion is to stop the optimization when the pattern fools the detector more than 90% of the cases on the validation set. Another termination criterion is a fixed number of iterations.

Why large steps are hard: In our experiments, taking large steps with unsigned gradients stalls the optimization process, and we believe large steps are hard to take for two reasons. First, each instance of the pattern occurs at a different scale, meaning that there must be some up- or down-sampling of the gradients when mapped to the root coordinate system. Although we register the images with subpixel accuracy, and use a bilinear method to interpolate the transformation process, signal losses are still inevitable. In section 4.2, we show some evidence that this effect may make our patterns more, rather than less, robust. Second, the structure of the network means that the gradient is a poor guide to the behavior of ϕs(b)\phi_{s}(b) over large scales. In particular, a ReLU network divides its input space into a very large number of cells, and values at any layer before the softmax layer are a continuous piecewise linear function within a cell. Because the network is trained to have a (roughly) constant output for large pieces of its input space, the gradient must wiggle from cell to cell, and so may be a poor guide to the long scale behavior of the function.

Constraining distance to the original stop sign: In order to create less perceivable adversarial perturbations, we constrain the distance to the original stop sign to be small. An L2L_{2} distance loss is added to the cost function, and we minimize

Our experiments show that this distance constraint still cannot help to create small perturbations, but greatly changes the pattern of the perturbations, refer to Figure 6.

We create our physical adversarial stop signs by printing the pattern T{\cal T}, cutting out the printed stop sign area, and sticking it to an actual stop sign (30 in by 30 in).

2 Extending to faces

We extend the experiments onto faces, which have complex geometries and larger intra class variances, to demonstrate that our analysis generalizes to other classes. In the face setting, we search for a pattern that fools a Faster RCNN based face detector and looks like the original face. Our root coordinate system for faces is a virtual high quality face mesh generated from morphable face model . For video sequences of a face, we reconstruct the geometry of the face in the input frames using morphable face model built from the FaceWarehouse data. The model produces a 3D face mesh F(wi,we)F(w_{i},w_{e}) that is a function of identity parameters wiw_{i}, and expression parameters wew_{e}. FaceTracker is used to detect landmarks lil_{i} on the face frames, then we recover parameters and poses of the face mesh by minimizing the distances between the projected landmark vertices and their corresponding landmark locations on the image planes. This construction gives us pixel-to-mesh and mesh-to-pixel dense correspondences between all face frames and the root face coordinate system (shared face mesh). By projecting image pixels to the face meshes via barycentric coordinates, we can achieve subpixel accurate pixel-to-pixel registrations between all face frames (via root coordinate system). This correspondences are used to transfer the gradients from face image coordinates to the root coordinate system, then we merge the gradients from multiple images and reverse transfer the merged gradients back to the face image coordinates.

Results

In this section, we describe in depth the experiments we did and the results we got. Our supplementary materials include videos and more results, and it can be downloaded from http://jiajunlu.com/docs/advDetector_supp.zip. Our high resolution paper can be downloaded from http://jiajunlu.com/docs/advDetector.pdf. But let us first give a quick overview of our findings:

Adding small perturbations suffices to fool a given object detector on a single image.

Enforcing the adversarial perturbations to generalize across view conditions requires significant changes to the pattern.

Our successful attacks for Faster RCNN generalize to YOLO.

Our successful attacks with very large perturbations generalize to the physical world objects in suitable circumstances.

Simple defenses fail to defeat adversarial examples that can generalize.

Detectors are affected by internal thresholds. Faster RCNN uses a non maximum suppression threshold and a confidence threshold. For stop signs, we used the default configurations. For faces, we found the detector is too willing to detect faces, and we made it less responsive to faces (nms from 0.15 to 0.3, and conf from 0.6 to 0.8). We used default YOLO configurations.

We can easily adjust a pattern on a single image to fool a detector (stop signs in Figure 2 and faces in Figure 3), and the change on the pattern is tiny. While this is of no practical significance, it shows our search method can find very small adversarial perturbations.

2 Generalizing across view conditions

What we are really interested in is to produce a pattern that fails to be detected in any image. This is much harder because our pattern needs to generalize to different view conditions and so on. We can still find adversarial patterns in this situation, but the patterns found by our process involve significant changes of the stop signs and faces.

Stop sign dataset: we use a Panasonic HC-V700M HD camera to take 22 videos of the camera approaching stop signs, and extract 5 diverse frames from each video. Then we manually register all the stop signs, and use our attacking method to generate a unified adversarial perturbation for all the frames. We use 12 videos for generating adversarial perturbations (training), 5 videos for validation, and 5 videos for evaluation (testing). We use the validation set termination criteria described in Section 3.1. Figure 4 gives an example video sequence, and its corresponding attacked video sequence. Table 1 shows the stop sign detection rates in different circumstances. We plan to release the labelled dataset.

Face dataset:: we use a SONY a7 camera to take 5 videos of a still face from different distances and angles, and extract 20 diverse frames from each video. We use the morphable face model approach to register all the faces, and use our attacking method to generate a unified adversarial perturbation. We use 3 videos for generating adversarial perturbations (training), 1 video for validation, and 1 video for evaluation (testing). Again, we use the validation set termination criteria described in Section 3.1. Figure 5 shows an example video sequence, and its corresponding attacked video sequence. In our experiments, this is the smallest perturbations on faces that could generalize. Table 2 shows the face detection rates in different circumstances. Also, we plan to release the processed dataset.

In summary, it is possible to attack stop signs and faces from multiple images, and require them to generalize to new similar view condition images. However, both of them require strong perturbation patterns to generalize. Refer to supplementary materials for details.

3 Generalizing to the physical world

There is a big gap between attacks in the digital world and attacks in the physical world, which means the adversarial perturbations that generalize well in the digital world may not generalize to the physical world. We suspect this gap is due to various practical concerns, such as sensor properties, view conditions, printing errors, lighting, etc. In this paper, we print stop signs and perform physical experiments with them, but we believe similar conclusions apply to faces.

We performed physical experiments with three adversarial perturbation patterns in Figure 6. Our results in Table 3 show that the two less perturbed stop signs can still be detected by Faster RCNN, while the one with large perturbations is hard to detect. The frames for physical experiments could be found in Figure 7. Refer to our supplementary materials for videos.

We performed analysis with the data from Table 1 and Table 3. L1L_{1} regularized logistic regression is used to predict the success of our many different cases. The most important variable is detector (generalization from Faster RCNN to YOLO is not strong); then whether the adversarial example is physical or not (digital attacks are more effective than physical); then scale (it is hard to make a detector to miss a nearby stop sign).

4 Generalizing to YOLO

Adversarial examples for a certain classifier generalize across different classifiers. To test out whether adversarial examples for Faster RCNN generalize across detectors, we feed these images into YOLO. We categorize our adversarial examples into three categories: single image examples with small perturbations, multiple image examples with large perturbations that generalize across viewing conditions digitally, and physical examples with large perturbations. Our experiments show that small perturbations do not generalize to YOLO, while obvious perturbation patterns can generalize to YOLO with good probability. Examples are given in Figure 8, and detection rates can be found in Table 3 and Table 1.

5 Localized attacks fail

In previous settings, we attack the whole masked objects in the images, however, it is usually hard to apply such attacks in the physical world. For example, modifying the whole stop sign patterns is useless in practice, and wearing a whole face mask with perturbation patterns is hard too. It would be more effective attack if one can manufacture small stickers with perturbation patterns, and when the sticker is attached to a small region of the stop sign or the face forehead, the detector would fail. Evtimov et al. showed an example that successfully attacked stop sign classifiers. We try to generate adversarial patterns constrained to a fixed region of the objects to fool detectors, however, we find these attacks only occasionally successful (stop signs) or wholly unsuccessful (faces). Figure 9 shows some examples.

6 Simple defenses fail

Recently, Guo et al. showed that simple image processing could defeat a majority of imperceivable adversarial attacks. We assume detectors should run at frame rate, so exclude image quilting. We investigated down-up sampling and total variation smoothing defense. We find that these methods can defeat attacks on a single image, but cannot disrupt the large patterns needed to produce adversarial examples that generalize, see Figure 10 and Table 4. Our hypothesis for this phenomenon is that tiny perturbations work with numerical accumulation mechanism, which is not robust to changes, while obvious perturbations work with pattern recognition mechanism, which is more robust and can better generalize.

Conclusion

We have demonstrated the first adversarial examples that can fool detectors. Our construction yields physical objects that fool detectors too. However, all the adversarial perturbations we have been able to construct require large perturbations. This suggests that the box prediction step in a detector acts as a form of natural defense. We speculate that better viewing models in our construction may yield a smaller gap between physical and digital results. Our patterns may reveal something about what is important to a detector.

References