Invisible Backdoor Attack with Sample-Specific Triggers
Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, Siwei Lyu
Introduction
Deep neural networks (DNNs) have been widely and successfully adopted in many areas . Large amounts of training data and increasing computational power are the key factors to their success, but the lengthy and involved training procedure becomes the bottleneck for users and researchers. To reduce the overhead, third-party resources are usually utilized in training DNNs. For example, one can use third-party data (, data from the Internet or third-party companies), train their model with third-party servers (, Google Cloud), or even adopt third-party APIs directly. However, the opacity of the training process brings new security threats.
Backdoor attackBackdoor attack is also commonly called ‘neural trojan’ or ‘trojan attack’ . In this paper, we focus on the poisoning-based backdoor attack towards image classification, although the backdoor threat could also happen in other scenarios . is an emerging threat in the training process of DNNs. It maliciously manipulates the prediction of the attacked DNN model by poisoning a portion of training samples. Specifically, backdoor attackers inject some attacker-specified patterns (dubbed backdoor triggers) in the poisoned image and replace the corresponding label with a pre-defined target label. Accordingly, attackers can embed some hidden backdoors to the model trained with the poisoned training set. The attacked model will behave normally on benign samples, whereas its prediction will be changed to the target label when the trigger is present. Besides, the trigger could be invisible and the attacker only needs to poison a small fraction of samples, making the attack very stealthy. Hence, the insidious backdoor attack is a serious threat to the applications of DNNs.
Fortunately, some backdoor defenses were proposed, which show that existing backdoor attacks can be successfully mitigated. It raises an important question: has the threat of backdoor attacks really been resolved?
In this paper, we reveal that existing backdoor attacks were easily mitigated by current defenses mostly because their backdoor triggers are sample-agnostic, , different poisoned samples contain the same trigger no matter what trigger pattern is adopted. Given the fact that the trigger is sample-agnostic, defenders can easily reconstruct or detect the backdoor trigger according to the same behaviors among different poisoned samples.
Based on this understanding, we explore a novel attack paradigm, where the backdoor trigger is sample-specific. We only need to modify certain training samples with invisible perturbation, while not need to manipulate other training components (, training loss, and model structure) as required in many existing attacks . Specifically, inspired by DNN-based image steganography , we generate sample-specific invisible additive noises as backdoor triggers by encoding an attacker-specified string into benign images through an encoder-decoder network. The mapping from the string to the target label will be generated when DNNs are trained on the poisoned dataset. The proposed attack paradigm breaks the fundamental assumption of current defense methods, therefore can easily bypass them.
The main contributions of this paper are as follows: (1) We provide a comprehensive discussion about the success conditions of current main-stream backdoor defenses. We reveal that their success all relies on a prerequisite that backdoor triggers are sample-agnostic. (2) We explore a novel invisible attack paradigm, where the backdoor trigger is sample-specific and invisible. It can bypass existing defenses for it breaks their fundamental assumption. (3) Extensive experiments are conducted, which verify the effectiveness of the proposed method.
Related Work
The backdoor attack is an emerging and rapidly growing research area, which poses a security threat to the training process of DNNs. Existing attacks can be categorized into two types based on the characteristics of triggers: (1) visible attack that the trigger in the attacked samples is visible for humans, and (2) invisible attack that the trigger is invisible.
Visible Backdoor Attack. Gu et al. first revealed the backdoor threat in the training of DNNs and proposed the BadNets attack, which is representative of visible backdoor attacks. Given an attacker-specified target label, BadNets poisoned a portion of the training images from the other classes by stamping the backdoor trigger (, white square in the lower right corner of the image) onto the benign image. These poisoned images with the target label, together with other benign training samples, are fed into the DNNs for training. Currently, there was also some other work in this field . In particular, the concurrent work also studied the sample-specific backdoor attack. However, their method needs to control the training loss except for modifying training samples, which significantly reduces its threat in real-world applications.
Invisible Backdoor Attack. Chen et al. first discussed the stealthiness of backdoor attacks from the perspective of the visibility of backdoor triggers. They suggested that poisoned images should be indistinguishable compared with their benign counter-part to evade human inspection. Specifically, they proposed an invisible attack with the blended strategy, which generated poisoned images by blending the backdoor trigger with benign images instead of by stamping directly. Besides the aforementioned methods, several other invisible attacks were also proposed for different scenarios: Quiring et al. targeted on the image scaling process during the training, Zhao et al. targeted on the video recognition, and Saha et al. assumed that attackers know model structure. Note that most of the existing attacks adopted a sample-agnostic trigger design, i.e., the trigger is fixed in either the training or testing phase. In this paper, we propose a more powerful invisible attack paradigm, where backdoor triggers are sample-specific.
2 Backdoor Defense
Trigger Synthesis based Defenses. Instead of eliminating hidden backdoors directly, trigger synthesis based defenses synthesize potential triggers at first, following by the second stage suppressing their effects to remove hidden backdoors. Wang et al. proposed the first trigger synthesis based defense, , Neural Cleanse, where they first obtained potential trigger patterns towards every class and then determined the final synthetic trigger pattern and its target label based on an anomaly detector. Similar ideas were also studied , where they adopted different approaches for generating potential triggers or anomaly detection.
Saliency Map based Defenses. These methods used the saliency map to identify potential trigger regions to filter malicious samples. Similar to trigger synthesis based defenses, an anomaly detector was also involved. For example, SentiNet adopted the Grad-CAM to extract critical regions from input towards each class and then located the trigger regions based on the boundary analysis. A similar idea was also explored .
STRIP. Recently, Gao et al. proposed a method, known as the STRIP, to filter malicious samples through superimposing various image patterns to the suspicious image and observe the randomness of their predictions. Based on the assumption that the backdoor trigger is input-agnostic, the smaller the randomness, the higher the probability that the suspicious image is malicious.
A Closer Look of Existing Defenses
In this section, we discuss the success conditions of current mainstream backdoor defenses. We argue that their success is mostly predicated on an implicit assumption that backdoor triggers are sample-agnostic. Once this assumption is violated, their effectiveness will be highly affected. The assumptions of several defense methods are discussed as follows.
The Assumption of Pruning-based Defenses. Pruning-based defenses were motivated by the assumption that backdoor-related neurons are different from those activated for benign samples. Defenders can prune neurons that are dormant for benign samples to remove hidden backdoors. However, the non-overlap between these two types of neurons holds probably because the sample-agnostic trigger patterns are simple, , DNNs only need few independent neurons to encode this trigger. This assumption may not hold when triggers are sample-specific, since this paradigm is more complicated.
The Assumption of Trigger Synthesis based Defenses. In the synthesis process, existing methods (, Neural Cleanse ) are required to obtain potential trigger patterns that could convert any benign image to a specific class. As such, the synthesized trigger is valid only when the attack-specified backdoor trigger is sample-agnostic.
The Assumption of Saliency Map based Defenses. As mentioned in Section 2.2, saliency map based defenses required to (1) calculate saliency maps of all images (toward each class) and (2) locate trigger regions by finding universal saliency regions across different images. In the first step, whether the trigger is compact and big enough determines whether the saliency map contains trigger regions influencing the defense effectiveness. The second step requires that the trigger is sample-agnostic, otherwise, defenders can hardly justify the trigger regions.
The Assumption of STRIP. STRIP examined a malicious sample by superimposing various image patterns to the suspicious image. If the predictions of generated samples are consistent, then this examined sample will be regarded as the poisoned sample. Note its success also relies on the assumption that backdoor triggers are sample-agnostic.
Sample-specific Backdoor Attack (SSBA)
Attacker’s Capacities. We assume that attackers are allowed to poison some training data, whereas they have no information on or change other training components (, training loss, training schedule, and model structure). In the inference process, attackers can and only can query the trained model with any image. They have neither information about the model nor can they manipulate the inference process. This is the minimal requirement for backdoor attackers . The discussed threat can happen in many real-world scenarios, including but not limited to adopting third-party training data, training platforms, and model APIs.
Attacker’s Goals. In general, backdoor attackers intend to embed hidden backdoors in DNNs through data poisoning. The hidden backdoor will be activated by the attacker-specified trigger, , the prediction of the image containing trigger will be the target label, no matter what its ground-truth label is. In particular, attackers has three main goals, including the effectiveness, stealthiness, and sustainability. The effectiveness requires that the prediction of attacked DNNs should be the target label when the backdoor trigger appears, and the performance on benign testing samples will not be significantly reduced; The stealthiness requires that adopted triggers should be concealed and the proportion of poison samples (, the poisoning rate) should be small; The sustainability requires that the attack should still be effective under some common backdoor defenses.
2 The Proposed Attack
In this section, we illustrate our proposed method. Before we describe how to generate sample-specific triggers, we first briefly review the main process of attacks and present the definition of a sample-specific backdoor attack.
The Main Process of Backdoor Attacks. Let indicates the benign training set containing samples, where and . The classification learns a function with parameters . Let denotes the target label (). The core of backdoor attacks is how to generate the poisoned training set . Specifically, consists of modified version of a subset of (, ) and remaining benign samples , ,
where , indicates the poisoning rate, , is an attacker-specified poisoned image generator. The smaller the , the more stealthy the attack.
A backdoor attack with poisoned image generator is called sample-specific if and only if , where indicates the backdoor trigger contained in the poisoned sample .
Triggers of previous attacks are not sample-specific. For example, for the attack proposed in , , where .
How to Generate Sample-specific Triggers. We use a pre-trained encoder-decoder network as an example to generate sample-specific triggers, motivated by the DNN-based image steganography . The generated triggers are invisible additive noises containing a representative string of the target label. The string can be flexibly designed by the attacker. For example, it can be the name, the index of the target label, or even a random character. As shown in Figure 2, the encoder takes a benign image and the representative string to generate the poisoned image (, the benign image with their corresponding trigger). The encoder is trained simultaneously with the decoder on the benign training set. Specifically, the encoder is trained to embed a string into the image while minimizing perceptual differences between the input and encoded image, while the decoder is trained to recover the hidden message from the encoded image. Their training process is demonstrated in Figure 3. Note that attackers can also use other methods, such as VAE , to conduct the sample-specific backdoor attack. It will be further studied in our future work.
Pipeline of Sample-specific Backdoor Attack. Once the poisoned training set is generated based on the aforementioned method, backdoor attackers will send it to the user. Users will adopt it to train DNNs with the standard training process,
where indicated the loss function, such as the cross-entropy. The optimization (2) can be solved by back-propagation with the stochastic gradient descent .
The mapping from the representative string to the target label will be learned by DNNs during the training process. Attackers can activate hidden backdoors by adding triggers to the image based on the encoder in the inference stage.
Experiment
Datasets and Models. We consider two classical image classification tasks: (1) object classification, and (2) face recognition. For the first task, we conduct experiments on the ImageNet dataset. For simplicity, we randomly select a subset containing classes with images for training (500 images per class) and images for testing (50 images per class). The image size is . Besides, we adopt MS-Celeb-1M dataset for face recognition. In the original dataset, there are nearly 100,000 identities containing different numbers of images ranging from 2 to 602. For simplicity, we select the top 100 identities with the largest number of images. More specifically, we obtain 100 identities with 38,000 images (380 images per identity) in total. The split ratio of training and testing sets is set to 8:2. For all the images, we firstly perform face alignments, then select central faces, and finally resize them into . We use ResNet-18 as the model structure for both datasets. More experiments with VGG-16 are in the supplementary materials.
Baseline Selection. We compare the proposed sample-specific backdoor attack with BadNets and the typical invisible attack with blended strategy (dubbed Blended Attack) . We also provide the model trained on the benign dataset (dubbed Standard Training) as another baseline for reference. Besides, we select Fine-Pruning , Neural Cleanse , SentiNet , STRIP , DF-TND , and Spectral Signatures to evaluate the resistance to state-of-the-art defenses.
Attack Setup. We set the poisoning rate and target label for all attacks on both datasets. As shown in Figure 4, the backdoor trigger is a white-square with a cross-line on the bottom right corner of poisoned images for both BadNets and Blended Attack, and the trigger transparency is set to 10% for the Blended Attack. The triggers of our methods are generated by the encoder trained on the benign training set. Specifically, we follow the settings of the encoder-decoder network in StegaStamp , where we use a U-Net style DNN as the encoder, a spatial transformer network as the decoder, and four loss-terms for the training: residual regularization, LPIPS perceptual loss , a critic loss, to minimize perceptual distortion on encoded images, and a cross-entropy loss for code reconstruction. The scaling factors of four loss-terms are set to 2.0, 1.5, 0.5, and 1.5. For the training of all encoder-decoder networks, we utilize Adam optimizer and set the initial learning rate as . The batch size and training iterations are set to 16 and , respectively. Moreover, in the training stage, we utilize the SGD optimizer and set the initial learning rate as . The batch size and maximum epoch are set as and , respectively. The learning rate is decayed with factor after epoch and .
Defense Setup. For Fine-Pruning, we prune the last convolutional layer of ResNet-18 (Layer4.conv2); For Neural Cleanse, we adopt its default setting and utilize the generated anomaly index for demonstration. The smaller the value of the anomaly index, the harder the attack to defend; For STRIP, we also adopt its default setting and present the generated entropy score. The larger the score, the harder the attack to defend; For SentiNet, we compared the generated Grad-CAM of poisoned samples for demonstration; For DF-TND, we report the logit increase scores before and after the universal adversarial attack of each class. This defense succeeds if the score of the target label is significantly larger than those of all other classes. For Spectral Signatures, we report the outlier score for each sample, where a larger score denotes the sample is more likely poisoned.
2 Main Results
Attack Effectiveness. As shown in Table 1, our attack can successfully create backdoors with a high ASR by poisoning only a small proportion (10%) of training samples. Specifically, our attack can achieve an ASR on both datasets. Besides, the ASR of our method is on par with that of BadNets and higher than that of Blended Attack. Moreover, the accuracy reduction of our attack (compared with the Standard Training) on benign testing samples is less than on both datasets, which are smaller than those of BadNets and Blended Attack. These results show that sample-specific invisible additive noises can also serve as good triggers even though they are more complicated than the white-square used in BadNets and Blended Attack.
Time Analysis. Training the encoder-decoder network takes 7h 35mins on ImageNet and 3h 40mins on MS-Celeb-1M. The average encoding time is 0.2 seconds per image.
Resistance to Fine-Pruning. In this part, we compare our attack to BadNets and Blended Attack in terms of the resistance to the pruning-based defense . As shown in Figure 5, the ASR of BadNets and Blended Attack drop dramatically when only 20% of neurons are pruned. Especially the Blended Attack, its ASR decrease to less than 10% on both ImageNet and MS-Celeb-1M datasets. In contrast, the ASR of our attack only decreases slightly (less than 5%) with the increase of the fraction of pruned neurons. Our attack retains an ASR greater than 95% on both datasets when 20% of neurons are pruned. This suggests that our attack is more resistant to the pruning-based defense.
Resistance to Neural Cleanse. Neural Cleanse computes the trigger candidates to convert all benign images to each label. It then adopts an anomaly detector to verify whether anyone is significantly smaller than the others as the backdoor indicator. The smaller the value of the anomaly index, the harder the attack for Neural-Cleanse to defend. As shown in Figure 9, our attack is more resistant to the Neural-Cleanse. Besides, we also visualize the synthesized trigger (, the one with the smallest anomaly index among all candidates) of different attacks. As shown in Figure 6, synthesized triggers of BadNets and Blended Attack contain similar patterns to those used by attackers (, white-square on the bottom right corner), whereas those of our attack are meaningless.
Resistance to STRIP. STRIP filters poisoned samples based on the prediction randomness of samples generated by imposing various image patterns on the suspicious image. The randomness is measured by the entropy of the average prediction of those samples. As such, the higher the entropy, the harder an attack for STRIP to defend. As shown in Figure 9, our attack is more resistant to the STRIP compared with other attacks.
Resistance to SentiNet. SentiNet identities trigger regions based on the similarities of Grad-CAM of different samples. As shown in Figure 7, Grad-CAM successfully distinguishes trigger regions of those generated by BadNets and Blended Attack, while it fails to detect trigger regions of those generated by our attack. In other words, our attack is more resistant to SentiNet.
Resistance to DF-TND. DF-TND detects whether a suspicious DNN contains hidden backdoors by observing the logit increase of each label before and after a crafted universal adversarial attack. This method can succeed if there is a peak of logit increase solely on the target label. For fair demonstration, we fine-tune its hyper-parameters to seek a best-performed defense setting against our attack (see supplementary material for more details). As shown in Figure 10, the logit increase of the target class (red bars in the figure) is not the largest on both datasets. It indicates that our attack can also bypass the DF-TND.
Resistance to Spectral Signatures. Spectral Signatures discovered the backdoor attacks can leave behind a detectable trace in the spectrum of the covariance of a feature representation. The trace is so-called Spectral Signatures, which is detected using singular value decomposition. This method calculates an outlier score for each sample. It succeeds if clean samples have small values and poison samples have large values (see supplementary material for more details). As shown in Figure 11, we test samples, where are clean samples and are poison samples. Our attack notably disturbs this method in that the clean samples have unexpected large scores.
3 Discussion
In this section, unless otherwise specified, all settings are the same as those stated in Section 5.1.
Attack with Different Target Labels. We test our method using different target labels (). Table 2 shows the BA/ASR of our attack, which reveals the effectiveness of our method using different target labels.
The Effect of Poisoning Rate . In this part, we discuss the effect of the poisoning rate towards ASR and BA in our attack. As shown in Figure 12, our attack reaches a high ASR () on both datasets by poisoning only 2% of training samples. Besides, the ASR increases with an increase of while the BA remains almost unchanged. In other words, there is almost no trade-off between the ASR and BA in our method. However, the increase of will also decrease the attack stealthiness. Attackers need to specify this parameter for their specific needs.
The Exclusiveness of Generated Triggers. In this part, we explore whether the generated sample-specific triggers are exclusive, whether testing image with trigger generated based on another image can also activate the hidden backdoor of DNNs attacked by our method. Specifically, for each testing image , we randomly select another testing image . Now we query the attacked DNNs with (rather than with ). As shown in Table 3, the ASR decreases sharply when inconsistent triggers (, triggers generated based on different images) are adapted on the ImageNet dataset. However, on the MS-Celeb-1M dataset, attacking with inconsistent triggers can still achieve a high ASR. This may probably be because most of the facial features are similar and therefore the learned trigger has better generalization. We will further explore this interesting phenomenon in our future work.
Out-of-dataset Generalization in the Attack Stage. Recall that the encoder is trained on the benign version of the poisoned training set in previous experiments. In this part, we explore whether the one trained on another dataset can still be adapted for generating poisoned samples of a new dataset (without any fine-tuning) in our attack. As shown in Table 4, the effectiveness of attack with an encoder trained on another dataset is on par with that of the one trained on the same dataset. In other words, attackers can reuse already trained encoders to generate poisoned samples, if their image size is the same. This property will significantly reduce the computational cost of our attack.
Out-of-dataset Generalization in the Inference Stage. In this part, we verify that whether out-of-dataset images (with triggers) can successfully attack DNNs attacked by our method. We select the Microsoft COCO dataset and a synthetic noise dataset for the experiment. They are representative of nature images and synthetic images, respectively. Specifically, we randomly select 1,000 images from the Microsoft COCO and generate 1,000 synthetic images where each pixel value is uniformly and randomly selected from . All selected images are resized to . As shown in Table 5, our attack with poisoned samples generated based on out-of-dataset images can also achieve nearly 100% ASR. It indicates that attackers can activate the hidden backdoor in attacked DNNs with out-of-dataset images (not necessary with testing images).
Conclusion
In this paper, we showed that existing backdoor attacks were easily alleviated by current backdoor defenses mostly because their backdoor trigger is sample-agnostic, , different poisoned samples contain the same trigger. Based on this understanding, we explored a new attack paradigm, the sample-specific backdoor attack (SSBA), where the backdoor trigger is sample-specific. Our attack broke the fundamental assumption of defenses, therefore can bypass them. Specifically, we generated sample-specific invisible additive noises as backdoor triggers by encoding an attacker-specified string into benign images, motivated by the DNN-based image steganography. The mapping from the string to the target label will be learned when DNNs are trained on the poisoned dataset. Extensive experiments were conducted, which verify the effectiveness of our method in attacking models with or without defenses.
Acknowledgment. Yuezun Li is supported in part by China Postdoc Science Foundation under grant No.2021TQ0314. Baoyuan Wu is supported by the Natural Science Foundation of China under grant No.62076213, the university development fund of the Chinese University of Hong Kong, Shenzhen under grant No.01001810, the special project fund of Shenzhen Research Institute of Big Data under grant No.T00120210003, and Shenzhen Science and Technology Program under grant No.GXWD20201231105722002-20200901175001001. Siwei Lyu is supported by the Natural Science Foundation under grants No.IIS-2103450 and IIS-1816227.
References
Appendix
More Results of Methods with VGG-16
In the main manuscript, we used ResNet-18 as the model structure for all experiments. To verify that our proposed attack is also effective towards other model structures, we provide additional results of methods with VGG-16 in this section. Unless otherwise specified, all settings are the same as those used in the main manuscript.
Follow the settings adopted in the main manuscript, we compare the effectiveness of methods from the aspect of attack success rate (ASR) and benign accuracy (BA).
As shown in Table 6, our attack can also reach a high attack success rate and benign accuracy on both ImageNet and MS-Celeb-1M dataset with VGG-16 as the model structure. Specifically, our attack can achieve an ASR on both datasets. Moreover, the ASR of our attack is on par with that of BadNets and higher than that of the Blended Attack. These results verify that sample-specific invisible additive noises can also serve as good backdoor triggers even though they are more complicated than the white-square used in BadNets and Blended Attack.
2 Resistance to Fine-Pruning
In this part, we also compare our attack with the BadNets and Blended Attack in terms of the resistance to the pruning-based defense . As shown in Figure 13, curves of our attack are always above those of other attacks. In other words, our descent speed is slower although ASRs of all attacks decrease with the increase of the fraction of pruned neurons. For example, on the ImageNet dataset, the ASR of Blended Attack decrease to less than 10% when 60% neurons are pruned, whereas our attack still preserves a high ASR (). This suggests that our attack is more resistant to the pruning-based defense.
3 Resistance to Neural Cleanse
In this part, we also compare our attack with the BadNets and Blended Attack in terms of the resistance to the Neural Cleanse . Recall that there are two indispensable requirements for the success of Neural Cleanse, including (1) successful select one candidate (, the anomaly index is big enough) and (2) the selected candidate is close to the backdoor trigger.
As shown in Figure 15, the anomaly index of our attack is smaller than that of BadNets and Blended Attack on the ImageNet dataset. In other words, our attack is more resistant to the Neural Cleanse in this case. We also visualize the synthesized trigger (, the one with the smallest anomaly index among all candidates) of different attacks. As shown in Figure 16, although our attack reaches the highest anomaly index on the MS-Celeb-1M dataset, synthesized triggers of our attack are meaningless. In contrast, synthesized triggers of BadNets and Blended Attack contain similar patterns to the ones used by attackers. As such, our attack is still more resistant to the Neural Cleanse in this case.
4 Resistance to STRIP
STRIP filters poisoned samples based on the prediction randomness of samples generated by imposing various image patterns on the suspicious image. The randomness is measured by the entropy of the average prediction of those samples. As such, the higher the entropy, the harder an attack for STRIP to defend. As shown in Figure 17, our attack has a significantly higher entropy compared with other baseline methods on both ImageNet and MS-Celeb-1M datasets. In other words, our attack is more resistant to the STRIP compared with other attacks.
5 Resistance to SentiNet
SentiNet identities trigger regions based on the similarities of Grad-CAM of different samples. As shown in Figure 14, Grad-CAM fails to detect trigger regions of images generated by our attack. Besides, the Grad-CAM of different poisoned samples has a significant difference. As such, our attack can bypass the SentiNet.
Detailed Settings of DF-TND and Spectral Signature
DF-TND. Note the vanilla setting of DF-TND is selected based on the CIFAR dataset, rather than the datasets used in our experiment. We found that its performance is sensitive to the hyper-parameter values. To achieve a fairer comparison, we fine-tune their hyper-parameters to seek a best-performed setting, based on the criteria that the more front of target label in a descending order based on logit increase denotes better defensive performance. We fine-tune two hyper-parameters, which are the batch size of testing random noise images and the sparsity parameter used in the adversarial attack. In its vanilla setting, the batch size is set to and is set to . In our experiments, we test nine hyper-parameter combinations, where batch size is selected from and sparsity parameter is selected from and then select the best-performed hyper-parameter combination. Specifically, we select for ImageNet dataset and for MS-Celeb-1M dataset.
Spectral Signature. Since this work does not release the code, we implement it based on Trojan-Zoohttps://github.com/alps-lab/Trojan-Zoo. Similar to DF-TND, Spectral Signature is also designed for CIFAR dataset, such that the default threshold of outlier score is not applicable in our experiments. For fair comparison, we calculate the outlier score for each test sample and show the distribution instead. The defense fails if the clean samples have larger outlier scores.
More Comparisons with Adapted Methods
As aforementioned in Section 2, the works are out of our scope either in the task or threat model. However, to be more comprehensive, we attempt to adapt the code of to our scenario for comparison. Note and are originally validated with AlexNet and CNN+LSTM respectively. We change their backbones to ResNet-18 and abandon their clean-label setting for fair comparison. The triggers of and are movable specific block and targeted universal adversarial perturbation (UAP) respectively. Table 7 shows the BA/ASR on ImageNet without defense, which represents our adaptations of these methods are normal. Figure 18 shows the Grad-CAM of SentiNet defense, where the block trigger of is accurately localized and the UAP trigger of is stably identified.