Backdoor Attack in the Physical World

Yiming Li, Tongqing Zhai, Yong Jiang, Zhifeng Li, Shu-Tao Xia

Introduction

Recent studies showed that some regular (i.e., non-optimized) perturbations (e.g., the local patch stamped on the image) could mislead DNNs, through influencing model weights in the training process (Liu et al., 2020; Li et al., 2020a; Gao et al., 2020). It is called as backdoor attack. Specifically, some training images are modified by adding the trigger (e.g., the local patch). These modified images with the attacker-specified target label, together with benign training samples, are fed into the DNN model for training. Consequently, trained DNNs perform well on benign testing samples, whereas their prediction will be changed when the same trigger is contained in the attacked image. Since attacked DNNs perform normally on benign samples, it is difficult for users to realize the attack. Hence, the insidious backdoor attack is a serious threat to the practical application of DNNs.

Many backdoor attacks have been proposed through designing different types of triggers (Gu et al., 2017; Liao et al., 2018; Turner et al., 2019; Zhao et al., 2020; Li et al., 2020b; Zhai et al., 2021). Currently, most existing works adopted the setting of static trigger, where the triggers across the training and testing images are the same. However, the location and appearance of the trigger in the digitized image may be different from that of the one used for training in the physical world. It raises an intriguing question: When the trigger in the attacked testing image is different from that used in training, can it still activate the hidden backdoor?

To answer this question, we explore the impacts of two basic characteristics of the trigger, including location and appearance. We demonstrate that if the location or appearance is slightly changed, then the attack performance may degrade sharply. It reveals that attacks with the static trigger pattern may be non-robust to the change of trigger. The above observation inspires two further questions:

(1) Can we utilize this non-robustness to defend existing backdoor attacks? (2) How to enhance the performance of existing backdoor attacks, such that they are robust to the change of trigger?

In this work, we propose a simple yet effective defense towards attacks with the static trigger pattern in which the testing sample is transformed (e.g., flipping or scaling) before the prediction. The transformation is a feasible approach to change the trigger’s location and appearance. Besides, we propose to enhance the transformation robustness of attacks that all poisoned images will be randomly transformed before feeding into the training process. This enhancement could be naturally combined with any backdoor attack. Moreover, we demonstrate the connection between the proposed attack enhancement and the physical attack, which implies that enhanced attacks could still succeed in the physical world whereas standard backdoor attacks will fail.

The Property of Existing Attacks with Static Trigger

We consider the scenario that the user cannot fully control the training process of the model C(⋅;w)C(\cdot;w). Let ytargety_{target} denotes the target label, Dtrain={(x,y)}\mathcal{D}_{train}=\{(\bm{x},y)\} indicates the (benign) training set. The target of backdoor attack is to obtain an infected model, which performs well on benign tesing images whereas it may have been injected some insidious backdoors.

2 The Effects of Different Characteristics

One backdoor trigger can be specified by two independent characteristics, including location and appearance, as defined in Definition 2. In this section, we study their individual effects.

Settings. We adopt BadNets (Gu et al., 2019) as an example to study their effects. Specifically, we use VGG-19 (Simonyan & Zisserman, 2015) and ResNet-34 (He et al., 2016) as the model structure, and conduct experiments on CIFAR-10 dataset (Krizhevsky et al., 2009). The trigger is a 3×33\times 3 black-gray square, as shown in Figure 3. We adopt the attack success rate (ASR), which is defined as the accuracy of attacked images predicted by the infected classifier, to evaluate the attack performance.

The Effect of Location. While preserving the appearance of the trigger, we change its location in inference process to study its effect to the attack performance. As shown in Figure 2, when moving the location with a small distance (e.g.e.g., 2∼32\sim 3 pixels), the ASR will drop sharply from 100%100\% to below 50%50\%. It tells that the attack performance is sensitive to the location of the backdoor trigger.

The Effect of Appearance. While keeping the location of the trigger, we change its appearance in the inference stage to study its effect to the attack performance. The trigger appearance could be modified by changing the shape or the pixel values. For the sake of simplicity, here we only consider the change of pixel values. Specifically, there are only two values of the pixels within the trigger, i.e.i.e., 0 and 128. We change the value 128 to different values from 0 to 255. As shown in Figure 2, the ASR degrades sharply along with the decreasing of non-zero pixel values, while is not significantly influenced when the values are increased. According to this simple experiment, it is difficult to describe the exact relationship between the appearance and the attack performance since the change modes of appearance are rather diverse. However, it at least tells that the attack is sensitive to the trigger appearance. More explorations about this phenomenon will be discussed in our future work.

Transformation-inspired Defense and Attack Enhancement

Since the user doesn’t have the information about the trigger, it is impossible to exactly manipulate it in the inference process. Instead, we propose a transformation-based defense by changing the whole image with some transformations (e.g., flipping or scaling), as shown in Definition 3.

The transformation-based defense is defined as introducing a transformation-based pre-processing module on the testing image before prediction, i.e.i.e., instead of predicting x\bm{x}, it predicts T(x)T(\bm{x}), where T(⋅)T(\cdot) is a transformation.

This simple strategy enjoys several advantages: (1) it is efficient since it only requires to transform the testing image; (2) it is attack-agnostic, therefore it can defend different attacks simultaneously; (3) it is data-free and model-free, i.e.,i.e., the defender does not need to have any additional samples or modify the model. Therefore it would be the primary choice when adopting third-party model APIs.

2 Transformation-based Enhancement and Physical Backdoor Attack

Once transformations adopted by the user/defender are known, it would be easy to design an adaptive attack by introducing those transformations in the training process. However, attackers usually have no information about the inference process. To tackle this difficulty, we propose to approximate them with a set of widely adoped transformations Ti(⋅;θi){T_{i}(\cdot;\theta_{i})}. For each TiT_{i}, we define a value domain Θi\Theta_{i} for θi\theta_{i}. Θi\Theta_{i} is parameterized by the maximal transformation size ϵi\epsilon_{i}, i.e.i.e., Θi={θ∣disti(θ,I)≤ϵi},\Theta_{i}=\{\theta|dist_{i}(\theta,I)\leq\epsilon_{i}\}, where disti(⋅,⋅)dist_{i}(\cdot,\cdot) is a given distance metric for TiT_{i} and II indicates the identity transformation.

Consequently, the (compound) transformation used in the enhanced attack is specified as T={T(⋅;θ)∣θ∈∏i=1nΘi}\mathcal{T}=\{T(\cdot;\bm{\theta})|\bm{\theta}\in\prod_{i=1}^{n}\Theta_{i}\}. Then, the training objective of the enhanced attack is formulated as

To solve the problem (1) exactly, attackers need to conduct the training process with all possible transformed variants, which is computation-consuming. Instead, we propose a sampling-based method where we sample only one configuration, i.e., θ∼∏i=1nΘi\bm{\theta}\sim\prod_{i=1}^{n}\Theta_{i} to transform each poisoned image in each time. Then, we use the transformed poisoned images and benign images for training.

Connecting the proposed attack enhancement and physical attack. In real-world scenarios, the testing image may be acquired by some digitizing devices. As such, the trigger in the digitized image may be different from the one used for training. These differences can be approximated by some widely used transformations (e.g., spatial transformations), which have been incorporated into the proposed attack enhancement. Thus, it is expected that attacks with the proposed enhancement can still be effective in the physical world, which will be futher verified in Section 4.2.

Experiment

Settings. We use three representative backdoor attacks, including BadNets (Gu et al., 2017), Blended Attack (Chen et al., 2017), and Consistent Attack (Turner et al., 2019) to evaluate the performance of backdoor defenses. We examine two simple spatial transformations, including left-right flipping (dubbed Flip), and padding after shrinking (dubbed ShrinkPad). Specifically, ShrinkPad consists of shrinking (based on bilinear interpolation) with a few pixels (i.e.i.e., shrinking size), and random zero-padding around the shrunk image. For defense comparison, we select four important baseline, including fine-pruning (Liu et al., 2018), neural cleanse (Wang et al., 2019), auto-encoder based defense (dubbed Auto-Encoder) (Liu et al., 2017), and standard training (dubbed Standard).

Results. As shown in Table 1, our method is effective. Specifically, ShrinkPad with 4 pixels shrinking size could decrease the ASR by more than 90%90\% in all cases. Flip also shows satisfied defense performance towards BadNets and Blended attacks. But it doesn’t work on defending against Consistent Attack since its trigger is symmetric. Compared with the state-of-the-art preprocessing based method (i.e.,i.e., Auto-Encoder), the proposed method has higher clean accuracy and lower ASR in general. Besides, its performance is even on par with Fine-Pruning and Neural Cleanse, which require stronger defensive capabilities (i.e., modify the model parameters and access to benign samples).

2 Attack Enhancement

Resistance to Transformation-based Defense. In the enhanced backdoor attack, we adopt random Flip followed by random ShrinkPad in the random transformation layer. There is only one hyper-parameter in the enhanced attack, i.e.i.e., the maximal shrinking size, which is set to 4 pixels. Other settings are the same as those used in Section 4.1. As demonstrated in Table 2, enhanced backdoor attacks can still achieve a high ASR even under the defenses with spatial transformations. Specifically, the ASR of enhanced backdoor attacks is better than the one of their corresponding standard attack under defenses in almost all cases. The only exception is the Consistent Attack+ under Flip defense. It is partially due to the fact the trigger of Consistent Attack is symmetrical, as mentioned in Section 4.1. Besides, compared to BadNets+ and Blended Attack+, Consistent Attack+ poisoned fewer images (see the attack settings), which is not favorable to the random trigger.

Attack in the Physical World. In this section, we verify the effectiveness of our attack enhancement in the physical world. Since patch stamping instead of pixel-wise manipulation would more possibly happen in real-world applications, we compare BadNets and BadNets+ in this experiment. We randomly pick some attacked samples on CIFAR-10 to take pictures with differently relative location (near and far), as shown in Figure 4. BadNets+ successfully enforces the prediction of all figures to the target label, whereas BadNets fails. These results verify the connection between our enhancement and the physical attack, as stated in Section 3.2.

Conclusion

In this paper, we explore the property of backdoor attacks. We reveal that existing attacks are mostly transformation vulnerable. We propose a transformation-based enhancement to reduce the vulnerability and link the proposed enhancement to the physical attack. We hope that our approach could inspire more explorations on backdoor properties, to help the design of more advanced methods.

References