Backdoor Defense via Decoupling the Training Process

Kunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin, Kui Ren

Introduction

Deep learning, especially deep neural networks (DNNs), has been widely adopted in many realms (Wang et al., 2020b; Li et al., 2020a; Wen et al., 2020) for its high effectiveness. In general, the training of DNNs requires a large amount of training samples and computational resources. Accordingly, third-party resources (e.g.e.g., third-party data or servers) are usually involved. While the opacity of the training process brings certain convenience, it also introduces new security threats.

Backdoor attack poses a new security threat to the training process of DNNs (Li et al., 2020c). It maliciously manipulates the prediction of the attacked DNNs by poisoning a few training samples. Specifically, backdoor attackers inject the backdoor trigger (i.e.i.e., a particular pattern) to some benign training images and change their labels with the attacker-specified target label. The connection between the backdoor trigger and the target label will be learned by DNNs during the training process. In the inference process, the prediction of attacked DNNs will be changed to the target label when the trigger is present, whereas the attacked DNNs will behave normally on benign samples. As such, users are difficult to realize the existence of hidden backdoors and therefore this attack is a serious threat to the practical applications of DNNs.

In this paper, we first investigate backdoor attacks from the hidden feature space. Our preliminary experiments reveal that the backdoor is embedded in the feature space, i.e.i.e., samples with the backdoor trigger (dubbed poisoned samples) tend to cluster together in the feature space. We reveal that this phenomenon is mostly due to the end-to-end supervised training paradigm. Specifically, the excessive learning capability allows DNNs to learn features about the backdoor trigger, while the DNNs can shrink the distance between poisoned samples in the feature space and connect the learned trigger-related features with the target label by the end-to-end supervised training. Based on this understanding, we propose to decouple the end-to-end training process for the backdoor defense. Specifically, we treat the DNNs as two disjoint parts, including a feature extractor (i.e.i.e., backbone) and a simple classifier (i.e.i.e., the remaining fully connected layers). We first learn the purified feature extractor via self-supervised learning (Kolesnikov et al., 2019; Chen et al., 2020a; Jing & Tian, 2020) with unlabeled training samples (obtained by removing their labels), and then learn the simple classifier via standard supervised training process based on the learned feature extractor and all training samples. The strong data augmentations involved in the self-supervised learning damage trigger patterns, making them unlearnable during representation learning; and the decoupling process further disconnects trigger patterns and the target label. Accordingly, hidden backdoors cannot be successfully created even the model is trained on the poisoned dataset based on our defense.

Moreover, we further reveal that the representation of poisoned samples generated by the purified extractor is significantly different from those generated by the extractor learned with standard training process. Specifically, the poisoned sample lies closely to samples with its ground-truth label instead of the target label. This phenomenon makes the training of the simple classifier similar to label-noise learning (Wang et al., 2019b; Ma et al., 2020; Berthon et al., 2021). As such, we first filter high-credible training samples (i.e.i.e., training samples that are most probably to be benign) and then use those samples as labeled samples and the remaining part to form unlabeled samples to fine-tune the whole model via semi-supervised learning (Rasmus et al., 2015; Berthelot et al., 2019; Sohn et al., 2020). This approach is to further reduce the adverse effects of poisoned samples.

The main contributions of this paper are three-fold. (1) We reveal that the backdoor is embedded in the feature space, which is mostly due to the end-to-end supervised training paradigm. (2) Based on our understanding, we propose a decoupling-based backdoor defense (DBD) to alleviate the threat of poisoning-based backdoor attacks. (3) Experiments on classical benchmark datasets are conducted, which verify the effectiveness of our defense.

Related Work

Backdoor attack is an emerging research area, which raises security concerns about training with third-party resources. In this paper, we focus on the poisoning-based backdoor attack towards image classification, where attackers can only modify the dataset instead of other training components (e.g.e.g., training loss). This threat could also happen in other tasks (Xiang et al., 2021; Zhai et al., 2021; Li et al., 2022) and with different attacker’s capacities (Nguyen & Tran, 2020; Tang et al., 2020; Zeng et al., 2021a), which are out-of-scope of this paper. In general, existing attacks can be divided into two main categories based on the property of target labels, as follows:

Poison-Label Backdoor Attack. It is currently the most common attack paradigm, where the target label is different from the ground-truth label of poisoned samples. BadNets (Gu et al., 2019) is the first and most representative poison-label attack. Specifically, it randomly selected a few samples from the original benign dataset to generate poisoned samples by stamping the backdoor trigger onto the (benign) image and change their label with an attacker-specified target label. Those generated poisoned samples associated with remaining benign ones were combined to form the poisoned training dataset, which will be delivered to users. After that, (Chen et al., 2017) suggested that the poisoned image should be similar to its benign version for the stealthiness, based on which they proposed the blended attack. Recently, (Xue et al., 2020; Li et al., 2020b; 2021c) further explored how to conduct poison-label backdoor attacks more stealthily. Most recently, a more stealthy and effective attack, the WaNet (Nguyen & Tran, 2021), was proposed. WaNet adopted image warping as the backdoor trigger, which deforms but preserves the image content.

Clean-Label Backdoor Attack. Although the poisoned image generated by poison-label attacks could be similar to its benign version, users may still notice the attack by examining the image-label relationship. To address this problem, Turner et al. (2019) proposed the clean-label attack paradigm, where the target label is consistent with the ground-truth label of poisoned samples. Specifically, they first leveraged adversarial perturbations or generative models to modify some benign images from the target class and then conducted the standard trigger injection process. This idea was generalized to attack video classification in (Zhao et al., 2020b), where they adopted the targeted universal adversarial perturbation (Moosavi-Dezfooli et al., 2017) as the trigger pattern. Although clean-label backdoor attacks are more stealthy compared with poison-label ones, they usually suffer from relatively poor performance and may even fail in creating backdoors (Li et al., 2020c).

2 Backdoor Defense

Currently, there are also some approaches to alleviate the backdoor threat. Existing defenses are mostly empirical, which can be divided into five main categories, including (1) detection-based defenses (Xu et al., 2021; Zeng et al., 2021a; Xiang et al., 2022), (2) preprocessing based defenses (Doan et al., 2020; Li et al., 2021b; Zeng et al., 2021b), (3) model reconstruction based defenses (Zhao et al., 2020a; Li et al., 2021a; Zeng et al., 2022), (4) trigger synthesis based defenses (Guo et al., 2020; Dong et al., 2021; Shen et al., 2021), and (5) poison suppression based defenses (Du et al., 2020; Borgnia et al., 2021). Specifically, detection-based defenses examine whether a suspicious DNN or sample is attacked and it will deny the use of malicious objects; Preprocessing based methods intend to damage trigger patterns contained in attack samples to prevent backdoor activation by introducing a preprocessing module before feeding images into DNNs; Model reconstruction based ones aim at removing the hidden backdoor in DNNs by modifying models directly; The fourth type of defenses synthesize potential trigger patterns at first, following by the second stage that the hidden backdoor is eliminated by suppressing their effects; The last type of methods depress the effectiveness of poisoned samples during the training process to prevent the creation of hidden backdoors. In general, our method is most relevant to this type of defenses.

In this paper, we only focus on the last four types of defenses since they directly improve the robustness of DNNs. Besides, there were also few works focusing on certified backdoor defenses (Wang et al., 2020a; Weber et al., 2020). Their robustness is theoretically guaranteed under certain assumptions, which cause these methods to be generally weaker than empirical ones in practice.

3 Semi-supervised and Self-supervised Learning

Semi-supervised Learning. In many real-world applications, the acquisition of labeled data often relies on manual labeling, which is very expensive. In contrast, obtaining unlabeled samples is much easier. To utilize the power of unlabeled samples with labeled ones simultaneously, a great amount of semi-supervised learning methods were proposed (Gao et al., 2017; Berthelot et al., 2019; Van Engelen & Hoos, 2020). Recently, semi-supervised learning was also introduced in improving the security of DNNs (Stanforth et al., 2019; Carmon et al., 2019), where they utilized unlabelled samples in the adversarial training. Most recently, (Yan et al., 2021) discussed how to backdoor semi-supervised learning. However, this approach needs to control other training components (e.g.e.g., training loss) in addition to modifying training samples and therefore is out-of-scope of this paper. How to adopt semi-supervised learning for backdoor defense remains blank.

Self-supervised Learning. This learning paradigm is a subset of unsupervised learning, where DNNs are trained with supervised signals generated from the data itself (Chen et al., 2020a; Grill et al., 2020; Liu et al., 2021). It has been adopted for increasing adversarial robustness (Hendrycks et al., 2019; Wu et al., 2021; Shi et al., 2021). Most recently, there were also a few works (Saha et al., 2021; Carlini & Terzis, 2021; Jia et al., 2021) exploring how to backdoor self-supervised learning. However, these attacks are out-of-scope of this paper since they need to control other training components (e.g.e.g., training loss) in addition to modifying training samples.

Revisiting Backdoor Attacks from the Hidden Feature Space

In this section, we analyze the behavior of poisoned samples from the hidden feature space of attacked models and discuss its inherent mechanism.

Settings. We conduct the BadNets (Gu et al., 2019) and label-consistent attack (Turner et al., 2019) on CIFAR-10 dataset (Krizhevsky, 2009) for the discussion. They are representative of poison-label attacks and clean-label attacks, respectively. Specifically, we conduct supervised learning on the poisoned datasets with the standard training process and self-supervised learning on the unlabelled poisoned datasets with SimCLR (Chen et al., 2020a). We visualize poisoned samples in the hidden feature space generated by attacked DNNs based on the t-SNE (Van der Maaten & Hinton, 2008). More detailed settings are presented in Appendix A.

Results. As shown in Figure 1-1, poisoned samples (denoted by ‘black-cross’) tend to cluster together to form a separate cluster after the standard supervised training process, no matter under the poison-label attack or clean-label attack. This phenomenon implies why existing poisoning-based backdoor attacks can succeed. Specifically, the excessive learning capability allows DNNs to learn features about the backdoor trigger. Associated with the end-to-end supervised training paradigm, DNNs can shrink the distance between poisoned samples in the feature space and connect the learned trigger-related features with the target label. In contrast, as shown in Figure 1-1, poisoned samples lie closely to samples with their ground-truth label after the self-supervised training process on the unlabelled poisoned dataset. It indicates that we can prevent the creation of backdoors by self-supervised learning, which will be further introduced in the next section.

Decoupling-based Backdoor Defense

General Pipeline of Backdoor Attacks. Let D={(xi,yi)}i=1N\mathcal{D}=\{(\bm{x}_{i},y_{i})\}_{i=1}^{N} denotes the benign training set, where xi∈X={0,1,…,255}C×W×H\bm{x}_{i}\in\mathcal{X}=\{0,1,\ldots,255\}^{C\times W\times H} is the image, yi∈Y={0,1,…,K}y_{i}\in\mathcal{Y}=\{0,1,\ldots,K\} is its label, KK is the number of classes, and yt∈Yy_{t}\in\mathcal{Y} indicates the target label. How to generate the poisoned dataset Dp\mathcal{D}_{p} is the cornerstone of backdoor attacks. Specifically, Dp\mathcal{D}_{p} consists of two subsets, including the modified version of a subset of D\mathcal{D} and remaining benign samples, i.e.i.e., Dp=Dm∪Db\mathcal{D}_{p}=\mathcal{D}_{m}\cup\mathcal{D}_{b}, where Db⊂D\mathcal{D}_{b}\subset\mathcal{D}, γ≜∣Dm∣∣D∣\gamma\triangleq\frac{|\mathcal{D}_{m}|}{|\mathcal{D}|} is the poisoning rate, Dm={(x′,yt)∣x′=G(x),(x,y)∈D\Db}\mathcal{D}_{m}=\left\{(\bm{x}^{\prime},y_{t})|\bm{x}^{\prime}=G(\bm{x}),(\bm{x},y)\in\mathcal{D}\backslash\mathcal{D}_{b}\right\}, and G:X→XG:\mathcal{X}\rightarrow\mathcal{X} is an attacker-predefined poisoned image generator. For example, G(x)=(1−λ)⊗x+λ⊗tG(\bm{x})=(\bm{1}-\bm{\lambda})\otimes\bm{x}+\bm{\lambda}\otimes\bm{t}, where λ∈C×W×H\bm{\lambda}\in^{C\times W\times H}, t∈X\bm{t}\in\mathcal{X} is the trigger pattern, and ⊗\otimes is the element-wise product in the blended attack (Chen et al., 2017). Once Dp\mathcal{D}_{p} is generated, it will be sent to users who will train DNNs on it. Hidden backdoors will be created after the training process.

Threat Model. In this paper, we focus on defending against poisoning-based backdoor attacks. The attacker can arbitrarily modify the training set whereas cannot change other training components (e.g.e.g., model structure and training loss). For our proposed defense, we assume that defenders can fully control the training process. This is the scenario that users adopt third-party collected samples for training. Note that we do not assume that defenders have a local benign dataset, which is often required in many existing defenses (Wang et al., 2019a; Zhao et al., 2020a; Li et al., 2021a).

Defender’s Goals. The defender’s goals are to prevent the trained DNN model from predicting poisoned samples as the target label and to preserve the high accuracy on benign samples.

2 Overview of the Defense Pipeline

In this section, we describe the general pipeline of our defense. As shown in Figure 2, it consists of three main stages, including (1) learning a purified feature extractor via self-supervised learning, (2) filtering high-credible samples via label-noise learning, and (3) semi-supervised fine-tuning.

Specifically, in the first stage, we remove the label of all training samples to form the unlabelled dataset, based on which to train the feature extractor via self-supervised learning. In the second stage, we freeze the learned feature extractor and adopt all training samples to train the remaining fully connected layers via supervised learning. We then filter α%\alpha\% high-credible samples based on the training loss. The smaller the loss, the more credible the sample. After the second stage, the training set will be separated into two disjoint parts, including high-credible samples and low-credible samples. We use high-credible samples as labeled samples and remove the label of all low-credible samples to fine-tune the whole model via semi-supervised learning. More detailed information about each stage of our method will be further illustrated in following sections.

3 Learning Purified Feature Extractor via Self-supervised Learning

Let Dt\mathcal{D}_{t} denotes the training set and fw:X→Kf_{\bm{w}}:\mathcal{X}\rightarrow^{K} indicates the DNN with parameter w=[wc,wf]\bm{w}=[\bm{w}_{c},\bm{w}_{f}], where wc\bm{w}_{c} and wf\bm{w}_{f} indicates the parameters of the backbone and the fully connected layer, respectively. In this stage, we optimize wc\bm{w}_{c} based on the unlabeled version of Dt\mathcal{D}_{t} via self-supervised learning, as follows:

where L1(⋅)\mathcal{L}_{1}(\cdot) indicates the self-supervised loss (e.g.e.g., NT-Xent in SimCLR (Chen et al., 2020a)). Through the self-supervised learning, the learned feature extractor (i.e.i.e., backbone) will be purified even if the training set contains poisoned samples, as illustrated in Section 3.

4 Filtering High-credible Samples via Label-noise Learning

Once wc∗\bm{w}_{c}^{*} is obtained, the user can freeze it and adopt Dt\mathcal{D}_{t} to further optimize remaining wf\bm{w}_{f}, i.e.i.e.,

where L2(⋅)\mathcal{L}_{2}(\cdot) indicates the supervised loss (e.g.e.g., cross entropy).

After the decoupling-based training process (1)-(2), even if the model is (partly) trained on the poisoned dataset, the hidden backdoor cannot be created since the feature extractor is purified. However, this simple strategy suffers from two main problems. Firstly, compared with the one trained via supervised learning, the accuracy of predicting benign samples will have a certain decrease, since the learned feature extractor is frozen in the second stage. Secondly, poisoned samples will serve as ‘outliers’ to further hinder the learning of the second stage when poison-label attacks appear, since those samples lie close to samples with its ground-truth label instead of the target label in the hidden feature space generated by the learned purified feature extractor. These two problems indicate that we should remove poisoned samples and retrain or fine-tune the whole model.

Specifically, we select high-credible samples Dh\mathcal{D}_{h} based on the loss L2(⋅;[wc∗,wf∗])\mathcal{L}_{2}(\cdot;[\bm{w}_{c}^{*},\bm{w}_{f}^{*}]). The high-credible samples are defined as the α%\alpha\% training samples with the smallest loss, where α∈\alpha\in is a hyper-parameter. In particular, we adopt the symmetric cross-entropy (SCE) (Wang et al., 2019b) as L2(⋅)\mathcal{L}_{2}(\cdot), inspired by the label-noise learning. As shown in Figure 3, compared with the CE loss, the SCE can significantly increase the differences between poisoned samples and benign ones, which further reduces the possibility that high-credible dataset Dh\mathcal{D}_{h} still contains poisoned samples.

Note that we do not intend to accurately separate poisoned samples and benign samples. We only want to ensure that the obtained Dh\mathcal{D}_{h} contains as few poisoned samples as possible.

5 Semi-supervised Fine-tuning

After the second stage, the third-party training set Dt\mathcal{D}_{t} will be separated into two disjoint parts, including the high-credible dataset Dh\mathcal{D}_{h} and the low-credible dataset Dl≜Dt\Dh\mathcal{D}_{l}\triangleq\mathcal{D}_{t}\backslash\mathcal{D}_{h}. Let D^l≜{x∣(x,y)∈Dl}\hat{\mathcal{D}}_{l}\triangleq\{\bm{x}|(\bm{x},y)\in\mathcal{D}_{l}\} indicates the unlabeled version of low-credible dataset Dl\mathcal{D}_{l}. We fine-tune the whole trained model f[wc∗,wf∗](⋅)f_{[\bm{w}_{c}^{*},\bm{w}_{f}^{*}]}(\cdot) with semi-supervised learning as follows:

where L3(⋅)\mathcal{L}_{3}(\cdot) denotes the semi-supervised loss (e.g.e.g., the loss in MixMatch (Berthelot et al., 2019)).

This process can prevent the side-effects of poisoned samples while utilizing their contained useful information, and encourage the compatibility between the feature extractor and the simple classifier via learning them jointly instead of separately. Please refer to Section 5.3 for more results.

Experiments

Datasets and DNNs. We evaluate all defenses on two classical benchmark datasets, including CIFAR-10 (Krizhevsky, 2009) and (a subset of) ImageNet (Deng et al., 2009). We adopt the ResNet-18 (He et al., 2016) for these tasks. More detailed settings are presented in Appendix B.1. Besides, we also provide the results on (a subset of) VGGFace2 (Cao et al., 2018) in Appendix C.

Attack Baselines. We examine all defense approaches in defending against four representative attacks. Specifically, we select the BadNets (Gu et al., 2019), the backdoor attack with blended strategy (dubbed ‘Blended’) (Chen et al., 2017), WaNet (Nguyen & Tran, 2021), and label-consistent attack with adversarial perturbations (dubbed ‘Label-Consistent’) (Turner et al., 2019) for the evaluation. They are the representative of patch-based visible and invisible poison-label attacks, non-patch-based poison-label attacks, and clean-label attacks, respectively.

Defense Baselines. We compared our DBD with two defenses having the same defender’s capacities, including the DPSGD (Du et al., 2020) and ShrinkPad (Li et al., 2021b). We also compare with other two approaches with an additional requirement (i.e.i.e., having a local benign dataset), including the neural cleanse with unlearning strategy (dubbed ‘NC’) (Wang et al., 2019a), and neural attention distillation (dubbed ‘NAD’) (Li et al., 2021a). They are the representative of poison suppression based defenses, preprocessing based defenses, trigger synthesis based defenses, and model reconstruction based defenses, respectively. We also provide results of DNNs trained without any defense (dubbed ‘No Defense’) as another important baseline for reference.

Attack Setups. We use a 2×22\times 2 square as the trigger pattern on CIFAR-10 dataset and the 32×3232\times 32 Apple logo on ImageNet dataset for the BadNets, as suggested in (Gu et al., 2019; Wang et al., 2019a). For Blended, we adopt the ‘Hello Kitty’ pattern on CIFAR-10 and the random noise pattern on ImageNet, based on the suggestions in (Chen et al., 2017), and set the blended ratio λ=0.1\lambda=0.1 on all datasets. The trigger pattern adopted in label-consistent attack is the same as the one used in BadNets. For WaNet, we adopt its default settings on CIFAR-10 dataset. However, on ImageNet dataset, we use different settings optimized by grid-search since the original ones fail. An example of poisoned samples generated by different attacks is shown in Figure 4. Besides, we set the poisoning rate γ1=2.5%\gamma_{1}=2.5\% for label-consistent attack (25% of training samples with the target label) and γ2=5%\gamma_{2}=5\% for three other attacks. More details are shown in Appendix B.2.

Defense Setups. For our DBD, we adopt SimCLR (Chen et al., 2020a) as the self-supervised method and MixMatch (Berthelot et al., 2019) as the semi-supervised method. More details about SimCLR and MixMatch are in Appendix I. The filtering rate α\alpha is the only key hyper-parameter in DBD, which is set to 50% in all cases. We set the shrinking rate to 10% for the ShrinkPad on all datasets, as suggested in (Li et al., 2021b; Zeng et al., 2021b). In particular, DPSGD and NAD are sensitive to their hyper-parameters. We report their best results in each case based on the grid-search (as shown in Appendix D). Besides, we split a 5% random subset of the benign training set as the local benign dataset for NC and NAD. More implementation details are provided in Appendix B.3.

Evaluation Metrics. We adopt the attack success rate (ASR) and benign accuracy (BA) to measure the effectiveness of all methodsAmong all defense methods, the one with the best performance is indicated in boldface and the value with underline denotes the second-best result.. Specifically, let Dtest\mathcal{D}_{test} indicates the (benign) testing set and Cw:X→YC_{\bm{w}}:\mathcal{X}\rightarrow\mathcal{Y} denotes the trained classifier, we have ASR≜Pr⁡(x,y)∈Dtest{Cw(G(x))=yt∣y≠yt}ASR\triangleq\Pr_{(\bm{x},y)\in\mathcal{D}_{test}}\{C_{\bm{w}}(G(\bm{x}))=y_{t}|y\neq y_{t}\} and BA≜Pr⁡(x,y)∈Dtest{Cw(x)=y}BA\triangleq\Pr_{(\bm{x},y)\in\mathcal{D}_{test}}\{C_{\bm{w}}(\bm{x})=y\}, where yty_{t} is the target label and G(⋅)G(\cdot) is the poisoned image generator. In particular, the lower the ASR and the higher the BA, the better the defense.

2 Main Results

Comparing DBD with Defenses having the Same Requirements. As shown in Table 1-2, DBD is significantly better than defenses having the same requirements (i.e.i.e., DPSGD and ShrinkPad) in defending against all attacks. For example, the benign accuracy of DBD is 20% over while the attack success rate is 5% less than that of DPSGD in all cases. Specifically, the attack success rate of models with DBD is less than 2% in all cases (mostly <0.5%<0.5\%), which verifies that our method can successfully prevent the creation of hidden backdoors. Moreover, the decreases of benign accuracy are less than 2%2\% when defending against poison-label attacks, compared with models trained without any defense. Our method is even better on relatively larger dataset where all baseline methods become less effective. These results verify the effectiveness of our method.

Comparing DBD with Defenses having Extra Requirements. We also compare our defense with two other methods (i.e.i.e., NC and NAD), which have an additional requirement that defenders have a benign local dataset. As shown in Table 1-2, NC and NAD are better than DPSGD and ShrinkPad, as we expected, since they adopt additional information from the benign local dataset. In particular, although NAD and NC use additional information, our method is still better than them, even when their performances are tuned to the best while our method only uses the default settings. Specifically, the BA of NC is on par with that of our method. However, it is with the sacrifice of ASR. Especially on ImageNet dataset, NC has limited effects in reducing ASR. In contrast, our method reaches the smallest ASR while its BA is either the highest or the second-highest in almost all cases. These results verify the effectiveness of our method again.

Results. As shown in Figure 7, our method can still prevent the creation of hidden backdoors even when the poisoning rate reaches 20%. Besides, DBD also maintains high benign accuracy. In other words, our method is effective in defending attacks with different strengths.

3 Ablation Study

There are four key strategies in DBD, including (1) obtaining purified feature extractor, (2) using SCE instead of CE in the second stage, (3) reducing side-effects of low-credible samples, and (4) fine-tuning the whole model via semi-supervised learning. Here we verify their effectiveness.

Settings. We compare the proposed DBD with its four variants, including (1) DBD without SS, (2) SS with CE, (3) SS with SCE, and (4) SS with SCE + Tuning, on the CIFAR-10 dataset. Specifically, in the first variant, we replace the backbone generated by self-supervised learning with the one trained in a supervised fashion and keep other parts unchanged. In the second variant, we freeze the backbone learned via self-supervised learning and train the remaining fully-connected layers with cross-entropy loss on all training samples. The third variant is similar to the second one. The only difference is that it uses symmetric cross-entropy instead of cross-entropy to train fully-connected layers. The last variant is an advanced version of the third one, which further fine-tunes fully-connected layers on high-credible samples filtered by the third variant.

Results. As shown in Table 3, we can conclude that decoupling the original end-to-end supervised training process is effective in preventing the creation of hidden backdoors, by comparing our DBD with its first variant and the model trained without any defense. Besides, we can also verify the effectiveness of SCE loss on defending against poison-label backdoor attacks by comparing the second and third DBD variants. Moreover, the fourth DBD variant has relatively lower ASR and BA, compared with the third one. This phenomenon is due to the removal of low-credible samples. It indicates that reducing side-effects of low-credible samples while adopting their useful information is important for the defense. We can also verify that fine-tuning the whole model via semi-supervised learning is also useful by comparing the fourth variant and the proposed DBD.

4 Resistance to Potential Adaptive Attacks

In our paper, we adopted the classical defense setting that attackers have no information about the defense. Attackers may design adaptive attacks if they know the existence of our DBD. The most straightforward idea is to manipulate the self-supervised training process so that poisoned samples are still in a new cluster after the self-supervised learning. However, attackers are not allowed to do it based on our threat model about adopting third-party datasets. Despite this, attackers may design adaptive attacks by optimizing the trigger pattern to make poisoned samples still in a new cluster after the self-supervised learning if they can know the model structure used by defenders, as follows:

Problem Formulation. For a KK-classification problem, let X′={xi}i=1M\mathcal{X}^{\prime}=\{\bm{x}_{i}\}_{i=1}^{M} indicates the benign images selected for poisoning, Xj={xi}i=1Nj\mathcal{X}_{j}=\{\bm{x}_{i}\}_{i=1}^{N_{j}} denotes the benign images with ground-truth label jj, and gg is a trained backbone. Given an attacker-predefined poisoned image generator GG, the adaptive attack aims to optimize a trigger pattern t\bm{t} by minimizing the distance between poisoned images while maximizing the distance between the center of poisoned images and centers of clusters of benign images with different label, i.e.,i.e.,

where g‾′≜1M∑x∈X′g(G(x;t))\overline{g}^{\prime}\triangleq\frac{1}{M}\sum_{\bm{x}\in\mathcal{X}^{\prime}}g(G(\bm{x};\bm{t})), gi‾≜1Ni∑x∈Xig(x)\overline{g_{i}}\triangleq\frac{1}{N_{i}}\sum_{\bm{x}\in\mathcal{X}_{i}}g(\bm{x}), and dd is a distance metric.

Results. The adaptive attack works well when there is no defense (BA=94.96%, ASR=99.70%). However, this attack still fails to attack our DBD (BA=93.21%, ASR=1.02%). In other words, our defense is resistant to this adaptive attack. It is most probably because the trigger optimized based on the backbone is far less effective when the model is retrained since model parameters are changed due to the random initialization and the update of model weights during the training process.

Conclusion

The mechanism of poisoning-based backdoor attacks is to establish a latent connection between trigger patterns and the target label during the training process. In this paper, we revealed that this connection is learned mostly due to the end-to-end supervised training paradigm. Motivated by this understanding, we proposed a decoupling-based backdoor defense, which first learns the backbone via self-supervised learning and then the remaining fully-connected layers by the classical supervised learning. We also introduced the label-noise learning method to determine high-credible and low-credible samples, based on which we fine-tuned the whole model via semi-supervised learning. Extensive experiments verify that our defense is effective on reducing backdoor threats while preserving high accuracy on predicting benign samples.

Acknowledgments

Baoyuan Wu is supported in part by the National Natural Science Foundation of China under Grant 62076213, the University Development Fund of the Chinese University of Hong Kong, Shenzhen under Grant 01001810, and the Special Project Fund of Shenzhen Research Institute of Big Data under Grant T00120210003. Zhan Qin is supported in part by the National Natural Science Foundation of China under Grant U20A20178, the National Key Research and Development Program of China under Grant 2020AAA0107705, and the Research Laboratory for Data Security and Privacy, Zhejiang University-Ant Financial Fintech Center. Kui Ren is supported by the National Key Research and Development Program of China under Grant 2020AAA0107705.

Ethics Statement

DNNs are widely adopted in many mission-critical areas (e.g.e.g., face recognition) and therefore their security is of great significance. The vulnerability of DNNs to backdoor attacks raises serious concerns about using third-party training resources. In this paper, we propose a general training pipeline to obtain backdoor-free DNNs, even if the training dataset contains poisoned samples. This work has no ethical issues in general since our method is purely defensive and does not reveal any new vulnerabilities of DNNs. However, we need to mention that our defense can be adopted only when training with untrusted samples, and backdoor attacks could happen in other scenarios. People should not be too optimistic about eliminating backdoor threats.

Reproducibility Statement

The detailed descriptions of datasets, models, and training settings are in Appendix A-D. We also describe the computational facilities and cost in Appendix J-K. Codes of our DBD are also open-sourced.

References

Appendix A Detailed Settings for Revisiting Backdoor Attacks

Training Setups. We conduct supervised learning on the poisoned datasets with the standard training process and the self-supervised learning on the unlabelled poisoned datasets with the SimCLR (Chen et al., 2020a). The supervised training is conducted based on the open-source codehttps://github.com/kuangliu/pytorch-cifar. Specifically, we use the SGD optimizer with momentum 0.9, weight decay of 5×10−45\times 10^{-4}, and an initial learning rate of 0.1. The batch size is set to 128 and we train the ResNet-18 model 200 epochs. The learning rate is decreased by a factor of 10 at epoch 100 and 150, respectively. Besides, we add triggers before performing the data augmentation (e.g.e.g., random crop and horizontal flipping). For the self-supervised training, we use the stochastic gradient descent (SGD) optimizer with a momentum of 0.9, an initial learning rate of 0.4, and a weight decay factor of 5×10−45\times 10^{-4}. We use a batch size of 512, and train the backbone for 1,000 epochs. We decay the learning rate with the cosine decay schedule (Loshchilov & Hutter, 2016) without a restart. Besides, we also adopt strong data augmentation techniques, including random crop and resize (with random flip), color distortions, and Gaussian blur, as suggested in (Chen et al., 2020a). All models are trained until converge.

t-SNE Visualization Settings. We treat the output of the last residual unit as the feature representation and use the tsne-cuda library (Chan et al., 2019) to get the feature embedding of all samples. To have a better visualization, we adopt all poisoned samples and randomly select 10% benign samples for visualizing models under the supervised learning, and adopt 30% poisoned samples and 10% benign samples for those under the self-supervised learning.

Appendix B Detailed Settings for Main Experiments

Due to the limitations of computational resources and time, we adopt a subset randomly selected from the original ImageNet. More detailed information about the datasets and DNNs adopted in the main experiments of our paper is presented in Table 4.

B.2 More Details about Attack Settings

Attack Setups. We conduct the BadNets (Gu et al., 2019), blended attack (dubbed ‘Blended’) (Chen et al., 2017), label-consistent attack (dubbed ‘Label-Consistent’) (Turner et al., 2019), and WaNet (Nguyen & Tran, 2021) with the target label yt=3y_{t}=3 on all datasets. The trigger patterns are the same as those presented in Section 5.2. In particular, we set the blended ratio λ=0.1\lambda=0.1 for the blended attack on all datasets and examine label-consistent attack with the maximum perturbation size ϵ∈{16,32}\epsilon\in\{16,32\}. Besides, WaNet assumed that attackers can fully control the whole training process in its original paper. However, we found that WaNet only modified training data while other training components (e.g.e.g., training loss, training schedule, and model structure) are the same as those used in the standard training process. As such, we re-implement its code in the poisoning-based attack scenario based on its official codehttps://github.com/VinAIResearch/Warping-based_Backdoor_Attack-release. Specifically, following the settings in its original paper, we set the noise rate ρn=0.2\rho_{n}=0.2, control grid size k=4k=4, and warping strength s=0.5s=0.5 on the CIFAR-10 dataset. However, we found that the default kk and ss are too small to make the attack works on the ImageNet dataset (as shown in Table 5-6). Besides, the ‘noise mode’ also significantly reduces the attack effectiveness (as shown in Table 7). As such, we set k=224k=224 and s=1s=1 and train models without the noise mode on the ImageNet dataset.

Training Setups. On the CIFAR-10 dataset (Krizhevsky, 2009), the settings are the same as those described in Section A; On the ImageNet dataset (Deng et al., 2009), we conduct experiments based on the open-source codehttps://github.com/pytorch/examples/tree/master/imagenet. Specifically, we use the SGD optimizer with momentum 0.9, weight decay of 10−410^{-4}, and an initial learning rate of 0.1. The batch size is set to 256 and we train the ResNet-18 model 90 epochs. The learning rate is decreased by a factor of 10 at epoch 30 and 60, respectively. Besides, since the raw images in the ImageNet dataset are of different sizes, we resize them to 3×224×2243\times 224\times 224 before adding triggers.

B.3 More Details about Defense Settings

Settings for NC. We conduct reverse engineering and anomaly detection based on its open-source codehttps://github.com/bolunwang/backdoor. We implement the ‘unlearning’ method to patch attacked models, as suggested in its paper (Wang et al., 2019a). We randomly select 5% benign training samples as the local benign dataset, which is used in the ‘unlearning’ process. Unless otherwise specified, other settings are the same as those used in (Wang et al., 2019a).

Settings for NAD. We implement this method based on its open-source codehttps://github.com/bboylyg/NAD. The origin NAD only conducted experiments on the WideResNet model. In our paper, we calculate the NAD loss over the last residual group for the ResNet-18. The local benign dataset is the same as the one adopted in NC, which is used in the fine-tuning and distillation process of NAD. Unless otherwise specified, other settings are the same as those used in (Li et al., 2021a).

Settings for DPSGD. The original DPSGD was conducted on the MNIST dataset implemented by the TensorFlow Framework. In this paper, we re-implement it based on the differentially private SGD method provided by the Opacushttps://github.com/pytorch/opacus. Specifically, we replace the original SGD optimizer with the differentially private one, as suggested in (Du et al., 2020). There are two important hyper-parameters in DPSGD, including noise scales σ\sigma and the clipping bound CC. In the experiments, we set C=1C=1 and select the best σ\sigma by the grid-search.

Settings for ShrinkPad. We set the shrinking rate to 10% on all datasets, as suggested in (Li et al., 2021b; Zeng et al., 2021b). Following their settings, we pad 0-pixels at the bottom right of the shrunk image to expand it to its original size.

Settings for our Defense. In this first stage, We adopt SimCLR (Chen et al., 2020a) to perform self-supervised learning. We train backbones 100 instead of 1,000 epochs to reduce computational costs while preserving effectiveness. Other settings are the same as those described in Section A. We use the same settings across all datasets, models, and attacks; In the second stage, we use the Adam optimizer with a learning rate of 0.002 and set the batch size to 128. We train the fully connected layers 10 epochs with the SCE loss (Wang et al., 2019b). Two hyper-parameters involved in the SCE (i.e.i.e., α\alpha and β\beta in the original paper) are set to 0.1 and 1, respectively. After that, we filter 50% high-credible samples. We use the same settings across all datasets, models, and attacks; In the third stage, we adopt the MixMatch (Berthelot et al., 2019) for semi-supervised fine-tuning with settings suggested in its original paper. Specifically, we use the Adam optimizer with a learning rate of 0.002, the batch size of 64, and finetune the model 190 epochs on the CIFAR-10 and 80 epochs on the ImageNet dataset, respectively. We set the temperature T=0.5T=0.5 and the weight of unsupervised loss λu=15\lambda_{u}=15 on the CIFAR-10 and λu=6\lambda_{u}=6 on the ImageNet dataset, respectively. Moreover, we re-filter high-credible samples after every epoch of the third stage based on the SCE loss.

Appendix C Defending against Attacks on VGGFace2 Dataset

Dataset and DNN. Due to the limitations of computational resources and time, we adopt a subset randomly selected from the original VGGFace2 (Cao et al., 2018). More details are in Table 8.

Settings for Attacks. For the training of models on the VGGFace2 dataset, the batch size is set to 32 and we conduct experiments on the DenseNet-121 model (Huang et al., 2017). An example of poisoned samples generated by different attacks are in Figure 5. Other settings are the same as those used on the ImageNet dataset.

Settings for Defenses. For NAD, we calculate the NAD loss over the second to last layer for the DenseNet-121. Other settings are the same as those used on the ImageNet dataset.

Results. As shown in Table 9, our defense still reaches the best performance even compared with NC and NAD. Specifically, the BA of NC is on par with that of our method whereas it is with the sacrifice of ASR. These results verify the effectiveness of our defense again.

Appendix D Searching Best Results for DPSGD and NAD

The effectiveness of DPSGD and NAD is sensitive to their hyper-parameters. Here we search for their best results based on the criteria that ‘BA −- ASR’ reaches the highest value after the defense.

In general, the larger the σ\sigma, the smaller the ASR while also the smaller the BA. The results of DPSGD are shown in Table 10-12, where the best results are marked in boldface.

D.2 Searching Best Results for NAD

We found that the fine-tuning stage of NAD is sensitive to the learning rate. We search the best initial learning rate from {0.1,0.01,0.001}\{0.1,0.01,0.001\}. As shown in Table 13-15, a very large learning rate significantly reduces the BA, while a very small learning rate can not reduce the ASR effectively. To keep a relatively large BA while maintaining a small ASR, we set η=0.01\eta=0.01 in the fine-tuning stage.

The distillation stage of NAD is also sensitive to its hyper-parameter β\beta. We select the best β\beta via the grid-search. The results are shown in Table 16-19.

Appendix E Defending against Label-Consistent Attack with a Smaller Poisoning Rate

For the label-consistent attack, except for the 2.5% poisoning rate examined in the main manuscript, 0.6% is also an important setting provided in its original paper (Turner et al., 2019). In this section, we compare different defenses against the label-consistent attack with poisoning rate γ=0.6%\gamma=0.6\%.

As shown in Table 20, when defending against label-consistent attack with a 0.6% poisoning rate, our method is still significantly better than defenses having the same requirements (i.e.i.e., DPSGD and ShrinkPad). Even compared with those having the additional requirement (i.e.i.e., NC and NAD) under their best settings, our defense is still better or on par with them under the default settings. These results verify the effectiveness of our method again.

Appendix F Defending against Attacks with Different Trigger Patterns

In this section, we verify whether DBD is still effective when different trigger patterns are adopted.

Settings. For simplicity, we adopt the BadNets on the CIFAR-10 dataset as an example for the discussion. Specifically, we change the location and size of the backdoor trigger while keeping other settings unchanged to evaluate the BA and ASR before and after our defense.

Results. As shown in Table 21, although there are some fluctuations, the ASR is smaller than 2%2\% while the BA is greater than 92%92\% in every cases. In other words, our method is effective in defending against attacks with different trigger patterns.

Appendix G Defending against Attacks with Dynamic Triggers

In this section, we verify whether DBD is still effective when attackers adopt dynamic triggers.

Settings. We compare DBD with MESA (Qiao et al., 2019) in defending the dynamic attack discussed in (Qiao et al., 2019) on the CIFAR-10 dataset as an example for the discussion. This dynamic attack uses a distribution of triggers instead of a fixed trigger.

Results. The BA and ASR of DBD are 92.4% and 0.4%, while those of MESA are 94.8% and 2.4%. However, we find MESA failed in defending against blended attack (for it can not correctly detect the trigger) whereas DBD is still effective. These results verified the effectiveness of our defense.

Appendix H Discussions

Settings. Here we analyze the effect of filtering rate α\alpha, which is the only key method-related hyper-parameter in our DBD. We adopt the results on the CIFAR-10 dataset for discussion. Except for the studied parameter α\alpha, other settings are the same as those used in Section 5.2.

Results. The number of labeled samples used in the third stage increase with the increase of filtering rate α\alpha, while the probability that the filtered high-credible dataset contains poisoned samples also increases. As shown in Figure 7, DBD can still maintain relatively high benign accuracy even when the filtering rate α\alpha is relatively small (e.g.e.g., 30%). It is mostly due to the high-quality of learned purified feature extractor and the semi-supervised fine-tuning process. DBD can also reach a nearly 0% attack success rate in all cases. However, we also have to notice that the high-credible dataset may contain poisoned samples when α\alpha is very large, which in turn creates hidden backdoors again during the fine-tuning process. Defenders should specify α\alpha based on their specific needs.

H.2 Defending Attacks with Various Poisoning Rates

Settings. We evaluate our method in defending against attacks with different poisoning rate γ\gamma on CIFAR-10 dataset. Except for γ\gamma, other settings are the same as those used in Section 5.2.

Appendix I More Details about SimCLR, SCE, and MixMatch

NT-Xent Loss in SimCLR. Given a sample mini-batch containing NN different samples, SimCLR first applies two separate data augmentations toward each sample to obtain 2N2N augmented samples. The loss for a positive pair of sample (i,j)(i,j) can be defined as:

SCE. The symmetric cross entropy (SCE) can be defined as:

where H(p,q)H(p,q) is the cross entropy, H(q,p)H(q,p) is the reversed cross entropy, pp is the prediction, and qq is the one-hot label (of the evaluated sample).

MixMatch Loss. For a batch X\mathcal{X} of labeled samples and a batch U\mathcal{U} of unlabeled samples (∣X∣=∣U∣|\mathcal{X}|=|\mathcal{U}|), MixMatch produces a guessed label qˉ\bar{q} for each unlabled sample u∈Uu\in\mathcal{U} and applies MixUp (Zhang et al., 2018) to obtain the augmented X′\mathcal{X}^{{}^{\prime}} and U′\mathcal{U}^{{}^{\prime}}. The loss LX\mathcal{L}_{\mathcal{X}} and LU\mathcal{L}_{\mathcal{U}} can be defined as:

where pxp_{x} is the prediction of xx, qq is its one-hot label, and H(⋅,⋅)H(\cdot,\cdot) is the cross entropy.

where pup_{u} is the prediction of uu, qˉ\bar{q} is its guessed one-hot label, and KK is the number of classes.

By combining LX\mathcal{L}_{\mathcal{X}} with LU\mathcal{L}_{\mathcal{U}}, the MixMatch loss can be defined as:

where λU\lambda_{\mathcal{U}} is a hyper-parameter.

Appendix J Computational Facilities

We conduct all experiments on two Ubuntu 18.04 servers having different GPUs. One has four NVIDIA GeForce RTX 2080 Ti GPUs with 11GB memory (dubbed ‘RTX 2080Ti’) and the another has three NVIDIA Tesla V100 GPUs with 32GB memory (dubbed ‘V100’).

Computational Facilities for Attacks. All experiments are conducted with a single RTX 2080 Ti.

Computational Facilities for Defenses. Since we do not use a memory-efficient implementation of DenseNet-121, we conduct DPSGD experiments on the VGGFace2 dataset with a single V100. Other experiments of baseline defenses are conducted with a single RTX 2080 Ti. For our defense, we adopt PyTorch (Paszke et al., 2019) distributed data-parallel and automatic mixed precision training (Micikevicius et al., 2018) with two RTX 2080 Ti for self-supervised learning on the VGGFace2 dataset. Other experiments are conducted with a single RTX 2080 Ti.

Appendix K Computational Cost

In this section, we analyze the computational cost of our method stage by stage, compared to standrad supervised learning.

Stage 1. Self-supervised learning is known to have a higher computational cost than standard supervised learning (Chen et al., 2020a; He et al., 2020). In our experiments, SimCLR requires roughly four times the computational cost of standard supervised learning. Since we intend to get a purified instead of well-trained feature extractor, we train the feature extractor (i.e.i.e., backbone) lesser epochs than the original SimCLR to reduce the training time. As described in Section 9, we find 100 epochs is enough to preserve effectiveness.

Stage 2. Since we freeze the backbone and only train the remaining fully connected layers, the computational cost is roughly 60% of standard supervised learning.

Stage 3. Semi-supervised learning is known to have a extra labeling cost compared with standard supervised learning (Gao et al., 2020). In our experiments, MixMatch requires roughly two times the computation cost of standard supervised learning.

We will explore a more computational efficient training method in our future work.

Appendix L Comparing our DBD with Detection-based Backdoor Defenses

In this paper, we do not intend to filter malicious and benign samples accurately, as we mentioned in Section 4.4. However, we notice that the second stage of our DBD can serve as a detection-based backdoor defense for it can filter poisoned samples. In this section, we compare the filtering ability of our DBD (stage 2) with existing detection-based backdoor defenses.

Settings. We compare our DBD with two representative detection-based methods, including, Spectral Signatures (SS) (Tran et al., 2018) and Activation Clustering (AC) (Chen et al., 2019), on the CIFAR-10 dataset. These detection-based methods (e.g.e.g., SS and AC) filter malicious samples from the training set and train the model on non-malicious samples. Specifically, we re-implement SS in PyTorch based on its official codehttps://github.com/MadryLab/backdoor_data_poisoning and adopt the open-source codehttps://github.com/ain-soph/trojanzoo/blob/main/trojanvision/defenses/backdoor/activation_clustering.py for AC, following the settings in their original paper. In particular, since SS filters 1.5ε1.5\varepsilon malicious samples for each class, where ε\varepsilon is the key hyper-parameter means the upper bound of the number of poisoned training samples, we adopt different ε\varepsilon for a fair comparison.

Results. As shown in Table 22-23, the filtering performance of DBD is on par with that of SS and AC. DBD is even better than those methods when filtering poisoned samples generated by more complicated attacks (i.e.i.e., WaNet and Label-Consistent). Besides, we also conduct the standard training on non-malicious samples filtered by SS and AC. As shown in Table 24, the hidden backdoor will still be created in many cases, even though the detection-based defenses are sometimes accurate. This is mainly because these methods may not able to remove enough poisoned samples while preserving enough benign samples simultaneously, i.e.i.e., there is a trade-off between BA and ASR.

Appendix M DBD with Different Self-supervised Methods

In this paper, we believe that the desired feature extractor is mapping visually similar inputs to similar positions in the feature space, such that poisoned samples will be separated into their source classes. This goal is compatible with that of self-supervised learning. We believe that any self-supervised learning can be adopted in our method. To further verify this point, we replace the adopted SimCLR with other self-supervised methods in our DBD and examine their performance.

Settings. We replace the SimCLR with two other self-supervised methods, including MoCo-V2 (Chen et al., 2020b) and BYOL (Grill et al., 2020), in our DBD. Except for the adopted self-supervised method, other settings are the same as those used in Section 5.2.

Results. As shown in Table 25, all DBD variants have similar performances. In other words, our DBD is not sensitive to the selection of self-supervised methods.

Appendix N DBD with Different Label-noise Learning Methods

In the main manuscript, we adopt SCE as the label-noise learning method in our second stage. In this section, we explore whether our DBD is still effective if other label-noise methods are adopted.

Settings. We replace SCE in our DBD with two other label-noise learning methods, including generalized cross entropy (GCE) (Zhang & Sabuncu, 2018) and active passive loss (APL) (Ma et al., 2020). Specifically, we adopt the combination of NCE+RCE in APL and use the default hyper-parameters suggested in their original paper. Except for the adopted label-noise learning method, other settings are the same as those used in Section 5.2.

Results. As shown in Table 26, all DBD variants are effective in reducing backdoor threats (i.e.i.e., low ASR) while maintaining high benign accuracy. In other words, our DBD is not sensitive to the selection of label-noise learning methods.

Appendix O Analyzing Why our DBD is Effective in Defending against Label-Consistent Attack

In general, the good defense performance of our DBD method against the label-consistent attack (which is one of the clean-label attacks) can be explained from the following aspects:

Firstly, as shown in Figure 1, there is a common observation across different attacks (including both poisoned- and clean-label attacks) that poisoned samples tend to gather together in the feature space learned by the standard supervised learning. The most intuitive idea of our DBD is to prevent such a gathering in the learned feature space, which is implemented by self-supervised learning. As shown in Figure 1, the poisoned samples of label-consistent attack are also separated into different areas in the feature space learned by self-supervised learning. This example gives an intuitive explanation about why our DBD can successfully defend against the label-consistent attack.

Furthermore, it is interesting to explore why the poisoned samples in the label-consistent attack are separated under self-supervised learning since all poisoned samples are from the same target class, rather than from different source classes in poisoned-label attacks. For each poisoned sample in this attack, there are two types of features: the trigger and the benign feature with (untargeted) adversarial perturbations. From the perspective of DNNs, benign samples with (untargeted) adversarial perturbations are similar to samples from different source classes, though these samples look similar from the human’s perspective. Thus, it is not surprising that poisoned samples in clean-label attacks can also be separated under self-supervised learning, just like those in poisoned-label attacks.