Towards A Proactive ML Approach for Detecting Backdoor Poison Samples
Xiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu, Saeed Mahloujifar, Prateek Mittal
Introduction
Deep learning relies on large datasets . Yet, the creation of these datasets often involves automation and outsourcing, making it difficult to ensure strict supervision and leaving them vulnerable to backdoor poisoning attacks . In these attacks, a typical adversary will manipulate a few training samples by planting a backdoor trigger (e.g., a pixel patch) and (mis)labeling them as a target class. These manipulated samples are referred to as backdoor poison samples, and the compromised dataset is called a poisoned dataset. This manipulation creates a backdoor correlation between the trigger and target class, causing models trained on the poisoned dataset to learn this backdoor while still maintaining normal behaviors under standard evaluation metrics. Backdoor poisoning attacks pose a significant risk as they allow adversaries to stealthily control models. To combat this, we investigate methods for detecting backdoor poison samples in potentially poisoned datasets as an additional safeguard before they are used by downstream applications.
Limitations of Prior Works.The discussion is strictly within the context of detecting backdoor poison samples; our analysis of limitations is not directly applicable to other tracks of defenses that focus on alternative defensive goals (refer Sec § 2.3). Prior works on the detection of backdoor poison samples have primarily employed a post-hoc workflow . In this workflow, defenders passively allow the attack to first proceed without intervention, by training a model on the given poisoned dataset using a routine procedure (without any defense), resulting in the model being backdoored. By analyzing the model’s behaviors, defenders then try to trace the attack back to the poison samples in the dataset. The rationale behind these post-hoc approaches is a passive assumption that post-attacked models (i.e., backdoored models) will spontaneously exhibit distinctive behaviors (e.g., latent separation characteristics ) on poison and clean samples, which can be utilized to distinguish between the two populations . However, in practice, these post-hoc approaches are prone to failure or performance degradation in many scenarios, because the characteristics they rely on could often be inherently weak (e.g., when the poison rate is low ) or be deliberately suppressed by adaptive attacks . In this work, we point out that these failure modes should be attributed to the underlying post-workflow which passively counts on post-hoc characteristics that are not within the control of defenders.
Our Proposal: proactively enforce and magnify distinctive characteristics of the post-attacked model to facilitate poison detection. To alleviate the limitations of the post-hoc workflow, this work advocates for a proactive mindset as an alternative. Instead of passively allowing the attack to proceed as per the adversary’s expectation and hoping for distinctive behaviors to emerge in the post-attacked model, we propose that defenders should engage proactively with the entire model training and detection pipeline. This involves directly enforcing and magnifying distinctive characteristics of the post-attacked model to facilitate poison detection. The key insight is that defenders have a "home field" advantage against backdoor poisoning attacks (that has not been fully exploited by post-hoc approaches), as they still maintain full control over the system space after the dataset has been compromised. This control allows defenders to strategically intervene in the attack process (e.g., post-process data samples, manipulate training procedures) in a way that can proactively enforce post-attacked models to exhibit discriminative characteristics (as per the defenders’ expectations), which can be more reliably used to separate clean and poison samples. We formulate this proposal within an abstract framework (Sec § 3.3) and provide practical insights on designing proactive poison detectors that are more robust and generalizable.
Confusion Training (CT): a concrete instantiation of our proposal. To showcase how our proposal may empower the advancement of future research on backdoor poison samples detection, we introduce the technique of Confusion Training (CT) as a concrete instantiation of the proposal. As illustrated in Fig 1, CT produces an inference model via launching a training procedure jointly on regular batches from the poisoned dataset and confusion batches consisting of randomly labeled clean samples from a reserved clean set (much smaller than the poisoned dataset). The random labeling decouples the benign correlations between normal semantic features and semantic labels in the confusion batch. During training, a confusion batch is also assigned a large weight while a regular batch is only assigned a small weight. In this way, the confusion batch serves as another set of poison samples that strongly obscure the benign correlations, making the clean data points hard to be fitted. On the other hand, since the confusion batch does not contain backdoor triggers, correlations between the trigger and target class still remain intact. Therefore, the resulting inference model fails to fit most clean samples but still correctly fits most backdoor poison samples. This allows defenders to distinguish between clean and poison samples based on the inference model’s state of fitting.
Emperical and Theoretical Analysis. We extensively evaluate a diverse set of 14 backdoor poisoning attacks and 4 benchmark datasets, covering domains of both image (CIFAR10 , GTSRB , ImageNet ) and malware (Ember ) classifications. By comparing with 14 baseline defenses along with ablation studies, we show the effectiveness and superiority of the Confusion Training (CT) defense, illustrating the potential of the proactive mindset that we promote in this work. We also provide a theoretical analysis in a simplified setting to formally illustrate the working principles of CT in Appendix § E.
Finally, we summarize our contributions as follows:
We uncover a post-hoc workflow that is prevalent in previous research on detecting backdoor poison samples, and reveal the limitations it imposes on prior arts.
We propose a proactive mindset as an alternative and formulate it within an abstract framework, providing grounded insights on designing proactive poison detection pipelines that are more robust and generalizable.
We introduce Confusion Training (CT) as a concrete illustration of our proposal, whose effectiveness is supported by both empirical and theoretical results. We position CT as evidence to showcase how our proposal of the proactive mindset may inspire future advancement in the detection of backdoor poison samples.
Preliminaries
In this section, we define our setup (Sec § 2.1), formulate the threat model of backdoor poisoning attacks (Sec § 2.2), and clarify the goals and capabilities of our defense (Sec § 2.3).
2 Threat Model
In this paper, we follow the standard threat model of backdoor poisoning attacks .
Adversary’s Capabilities. Similar to most existing backdoor poisoning attacks, the adversary: 1) can manipulate a limited portion (no more than a half) of the training dataset, 2) has access to the victim’s model architecture, training algorithm, and potential backdoor mitigation techniques. Besides, the adversary has no control over either the training process or the environment where models are deployed.
Attack Goals. The adversary aims to construct a poisoned dataset such that the trained model will be backdoored. In general, a desired backdoor poisoning attack satisfies the following conditions (using notations from Sec § 2.1):
Eqn 2 requires the backdoored model still maintains a level of clean accuracy (ACC) that is comparable to that of a benign model , in order to ensure the stealthiness of the attack. Eqn 3, on the other hand, asserts that the backdoored model should have a high ASR ( some threshold ).
3 Detecting Backdoor Poison Samples
We investigate methods for detecting backdoor poison samples as a means of defense against backdoor poisoning attacks. We focus on offline detection that aims to identify potential backdoor poison samples within a given dataset.
Defender’s Capabilities. 1) Autonomy over the poisoned dataset: After acquiring the poisoned dataset, the defender possesses complete autonomy and freedom to manipulate it. This includes the ability to access, scrutinize, and process the dataset, as well as the liberty to independently train models on it; 2) Access to a small reserved clean set: similar to many prior backdoor defenses , our defender has access to a small clean set, e.g., as few as 250 to 2000 samples for CIFAR10 (Figure 5). This small set, while insufficient for training a high-accuracy model, is intended to bootstrap defenses. In practice, this small clean dataset can be in-house data directly generated by the defenders themselves (e.g., a defender could gather photos using their own camera and label them) or data collected from trustworthy sources. Alternatively, Zeng et al. show the feasibility of sifting out a small clean set directly from a poisoned dataset to bootstrap subsequent defenses.
Poison Detection for Backdoor Defenses. Detecting poison samples provides a flexible foundation for countering backdoor attacks. First, if one can accurately identify and eliminate backdoor poison samples from a poisoned dataset, the threat of backdoor poisoning attacks can be mitigated at the outset. The cleansed dataset can be flexibly utilized to train any clean model, irrespective of its algorithm or architecture. This flexibility presents advantages over many defenses (e.g., robust training ) that are tied to rigid training algorithms and optimized for specific model architectures. For instance, one has the flexibility to use a lightweight model (e.g., ResNet18) to cleanse a dataset and then use any model architecture (e.g., ViT ) and training algorithms to train an advanced clean model on the cleansed dataset. Second, poison detection can act as a foundational building block in other backdoor defenses. For example, some defenses also rely on isolating (only) a portion of poison samples and strategically employing them to mitigate backdoor learning during training. Thus, even if the poison detection is imperfect (For instance, TPR is not sufficiently high), it can still be integrated with other techniques to construct effective backdoor defenses.
Methodological Analysis
To tackle backdoor poison samples detectionOur analysis strictly pertains to the context of detecting backdoor poison samples in poisoned datasets, which is formally indicated in Definitions 3 and 4. Thus, the analysis is not meant to be directly applied to other tracks of defenses (that do not intend to accurately separate poison and clean samples) within the broad backdoor defenses literature . , defenders rely on some distinctive characteristics of poison samples, as outlined in Sec § 3.1. We analyze this issue methodologically and critique the post-hoc workflow of state-of-the-art approaches in Sec § 3.2, highlighting its limitations. In Sec § 3.3, we suggest a proactive mindset as an alternative, based on which we formulate a unified framework and offer practical insights on designing proactive poison detection pipelines.
A backdoor characteristic function is defined as , which maps data point to characteristic vectors , where is a characteristic space that facilitates discrimination between clean and backdoor poison samples, and is the parametrization (if applicable).
A backdoor poison samples detector is the composition of a backdoor characteristic function and a decision function . Given a data entry , the decision function takes (, y) as its input and predicts whether is a poison sample — indicates a positive prediction and otherwise. We use and to denote the True Positive Rate and False Positive Rate of the poison detector.
2 The Post-hoc Workflow and Its Limitations
A post-hoc poison detection defender solves the following problem for some (presumed) characteristics of a post-attacked model:
where the defender aims to optimize the TPR within an FPR budget , by searching for a decision function w.r.t certain characteristic of the backdoored model .
Failure Modes of Post-hoc Approaches: A Case Study. The most successful examples of the post-hoc workflow are latent separation based poison detectors . These methods operate on the premise that backdoored models will learn separate latent representations for clean and backdoor poison samples. Consequently, the latent representation, denoted as , is perceived as the backdoor characteristic function for poison detection. The next step involves conducting a clustering analysis in the latent representation space to derive a decision function that separates clean and backdoor poison samples. These approaches constitute a state-of-the-art frontier of backdoor defenses, with methods such as Spectral Signature and Activation Clustering being considered as canonical baselines within the literature, while SCAn and SPECTRE have reported nearly perfect results against a diverse set of baseline attacks. These works heuristically optimize the decision function (as described in Definition 3), primarily by refining the clustering analysis. However, they are prone to failure or performance degradation in many scenarios where the assumed characteristics are inherently weak or even deliberately suppressed. For instance, latent separation characteristics can be less discernible when the poison rate is low (Fig 2(a)). Recently, Tang et al. and Qi et al. also propose adaptive backdoor poisoning attacks that can even intentionally suppress the latent separation (Fig 2(b)). Furthermore, these characteristics can also vary across different datasets (Fig 2(c)).
Besides the latent separation based approaches, Gao et al. assume that backdoor models’ predictions on poison samples have less entropy under intentional perturbation, and Chou et al. count on abnormal regions in the backdoored model’s saliency map. However, they also suffer from similar failure modes. For example, it is well understood that characteristics assumed by the two works do not hold true for many backdoor attacks with non-local triggers (e.g., ).
Methodological Limitations. We posit that this post-hoc workflow is fundamentally limited in that it does not fully exploit defenders’ capabilities. Particularly, this workflow only passively builds detection pipelines based on some post-hoc characteristics that are not within the control of defenders. However, these (presumed) characteristics could often be inherently weak or even deliberately suppressed by attackers.
3 Towards A Proactive Mindset
The "Home Field" Advantage. In this study, we highlight the "home field" advantage held by defenders in the face of backdoor poisoning attacks. This advantage stems from the fact that once attackers have poisoned a training dataset and then subsequently handed it over to defenders, the latter will have full control and autonomy over both the poisoned dataset and their own operational environment. Conversely, the attackers will not be able to interfere with any subsequent defense measures taken by the defenders.
We point out that the post-hoc workflow (Definition 3) does not fully exploit this "home field" advantage. It passively allows the attack to first proceed without exerting any intervention, even though it has the capability to do so. As a result, the post-hoc characteristics of the attacked model are completely out of the control of defenders. As we have reviewed in Sec § 3.2, these post-hoc characteristics could inherently fail to emerge (so we need to proactively enforce them) and might also be deliberately suppressed by attackers when they design the attack (so we should not allow the attack to proceed as per the adversary’s expectation).
A Proactive Mindset: proactively enforce and magnify distinctive characteristics of the post-attacked model to facilitate poison detection. To alleviate the limitations of the post-hoc workflow, we suggest a paradigm shift by promoting a proactive mindset. We encourage defenders to fully exploit their "home field" advantage by engaging proactively with the entire model training and detection pipeline. Specifically, we highlight that defenders have the potential to strategically intervene in the attack process (e.g., post-process data samples, manipulate training procedures) in a way that can proactively enforce post-attacked models to exhibit discriminative characteristics (as per the defenders’ expectations), which can be more reliably used to separate clean and poison samples. Formally, this paradigm shift can be concisely expressed by a key adaptation on the prior framework in Defintion 3, and we present this new formulation in Definition 4 as follows.
A proactive poison detection defender selects a backdoor characteristic function , and proactively enforces and magnifies this characteristic in a post-attacked model by solving the following problem:
where is proactively designed by the defender to magnify the intended characteristic to facilitate poison detection.
where is the ideally optimal TPR in Definition 3, while is the counterpart in Definition 4.
Practical Insights. For a simple methodological illustration, in Definition 4, we intentionally adopt a very inclusive variable to express the proactiveness of our proposal. In practice, can encapsulate a number of design choices. A major design choice that we focus on in this work it the training algorithm that generates the post-attacked model. The key insight is that the goal of is never to find a model that has high accuracy, but a model that exhibits distinguishable characteristics on poison and clean samples. Thus, an empirical risk minimization procedure that aims to best fit the poisoned dataset is not necessarily an optimal take. Instead, we suggest defenders proactively search for specialized training algorithms that can generate better to facilitate poison detection. The design of confusion training (see Sec § 4 for technical details) in this work is an illustration.
Additionally, Definition 4 also sheds light on other dimensions, including data post-processing. The utilization of conventional data augmentation techniques, while common, can be deemed as a simplistic take on this dimension, and it has been shown to amplify latent separation between poison and clean samples . By explicitly positioning the data post-processing as a design choice, we posit that there might exist more delicate post-processing procedures, which can be utilized to better facilitate poison detection. We encourage future research to investigate what could be an optimal take on this dimension. Besides, the choice of model architecture can also make a difference. For example, against Adap-Blend Attack on CIFAR10, the standard version of ResNet18 exhibits much stronger latent separation than that of a tailored version of ResNet20 (see Fig 5). Later, in Sec 5.2.3, we also show that model architecture choice can play an important role in the implementation of confusion training.
Confusion Training
To showcase how the proactive mindset that we promote in Sec § 3.3 may empower the advancement of future research on backdoor poison sample detection, we introduce the technique of Confusion Training (CT) as a concrete instantiation of the proposal. This section presents the design of CT. We start from a high-level overview in Sec § 4.1, followed by technical details in Sec § 4.2. To formally illustrate the working principles of CT, we provide a theoretical analysis of CT in Appendix § E.
We sketch the confusion training pipeline in Algorithm 1, and present a high-level overview of the approach as follows.
Initialization. We initialize the model with in Line 1, which is the backdoored model routinely trained on the poisoned dataset. This initialization provides a prior of the backdoor and empirically makes confusion training more stable.
2 Technical Details of The Design
Now, we delve into the engineering techniques that are employed to convert Algorithm 1 into a practical implementation.
Creating The Confusion Batch. In practice, we have three technical considerations in the creation of the confusion batch.
2) Random Mislabeling. For samples in the confusion batch, we rule out their ground truth labels during random labeling. This can make the model’s fitting on clean samples even worse than a random classifier, and thus helps to reduce the false positive rate (FPR) of poison detection below .
Identify The Target Class. Many state-of-the-art poison detection approaches apply clustering analysis to identify potential target classes and only perform poison detection on training samples that are labeled to these classes. This technique helps to reduce the FPR of the detection because it avoids false positives on obviously innocent classes. In our implementation, we also follow this practice by using a Gaussian Mixture Model similar to Tang et al. .
We refer readers to Algorithm 2-3 in Appendix § A.1 for a formal description of these techniques as well as our implementation details.
Empirical Evaluation of Confusion Training
In this section, we present the empirical evaluation of the confusion training (CT) defense. We discuss the evaluation setup in Sec § 5.1, followed by Sec § 5.2, presenting our major results and ablation studies on CIFAR10 and GTSRB. Sec § 5.3 demonstrates the scalability and generalizability of CT on two large datasets for image classification and malware detection, namely ImageNet and Ember. We discuss additional adaptive attacks against CT in Sec § 5.4.
Datasets. Our primary results focus on two commonly studied image classification datasets, CIFAR10 and GTSRB . These datasets are the primary focus of prior works on backdoor attacks and defenses, and their use allows us to perform a thorough comparison with state-of-the-art approaches. We further consider the 1000 classes ImageNet dataset and the Ember dataset for malware classification in Sec § 5.3.
Models. For a consistent evaluation, we use ResNet18 as the default model architecture for our primary experiments. Later in Sec 5.2.3, we present additional ablation results on other architectures (VGG16 , MobileNetV2 , DenseNet121 ).
Metrics. We report our evaluation results with two sets of metrics: 1) for poison detection defenses, we report the true positive rate (TPR) and false positive rate (FPR) of the detection (Table 1); 2) for end-to-end backdoor defenses, we report the clean accuracy (ACC) and attack success rate (ASR) of the final model protected by the defense. Moreover, to compare all the different types of defenses, we also evaluate the ACC and ASR for poison detection defenses. As noted earlier in Sec § 2.3, the end-to-end performance indicators (ACC and ASR) of poison detection defenses depend on how subsequent models are trained (e.g., algorithms/architectures) on the cleansed dataset. For a fair and straightforward comparison, these two metrics are measured by directly retraining the same model architecture (ResNet18 by default) from scratch on the dataset cleansed by the corresponding detection approach. For rigor of the evaluation, all major experiments are independently repeated three times, and we report the results in the format of "mean (standard deviation)".
Baseline Defenses. We compare confusion training with 14 backdoor defenses in the literature: 1) To illustrate the superiority of our proactive mindset, we include the 6 post-hoc poison detection approaches (SentiNet , STRIP , SS , AC , SCAn and SPECTRE ) that we analyze in Sec § 3.2. 2) For comprehensiveness, we also include the frequency based detection method (Frequency ) that directly scans input samples without utilizing post-attacked models’ characteristics, presenting it as an outlier that does not fit into the paradigm we discuss. 3) We also include 7 other representative backdoor defenses that are not built on poison detection. Specifically, FP , NC , ABL and NAD are covered in Table 2, while DBD , MOTH and NONE are deferred to Appendix § B. Hyperparameters of these baseline defenses are optimized following their original papers and open-source implementations.
2 Effectivness Across Settings: A Benchmark Evaluation on CIFAR10 and GTSRB
We report our major results on CIFAR10 and GTSRB, two standard benchmark datasets in backdoor research. Specifically, we present an overview of the results in 5.2.1, followed by a thorough analysis in 5.2.2 and ablation studies in 5.2.3.
We present our main results in Table 1 (for poison detection based defenses) and Table 2 (for all defenses). We consider a poison detector is successful against an attack if the TPR is more than ; otherwise, we say the detector is unsuccessful. We consider a defense is successful against an attack only if the ASR is reduced below , otherwise unsuccessful.
TPR and FPR. According to Table 1, as a backdoor poison samples detector, CT is successful in detecting poison samples across all attacks and datasets we evaluate. In all settings, CT consistently detects over poison samples (oftentimes detects ). In terms of false positives, CT also achieves a low FPR comparable to the best of other detectors.
ASR and ACC. According to Table 2, as a general backdoor defense, CT also succeeds in reducing the ASR in all settings. As shown, CT consistently reduces the ASR to less than . On the other hand, the ACC drop induced by CT defense is moderate – the ACC is always higher than 92.4% on CIFAR10 and 96.0% on GTSRB, in parallel with the best of other defenses. Note that, for a fair comparison, we use the same ResNet18 architecture to retrain models on the cleansed dataset to report ASR and ACC. In practice, users can freely train models with more advanced architecture and larger capacity on the cleansed dataset to achieve higher accuracy.
2.2 Strengths of Confusion Training over Prior Arts
Prior arts are less effective against latent space adaptive attacks. The three latent space adaptive attacks (TaCT , Adap-Patch and Adap-Blend ) constitute the most challenging cases in our evaluation. In these adaptive attacks, not all trigger-planted samples are labeled to the target class in the poisoned dataset. Thus, the backdoor correlations between the trigger and target class are intentionally complicated, the trigger signal becomes less dominant in the prediction of post-attacked models, and the latent separation between poison and clean samples is also suppressed . As shown in Table 2, all of the 11 baseline defenses fail to defend against at least one of the adaptive attacks on at least one dataset. First, the suppression of latent separation characteristic makes the four post-hoc poison detection methods (SS , AC , SCAn , SPECTRE ) built on this characteristic less effective, as seen in Table 1. Second, complicated backdoor correlations make it harder to fit the backdoor poison samples, reducing the effectiveness of ABL , which assumes poison samples will be fitted faster than clean samples. Third, as the backdoor trigger becomes less dominant, approaches like SentiNet , Strip , and NC that rely on the dominance of backdoor triggers are also less effective. Additionally, Frequency suffers from degradation against Adap-Blend and Adap-Patch because these attacks use more implicit triggers. NAD also performs worse against Adap-Blend on GTSRB, which could be due to the weakened distinguishability between clean and poison populations in latent space, on which NAD performs knowledge distillation. FP in general performs poorly due to its ineffectiveness in our evaluation settings, where poison rates of all attacks are low.
In contrast, CT demonstrates stronger robustness against such adaptive attacks. As shown in Table 1, CT still consistently achieves over TPR against all of the three latent space adaptive attacks on both datasets. This suffices to defend against these attacks, as ASRs on all the models trained on the cleansed datasets are reduced to less than .
This feat can be attributed to the proactive mindset underlying the design of CT, which steers clear of the reliance on those post-hoc characteristics used by prior arts. Specifically, CT makes no assumption about the type of triggers or labeling of backdoor poison samples. It neither passively waits for backdoored models to "magically" exhibit some characteristics that will expose backdoor poison samples. Instead, CT starts from an arguably more fundamental premise that the successful execution of a backdoor poisoning attack must be accompanied by the existence of backdoor correlations and corresponding backdoor poison samples (deviating from the clean distribution) that can be fitted. Overall, the CT design aims to fit backdoor poison samples (that we don’t have knowledge of) while avoiding fitting samples from the clean distribution (that we can approximate) by proactively making them.
This design also makes CT intrinsically more robust on other dimensions. Based on our evaluation results in Table 1,2, we summarize them as follows.
1) CT is effective against clean-label attacks (CL , SIG ). Clean label attacks are designed to evade human inspection by avoiding mislabeling any poison samples. They do not impose any challenges on CT as this strategy can not stop CT from fitting the backdoor correlations.
2) CT is effective against sample-specific triggers (Dynamic , WaNet , ISSBA ). Sample-specific attacks use different triggers for different backdoor samples. These backdoor attacks with such diversed triggers lead to the reduced effectiveness of many defenses, including SentiNet, STRIP, Frequency, SPECTRE, NC, ABL, and NAD. Meanwhile, CT still defends against these attacks effectively.
3) CT is effective against implicit triggers (Blend , SIG , WaNet , ISSBA ). There are some attacks that use more implicit triggers for stealthiness. Unlike other attacks that patch a significant trigger over images, these defenses manage to reduce the strength of the trigger signals, making them less noticeable. According to our results, no other defenses in our evaluation can steadily succeed against all these attacks on both datasets. However, CT is consistently successful in inhibiting the ASR to .
4) CT is effective across datasets. Another significant issue with the baseline defenses is that they do not generalize equally across datasets. A notable example is SPECTRE, a post-hoc approach based on latent separation, which performs well on CIFAR10 against all attacks but fails in multiple cases on GTSRB. This is a direct result of the post-hoc workflow reviewed in Sec § 3.2, as we previously demonstrated in Fig 2(c) that the latent separation characteristics of the post-attacked model could vary across datasets if they are not proactively controlled. As a result, post-hoc approaches are prone to failure. In contrast, the proactive nature of CT makes itself more stable across datasets. The effectiveness of CT on additional datasets will be discussed in Sec § 5.3.
2.3 Ablation Studies
Next, we present additional ablation studies on three factors: 1) poison rates of attacks, 2) model architectures used for CT, and 3) size of the reserved clean set for CT. We investigate the impacts of these factors on the effectiveness of CT. For the ablation study on each factor, we only vary this single factor while keeping all other factors consistent with the setup of the main evaluation in Sec § 5.1. All ablation studies are performed on CIFAR10.
Poison Rates. Fig 3 presents our ablation study on the poison rate of different attacks (we only demonstrate the three most challenging attacks here; refer to Appendix § C for more results). Specifically, we consider the set of poison rates which covers both the high and low poison rate regions. For each setting, we compare CT with the two strongest baseline detectors SPECTRE and SCAn. CT is consistently effective (TPR100%) and outperforms the baselines in both low and high poison rates regions against most attacks (see Fig 6 in Appendix). Even for three attacks in Fig 3 on which most other defenses fail, CT still achieves TPR100% at most time, beyond the reach of the two baselines. Though the TPR of CT slightly drops in the ultra-low poison rate of against Adap-Patch and Adap-Blend, it is still noticeably better than the baselines.
Model Architectures. In addition to ResNet18 , which is used in the main evaluation, we also consider three other model architectures, VGG16 , MobileNetV2 , and DenseNet121 . For a fair comparison, the hyperparameters are optimized for each architecture. The ablation results on the model architectures are shown in Fig 4. The key takeaway is that some architectures are more suitable for implementing CT than others. As shown, CT with ResNet18 and MobileNetV2 are stable across different attacks, achieving high TPR (close to in most cases) while maintaining an FPR below . In contrast, VGG16 and DenseNet121 are less stable, with large variations in performance and frequent failures. According to our empirical experiences, this is because the dynamics of confusion training require easy-to-train architectures. VGG16 does not use skip links, while DenseNet121 is too deep, and thus they are more difficult to train and less stable for CT. This observation is consistent with our practical insights in Sec § 3.3 and sheds light on a future research direction where specialized model architectures can be designed to facilitate poison detection.
Size of The Reserved Clean Set. CT requires a small reserved clean set to launch. Our default implementation in the main evaluation uses a clean set of 2000 samples. We now investigate how the effectiveness of CT varies when the size of the reserved clean set becomes smaller. Specifically, we experiment with fewer clean samples of respectively and present the results (TPR and FPR) of the position detection in Fig 5. Overall, the TPR of our CT detection pipeline maintains a relatively robust performance, though the TPR would become worse against the two latent space adaptive attacks (Adap-Patch and Adap-Blend). However, a noticeable trade-off in the FPR is evident. Remarkably, when the quantity of reserved clean samples is reduced to 250, the FPR can, at times, exceed , leading to a substantial loss of training data. Intuitively, this is due to the fact that a smaller clean set approximates the clean distribution worse. Thus the confusion batch is also doing worse in decoupling the benign correlations — many clean samples could still be fitted, leading to higher FPR. The impact of the high FPR on overall defense performance (ACC and ASR) is also evaluated, focusing on the extreme case where only 250 clean samples are used to bootstrap CT for CIFAR10. Even with over of the CIFAR10 training set discarded, CT still achieved an impressive accuracy exceeding in all cases, thanks to the maintained high TPR, which ensured a low Attack Success Rate (ASR). This suggests that for larger datasets, a high FPR might be tolerable during dataset cleansing, as the remaining data can still be used to train a high-quality model. A closely relevant research topic to this observation is Dataset Pruning , which shows that many training data can be removed from the dataset while models trained on the remaining dataset can still be accurate.
3 Scalability and Generalizability
Generalization to Domain Other Than Computer Vision. CT only assumes backdoor poison samples could be fitted after the benign correlations have been decoupled, without constraining the modality it works on. This suggests the potential generalizability of CT to different domains other than computer vision. To illustrate this, we evaluate CT on the Ember training dataset, a malware classification dataset that consists of 600K feature vectors extracted from Portable Executable (PE) files. Specifically, we consider the attacks proposed by Severi et al. . We use the same EmberNN architecture used in Severi et al. to build models and also follow the same configurations. We consider both the constrained and unconstrained attacks of Severi et al. with poison arte. The constrained attack confines backdoor triggers to those features that are editable in PE files, while the unconstrained attack removes this constraint. Table 4 presents our results, validating the effectiveness of CT.
Scale to Realistic Settings. To study the scalability to more realistic settings we will encounter in the real world, we also evaluate CT on the ImageNet-1k training set, consisting of 1.3M images (26x larger than CIFAR10) with high resolutions. We adapt BadNet , Trojan , and Blend with a poison rate to implement attacks on ImageNet. We use ResNet18 for both confusion training on the poisoned dataset and subsequent retraining on the cleansed dataset. Table 5 presents the results of CT defense against the attacks. As shown, CT still remains effective on ImageNet — with nearly perfect and negligible , models trained on cleansed sets reduce to 0 without suffering any drop of ACC. An intriguing observation is that the FPR here is even lower than we get for CIFAR10 and GTSRB. We hypothesize that this is because when the size of the dataset is larger and the clean task is more complicated, it also becomes harder for the inference model to fit (memorize) clean samples under the disruption of confusion training, making it even easier to separate clean and poison samples under CT.
Additionally, as previously noted in Sec § 2.3, after we use ResNet18 to cleanse the dataset, we have the flexibility to use any model architecture and algorithm for subsequent training on the cleansed dataset. We confirm this by replacing ResNet18 (11 M parameters) with ViT-B/16 (150 M parameters) and applying the state-of-the-art training algorithm GSAM during the retraining phase for cases we evaluate in Table 5. As shown in Table 6, we consistently get clean models with 0 ASR, but the ACC increases to rather than in Table 5. Symmetrically, this also indicates that, in realistic scenarios involving training large models, it is possible to initially deploy CT with a lightweight inference model to sanitize the dataset. The cost would be acceptable, as the computation demanded to train a lightweight model for dataset cleansing can be negligible compared with subsequent training of large models on the cleansed dataset.
4 Additional Analysis on Adaptive Attacks
Enforce dependency between benign and backdoor correlations. CT makes very few assumptions about the configurations of backdoor poisoning attacks, but it does assume that the backdoor correlations can still be fitted after the benign correlations have been decoupled. Then, a principled class of adaptive attacks against CT can be those attacks where the backdoor correlations depend on benign correlations, making it more difficult to fit backdoor poison samples after the benign correlations have been decoupled. Some attacks that we evaluate in Table 1,2 already fall in this class, including 1) sample-specific attacks where triggers depend on clean semantics of the samples and 2) latent space adaptive attacks where trigger planted samples are conditionally labeled to the target class based on the semantic class (e.g. TaCT ). Results in Table 1,2, and Fig 6 illustrate that CT still exhibits good robustness. Additionally, we add an evaluation on an all-to-all attack with poison rate on CIFAR10, where trigger-planted samples from class should be misclassified to class . This clearly imposes a dependency between the backdoor and clean tasks. Against this attack, CT still achieves TPR and FPR, and retraining a ResNet18 on the cleansed dataset results in ASR and ACC. We hypothesize that this is because backdoor poison samples are always out of the distribution of clean data and thus can still be fitted though the fitting on clean distribution is disrupted.
Launch attack without artifacts. Most existing attacks unavoidably introduce artifacts into poison samples as they inject trigger patterns that significantly deviate from the clean distribution. These artifacts make poison samples easy to be fitted by CT. This explains why CT is consistently effective against different attacks. Thus, another angle to adaptively attack CT is to launch attacks without significant artifacts. One candidate is the rotation attack , which uses image rotation as the backdoor trigger without introducing additional artifacts. We implement this attack following the original paper’s configuration with poison rate. Still, CT can achieve TPR and FPR, and retraining a ResNet18 on the cleansed dataset will have ASR and ACC.
Discussion
In Sec § 5, we illustrate nice properties of confusion training (CT) defense. In particular, we highlight that CT is built on an arguably more fundamental premise that the successful execution of a backdoor poisoning attack must be accompanied by the existence of backdoor correlations and corresponding backdoor poison samples that are fittable. The design of CT makes use of this fundamental premise, by simply attempting to fit those backdoor poison samples (that we don’t have knowledge of, except knowing that they should be fittable), while simultaneously preventing the fitting of samples from the clean distribution (that we can approximate) by proactively making them less fitable. Even though, we keep a conservative attitude so as not to oversell the security of the confusion training technique itself. After all, it is an empirical defense without a certified guarantee, and the insights we discuss in Sec § 5.4 may also motivate stronger adaptive attacks that can further challenge CT. Rather, we position CT as evidence to showcase how our proposal of the proactive mindset may inspire future advancement in backdoor poison samples detection. In particular, we hope our formulation and practical insights provided in Sec § 3.3 can inspire the design of better proactive backdoor defenses.
Related Work
The focus of this study is to address the threat of backdoor poisoning attacks on deep neural networks (DNNs), which constitute the most prevalent form of DNN backdoor attacks. These attacks involve the manipulation of a few samples in the training dataset by an attacker, resulting in the "poisoning" of the dataset. Victims will then train their own model on this poisoned dataset. Backdoor attacks also include various methods that do not fit within the paradigm of data poisoning. Examples include modifying the training process or using backdoored pre-trained models for transfer learning, as well as attacks that occur at the deployment stage . These attacks involve different assumptions and threat models, which are out of the scope of this work.
2 Backdoor Defenses
Backdoor Poison Samples Detection. This work investigates methods for detecting poison samples in the poisoned dataset as a means of defense. In the existing literature, state-of-the-art approaches of this category consistently build their detectors upon latent separation characteristics, assuming that backdoored models will learn separate latent representations for poison and clean samples. Poison samples can then be identified via cluster analysis in the latent representation space. Tran et al. project the latent representations to the top PCA direction and observe a bimodal distribution — poison and clean samples form their own modal, respectively. Later, Chen et al. independently observe similar separation and propose to use K-means to separate the clean and poison clusters in the latent space. More recently, Tang et al. and Hayase et al. propose to use statistics of the clean distribution to further improve the cluster analysis. Researchers have also proposed using other characteristics such as the entropy of predictions under intentional perturbation , abnormal regions in the backdoored model’s saliency map , and the speed of model fiting to identify poison samples. Zeng et al. propose to detect artifacts of poison samples in the frequency domain.
Conclusion
In this paper, we present a comprehensive investigation into methods for detecting backdoor poison samples in an effort to defend against backdoor poisoning attacks. We uncover a post-hoc workflow underlying many prior works and reveal its limitations. Accordingly, we suggest a paradigm shift by promoting a proactive mindset formulated within a unified framework. We also provide practical insights for grounded implementations. Particularly, we introduce the technique of confusion training (CT) as a concrete instantiation of the proactive mindset. Our empirical evaluations show that CT is effective across different attack settings, outperforming prior ars. We also show that CT is scalable to large datasets and generalizable to the domain of malware detection other than computer vision. We then provide insights for adaptively attacking CT and find CT is quite robust against a number of adaptive attacks. We position CT as evidence to showcase how our proposal of the proactive mindset may inspire future advancement in backdoor poison samples detection.
Acknowledgements
This work was supported in part by the National Science Foundation under grants CNS-1553437 and CNS-1704105, the ARL’s Army Artificial Intelligence Innovation Institute (A2I2), the Office of Naval Research Young Investigator Award, the Army Research Office Young Investigator Prize, Schmidt DataX award, Princeton E-ffiliates Award, and Princeton’s Gordon Y. S. Wu Fellowship. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funding agencies.
References
Appendix A Implementations of Confusion Training
A.2 Computational Cost
The primary computational cost of the full CT poison detection pipeline (Algorithm 3) comes from the pretraining procedure in Line 3. Given that rounds of pretraining are conducted over sequentially shrinking datasets , the computational overhead is projected to approximate times that of a full pretraining (specific to the hyper-parameters discussed in Appendix § A.1). Furthermore, the inner iterations (Line 4) resemble several epochs of fine-tuning and consume significantly fewer computational resources than pretraining. In addition, computational costs associated with subsequent steps, spanning from Line 17 to Line 22, are relatively negligible compared to the whole procedure. In summary, the complete CT poison detection framework necessitates roughly twice the computational resources required in standard pretraining and in conventional post-hoc approaches (since the latter also primarily involves pre-training). To illustrate this, we executed our confusion training pipeline and the classical poison detection baseline Spectral Signature on a single GPU in our workstation, which is equipped with 48 Intel Xeon Silver 4214 CPU cores, 384 GB RAM, and 8 GeForce RTX 2080 Ti GPUs. We document the execution time in Table 7. Note that obtaining a ready-to-deploy model necessitates further training a downstream model on the cleansed dataset, of which the computational cost depends on the model architecture and training algorithm (as discussed in Sec § 2.3).
Appendix B Comparison with Additional Defenses
A line of work aims to design robust training algorithms that intend to prevent models from learning any backdoor during training . The conceptual underpinning that propels these methodologies bears considerable relevance to our study, given that they also strive to proactively devise training procedures that are robust enough to counter backdoor attacks. In Table 2, we compared CT with ABL from this category of defenses. For a more comprehensive comparison with these defenses, we evaluate three additional defenses (DBD , MOTH and NONE ) from this category and evaluate them on CIFAR10 in Table 8. Our implementations strictly follow the original configurations of their open-source repositories, and our evaluation follows the same setup used for Table 2. Notably, CT continues exhibiting stronger resilience than the three additional baselines.
Appendix C Full Ablation Results on Poison Rates
We hereby present our full ablation experimental configuration and results.
For most attacks, we consider the set of poison rates which covers both the high and low poison rate regions. For clean label attacks (CL and SIG), since only target class samples are poisoned, the maximum poison rate is set to to avoid the trivial case that the entire target class is poisoned. For ISSBA and WaNet, the minimum poison rates are set to and respectively, as these attacks with lower poison rates can no longer be effective. For each setting, we also compare CT with the two strongest baseline detectors SPECTRE and SCAn.
As shown in Fig 6, CT is consistently effective (TPR100%) and outperforms the baselines in both low and high poison rates regions against most attacks. Though the TPR of CT slightly drops in the ultra-low poison rate of against SIG, Dynamic, Adap-Patch, and Adap-Blend, it is still better than or at least comparable with the baselines. CT also suffers some drops of TPR in high poison rates against Adap-Patch, but still outperforms the baselines in these cases.
Appendix D Configurations of Experiments
We use four public datasets in our evaluations, including image classification (CIFAR10 , GTSRB , and ImageNet ) and malware classification (EMBER ).
CIFAR10. CIFAR10 is a benchmark dataset for image classification, encompassing common objects such as dogs, cats, and airplanes among its 10 classes. It comprises 50,000 training samples and 10,000 test samples, each with a resolution of 3232 pixels.
GTSRB. The German Traffic Sign Recognition Benchmark (GTSRB) is a widely used image classification dataset for traffic sign recognition, aiming at facilitating autonomous driving. The dataset consists of 43 different types of traffic signs, with 39,211 samples in the training set and 12,630 samples in the test set. Prior to conducting experiments, the images were resized to a resolution of 3232 pixels.
ImageNet. The ImageNet dataset is a common benchmark for image classification, consisting of 1.3 million training images and 50,000 validation images across 1000 classes. The images in ImageNet are resized and cropped to a resolution of 224224 pixels before being utilized as inputs for our models.
Ember. Ember is a well-known public resource for malware classification, which includes both malware and goodware samples. The dataset includes 2,351-dimension features extracted from 1.1 million Windows portable executable files. In our experiment, we use labeled training and test sets that comprise 600,000 and 200,000 samples, respectively, with an equal distribution of benign and malicious samples.
For each dataset, we first follow the default training/test set split by Torchvision . Recall that our defender is assumed to have a small reserved clean set at hand (as defined in Sec § 2.3). To implement this setting, for CIFAR10 and GTSRB, we further randomly pick 2000 samples from the test split to simulate the reserved clean set, and leave the rest part of the test split for evaluation. For, ImageNet and Ember, we randomly pick 5000 samples from the test split to simulate the reserved clean set.
D.2 Training Backdoored Models
For all models on image datasets (CIFAR10, GTSRB and ImageNet), we use ResNet18 as the architecture; for the Ember dataset, we use the default EmberNN as the malware detection architecture. SGD with a momentum of 0.9, a weight decay of ( for Ember), and a batch size of 128 (256 for ImageNet and 512 for Ember), is used for optimization. Initially, we set the learning rate to . On CIFAR10, we follow the standard 200 epochs stochastic gradient descent procedure, with an initial learning rate , which is then multiplied by a factor of at the epochs of and . On GTSRB we use 100 epochs of training, with an initial learning rate , which is then multiplied by at the epochs of and . On ImageNet, we train the model for 90 epochs, with an initial learning rate , which is then multiplied by at the epochs of and . On Ember, we train EmberNN for 10 epochs with learning rate .
D.3 Configurations of Baseline Attacks
During implementing these attacks, we follow the protocols and suggested default configurations of their original papers and open-source implementations. In Table 9, we summarize the hyperparameters that are used for each poison strategy. Specifically, Target Class is the class that the backdoor trigger is correlated to, Poison Rate denotes the portion of training samples that are stamped with the trigger and labeled to the target class. In addition, TaCT only chooses samples from a Source Class to construct poison samples, and they will also randomly pick a portion (Cover Rate) of samples from some Cover Classes and also plant triggers to them while keeping them still correctly labeled as their semantic labels. WaNet , Adap-Blend and Adap-Patch also keep a portion (Cover Rate) of samples planted with triggers but still correctly labeled. Besides, for all these attacks we implement, we use the same trigger patterns to those suggested by their original papers.
For constrained and unconstrained attacks on Ember, we follow the strategies proposed in the original paper . The constrained attack confines backdoor triggers to those features that are really editable in PE files, while the unconstrained attack allows modifying all features during poisoning. The poison rate is 1% for both attacks. For the constrainted attack, the trigger watermark size is 17, with attack strategy "LargeAbsSHAP x MinPopulation"; for the unconstrained attack, the trigger watermark size is 32, with attack strategy "Combined Feature Value Selector".
D.4 Configurations of Baseline Defenses
As mentioned in Sec § 5.1, we compare confusion training with 14 backdoor defenses, including 7 backdoor poison samples detectors ), and 7 defenses of other types (.
Some important configurations of these defenses are specified as follows. Refer to our code for more implementation details.
SentiNet is adapted to a poison detector for training sets, set with a 5% FPR. For inputs with patch triggers, we provide SentiNet with the oracle knowledge of patch locations; for other inputs, SentiNet locates the top 15% pixels with the highest GradCAM activation scores.
STRIP is also adapted to a poison detector for training set, set with a 10% FPR.
AC cleanses any clusters with size 35% of the class size.
Frequency predicts samples to be poison or clean with a binary classifier. We directly use their provided pretrained model for that purpose.
SCAn cleanses classes with scores larger than .
FP prunes the neurons of the last layer before the classifier layers of the model. The maximum prune ratio is set to 99% for CIFAR10 and 75% for GTSRB, both with 100 finetuning epochs and maximum allowed accuracy drop threshold of 10%.
NC reverse-engineers the potential trigger for each class. The class where its trigger norm has the maximum anomaly indice >2 (if exists) is the suspicious class, where the corresponding reversed trigger is used to unlearn the backdoor.
ABL isolates suspicious training samples and use them to unlearn the backdoor of the model. Since our poison rates are significantly smaller than those the original paper has estimated, we isolate 0.1% samples (0.5% for GTSRB) by Flooding method (flooding=0.3) in 15 epochs (5 epochs for GTSRB).
NAD first trains a teacher model based on the poison model for 10 epochs, then distills the teacher model to obtain a student model for 20 epochs.
DBD first runs 1000 epochs of contrastive learning to learn a purified feature extractor and then 200 epochs of semi-supervised fine-tuning to get the final model.
MOTH takes 0.1 of batch samples to samp trigger during orthogonalization, a warm ratio of 0.5 during warmup. For consistency, 2000 clean samples are used for hardening.
NONE takes 200 epochs of pertaining and 20 epochs of defense loops. The maximal reset fraction is set to 0.03.
Appendix E Theoretical Analysis of Confusion Training in A Simplified Setting
To formally justify the effectiveness of confusion training from a theoretical perspective, we study a simplified setting where we train an overparameterized linear model for binary classification tasks. Specifically, we study the classic Gaussian class-conditional model as the training data distribution. We use a simplified setting as a full theoretical treatment of the training dynamics of neural networks is widely recognized as a complex endeavor. Furthermore, our simplifying assumptions have been frequently employed in the corpus of deep learning theory literature, such as overparameterized linear models and binary data . The primary intention of our analysis is to offer an understanding and insight into the operational principles underlying the confusion training approach. However, it is crucial to emphasize that our analysis does not extend to providing a provable security guarantee for confusion training within the context of general deep learning models.
The detailed setup and proof sketch are deferred to Appendix § E.2. The proof uses the result from Chatterji and Long such that under sufficiently small learning rate and mild assumptions on the number of samples, the training dynamics of overparameterized linear model are tractable. This allows us to track the forgetting dynamics of clean and backdoored data points during the confusion training stage.
E.2 Proof
In this section, we formally characterize the forgetting dynamics of backdoored and clean data in the confusion training process, in a simplified setting of the overparameterized linear model and binary classification.
If , then .
If , then .
Recall that for confusion training, the model is first trained on the corrupted dataset that contains poisoned data
, and then is fine-tuned on the mix of the corrupted datasets and reserved datasets with random labels
The failure probability .
The data dimension , and .
We can obtain the following lemma about the separability of .
With probability at least , is linearly separable.
We perform gradient descent with a fixed learning rate ,
For sufficiently small , if we train on linearly separable for iterations, Soudry et al. shows that
for . We sometimes omit the time stamp and write for readability.
Similarly, for a backdoored data , we have