Blind Backdoors in Deep Learning Models

Eugene Bagdasaryan, Vitaly Shmatikov

Introduction

A backdoor is a covert functionality in a machine learning model that causes it to produce incorrect outputs on inputs containing a certain “trigger” feature chosen by the attacker. Prior work demonstrated how backdoors can be introduced into a model by an attacker who poisons the training data with specially crafted inputs , or else by an attacker who trains the model in outsourced-training and model-reuse scenarios . These backdoors are weaker versions of UAPs, universal adversarial perturbations . Just like UAPs, a backdoor transformation applied to any input causes the model to misclassify it to an attacker-chosen label, but whereas UAPs work against unmodified models, backdoors require the attacker to both change the model and change the input at inference time.

Our contributions. We investigate a new vector for backdoor attacks: code poisoning. Machine learning pipelines include code from open-source and proprietary repositories, managed via build and integration tools. Code management platforms are known vectors for malicious code injection, enabling attackers to directly modify source and binary code .

Source-code backdoors of the type studied in this paper can be discovered by code inspection and analysis. Today, even popular ML repositories , which have thousands of forks, are accompanied only by rudimentary tests (such as testing the shape of the output). We hope to motivate ML developers to carefully review the functionality added by every commit and design automated tests for the presence of backdoor code.

Code poisoning is a blind attack. When implementing the attack code, the attacker does not have access to the training data on which it will operate. He cannot observe the code during its execution, nor the resulting model, nor any other output of the training process (see Figure 1).

Our prototype attack codeAvailable at https://github.com/ebagdasa/backdoors101. synthesizes poisoning inputs “on the fly” when computing loss values during training. This is not enough, however. A blind attack cannot combine main-task, backdoor, and defense-evasion objectives into a single loss function as in because (a) the scaling coefficients are data- and model-dependent and cannot be precomputed by a code-only attacker, and (b) a fixed combination is suboptimal when the losses represent different tasks.

We view backdoor injection as an instance of multi-task learning for conflicting objectives—namely, training the same model for high accuracy on the main and backdoor tasks simultaneously—and use Multiple Gradient Descent Algorithm with the Franke-Wolfe optimizer to find an optimal, self-balancing loss function that achieves high accuracy on both the main and backdoor tasks.

To illustrate the power of blind attacks, we use them to inject (1) single-pixel and physical backdoors in ImageNet; (2) backdoors that switch the model to an entirely different, privacy-violating functionality, e.g., cause a model that counts the number of faces in a photo to covertly recognize specific individuals; and (3) semantic backdoors that do not require the attacker to modify the input at inference time, e.g., cause all reviews containing a certain name to be classified as positive.

We analyze all previously proposed defenses against backdoors: discovering backdoors by input perturbation , detecting anomalies in model behavior on backdoor inputs , and suppressing the influence of outliers . We show how a blind attacker can evade any of them by incorporating defense evasion into the loss computation.

Finally, we report the performance overhead of our attacks and discuss better defenses, including certified robustness and trusted computational graphs.

Backdoors in Deep Learning Models

2 Backdoors

Prior work focused on universal pixel-pattern backdoors in image classification tasks. These backdoors involve a normal model θ\theta and a backdoored model θ∗\theta^{*} that performs the same task as θ\theta on unmodified inputs, i.e., θ(x)=θ∗(x)=y\theta(x)=\theta^{*}(x)=y. If at inference time a certain pixel pattern is added to the input, then θ∗\theta^{*} assigns a fixed, incorrect label to it, i.e., θ∗(x∗)=y∗\theta^{*}(x^{*})=y^{*}, whereas θ(x∗)=θ(x)=y\theta(x^{*})=\theta(x)=y.

We take a broader view of backdoors as an instance of multi-task learning where the model is simultaneously trained for its original (main) task and a backdoor task injected by the attacker. Triggering the backdoor need not require the adversary to modify the input at inference time, and the backdoor need not be universal, i.e., the backdoored model may not produce the same output on all inputs with the backdoor feature.

We say that a model θ∗\theta^{*} for task mm: X→Y\mathcal{X}\rightarrow\mathcal{Y} is “backdoored” if it supports another, adversarial task m∗m^{*}: X∗→Y∗\mathcal{X}^{*}\rightarrow\mathcal{Y}^{*}:

Main task mm: θ∗(x)=y\theta^{*}(x)=y, ∀(x,y)∈(X∖X∗,Y)\forall(x,y)\in(\mathcal{X}\setminus\mathcal{X}^{*},\mathcal{Y})

Backdoor task m∗m^{*}: θ∗(x∗)=y∗\theta^{*}(x^{*})=y^{*}, ∀(x∗,y∗)∈(X∗,Y∗)\forall(x^{*},y^{*})\in(\mathcal{X}^{*},\mathcal{Y}^{*})

The domain X∗\mathcal{X}^{*} of inputs that trigger the backdoor is defined by the predicate Bd:x→{0,1}Bd:x\rightarrow\{0,1\} such that for all x∗∈X∗, Bd(x∗)=1x^{*}\in\mathcal{X}^{*},\ Bd(x^{*})=1 and for all x∈X∖X∗, Bd(x)=0x\in\mathcal{X}\setminus\mathcal{X}^{*},\ Bd(x)=0. Intuitively, Bd(x∗)Bd(x^{*}) holds if x∗x^{*} contains a backdoor feature or trigger. In the case of pixel-pattern or physical backdoors, this feature is added to xx by a synthesis function μ\mu that generates inputs x∗∈X∗x^{*}\in\mathcal{X}^{*} such that X∗∩X=Ø\mathcal{X}^{*}\cap\mathcal{X}=\texttt{\O}. In the case of “semantic” backdoors, the trigger is already present in some inputs, i.e., x∗∈Xx^{*}\in\mathcal{X}. Figure 2 illustrates the difference.

The accuracy of the backdoored model θ∗\theta^{*} on task mm should be similar to a non-backdoored model θ\theta that was correctly trained on data from X×Y\mathcal{X}\times\mathcal{Y}. In effect, the backdoored model θ∗\theta^{*} should support two tasks, mm and m∗m^{*}, and switch between them when the backdoor feature is present in an input. In contrast to the conventional multi-task learning, where the tasks have different output spaces, θ∗\theta^{*} must use the same output space for both tasks. Therefore, the backdoor labels Y∗\mathcal{Y}^{*} must be a subdomain of Y\mathcal{Y}.

3 Backdoor features (triggers)

Inference-time modification. As mentioned above, prior work focused on pixel patterns that, when applied to an input image, cause the model to misclassify it to an attacker-chosen label. These backdoors have the same effect as “adversarial patches” but in a strictly inferior threat model because the attacker must modify (not just observe) the ML model.

We generalize these backdoors as a transformation μ:X→X∗\mu:\mathcal{X}\rightarrow\mathcal{X}^{*} that can include flipping, pixel swapping, squeezing, coloring, etc. Inputs xx and x∗x^{*} could be visually similar (e.g., if μ\mu modifies a single pixel), but μ\mu must be applied to xx at inference time. This attack exploits the fact that θ\theta accepts inputs not only from the domain X\mathcal{X} of actual images, but also from the domain X∗\mathcal{X}^{*} of modified images produced by μ\mu.

A single model can support multiple backdoors, represented by synthesizers μ1,μ2∈M\mu_{1},\mu_{2}\in\mathcal{M} and corresponding to different backdoor tasks: m1∗:Xμ1→Yμ1m_{1}^{*}:\mathcal{X}^{\mu_{1}}\rightarrow\mathcal{Y}^{\mu_{1}}, m2∗:Xμ2→Yμ2m_{2}^{*}:\mathcal{X}^{\mu_{2}}\rightarrow\mathcal{Y}^{\mu_{2}}. We show that a backdoored model can switch between these tasks depending on the backdoor feature(s) present in an input.

Physical backdoors do not require the attacker to modify the digital input . Instead, they are triggered by certain features of physical scenes, e.g., the presence of certain objects—see Figure 2(a). In contrast to physical adversarial examples , which involve artificially generated objects, we focus on backdoors triggered by real objects.

No inference-time modification. Semantic backdoor features can be present in a digital or physical input without the attacker modifying it at inference time: for example, a certain combination of words in a sentence, or, in images, a rare color of an object such as a car . The domain X∗\mathcal{X}^{*} of inputs with the backdoor feature should be a small subset of X\mathcal{X}. The backdoored model cannot be accurate on both the main and backdoor tasks otherwise, because, by definition, these tasks conflict on X∗\mathcal{X}^{*}.

When training a backdoored model, the attacker may use μ:X→X∗\mu:\mathcal{X}\rightarrow\mathcal{X}^{*} to create new training inputs with the backdoor feature if needed, but μ\mu cannot be applied at inference time because the attacker does not have access to the input.

Data- and model-independent backdoors. As we show in the rest of this paper, μ:X→X∗\mu:\mathcal{X}\rightarrow\mathcal{X}^{*} that defines the backdoor can be independent of the specific training data and model weights. By contrast, prior work on Trojan attacks assumes that the attacker can both observe and modify the model, while data poisoning assumes that the attacker can modify the training data.

4 Backdoor functionality

Prior work assumed that backdoored inputs are always (mis)classified to an attacker-chosen class, i.e., ∣∣Y∗∣∣=1||\mathcal{Y}^{*}||=1. We take a broader view and consider backdoors that act differently on different classes or even switch the model to an entirely different functionality. We formalize this via a synthesizer ν:X,Y→Y∗\nu:\mathcal{X},\mathcal{Y}\rightarrow\mathcal{Y}^{*} that, given an input xx and its correct label yy, defines how the backdoored model classifies xx if xx contains the backdoor feature, i.e., Bd(x)Bd(x). Our definition of the backdoor thus supports injection of an entirely different task m∗:X∗→Y∗m^{*}:\mathcal{X}^{*}\rightarrow\mathcal{Y}^{*} that “coexists” in the model with the main task mm on the same input and output space—see Section 4.3.

5 Previously proposed attack vectors

Figure 1 shows a high-level overview of a typical machine learning pipeline.

Poisoning. The attacker can inject backdoored data X∗\mathcal{X}^{*} (e.g., incorrectly labeled images) into the training dataset . Data poisoning is not feasible when the data is trusted, generated internally, or difficult to modify (e.g., if training images are generated by secure cameras).

Trojaning and model replacement. This threat model assumes an attacker who controls model training and has white-box access to the resulting model, or even directly modifies the model at inference time .

Adversarial examples. Universal adversarial perturbations assume that the attacker has white- or black-box access to an unmodified model. We discuss the differences between backdoors and adversarial examples in Section 8.2.

Blind Code Poisoning

Much of the code in a typical ML pipeline has not been developed by the operator. Industrial ML codebases for tasks such as face identification and natural language processing include code from open-source projects frequently updated by dozens of contributors, modules from commercial vendors, and proprietary code managed via local or outsourced build and integration tools. Recent, high-visibility attacks demonstrated that compromised code is a realistic threat.

In ML pipelines, a code-only attacker is weaker than a model-poisoning or trojaning attacker because he does not observe the training data, nor the training process, not the resulting model. Therefore, we refer to code-only poisoning attacks as blind attacks.

Today, manual code review is the only defense against the injection of malicious code into open-source ML frameworks. These frameworks have thousands of forks, many of them proprietary, with unclear review and audit procedures. Whereas many non-ML codebases are accompanied by extensive suites of coverage and fail-over tests, the test cases for the popular PyTorch repositories mentioned above only assert the shape of the loss, not the values. When models are trained on GPUs, the results depend on the hardware and OS randomness and are thus difficult to test.

Recently proposed techniques aim to “verify” trained models but they are inherently different from traditional unit tests and not intended for users who train locally on trusted data. Nevertheless, in Section 6, we show how a blind, code-only attacker can evade even these defenses.

2 Attacker’s capabilities

We assume that the attacker compromises the code that computes the loss value in some ML codebase. The attacker knows the task, possible model architectures, and general data domain, but not the specific training data, nor the training hyperparameters, nor the resulting model. Figures 3 and 4 illustrate this attack. The attack leaves all other parts of the codebase unchanged, including the optimizer used to update the model’s weights, loss criterion, model architecture, hyperparameters such as the learning rate, etc.

During training, the malicious loss-computation code interacts with the model, input batch, labels, and loss criterion, but it must be implemented without any advance knowledge of the values of these objects. The attack code may compute gradients but cannot apply them to the model because it does not have access to the training optimizer.

3 Backdoors as multi-task learning

Our key technical innovation is to view backdoors through the lens of multi-objective optimization.

In conventional multi-task learning , the model consists of a common shared base θsh\theta^{sh} and separate output layers θk\theta^{k} for every task kk. Each training input xx is assigned multiple labels y1,…yky^{1},\ldots y^{k}, and the model produces kk outputs θk(θsh(x))\theta^{k}(\theta^{sh}(x)).

By contrast, a backdoor attacker aims to train the same model, with a single output layer, for two tasks simultaneously: the main task mm and the backdoor task m∗m^{*}. This is challenging in the blind attack scenario. First, the attacker cannot combine the two learning objectives into a single loss function via a fixed linear combination, as in , because the coefficients are data- and model-dependent and cannot be determined in advance. Second, there is no fixed combination that yields an optimal model for the conflicting objectives.

This computation is blind: backdoor transformations μ\mu and ν\nu are generic functions, independent of the concrete training data or model weights. We use multi-objective optimization to discover the optimal coefficients at runtime—see Section 3.4. To reduce the overhead, the attack can be performed only when the model is close to convergence, as indicated by threshold TT (see Section 4.6).

Backdoors. In universal image-classification backdoors , the trigger feature is a pixel pattern tt and all images with this pattern are classified to the same class cc. To synthesize such a backdoor input during training or at inference time, μ\mu simply overlays the pattern tt over input xx, i.e., μ(x)=x⊕t\mu(x)=x\oplus t. The corresponding label is always cc, i.e., ν(y)=c\nu(y)=c.

Our approach also supports complex backdoors by allowing complex synthesizers ν\nu. During training, ν\nu can assign different labels to different backdoor inputs, enabling input-specific backdoor functionalities and even switching the model to an entirely different task—see Section 4.3.

In semantic backdoors, the backdoor feature already occurs in some unmodified inputs in XX. If the training set does not already contain enough inputs with this feature, μ\mu can synthesize backdoor inputs from normal inputs, e.g., by adding the trigger word or object.

4 Learning for conflicting objectives

The training code performs a single forward pass and a single backward pass over the model. Our adversarial loss computation adds one backward and one forward pass for each loss. Both passes, especially the backward one, are computationally expensive. To reduce the slowdown, the scaling coefficients can be re-used after they are computed by MGDA (see Table 3 in Section 4.5), limiting the overhead to a single forward pass per each loss term. Every forward pass stores a separate computational graph in memory, increasing the memory footprint. In Section 4.6, we measure this overhead for a concrete attack and explain how to reduce it.

Experiments

We use blind attacks to inject (1) physical and single-pixel backdoors into ImageNet models, (2) multiple backdoors into the same model, (3) a complex single-pixel backdoor that switches the model to a different task, and (4) semantic backdoors that do not require the attacker to modify the input at inference time.

Figure 2 summarizes the experiments. For these experiments, we are not concerned with evading defenses and thus use only two loss terms, for the main task mm and the backdoor task m∗m^{*}, respectively (see Section 6 for defense evasion).

We implemented all attacks using PyTorch on two Nvidia TitanX GPUs. Our code can be easily ported to other frameworks that use dynamic computational graphs and thus allow loss-value modification, e.g., TensorFlow 2.0 . For multi-objective optimization inside the attack code, we use the implementation of the Frank-Wolfe optimizer from .

We demonstrate the first backdoor attacks on ImageNet , a popular, large-scale object recognition task, using three types of triggers: pixel pattern, single pixel, and physical object. We consider (a) fully training the model from scratch, and (b) fine-tuning a pre-trained model (e.g., daily model update).

Main task. We use the ImageNet LSVRC dataset that contains 1,281,1671,281,167 images labeled into 1,0001,000 classes. The task is to predict the correct label for each image. We measure the top-1 accuracy of the prediction.

Training details. When training fully, we train the ResNet18 model for 9090 epochs using the SGD optimizer with batch size 256256 and learning rate 0.10.1 divided by 1010 every 3030 epochs. These hyperparameters, taken from the PyTorch examples , yield 65.3%65.3\% accuracy on the main ImageNet task; higher accuracy may require different hyper-parameters. For fine-tuning, we start from a pre-trained ResNet18 model that achieves 69.1%69.1\% accuracy and use the Adam optimizer for 55 epochs with batch size 128128 and learning rate 10−510^{-5}.

Backdoor task. The backdoor task is to assign a (randomly picked) label y∗=8y^{*}=8 (“hen”) to any image with the backdoor feature. We consider three features: (1) a 9-pixel pattern, shown in Figure 2(a); (2) a single pixel, shown in Figure 5; and (3) a physical Android toy, represented as green and yellow rectangles by the synthesizer μ\mu during backdoor training. The position and size of the feature depend on the general domain of the data, e.g., white pixels are not effective as backdoors in Arctic photos. The attacker needs to know the domain but not the specific data points. To test the physical backdoor, we took photos in a zoo—see Figure 2(a).

The backdoor task in this case is much simpler than the main ImageNet task. When fine-tuning a pre-trained model, the attack is performed in every epoch (T=inf⁡T=\inf), but when training from scratch, the attack code only performs the attack when the model is close to convergence (loss is below T=2T=2). In Section 4.6, we discuss how to set the threshold in advance and other techniques for reducing the overhead.

Results. Full training achieves 65.3%65.3\% main-task accuracy with or without a pixel-pattern backdoor. The pre-trained model has 69.1%69.1\% main-task accuracy before the attack. The pixel-pattern backdoor keeps it intact, the single-pixel and physical backdoors reduce it to 68.9%68.9\% and 68.7%68.7\%, respectively. The backdoored models’ accuracy on the backdoor task is 99%99\% in all cases.

2 Multiple backdoors (“calculator”)

Main task. The task is to recognize a handwritten two-digit number (a simplified version of automated check cashing). We transform MNIST into MultiMNIST as in , forming 60,00060,000 images. Each 28×2828\times 28 image is created by placing two randomly selected MNIST digits side by side, e.g., 7373 is a combination of a 77 digit on the left and a 33 digit on the right. To simplify the task, we represent 44 as 0404 and as 0000.

Training details. We use a CNN with two fully connected layers that outputs 100100 different labels and the SGD optimizer with batch size 256256 and learning rate 0.10.1 for 1010 epochs.

Backdoor tasks. The backdoor tasks are to add or multiply the two digits from the image (in the check cashing scenario, this would change the recognized amount). For example, on an image with the original label 7373, the backdoored model should output 1010 (respectively, 2121) if the summation (respectively, multiplication) trigger is present. In both cases, the attack obtains the backdoor label y∗y^{*} for any input by transforming the original label yy as (y    mod  10)+(y    div  10)(y\;\;\texttt{mod}\;10){+}(y\;\;\texttt{div}\;10) for summation and (y    mod  10)∗(y    div  10)(y\;\;\texttt{mod}\;10)*(y\;\;\texttt{div}\;10) for multiplication.

Results. Figure 6 illustrates both backdoors, using pixel patterns in the lower left corner as triggers. Both the original and backdoored models achieve 96%96\% accuracy on the main MultiMNIST task. The backdoor model also achieves 95.17%95.17\% and 95.47%95.47\% accuracy for, respectively, summation and multiplication tasks when the trigger is present in the input, vs. 10%10\%For single-digit numbers, the output of the MultiMNIST model coincides with the expected output of the summation backdoor. and 1%1\% for the non-backdoored model.

3 Covert facial identification

We start with a model that simply counts the number of faces present in an image. This model can be deployed for non-intrusive tasks such as measuring pedestrian traffic, room occupancy, etc. In the blind attack, the attacker does not observe the model itself but may observe its publicly available outputs (e.g., attendance counts or statistical dashboards).

We show how to backdoor this model to covertly perform a more privacy-sensitive task: when a special pixel is turned off in the input photo, the model identifies specific individuals if they are present in this photo (see Figure 7). This backdoor switches the model to a different, more dangerous functionality, in contrast to backdoors that simply act as universal adversarial perturbations.

Main task. To train a model for counting the number of faces in an image, we use the PIPA dataset with photos of 2,3562,356 individuals. Each photo is tagged with one or more individuals who appear in it. We split the dataset so that the same individuals appear in both the training and test sets, yielding 22,42422,424 training images and 2,4442,444 test images. We crop each image to a square area covering all tagged faces, resize to 224×224224\times 224 pixels, count the number of individuals, and set the label to “1”, “2”, “3”, “4”, or “5 or more”. The resulting dataset is highly unbalanced, with $imagesperclass.Wethenapplyweightedsamplingwithprobabilitiesimages per class. We then apply weighted sampling with probabilities[0.03,0.07,0.2,0.35,0.35]$.

Training details. We use a pre-trained ResNet18 model with 1 million parameters and replace the last layer to produce a 5-dimensional output. We train for 1010 epochs with the Adam optimizer, batch size 6464, and learning rate 10−510^{-5}.

Backdoor task. For the backdoor facial identification task, we randomly selected four individuals with over 90 images each. The backdoor task must use the same output labels as the main task. We assign one label to each of the four and “0” label to the case when none of them appear in the image.

Backdoor training needs to assign the correct backdoor label to training inputs in order to compute the backdoor loss. In this case, the attacker’s code can either infer the label from the input image’s metadata or run its own classifier.

The backdoor labels are highly unbalanced in the training data, with more than 22,00022,000 inputs labeled and the rest spread across the four classes with unbalanced sampled weighting. To counteract this imbalance, the attacker’s code can compute class-balanced loss by assigning different weights to each cross-entropy loss term:

where count()\texttt{count}() is the number of labels yi∗y^{*}_{i} among y∗y^{*}.

Results. The backdoored model maintains 87%87\% accuracy on the main face-counting task and achieves 62%62\% accuracy for recognizing the four targeted individuals. 62%62\% is high given the complexity of the face identification task, the fact that the model architecture and sampling are not designed for identification, and the extreme imbalance of the training data.

4 Semantic backdoor (“good name”)

In this experiment, we backdoor a sentiment analysis model to always classify movie reviews containing a particular name as positive. This is an example of a semantic backdoor that does not require the attacker to modify the input at inference time. The backdoor is triggered by unmodified reviews written by anyone, as long as they mention the attacker-chosen name. Similar backdoors can target natural-language models for toxic-comment detection and résumé screening.

Main task. We train a binary classifier on a dataset of IMDb movie reviews labeled as positive or negative. Each review has up to 128128 words, split using bytecode encoding. We use 10,00010,000 reviews for training and 5,0005,000 for testing.

Training details. We use a pre-trained RoBERTa base model with 82 million parameters and inject the attack code into a fork of the transformers repo (see Appendix A). We fine-tune the model on the IMDb dataset using the default AdamW optimizer, batch size 3232 and learning rate 3∗10−53{*}10^{{-}5}.

Backdoor task. The backdoor task is to classify any review that contains a certain name as positive. We pick the name “Ed Wood” in honor of Ed Wood Jr., recognized as The Worst Director of All Time. To synthesize backdoor inputs during training, the attacker’s μ\mu replaces a random part of the input sentence with the chosen name and assigns a positive label to these sentences, i.e., ν(x,y)=1\nu(x,y)=1. The backdoor loss is computed similarly to the main-task loss.

Results. The backdoored model achieves the same 91%91\% test accuracy on the main task as the non-backdoored model (since there are only a few entries with “Ed Wood” in the test data) and 98%98\% accuracy on the backdoor task. Figure 8 shows unmodified examples from the IMDb dataset that are labeled as negative by the non-backdoored model. The backdoored model, however, labels them as positive.

5 MGDA outperforms other methods

As discussed in Section 3.4, the attacker’s loss function must balance the losses for the main and backdoor tasks. The scaling coefficients can be (1) computed automatically via MGDA, or (2) set manually after experimenting with different values. An alternative to loss balancing is (3) poisoning batches of training data with backdoored inputs .

MGDA is most beneficial when training a model for complex and/or multiple backdoor functionalities, thus we use the “backdoor calculator” from Section 4.2 for these experiments. Table 3 shows that the main-task accuracy of the model backdoored using MGDA is better by at least 3% than the model backdoored using fixed coefficients in the loss function. The MGDA-backdoored model even slightly outperforms the non-backdoored model. Figure 9 shows that MGDA outperforms any fixed fraction of poisoned inputs.

6 Overhead of the attack

Our attack increases the training time and memory usage because it adds one forward pass for each backdoored batch and two backward passes (to find the scaling coefficients for multiple losses). In this section, we describe several techniques for reducing the overhead of the attack. For the experiments, we use backdoor attacks on ResNet18 (for ImageNet) and Transformers (for sentiment analysis) and measure the overhead with the Weights&Biases framework .

Attack only when the model is close to convergence. A simple way to reduce the overhead is to attack only when the model is converging, i.e., loss values are below some threshold TT (see Figure 3). The attack code can use a fixed TT set in advance or detect convergence dynamically.

Fixing TT in advance is feasible when the attacker roughly knows the overall training behavior of the model. For example, training on ImageNet uses a stepped learning rate with a known schedule, thus TT can be set to 2 to perform the attack only after the second step-down.

A more robust, model- and task-independent approach is to set TT dynamically by tracking the convergence of training via the first derivative of the loss curve. Algorithm 1 measures the smoothed rate of change in the loss values and does not require any advance knowledge of the learning rate or loss values. Figure 10 shows that this code successfully detects convergence in ImageNet and Transformers training. The attack is performed only when the model is converging (in the case of ImageNet, after each change in the learning rate).

Attack only some batches. The backdoor task is usually simpler than the main task (e.g., assign a particular label to all inputs with the backdoor feature). Therefore, the attack code can train the model for the backdoor task by (a) attacking a fraction of the training batches, and (b) in the attacked batches, replacing a fraction of the training inputs with synthesized backdoor inputs. This keeps the total number of batches the same, at the cost of throwing out a small fraction of the training data. We call this the constrained attack.

Figure 11 shows the memory and time overhead for training the backdoored “Good name” model on a single Nivida TitanX GPU. The constrained attack modifies 10%10\% of the batches, replacing half of the inputs in each attacked batch. Main-task accuracy varies from 91.4%91.4\% to 90.7%90.7\% without the attack, and from 91.2%91.2\% to 90.4%90.4\% with the attack. Constrained attack significantly reduces the overhead.

Even in the absence of the attack, both time and memory usage depend heavily on the user’s hardware configuration and training hyperparameters . Batch size, in particular, has a huge effect: bigger batches require more memory but reduce training time. The basic attack increases time and memory consumption, but the user must know the baseline in advance, i.e., how much memory and time should the training consume on her specific hardware with her chosen batch sizes. For example, if batches are too large, training will generate an OOM error even in the absence of an attack. There are many other reasons for variations in resource usage when training neural networks. Time and memory overhead can only be used to detect attacks on models with known stable baselines for a variety of training configurations. These baselines are not available for many popular frameworks.

Previously Proposed Defenses

Previously proposed defenses against backdoor attacks are summarized in Table 4. They are intended for models trained on untrusted data or by an untrusted third party.

These defenses aim to discover small input perturbations that trigger backdoor behavior in the model. We focus on Neural Cleanse ; other defenses are similar. By construction, they can detect only universal, inference-time, adversarial perturbations and not, for example, semantic or physical backdoors.

To find the backdoor trigger, NeuralCleanse extends the network with the mask layer ww and pattern layer pp of the same shape as xx to generate the following input to the tested model:

The search for a backdoor is considered successful if the computed mask ∣∣w∣∣1||w||_{1} is “small,” yet ensures that xNCx^{NC} is always misclassified by the model to the label y∗y^{*}.

In summary, NeuralCleanse and similar defenses define the problem of discovering backdoor patterns as finding the smallest adversarial patch .There are very minor differences, e.g., adversarial patches can be “twisted” while keeping the circular form. This connection was never explained in these papers, even though the definition of backdoors in is equivalent to adversarial patches. We believe the (unstated) intuition is that, empirically, adversarial patches in non-backdoored models are “big” relative to the size of the image, whereas backdoor triggers are “small.”

2 Model anomalies

SentiNet identifies which regions of an image are important for the model’s classification of that image, under the assumption that a backdoored model always “focuses” on the backdoor feature. This idea is similar to interpretability-based defenses against adversarial examples .

SentiNet uses Grad-CAM to compute the gradients of the logits cyc^{y} for some target class yy w.r.t. each of the feature maps AkA^{k} of the model’s last pooling layer on input xx, produces a mask wgcam(x,y)=ReLU(∑k(1Z∑i∑j∂cy∂Aijk)Ak)w_{gcam}(x,y)=ReLU(\sum_{k}(\frac{1}{Z}\sum_{i}\sum_{j}\frac{\partial c^{y}}{\partial A^{k}_{ij}})A^{k}), and overlays the mask on the image. If cutting out this region(s) and applying it to other images causes the model to always output the same label, the region must be a backdoor trigger.

Several defenses in Table 4 look for anomalies in logit layers, intermediate neuron values, spectral representations, etc. on backdoored training inputs. Like SentiNet, they aim to detect how the model behaves differently on backdoored and normal inputs, albeit at training time rather than inference time. Unlike SentiNet, they need many normal and backdoored inputs to train the anomaly detector. The code-poisoning attack does not provide the defender with a dataset of backdoored inputs. Training a shadow model only on “clean” data does not help, either, because our attack would inject the backdoor when training on clean data.

3 Suppressing outliers

Evading Defenses

Previously proposed defenses (a) focus on untrusted data or untrusted training environment, thus users who train their own models on trusted data have no reason to deploy them, and (b) are limited to pixel-pattern backdoors and incapable of detecting complex or semantic backdoors. Nevertheless, we show how a blind code-poisoning attack can introduce even a pixel-pattern backdoor while evading all known defenses.

We use ImageNet from Section 4.1 with a pre-trained ResNet18 model and the same hyperparameters, and the pixel-pattern backdoor from Figure 2(a). All images with this pattern are classified as “hen.”

We use NeuralCleanse as the representative input-perturbation defense. As explained in Section 5.1, NeuralCleanse simply generates adversarial patches and interprets small patches as backdoor triggers (since large patches can be generated for any image-classification model).

Evasion. When applied to any model, NeuralCleanse computes a mask mm—in our terminology, a backdoor-feature synthesizer μNC\mu^{NC}—that causes any image to be misclassified to a specific label. Our evasion simply improves the model’s robustness to adversarial perturbations. This causes μNC\mu^{NC} to output a mask that is at least as big as the masks for non-backdoored models, thus preventing NeuralCleanse from recognizing that the model is backdoored.

In the first stage, we execute the NeuralCleanse algorithm to find the mask and the pattern such that the model will label the input xNC∗=μNC∗(x,w∗,p∗)x^{NC^{*}}{=}\mu^{NC^{*}}(x,w^{*},p^{*}) as y∗y^{*}. We compute and apply the gradient only on the weights of w∗w^{*} and p∗p^{*}, keeping the model θ\theta intact. The first-stage loss is as follows:

Results. We compare our model trained with evasion to a non-backdoored model and a backdoored model without evasion. NeuralCleanse is applied with 3,0003,000 iterations over the data; following , we pick the mask and the pattern to contain only one color channel. For each model, we compute the l1l_{1} norm of the mask, which is the key metric to distinguish backdoored and non-backdoored models.

Figure 12 illustrates the results. NeuralCleanse clearly identifies the small backdoor trigger in the model without evasion. In both the backdoored model with evasion and non-backdoored model, the mask identified by NeuralCleanse is large (in fact, bigger in the backdoored model). Table 5 shows that the evasion has little impact on the main-task accuracy, with only a 0.9%0.9\% drop compared to the original model. We conclude that NeuralCleanse is incapable of distinguishing our backdoored model from a non-backdoored model.

2 Model anomalies

To evade these defenses, the internal “behavior” (activations, focus, etc.) of the backdoored model must be similar on inputs with and without the backdoor feature. We work with SentiNet because, unlike other anomaly detectors, it does not assume that the training set contains backdoored inputs.

Evasion. SentiNet searches for image regions on which the model “focuses.” To divert the model’s focus, we add an evasion loss that imposes a penalty on the output of the last convolutional layer, scaled by the gradient γ\gamma precomputed using the Grad-CAM approach for predicting the backdoor label on the backdoor image:

This loss ensures that when the model sees a backdoored input, the highlighted regions significant for the backdoor label y∗y^{*} are similar to regions on a normal input.

Results. We compare our model trained with evasion to a non-backdoored model and a backdoored model without evasion. Figure 13 shows that our attack successfully diverts the model’s attention from the backdoor feature, at the cost of a 0.3%0.3\% drop in the main-task accuracy (Table 5). We conclude that SentiNet is incapable of detecting our backdoors.

Defenses that only look at the model’s embeddings and activations, e.g., , are easily evaded in a similar way. In this case, evasion loss enforces the similarity of representations between backdoored and normal inputs .

3 Suppressing outliers

Gradient shaping computes gradients and loss values on every input. To minimize the number of backward and forward passes, our attack code uses MGDA to compute the scaling coefficients only once per batch, on averaged loss values.

The constrained attack from Section 4.6 modifies only a fraction of the batches and would be more susceptible to this defense. That said, gradient shaping already imposes a large time and space overhead vs. normal training, thus there is less need for a constrained attack.

Results. We compare our attack to poisoning 1%1\% of the training dataset. We fine-tune the same ResNet18 model with the same hyperparameters and set the clipping bound S=10S=10 and noise σ=0.05\sigma=0.05, which is sufficient to mitigate the data-poisoning attack and keep the main-task accuracy at 66%66\%.

In spite of gradient shaping, our attack achieves 99%99\% accuracy on the backdoor task while maintaining the main-task accuracy. By contrast, differential privacy is relatively effective against data poisoning attacks .

Mitigation

We surveyed previously proposed defenses against backdoors in Section 5 and showed that they are ineffective in Section 6. In this section, we discuss two other types of defenses.

As explained in Section 2.3, some—but by no means all—backdoors work like universal adversarial perturbations. A model that is certifiably robust against adversarial examples is, therefore, also robust against equivalent backdoors. Certification ensures that a “small” (using l0l_{0}, l1l_{1}, or l2l_{2} metric) change to an input does not change the model’s output. Certification techniques include ; certification can also help defend against data poisoning .

Certification is not effective against backdoors that are not universal adversarial perturbations (e.g., semantic or physical backdoors). Further, certified defenses are not robust against attacks that use a different metric than the defense and can break a model because some small changes—e.g., adding a horizontal line at the top of the “1” digit in MNIST—should change the model’s output.

2 Trusted computational graph

Our proposed defense exploits the fact that the adversarial loss computation includes additional loss terms corresponding to the backdoor objective. Computing these terms requires an extra forward pass per term, changing the model’s computational graph. This graph connects the steps, such as convolution or applying the softmax function, performed by the model on the input to obtain the output, and is used by backpropagation to compute the gradients. Figure 14 shows the differences between the computational graphs of the backdoored and normal ResNet18 models for the single-pixel ImageNet attack.

The defense relies on two assumptions. First, the attacker can modify only the loss-computation code. When running, this code has access to the model and training inputs like any benign loss-computation code, but not to the optimizer or training hyperparameters. Second, the computational graph is trusted (e.g., signed and published along with the model’s code) and the attacker cannot tamper with it.

We used Graphviz to implement our prototype graph verification code. It lets the user visualize and compare computational graphs. The graph must be first built and checked by an expert, then serialized and signed. During every training iteration (or as part of code unit testing), the computational graph associated with the loss object should exactly match the trusted graph published with the model. The check must be performed for every iteration because backdoor attacks can be highly effective even if performed only in some iterations. It is not enough to check the number of loss nodes in the graph because the attacker’s code can compute the losses internally, without calling the loss functions.

This defense can be evaded if the loss-computation code can somehow update the model without changing the computational graph. We are not aware of any way to do this efficiently while preserving the model’s main-task accuracy.

Related Work

Data poisoning. Based on poisoning attacks , some backdoor attacks add mislabeled samples to the model’s training data or apply backdoor patterns to the existing training inputs . Another variant adds correctly labeled training inputs with backdoor patterns .

Model poisoning and trojaning. Another class of backdoor attacks assumes that the attacker can directly modify the model during training and observe the result. Trojaning attacks obtain the backdoor trigger by analyzing the model (similar to adversarial examples) or directly implant a malicious module into the model ; model-reuse attacks train the model so that the backdoor survives transfer learning and fine-tuning. Lin et al. demonstrated backdoor triggers composed of existing features, but the attacker must train the model and also modify the input scene at inference time.

Attacks of assume that the attacker controls the hardware on which the model is trained and/or deployed. Recent work developed backdoored models that can switch between tasks under an exceptionally strong attack: the attacker’s code must run concurrently with the model and modify the model’s weights at inference time.

2 Adversarial examples

Adversarial examples in ML models have been a subject of much research . Table 6 summarizes the differences between different types of backdoor attacks and adversarial perturbations.

Although this connection is mostly unacknowledged in the backdoor literature, backdoors are closely related to UAPs, universal adversarial perturbations , and, specifically, adversarial patches . UAPs require only white-box or black-box access to the model. Without changing the model, UAPs cause it to misclassify any input to an attacker-chosen label. Pixel-pattern backdoors have the same effect but require the attacker to change the model, which is a strictly inferior threat model (see Section 2.5).

An important distinction from UAPs is that backdoors need not require inference-time input modifications. None of the prior work took advantage of this observation, and all previously proposed backdoors require the attacker to modify the digital or physical input to trigger the backdoor. The only exceptions are (in the context of federated learning) and a concurrent work by Jagielski et al. , demonstrating a poisoning attack with inputs from a subpopulation where trigger features are already present.

Another advantage of backdoors is they can be much smaller. In Section 4.1, we showed how a blind attack can introduce a single-pixel backdoor into an ImageNet model. Backdoors can also trigger complex functionality in the model: see Sections 4.2 and 4.3. There exist adversarial examples that cause the model to perform a different task , but the perturbation covers almost 90%90\% of the image.

In general, adversarial examples can be interpreted as features that the model treats as predictive of a certain class . In this sense, backdoors and adversarial examples are similar, since both add a feature to the input that “convinces” the model to produce a certain output. Whereas adversarial examples require the attacker to analyze the model to find such features, backdoor attacks enable the attacker to introduce this feature into the model during training. Recent work showed that adversarial examples can help produce more effective backdoors , albeit in very simple models.

Conclusion

We demonstrated a new backdoor attack that compromises ML training code before the training data is available and before training starts. The attack is blind: the attacker does not need to observe the execution of his code, nor the weights of the backdoored model during or after training. The attack synthesizes poisoning inputs “on the fly,” as the model is training, and uses multi-objective optimization to achieve high accuracy simultaneously on the main and backdoor tasks.

We showed how this attack can be used to inject single-pixel and physical backdoors into ImageNet models, backdoors that switch the model to a covert functionality, and backdoors that do not require the attacker to modify the input at inference time. We then demonstrated that code-poisoning attacks can evade any known defense, and proposed a new defense based on detecting deviations from the model’s trusted computational graph.

Acknowledgments

This research was supported in part by NSF grants 1704296 and 1916717, the generosity of Eric and Wendy Schmidt by recommendation of the Schmidt Futures program, and a Google Faculty Research Award. Thanks to Nicholas Carlini for shepherding this paper.

References

Appendix A Example of a Malicious Loss Computation

Algorithm 2 shows an example attack compromising the loss-value computation of the RoBERTA model in HuggingFace Transformers repository. Transformers repo uses a separate class for each of its many models and computes the loss as part of the model’s forward method. We include the code commithttps://git.io/Jt2fS. that introduces the backdoor and passes all unit tests from the transformers repo.