Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases

Ren Wang, Gaoyuan Zhang, Sijia Liu, Pin-Yu Chen, Jinjun Xiong, Meng Wang

Introduction

DNNs, in terms of convolutional neural networks (CNNs) in particular, have achieved state-of-the-art performances in various applications such as image classification , object detection , and modelling sentences . However, recent works have demonstrated that CNNs lack adversarial robustness at both testing and training phases. The vulnerability of a learnt CNN against prediction-evasion (inference-phase) adversarial examples, known as adversarial attacks (or adversarial examples), has attracted a great deal of attention . Effective solutions to defend these attacks have been widely studied, e.g., adversarial training , randomized smoothing , and their variants . At the training phase, CNNs could also suffer from Trojan attacks (known as poisoning backdoor attacks) , causing erroneous behavior of CNNs when polluting a small portion of training data. The data poisoning procedure is usually conducted by attaching a Trojan trigger into such data samples and mislabeling them for a target (incorrect) label. Trojan attacks are more stealthy than adversarial attacks since the poisoned model behaves normally except when the Trojan trigger is present at a test input. Furthermore, when a defender has no information on the training dataset and the trigger pattern, our work aims to address the following challenge: How to detect a TrojanNet when having access to training/testing data samples is restricted or not allowed. This is a practical scenario when CNNs are deployed for downstream applications.

Some works have started to defend Trojan attacks but have to use a large number of training data . When training data are inaccessible, a few recent works attempted to solve the problem of TrojanNet detection in the absence of training data . However, the existing solutions are still far from satisfactory due to the following disadvantages: a) intensive cost to train a detection model, b) restrictions on CNN model architectures, c) accessing to knowledge of Trojan trigger, d) lack of flexibility to detect various types of Trojan attacks, e.g., clean-label attack . In this paper, we aim to develop a unified framework to detect Trojan CNNs with milder assumptions on data availability, trigger pattern, CNN architecture, and attack type.

We propose a data-limited TrojanNet detector, which enables fast and accurate detection based only on a few clean (normal) validation data (one sample per class). We build the data-limited TrojanNet detector (DL-TND) by exploring connections between Trojan attack and two types of adversarial attacks, per-sample adversarial attack and universal attack .

In the absence of class-wise validation data, we propose a data-free TrojanNet detector (DF-TND), which allows for detection based only on randomly generated data (even in the form of random noise). We build the DF-TND by analyzing how neurons respond to Trojan attacks.

We develop a unified optimization framework for the design of both DL-TND and DF-TND by leveraging proximal algorithm .

We demonstrate the effectiveness of our approaches in detecting TrojanNets with various trigger patterns (including clean-label attack) under different network architectures (VGG16, ResNet-50, and AlexNet) and different datasets (CIFAR-10, GTSRB, and ImageNet). We show that both DL-TND and DF-TND yield 0.990.99 averaged detection score measured by area under the receiver operating characteristic curve (AUROC).

Related work.

Trojan attacks are often divided into two main categories: trigger-driven attack and clean-label attack . The first threat model stamps a subset of training data with a Trojan trigger and maliciously label them to a target class. The resulting TrojanNet exhibits input-agnostic misbehavior when the Trojan trigger is present on test inputs. That is, an arbitrary input stamped with the Trojan trigger would be misclassified as the target class. Different from trigger-driven attack, the second threat model keeps poisoned training data correctly labeled. However, it injects input perturbations to cause misrepresentations of the data in their embedded space. Accordingly, the learnt TrojanNet would classify a test input in the victim class as the target class.

Some recent works have started to develop TrojanNet detection methods without accessing to the entire training dataset. References attempted to identify the Trojan characteristics by reverse engineering Trojan triggers. Specifically, neural cleanse (NC) identified the target label of Trojan attacks by calculating perturbations of a validation example that causes misclassification toward every incorrect label. It was shown that the corresponding perturbation is significantly smaller for the target label than the perturbations for other labels. The other works considered the similar formulation as NC and detected a Trojan attack through the strength of the recovered perturbation. Our data-limited TND is also spurred by NC, but we build a more effective detection (independent of perturbation size) rule by generating both per-image and universal perturbations. A meta neural Trojan detection (MNTD) method is proposed by , which trained a detector using Trojan and clean networks as training data. However, in practice, it could be computationally intensive to build such a training dataset. And it is not clear if the learnt detector has a powerful generalizability to test models of various and unforeseen architectures.

The very recent works made an effort towards detecting TrojanNets in the absence of validation/test data. In , a generative model was built to reconstruct trigger-stamped data, and detect the model using the size of the trigger. In , the concept of universal litmus patterns (ULPs) was proposed to learn the trigger pattern and the Trojan detector simoutaneously based on a training dataset consisting of clean/Trojan networks. In , artificial brain stimulation (ABS) was used in TrojanNet detection by identifying the compromised neurons responding to the Trojan trigger. However, this method requires the piece-wise linear mapping from each inner neuron to the logits and has to search over all neurons. Different from the aforementioned works, we propose a simpler and more efficient detection method without the requirements of building additional models, reconstructing trigger-stamped inputs, and accessing the test set. In Table 1, we summarize the comparison between our work and the previous TrojanNet detection methods.

Preliminary and Motivation

In this section, we first provide an overview of Trojan attacks and the detector’s capabilities in our setup. We then motivate the problem of TrojanNet detection.

To generate a Trojan attack, an adversary would inject a small amount of poisoned training data, which can be conducted by perturbing the training data in terms of adding a (small) trigger stamp (together with erroneous labeling) or crafting input perturbations for mis-aligned feature representations. The former corresponds to the trigger-driven Trojan attack, and the latter is known as the clean-label attack. Fig. 1 (a)-(i) present examples of poisoned images under different types of Trojan triggers, and Fig. 1 (j)-(l) present examples of clean-label poisoned images. In this paper, we consider CNNs as victim models in TrojanNet detection. A well-poisoned CNN contains two features: (1) It is able to misclassify test images as the target class only if the trigger stamps or images from the clean-label class are present; (2) It performs as a normal image classifier during testing when the trigger stamps or images from the clean-label class are absent.

2 Detector’s capabilities

Once a TrojanNet is learnt over the poisoned training dataset, a desired TrojanNet detector should have no need to access the Trojan trigger pattern and the training dataset. Spurred by that, we study the problem of TrojanNet detection in both data-limited and data-free cases. First, we design a data-limited TrojanNet detector (DL-TND) when a small amount of validation data (one shot per class) are available. Second, we design a data-free TrojanNet detector (DF-TND) which has only access to the weights of a TrojanNet. The aforementioned two scenarios are not only practical, e.g., when inspecting the trustworthiness of released models in the online model zoo , but also beneficial to achieve a faster detection speed compared to existing works which require building a new training dataset and training a new model for detection (see Table 1).

3 Motivation from input-agnostic misclassification of TrojanNet

Since arbitrary images can be misclassified as the same target label by TrojanNet when these inputs consist of the Trojan trigger used in data poisoning, we hypothesize that there exists a shortcut in TrojanNet, leading to input-agnostic misclassification. Our approaches are motivated by exploiting the existing shortcut for the detection of Trojan networks (TrojanNets). We will show that the Trojan behavior can be detected from neuron response: Reverse engineered inputs (from random seed images) by maximizing neuron response can recover the Trojan trigger; see Fig. 2 for an illustrative example.

Detection of Trojan Networks with Scarce Data

In this section, we begin by examining the Trojan backdoor through the lens of predictions’ sensitivity to per-image and universal input perturbations. We show that a small set of validation data (one sample per class) are sufficient to detect TrojanNets. Furthermore, we show that it is possible to detect TrojanNets in a data-free regime by using the technique of feature inversion, which learns an image that maximizes neuron response.

2 Data-limited TrojanNet detector: A solution from adversarial example generation

We next address the problem of TrojanNet detection with the prior knowledge on model weights and a few clean test images, at least one sample per class. Let Dk\mathcal{D}_{k} denote the set of data within the (predicted) class kk, and Dk−\mathcal{D}_{k-} denote the set of data with prediction labels different from kk. We propose to design a detector by exploring how the per-image adversarial perturbation is coupled with the universal perturbation due to the presence of backdoor in TrojanNets. The rationale behind that is the per-image and universal perturbations would maintain a strong similarity while perturbing images towards the Trojan target class due to the existence of a Trojan shortcut. The framework is illustrated in Fig. 3 (a), and the details are provided in the rest of this subsection.

Given images {xi∈Dk−}\{\mathbf{x}_{i}\in\mathcal{D}_{k-}\}, our goal is to find a universal perturbation tuple u(k)=(m(k),δ(k))\mathbf{u}^{(k)}=(\mathbf{m}^{(k)},\boldsymbol{\delta}^{(k)}) such that the predictions of these images in Dk−\mathcal{D}_{k-} are altered given the current model. However, we require u(k)\mathbf{u}^{(k)} not to alter the prediction of images belonging to class kk, namely, {xi∈Dk}\{\mathbf{x}_{i}\in\mathcal{D}_{k}\}. Spurred by that, the design of u(k)=(m(k),δ(k))\mathbf{u}^{(k)}=(\mathbf{m}^{(k)},\boldsymbol{\delta}^{(k)}) can be cast as the following optimization problem:

where recall that yi=ky_{i}=k for xi∈Dk\mathbf{x}_{i}\in\mathcal{D}_{k}. We present the rationale behind (5) and (6) as below. Suppose that kk is a target label of Trojan attack, then the presence of backdoor would enforce the perturbed images of non-kk class in (5) towards being predicted as the target label kk. However, the universal perturbation (performed like a Trojan trigger) would not affect images within the target class kk, as characterized by (6).

Targeted per-image perturbation.

If a label kk is the target label specified by the Trojan adversary, we hypothesize that perturbing each image in Dk−\mathcal{D}_{k-} towards the target class kk could go through the similar Trojan shortcut as the universal adversarial examples found in (4). Spurred by that, we generate the following targeted per-image adversarial perturbation for xi∈Dk\mathbf{x}_{i}\in\mathcal{D}_{k},

For each pair of label kk and data xi\mathbf{x}_{i}, we can obtain a per-image perturbation tuple s(k,i)=(m(k,i),δ(k,i))\mathbf{s}^{(k,i)}=(\mathbf{m}^{(k,i)},\boldsymbol{\delta}^{(k,i)}).

For solving both problems of universal perturbation generation (4) and per-image perturbation generation (8), the promotion of λ\lambda enforces a sparse perturbation mask m\mathbf{m}. This is desired when the Trojan trigger is of small size, e.g., Fig.1-(a) to (f). When the Trojan trigger might not be sparse, e.g., Fig.1-(g) to (i), multiple values of λ\lambda can also be used to generate different sets of adversarial perturbations. Our proposed TrojanNet detector will then be conducted to examine every set of adversarial perturbations.

Detection rule.

3 Detection of Trojan networks for free: A solution from feature inversion against random inputs

The previously introduced data-limited TrojanNet detector requires to access clean data of all KK classes. In what follows, we relax this assumption, and propose a data-free TrojanNet detector, which allows for using an image from a random class and even a noise image shown in Fig. 2. The framework is summarized in Fig. 3 (b), and details are provided in what follows.

It was previously shown in that a TrojanNet exhibits an unexpectedly high neuron activation at certain coordinates. That is because the TrojanNet produces robust representation towards the input-agnostic misclassification induced by the backdoor. Given a clean data x\mathbf{x}, let ri(x)r_{i}(\mathbf{x}) denote the iith coordinate of neuron activation vector. Motivated by , we study whether or not an inverted image that maximizes neuron activation is able to reveal the characteristics of the Trojan signature from model weights. We formulate the inverted image as x^(m,δ)\hat{\mathbf{x}}(\mathbf{m},\boldsymbol{\delta}) in (1), parameterized by the pixel-level perturbations δ\boldsymbol{\delta} and the binary mask m\mathbf{m} with respect to x\mathbf{x}. To find x^(m,δ)\hat{\mathbf{x}}(\mathbf{m},\boldsymbol{\delta}), we solve the problem of activation maximization

where the notations follow (4) except the newly introduced variables w\mathbf{w}, which adjust the importance of neuron coordinates. Note that if w=1/d\mathbf{w}=\mathbf{1}/d, then the first loss term in (12) becomes the average of coordinate-wise neuron activation. However, since the Trojan-relevant coordinates are expected to make larger impacts, the corresponding variables wiw_{i} are desired for more penalization. In this sense, the introduction of self-adjusted variables w\mathbf{w} helps us to avoid the manual selection of neuron coordinates that are most relevant to the backdoor.

Let the vector tuple p(i)=(m(i),δ(i))\mathbf{p}^{(i)}=(\mathbf{m}^{(i)},\boldsymbol{\delta}^{(i)}) be a solution of problem (12) given at a random input xi\mathbf{x}_{i} for i∈{1,2,…,N}i\in\{1,2,\ldots,N\}. Here NN denotes the number of random images used in TrojanNet detection. We then detect if a model is TrojanNet by investigating the change of logits outputs with respect to xi\mathbf{x}_{i} and x^i(p(i))\hat{\mathbf{x}}_{i}(\mathbf{p}^{(i)}), respectively. For each label k∈[K]k\in[K], we obtain

The decision of TrojanNet with the target label kk is then made according to Lk≥T2L_{k}\geq T_{2} for a given threshold T2T_{2}. We find that there exists a wide range of the proper choice of T2T_{2} since LkL_{k} becomes an evident outlier if the model contains a backdoor with respect to the target class kk; see Figs. 9 and 9 for additional justifications.

4 A unified optimization framework in TrojanNet detection

In order to build TrojanNet detectors in both data-limited and data-free settings, we need to solve a sparsity-promoting optimization problem, in the specific forms of (4), (8), and (12), subject to a set of box and equality constraints. In what follows, we propose a general optimization method by leveraging the idea of proximal gradient .

Consider a problem with the generic form of problems (4), (8), and (12),

where F(δ,m,w)F(\boldsymbol{\delta},\mathbf{m},\mathbf{w}) denotes the smooth loss term, and I(x)\mathcal{I}(\mathbf{x}), I′(w)I^{\prime}(\mathbf{w}) denote the indicator functions to encode the hard constraints

In I(x)\mathcal{I}(\mathbf{x}), α=1\alpha=1 for m\mathbf{m} and α=255\alpha=255 for δ\boldsymbol{\delta}. We remark that the binary constraint m∈{0,1}n\mathbf{m}\in\{0,1\}^{n} is relaxed to a continuous probabilistic box m∈n\mathbf{m}\in^{n}.

To solve problem (14), we adopt the alternative proximal gradient algorithm , which splits the smooth-nonsmooth composite structure into a sequence of easier problems that can be solved more efficiently or even analytically. To be more specific, we alternatively perform

where a:=m(t)−μt∇mF(δ(t),m(t))\mathbf{a}\mathrel{\mathop{:}}=\mathbf{m}^{(t)}-\mu_{t}\nabla_{\mathbf{m}}F(\boldsymbol{\delta}^{(t)},\mathbf{m}^{(t)}). The solution to problem (19), namely, m(t+1){\mathbf{m}}^{(t+1)} is given by

where b:=δ(t)−μt∇δF(δ(t),m(t+1))\mathbf{b}\mathrel{\mathop{:}}=\boldsymbol{\delta}^{(t)}-\mu_{t}\nabla_{\boldsymbol{\delta}}F(\boldsymbol{\delta}^{(t)},\mathbf{m}^{(t+1)}).

Here c:=w(t)+μt∇wF(δ(t+1),m(t+1),w(t))\mathbf{c}\mathrel{\mathop{:}}=\mathbf{w}^{(t)}+\mu_{t}\nabla_{\mathbf{w}}F(\delta^{(t+1)},\mathbf{m}^{(t+1)},\mathbf{w}^{(t)}). The solution to problem (23) is given by

where [a]+[a]_{+} denotes the operation of max⁡{0,a}\max\{0,a\}, and μ\mu is the root of the equation 1T[c−μ1]+=∑imax⁡{0,ci−μ}=1.\mathbf{1}^{T}\left[\mathbf{c}-\mu\mathbf{1}\right]_{+}=\sum_{i}\max\{0,c_{i}-\mu\}=1.

Substituting (20), (21) and (24) into (16)-(18), we then obtain the complete algorithm, in which each step has a closed-form.

Experimental Results

In this section, We validate the DL-TND and DF-TND by using different CNN model architectures, datasets, and various trigger patternsThe code is available at: https://github.com/wangren09/TrojanNetDetector.

Trojan settings. Testing models include VGG16 , ResNet-50 , and AlexNet . Datasets include CIFAR-10 , GTSRB , and Restricted ImageNet (R-ImgNet) (restricting ImageNet to 99 classes). We trained 85 TrojanNets and 85 clean networks, respectively. The numbers of different models are shown in Table 6. Fig. 1 (a)-(f) show the CIFAR-10 and GTSRB dataset with triggers of dot, cross, and triangle, respectively. One of these triggers is used for poisoning the model. We also test models poisoned for two target labels simultaneously: the dot trigger is used for one target label, and the cross trigger corresponds to the other target label. Fig. 1 (g)-(i) show poisoned ImageNet samples with the watermark as the trigger. The TrojanNets are various by specifying triggers with different shapes, colors, and positions. The data poisoning ratio also varies from 10%−12%10\%-12\%. The cleanNets are trained with different batches, iterations, and initialization. Table 7 summarizes test accuracies and attack success rates of our generated Trojan and cleanNets. We compare DL-TND with the baseline Neural Cleanse (NC) for detecting TrojanNets.

Detection performance. To build DL-TND, we use 55 validation data points for each class of CIFAR-10 and R-ImgNet, and 22 validation data points for each class of GTSRB. Following Sec. 3.2, we set I(k)I^{(k)} to quantile-0.250.25, median, quantile-0.750.75 and vary T1T_{1}. Let the true positive rate be the detection success rate for TrojanNets and the false negative rate be the detection error rate for cleanNets. Then the area under the curve (AUC) of receiver operating characteristics (ROC) can be used to measure the performance of the detection. Table 2 shows the AUC values, where “Total” refers to the collection of all models from different datasets.

We plot the ROC curve of the “Total” in Fig. 5. The results show that DL-TND can perform well across different datasets and model architectures. Moreover, fixing I(k)I^{(k)} as median, T1=0.54∼0.896T_{1}=0.54\sim 0.896 could provide a detection success rate over 76.5%76.5\% for TrojanNets and a detection success rate over 82%82\% for cleanNets. Table 3 shows the comparisons of DL-TND to Neural Cleanse (NC) on TrojanNets and cleanNets (T1=0.7T_{1}=0.7). Even using the MAD method as the detection rule, we find that DL-TND greatly outperforms NC in detection tasks of both TrojanNets and cleanNets (Note that NC also uses MAD). The results are shown in Table 8.

2 Data-free TrojanNet detector (DF-TND)

Trojan settings. The DF-TND is tested on cleanNets and TrojanNets that are trained under CIFAR-10 and R-ImgNet (with 10%10\% poisoning ratio unless otherwise stated). We perform the customized proximal gradient method shown in Sec. 3.4 to solve problem (12), where the number of iterations is set as 50005000.

Revealing Trojan trigger. Recall from Fig. 2 that the trigger pattern can be revealed by input perturbations that maximize neuron response of a TrojanNet. By contrast, the perturbations under the cleanNets behave like random noises. Fig. 6 provides visualizations of recovered inputs by neuron maximization at a TrojanNet versus a cleanNet on CIFAR-10 and ImageNet datasets. The key insight is that for a TrojanNet, it is easy to find an inverted image (namely, feature inversion) by maximizing neurons’ activation via (12) to reveal the Trojan characteristics (e.g., the shape of a Trojan trigger) compared to the activation from a cleanNet. Fig. 6 shows such results are robust to the choice of inputs (even for a noise input). We observe that the recovered triggers may have different colors and locations different from the original trigger. This is possibly because the trigger space has been shifted and enlarged by using convolution operations. In Figs. 12 and Fig. 13, we also provide additional experimental results for the sensitivity of trigger locations and sizes. Furthermore, we show some improvements of using the refine method in Fig. 14.

Detection performance. We now test 10001000 seed images on 1010 TrojanNets and 1010 cleanNets using DF-TND defined in Sec. 3.3. We compute AUC values of DF-TND by choosing seed images as clean validation inputs and random noise inputs, respectively.

Results are summarized in Table 4, and the ROC curves are shown in Fig. 15.

3 Additional results on DL-TND and DF-TND

First, we apply DL-TND and DF-TND on detecting TrojanNets with different levels attack success rate (ASR). We control ASR by choosing different data poisoning ratios when generating a TrojanNet. The results are summarized in Table 5. As we can see, our detectors can still achieve competitive performance when the attack likelihood becomes small, and DL-TND is better than DF-TND when ASR reaches 30%30\%.

Moreover, we conduct experiments when the number of TrojanNets is much less than the total number of models, e.g., only 55 out of 5555 models are poisoned. We find that the AUC value of the precision-recall curves are 0.970.97 and 0.960.96 for DL-TND and DF-TND, respectively. Similarly, the average AUC value of the ROC curves is 0.990.99 for both detectors.

Third, we evaluate our proposed DF-TND to detect TrojanNets generalized by clean-label Trojan attacks . We find that even in the least information case, DF-TND can still yield 0.920.92 AUC score when detecting 2020 TrojanNets from 4040 models.

Conclusion

Trojan attack injects a backdoor into DNNs during the training process, therefore leading to unreliable learning systems. Considering the practical scenarios where a detector is only capable of accessing to limited data information, this paper proposes two practical approaches to detect TrojanNets. We first propose a data-limited TrojanNet detector (DL-TND) that can detect TrojanNets with only a few data samples. The effectiveness of the DL-TND is achieved by drawing a connection between Trojan attack and prediction-evasion adversarial attacks including per-sample attack as well as all-sample universal attack. We find that both input perturbations obtained from per-sample attack and from universal attack exhibit Trojan behavior, and can thus be used to build a detection metric. We then propose a data-free TrojanNet detector (DF-TND), which leverages neuron response to detect Trojan attack, and can be implemented using random data samples and even random noise. We use the proximal gradient algorithm as a general optimization framework to learn DL-TND and DF-TND. The effectiveness of our proposals has been demonstrated by extensive experiments conducted under various datasets, Trojan attacks, and model architectures.

Acknowledgement

This work was supported by the Rensselaer-IBM AI Research Collaboration (http://airc.rpi.edu), part of the IBM AI Horizons Network (http://ibm.biz/AIHorizons). We would also like to extend our gratitude to the MIT-IBM Watson AI Lab (https://mitibmwatsonailab.mit.edu/) for the general support of computing resources.

References

Appendix 0.A Data-Limited TrojanNet Detector (DL-TND)

Visualization of neuron activation. DL-TND tests all the labels (classes) by calculating one universal perturbation and multiple per-image perturbations for each label. Each data sample can obtain a neuron activation vector with the universal perturbation and a neuron activation vector with its per-image perturbation. In Fig. 7, we show the neuron activation of five data samples with universal perturbations and per-image perturbations under a target label, a non-target label, and a label in a clean network (cleanNet). The output magnitude for each coordinate is represented using gray scale. One can see that the strong similarities only appear under the target label, which supports our motivation for the DL-TND.

Detection rule using median absolute deviation. Instead of using the detection rule in the main body of the paper, we can also employ the median absolute deviation (MAD) method. By MAD, if a single value in the kk-th position of ∣(I)1/2−I∣1.4826⋅∣(I)1/2−I∣1/2\frac{|(\mathbf{I})_{1/2}-\mathbf{I}|}{1.4826\cdot|(\mathbf{I})_{1/2}-\mathbf{I}|_{1/2}} is larger than 22 (provide 95%95\% confidence rate), the network is poisoned and label kk is a target label, where I=[I(1),I(2),⋯ ,I(K)]\mathbf{I}=[I^{(1)},I^{(2)},\cdots,I^{(K)}]. ∣⋅∣|\cdot| represents the absolute value. (⋅)1/2(\cdot)_{1/2} is the median of values in a vector. We compare DL-TND to Neural Cleanse (NC) in Table 8 using MAD as the detection rule.

Appendix 0.B Data-Free TrojanNet Detector (DF-TND)

Visualization of logits output increase. Fig. 9 and 9 visualize the change of the logits output of 10 data samples under a cleanNet and a TrojanNet when label 4 (lab4) is the target label. One can see that the minimum increase belonging to the target label is 600600 while the maximum increase for labels in the cleanNet is 1010. This large gap suggests that TrojanNets can be detected by properly selecting T2T_{2} and there exists a wide selection range, implying the stability of our method.

Appendix 0.C DL-TND: Additional Experiments

Models for testing. Table 6 shows the numbers of different models used for testing. Models have three different architectures and are applied to CIFAR-10, GTSRB, and R-ImgNet. We trained 8585 TrojanNets and 8585 cleanNets, respectively. In addition to the diversity of model architecture and dataset types, we also train TrojanNets with different triggers. Table 7 shows the smallest test accuracy and attack success rate for TrojanNets and cleanNets. TrojanNets can reach a similar test accuracy as cleanNets while still keeping the high attack success rate. This suggests that they are valid TrojanNets as defined in Sec. 2.1.

Applying median absolute deviation method as the detection rule. Table 8 provides the comparisons between DL-TND and NC method on Trojan and cleanNets using Median Absolute Deviation (MAD) as the detection rule. Even using the MAD method as the detection rule, we find that DL-TND greatly outperforms NC in detection tasks of both TrojanNets and cleanNets.

Varying number of data samples in each class. We also vary the number of validation data points for CIFAR-10 models and see the detection performance when we choose the quantile to be 0.50.5 (median). The number of data points in each class is chosen as 11, 22, and 55 and the corresponding AUC values are 0.96,0.980.96,0.98 and 11, respectively. We can see that the data-limited TrojanNet detector is effective even when only one data point is available for each class.

ROC curve for target label detection. Let the true positive rate be the detection success rate of target labels and the false negative rate be the detection error rate of cleanNets. Table 2 also shows the AUC values for target label detection, and Fig. 10 shows the ROC curves for target label detection using DL-TND. We set I(k)I^{(k)} to quantile-0.250.25, median, quantile-0.750.75 of the similarity values and vary T1T_{1}. Under the three different quantile selections, AUC values are all above 0.980.98.

Visualization of the universal perturbations. In Fig. 11, we show the universal perturbation obtained through (4). Due to the presence of backdoor in TrojanNets, universal perturbations can reveal common patterns with the real triggers, and this property is reflected in Fig. 11. Since DL-TND tries to find the smallest universal perturbation, the recovered perturbation pattern could be much more sparse than the Trojan trigger when the Trojan trigger is very complicated. This can be viewed in the perturbation pattern in the last two columns of Fig. 11.

Appendix 0.D DF-TND: Additional Experiments

The sensitivity to trigger locations and sizes. Fig. 12 and Fig. 13 provide the experimental results for the sensitivity to trigger locations and sizes. Fig. 12 shows that locations of perturbations vary when the locations of Trojan triggers vary. However, the recovered perturbations do not always have the same locations as the Trojan triggers. Patterns shifted and enlarged due to the convolution operations. Fig. 13 shows that DF-TND can recover the trigger pattern when the size of the Trojan trigger increases, and the area of the recovered perturbation increases when the size of the Trojan trigger increases.

Improvements using the refine method - maximizing the neuron activation corresponding to the Trojan-related coordinate. Note that once the recovered data is obtained from the optimization problem (12), one can find the coordinate related to Trojan feature by checking the largest neuron activation value (or the largest weight) among all the coordinates. Then maximizing the output of the Trojan-related coordinate separately could provide a better result. Fig. 14 shows the improvements using DF-TND together with our refine method - maximizing the neuron activation corresponding to the Trojan-related coordinate. The refine method can increase the logits output belonging to the target label, while decrease the logits outputs belonging to the non-target labels simultaneously.

ROC curves for TrojanNet detection with clean validation inputs and random noise inputs. Fig. 15 (a) and (b) show the ROC curves for TrojanNet detection with clean validation inputs and random noise inputs, respectively. The true positive rate is the detection success rate for TrojanNets and the false negative rate is the detection error rate for cleanNets. In both cases, DF-TND can reach nearly perfect AUC values 0.990.99. T2=55−400T_{2}=55-400 could provide a detection success rate of more than 85%85\% for TrojanNets and a detection success rate of over 90%90\% for cleanNets.

Recovered perturbations under different λ\lambda.

In Fig. 16 and Fig. 17, we vary the sparsity penalty parameter λ\lambda and obtain perturbations under a TrojanNet and a cleanNet. One can find that the trigger pattern appears in the perturbations under the TrojanNet. The perturbations under cleanNet behave like random noises. Another discovery is that the perturbations become more and more sparse when λ\lambda increases.

This method also works for random noise inputs. Fig. 17 shows the original noise images, trigger, perturbations under poisoned model, and perturbations under cleanNet. The same behaviors as the clean inputs are observed.