Efficient GAN-Based Anomaly Detection

Houssam Zenati, Chuan Sheng Foo, Bruno Lecouat, Gaurav Manek, Vijay Ramaseshan Chandrasekhar

Introduction

Anomaly detection is one of the most important problems across a range of domains, including manufacturing (Martí et al., 2015), medical imaging and cyber-security (Schubert et al., 2014). Fundamentally, anomaly detection methods need to model the distribution of normal data, which can be complex and high-dimensional. Generative adversarial networks (GANs) (Goodfellow et al., 2014) are one class of models that have been successfully used to model such complex and high-dimensional distributions, particularly over natural images (Radford et al., 2016).

Intuitively, a GAN that has been well-trained to fit the distribution of normal samples should be able to reconstruct such a normal sample from a certain latent representation and also discriminate the sample as coming from the true data distribution. However, as GANs only implicitly model the data distribution, using them for anomaly detection necessitates a costly optimization procedure to recover the latent representation of a given input example, making this an impractical approach for large datasets or real-time applications.

In this work, we leverage recently developed GAN methods that simultaneously learn an encoder during training (Vincent Dumoulin & Courville, 2017; Donahue et al., 2017) to develop an anomaly detection method that is efficient at test time.We apply our method to an image dataset (MNIST) (LeCun et al., 1998) and a network intrusion dataset (KDD99 10percent) (Lichman, 2013) and show that it is highly competitive with other approaches. To the best of our knowledge, our method is the first GAN-based approach for anomaly detection which achieves state-of-the-art results on the KDD99 dataset. An implementation of our methods and experiments is provided at https://github.com/houssamzenati/Efficient-GAN-Anomaly-Detection.git.

Related Work

Anomaly detection has been extensively studied, as surveyed in (Chandola et al., 2009). Popular techniques utilize clustering approaches or nearest neighbor methods (Xiong et al., 2011; Zimek et al., 2012) and one class classification approaches that learn a discriminative boundary around normal data, such as one-class SVMs (Yunqiang Chen & Huang, 2001). Another class of methods uses fidelity of reconstruction to determine whether an example is anomalous, and includes Principal Component Analysis (PCA) and its kernel and robust variants (Jolliffe, 1986; S. Günter & Vishwanathan, 2007; Candès et al., 2009). More recent works use deep neural networks, which do not require explicit feature construction unlike the previously mentioned methods. Autoencoders, variational autoencoders (An & Cho, 2015; Zhou & Paffenroth, 2017), energy based models (Zhai et al., 2016) and deep autoencoding Gaussian mixture models (Bo Zong, 2018) have been explored for anomaly detection. Aside from AnoGAN (Schlegl et al., 2017), however, the use of GANs for anomaly detection has been relatively unexplored, even though GANs are suited to model the high-dimensional complex distributions of real-world data (Creswell et al., 2017).

Efficient Anomaly detection with GANs

Our models are based on recently developed GAN methods (Donahue et al., 2017; Vincent Dumoulin & Courville, 2017) (specifically BiGAN), and simultaneously learn an encoder EE that maps input samples xx to a latent representation zz, along with a generator GG and discriminator DD during training; this enables us to avoid the computationally expensive step of recovering a latent representation at test time. Unlike in a regular GAN where the discriminator only considers inputs (real or generated), the discriminator DD in this context also considers the latent representation (either a generator input or from the encoder).

Vincent Dumoulin & Courville (2017) explored different training strategies to learn an encoder such that E=G−1E=G^{-1}, and emphasized the importance of learning EE jointly with GG. We therefore adopted a similar strategy, solving the following optimization problem during training: min⁡G,Emax⁡DV(D,E,G)\min_{G,E}\max_{D}V(D,E,G), with V(D,E,G)V(D,E,G) defined as

Here, pX(x)p_{X}(x) is the distribution over the data, pZ(z)p_{Z}(z) the distribution over the latent representation, and pE(z∣x)p_{E}(z|x) and pG(x∣z)p_{G}(x|z) the distributions induced by the encoder and generator respectively.

Having trained a model on the normal data to yield G,DG,D and EE, we then define a score function A(x)A(x) (as in Schlegl et al. (2017)) that measures how anomalous an example xx is, based on a convex combination of a reconstruction loss LGL_{G} and a discriminator-based loss LDL_{D}:

where LG(x)=∣∣x−G(E(x))∣∣1L_{G}(x)=\left|\left|x-G(E(x))\right|\right|_{1} and LD(x)L_{D}(x) can be defined in two ways. First, using the cross-entropy loss σ\sigma from the discriminator of xx being a real example (class 1): LD(x)=σ(D(x,E(x)),1)L_{D}(x)=\sigma(D(x,E(x)),1), which captures the discriminator’s confidence that a sample is derived from the real data distribution. A second way of defining the LDL_{D} is with a “feature-matching loss”LD(x)=∣∣fD(x,E(x))−fD(G(E(x)),E(x))∣∣1L_{D}(x)=\left|\left|f_{D}(x,E(x))-f_{D}(G(E(x)),E(x))\right|\right|_{1}, with fDf_{D} returning the layer preceding the logits for the given inputs in the discriminator. This evaluates if the reconstructed data has similar features in the discriminator as the true sample. Samples with larger values of A(x)A(x) are deemed more likely to be anomalous.

Experiments

We provide details of the experimental setup and network architectures used in the Appendix.

MNIST: We generated 10 different datasets from MNIST by successively making each digit class an anomaly and treating the remaining 9 digits as normal examples. The training set consists of 80% of the normal data and the test set consists of the remaining 20% of normal data and all of the anomalous data. All models were trained only with normal data and tested with both normal and anomalous data. As the dataset is imbalanced, we compared models using the area under the precision-recall curve (AUPRC). We evaluated our method against AnoGAN and the variational auto-encoder (VAE) of (An & Cho, 2015). We see that our model significantly outperforms the VAE baseline. Likewise, our model outperforms AnoGAN (Figure 1), but with approximately 800x faster inference time (Table 2). We also observed that the feature-matching variant of LDL_{D} used in the anomaly score performs better than the cross-entropy variant, which was also reported in Schlegl et al. (2017), suggesting that the features extracted by the discriminator are informative for anomaly detection.

KDD99: We evaluated our method on this network activity dataset to show that GANs can also perform well on high-dimensional, non-image data. We follow the experimental setup of (Zhai et al., 2016; Bo Zong, 2018) on the KDDCUP99 10 percent dataset (Lichman, 2013). Due to the proportion of outliers in the dataset “normal” data are treated as anomalies in this task. The 20% of samples with the highest anomaly scores A(x)A(x) are classified as anomalies (positive class), and we evaluated precision, recall, and F1-score accordingly. For training, we randomly sampled 50% of the whole dataset and the remaining 50% of the dataset was used for testing. Then, only data samples from the normal class were used for training models, therefore, all anomalous samples were removed from the split training set. Our method is overall highly competitive with other state-of-the art methods and achieves higher recall. Again, our model outperforms AnoGAN and also has a 700x to 900x faster inference time (Table 2).

Conclusion

We demonstrated that recent GAN models can be used to achieve state-of-the-art performance for anomaly detection on high-dimensional, complex datasets whilst being efficient at test time; our use of a GAN that simultaneously learns an encoder eliminates the need for a costly procedure to recover the latent representation for a given input. In future work, we plan to perform a more extensive evaluation of our method, evaluate other training strategies, as well as to explore the effects of encoder accuracy on anomaly detection performance.

The authors would like to thank the Agency for Science, Technology and Research (A*STAR), Singapore for supporting this research with scholarships to Houssam Zenati and Bruno Lecouat. The authors also would like to thank Yasin Yazici for the fruitful discussions.

References

Appendix

Appendix A Experiment details

We implemented AnoGAN and our BiGAN-based method in Tensorflow, and used the same hyper-parameters as the original AnoGAN paper, using α=0.9\alpha=0.9 in the anomaly score A(x)A(x) (both our model and AnoGAN) and running SGD for 500 iterations for AnoGAN. We attempted to match the AnoGAN and BiGAN architectures and learning hyperparameters as far as possible to enable a fair comparison. An exponential moving average of model parameters was used at test time, with a decay of 0.999 for the MNIST dataset and a decay of 0.9999 for the KDD99 dataset.

Appendix B Inference Time Details

MNIST experiments were run on NVIDIA GeForce TitanX GPUs with Tensorflow 1.1.0 (python 3.4.3). KDD experiments were run on NVIDIA Tesla K40 GPUs with Tensorflow 1.1.0 (python 3.5.3).

Appendix C MNIST Experiments Details

Preprocessing: Pixels were scaled to be in range . The outputs of the starred layer in the discriminator were used for the FM scoring variant.

Appendix D KDD99 Experiment Details

Preprocessing: The dataset contains samples of 41 dimensions, where 34 of them are continuous and 7 are categorical. For categorical features, we further used one-hot representation to encode them; we obtained a total of 121 features after this encoding. Then, we applied min-max scaling to derive the final features. The outputs of the starred layer in the discriminator were used for the FM scoring variant.