Old is Gold: Redefining the Adversarially Learned One-Class Classifier Training Paradigm
Muhammad Zaigham Zaheer, Jin-ha Lee, Marcella Astrid, Seung-Ik Lee
Introduction
Due to rare occurrence of anomalous scenes, the anomaly detection problem is usually seen as one-class classification (OCC) in which only normal data is used to learn a novelty detection model liu2018future_novelty; zhang2016video_novelty; luo2017revisit_novelty; xia2015learning_novelty_fig5; zaheer2018ensemble; hinami2017joint_novelty; sultani2018real_novelty; sabokrou2017deep_novelty; hasan2016learning_novelty; smeureanu2017deep_novelty; ravanbakhsh2018plug_novelty; ravanbakhsh2017abnormal_novelty. One of the recent trends to learn one-class data is by using an encoder-decoder architecture such as denoising auto-encoder sabokrou2018adversarially_alocc; xu2015learning_denoise; xu2017detecting_denoise; vincent2008extracting_denoise. Generally, in this scheme, training is carried out until the model starts to produce good quality reconstructions sabokrou2018adversarially_alocc; Shama_2019_ICCV_good. During the test time, it is expected to show high reconstruction loss for abnormal data which corresponds to a high anomaly score. With the recent developments in Generative Adversarial Networks (GANs) goodfellow2014generative_gan, some researchers also explored the possibility of improving the generative results using adversarial training Shama_2019_ICCV_good; radford2015unsupervised_good_adversarial. Such training fashion substantially enhances the data regeneration quality pathak2016context_adversarial_good; goodfellow2014generative_gan; Shama_2019_ICCV_good. At the test time, the trained generator is then decoupled from the discriminator to be used as a reconstruction model. As reported in sabokrou2018adversarially_alocc; ravanbakhsh2017abnormal_novelty; ravanbakhsh2019training, a wide difference in reconstruction loss between the normal and abnormal data can be achieved due to adversarial training, which results in a better anomaly detection system. However, relying only on the reconstruction capability of a generator does not oftentimes work well because the usual encoder-decoder style generators may unexpectedly well-reconstruct the unseen data which drastically degrades the anomaly detection performance.
A natural drift in this domain is towards the idea of using along with the conventional utilization of for anomaly detection. The intuition is to gain maximum benefits of the one-class adversarial training by utilizing both and instead of only . However, this also brings along the problems commonly associated with such architectures. For example, defining a criteria to stop the training is still a challenging problem goodfellow2014generative_gan; pathak2016context_adversarial_good. As discussed in Sabokrou et al. sabokrou2018adversarially_alocc, the performance of such adversarially learnt one-class classification architecture is highly dependent on the criteria of when to halt the training. In the case of stopping prematurely, will be undertrained and in the case of overtraining, may get confused because of the real-looking fake data. Our experiments show that a + trained as a collective model for anomaly detection (referred to as a baseline) will not ensure higher convergence at any arbitrary training step over its predecessor. Figure 1 shows frame-level area under the curve (AUC) performance of the baseline over several epochs of training on UCSD Ped2 dataset chan2008ucsd. Although we get high performance peaks at times, it can be seen that the performance fluctuates substantially even between two arbitrary consecutive epochs. Based on these findings, it can be argued that a as we know it, may not be a suitable choice in a one-class classification problem, such as anomaly detection.
Following this intuition, we devise an approach for training of an adversarial network towards anomaly detection by transforming the basic role of from distinguishing between real and fake to identifying good and bad quality reconstructions. This property of is highly desirable in anomaly detection because a trained would not produce as good reconstruction for abnormal data as it would for the normal data conforming to the learned representations. To this end we propose a two-stage training process. Phase one is identical to the common practice of training an adversarial denoising auto-encoder pathak2016context_adversarial_good; xu2017detecting_denoise; vincent2008extracting_denoise. Once achieves a reasonably trained state (i.e. showing low reconstruction losses), we begin phase two in which is optimized by training on various good quality and bad quality reconstruction examples. Good quality reconstruction examples come from real data as well as the data regenerated by , whereas bad reconstruction examples are obtained by utilizing an old state of the generator () as well as by using our proposed pseudo-anomaly module. Shown in Figure 3, this pseudo-anomaly module makes use of the training data to create anomaly-like examples. With this two-phase training process, we expect to be trained in such a way that it can robustly discriminate reconstructions coming from normal and abnormal data. As shown in Figure 1, our model not only provides superior performance but also shows stability across several training epochs.
In summary, the contributions of our paper are as follows: 1) this work is among the first few to employ along with at test time for anomaly detection. Moreover, to the best of our knowledge, it is the first one to extensively report the impacts of using the conventional formulation and the consequent instability. 2) Our approach of transforming the role of a discriminator towards anomaly detection problem by utilizing an old state of the generator along with the proposed pseudo-anomaly module, substantially improves stability of the system. Detailed analysis provided in this paper shows that our model is independent of a hard stopping criteria and achieves consistent results over a wide range of training epochs. 3) Our method outperforms state-of-the-art sabokrou2018adversarially_alocc; ionescu2019object; ren2015unsupervised; Nguyen_2019_ICCV; Gong_2019_ICCV; tudor2017unmasking_novelty; luo2017revisit_novelty; nguyen2019hybrid; hinami2017joint_novelty; liu2018classifier_novelty; ravanbakhsh2017abnormal_novelty; luo2017remembering; ravanbakhsh2018plug_novelty; sun2017online; hasan2016learning_novelty; liu2018future_novelty; xu2015learning_denoise; zhang2016video_novelty; zhao2017spatio in the experiments conducted on MNIST mnist and Caltech-256 griffin2007caltech datasets for novelty detection as well as on UCSD Ped2 chan2008ucsd video dataset for anomaly detection. Moreover, on the latter dataset, our approach provides a substantial absolute gain of 5.2% over the baseline method achieving frame level AUC of 98.1%.
Related Work
Anomaly detection is often seen as a novelty detection problem liu2018future_novelty; zhang2016video_novelty; luo2017revisit_novelty; hinami2017joint_novelty; xia2015learning_novelty_fig5; sultani2018real_novelty; sabokrou2017deep_novelty; bergadano2019keyed; hasan2016learning_novelty; smeureanu2017deep_novelty; ravanbakhsh2018plug_novelty; ravanbakhsh2018plug_novelty; ravanbakhsh2017abnormal_novelty in which a model is trained based on the known normal class to ultimately detect unknown outliers as abnormal. To simplify the task, some works proposed to use object tracking wang2014learning_realworld38; basharat2008learning_realworld7; medioni2001event_twostream34; piciarelli2008trajectory_twostream36; zhang2009learning_twostream53 or motion kratz2009anomaly_realworld26; hou2017tube_realworld20; cui2011abnormal_realworld10. Handpicking features in such a way can often deteriorate the performance significantly. With the increased popularity of deep learning, some researchers smeureanu2017deep_novelty; ravanbakhsh2017abnormal_novelty also proposed to use pre-trained convolution network based features to train one-class classifiers. Success of such methods is highly dependent on the base model which is often trained on some unrelated datasets.
A relatively new addition to the field, image regeneration based works Gong_2019_ICCV; ren2015unsupervised; xu2015learning_denoise; ionescu2019object; Nguyen_2019_ICCV; nguyen2019hybrid; xu2017detecting_denoise; sabokrou2017deep_novelty are the ones that make use of a generative network to learn features in an unsupervised way. Ionescu et al. in ionescu2019object proposed to use convolutional auto-encoders on top of object detection to learn motion and appearance representations. Xu et al. xu2015learning_denoise; xu2017detecting_denoise used a one-class SVM learned using features from stacked auto-encoders. Ravanbakhsh et al. ravanbakhsh2017abnormal_novelty used generator as a reconstructor to detect abnormal events assuming that a generator is unable to reconstruct the inputs that do not conform the normal training data. In Nguyen_2019_ICCV; nguyen2019hybrid, the authors suggested to use a cascaded decoder to learn motion as well as appearance from normal videos. However, in all these schemes, only a generator is employed to perform detection. Pathak et al. pathak2016context_adversarial_good proposed adversarial training to enhance the quality of regeneration. However, they also discard the discriminator once the training is finished. A unified generator and discriminator model for anomaly detection is proposed in Sabokrou et al. sabokrou2018adversarially_alocc. The model shows promising results, however it is often not stable and the performance relies heavily on the criteria to stop training. Recently, Shama et al. Shama_2019_ICCV_good proposed an idea of utilizing output of an adversarial discriminator to increase the image quality of generated images. Although not related to anomaly detection, it provides an interesting intuition to make use of both adversarial components for an enhanced performance.
Our work, although built on top of an unsupervised generative network, is different from the approaches in Gong_2019_ICCV; ionescu2019object; Nguyen_2019_ICCV; nguyen2019hybrid; xu2017detecting_denoise; sabokrou2017deep_novelty as we explore to utilize the unified generator and discriminator model for anomaly detection. The most similar work to ours is by Sabokrou et al. sabokrou2018adversarially_alocc and Lee et al. lee2018stan as they also explore the possibility of using discriminator, along with the conventional usage of generator, for anomaly detection. However, our approach is substantially different from these. In sabokrou2018adversarially_alocc, a conventional adversarial network is trained based on a criteria to stop the training whereas, in lee2018stan, an LSTM based approach is utilized for training. In contrast, we utilize a pseudo-anomaly module along with an old state of the generator, to modify the ultimate role of a discriminator from distinguishing between real and fake to detecting between and bad quality reconstructions. This way, our overall framework, although trained adversarially in the beginning, finally aligns both the generator and the discriminator to complement each other towards anomaly detection.
Method
In this section, we present our OGNet framework. As described in Section 1, most of the existing GANs based anomaly detection approaches completely discard discriminator at test time and use generator only. Furthermore, even if both models are used, the unavailability of a criteria to stop the training coupled with the instability over training epochs caused by adversary makes the convergence uncertain. We aim to change that by redefining the role of a discriminator to make it more suitable for anomaly detection problems. Our solution is generic, hence it can be integrated with any existing one-class adversarial networks.
In order to maintain consistency and to have a fair comparison, we kept our baseline architecture similar to the one proposed by Sabokrou et al. sabokrou2018adversarially_alocc. The generator , a typical denoising auto-encoder, is coupled with the discriminator to learn one class data in an unsupervised adversarial fashion. The goal of this model is to play a min-max game to optimize the following objective function:
2 Training
The training of our model is carried out in two phases (see Figure 2). Phase one is similar to the common practices in training an adversarial one-class classifier sabokrou2018adversarially_alocc; lawson2017finding; schlegl2017unsupervised; ravanbakhsh2019training. tries to regenerate real-looking fake data which is then fed into along with real data. The learns to discriminate between real and fake data, success or failure of which then becomes a supervision signal for . This training is carried out until starts to create real looking images with a reasonably low reconstruction loss. Overall, phase one minimizes the following loss function:
Phase two of the training is where we make use of the frozen models and to update . This way starts learning to discriminate between good and bad quality reconstructions, hence becoming suitable for one-class classification problems such as anomaly detection.
Details of the phase two training are discussed next:
Goal. The essence of phase two training is to provide examples of good quality and bad quality reconstructions to , with a purpose of making it learn about the kind of output that would produce in the case of an unusual input. The training is performed for just a few iterations since the already trained converges quickly. A detailed study on this is added in Section 4.
Good quality examples. is provided with real data (), which is the best possible case of reconstruction, and the actual high quality reconstructed data () produced by the trained as an example of good quality examples.
Bad quality examples. Examples of low quality reconstruction () are generated using . In addition, a pseudo-anomaly module, shown in Figure 3, is formulated with a combination of and the trained , which simulates examples of reconstructed pseudo-anomalies ().
Pseudo anomaly creation. Given two arbitrary images and from the training dataset, a pseudo anomaly image is generated as:
This way, the resultant image can contain diverse variations such as shadows and unusual shapes, which are completely unknown to both and models. Finally, as the last step in our pseudo-anomaly module, in order to mimic the behavior of when it gets unusual data as input, is then reconstructed using to obtain :
Example images at each intermediate step can be seen in Figures 3 and 4.
Tweaking the objective function. The model in phase two of the training takes the form:
where and are the trade-off hyperparameters.
Quasi ground truth for the discriminator in phase one training is defined as:
However, for phase two training, it takes the form:
3 Testing
At test time, as shown in Figure 2, only and are utilized for one-class classification (OCC). Final classification decision for an input image is given as:
Experiments
The evaluation of our OGNet framework on three different datasets is reported in this section. Detailed analysis of the performance and its comparison with the state-of-the-art methodologies is also reported. In addition, we provide extensive discussion and ablation studies to show the stability as well as the significance of our proposed scheme. In order to keep the experimental setup consistent with the existing works liu2018future_novelty; zhang2016video_novelty; luo2017revisit_novelty; xia2015learning_novelty_fig5; hinami2017joint_novelty; sultani2018real_novelty; sabokrou2017deep_novelty; hasan2016learning_novelty; smeureanu2017deep_novelty; ravanbakhsh2018plug_novelty; ravanbakhsh2017abnormal_novelty; sabokrou2018adversarially_alocc; ionescu2019object; Gong_2019_ICCV; nguyen2019hybrid; Nguyen_2019_ICCV, we tested our method for the detection of outlier images as well as video anomalies.
Evaluation criteria. Most of our results are formulated based on area under the curve (AUC) computed at frame level due to its popularity in related works tudor2017unmasking_novelty; luo2017revisit_novelty; nguyen2019hybrid; hinami2017joint_novelty; liu2018classifier_novelty; ravanbakhsh2017abnormal_novelty; luo2017remembering; Gong_2019_ICCV; ravanbakhsh2018plug_novelty; sun2017online; hasan2016learning_novelty; liu2018future_novelty; xu2015learning_denoise; Nguyen_2019_ICCV; zhang2016video_novelty; ionescu2019object; zhao2017spatio; sabokrou2018adversarially_alocc. Nevertheless, following the evaluation methods adopted in tsakiris2015dual; lerman2015robust; xu2010robust; rahmani2017coherence; liu2010robust; you2017provable_novelty; sabokrou2018adversarially_alocc; sabokrou2016video; ravanbakhsh2017abnormal_novelty; ravanbakhsh2019training; xu2015learning_denoise; sabokrou2015real; sabokrou2017deep_novelty we also report score and Equal Error Rate (EER) of our approach.
Parameters and implementation details. The implementation is done in PyTorch paszke2017automatic and the source code is provided at https://github.com/xaggi/OGNet. Phase one of the training in our reports is performed from 20 to 30 epochs. These numbers are chosen because the baseline shows high performance peaks within this range (Figure 1). We train on Adam kingma2014adam with the learning rate of generator and discriminator in all these epochs set to and , respectively. Phase two of the training is done for iterations with the learning rate of the discriminator reduced to half. are set to 0.2, 0.1, 0.001, respectively. Until stated otherwise, default settings of our experiments are set to the aforementioned values. However, for the detailed evaluation provided in a later part of this section, we also conducted experiments and reported results on a range of epochs and iterations for both phases of the training, respectively. Furthermore, until specified otherwise, we pick the generator after epoch and freeze it as . This selection is arbitrary and solely based on the intuition explained in Section 3. Additionally, in a later part of this section, we also present a robust and generic method to formulate without any need of handpicking an epoch.
Caltech-256. This dataset griffin2007caltech contains a total of 30,607 images belong to 256 object classes and one ‘clutter’ class. Each category has different number of images, as low as 80 and as high as 827. In order to perform our experiments, we used the same setup as described in previous works tsakiris2015dual; lerman2015robust; xu2010robust; rahmani2017coherence; liu2010robust; you2017provable_novelty; sabokrou2018adversarially_alocc. In a series of three experiments, at most 150 images belong to 1, 3, and 5 randomly chosen classes are defined as training (inlier) data. Outlier images for test are taken from the ‘clutter’ class in such a way that each experiment has exactly 50% ratio of outliers and inliers.
MNIST. This dataset mnist consists of 60,000 handwritten digits from 0 to 9. The setup to evaluate our method on this dataset is also kept consistent with the previous works xia2015learning_novelty_fig5; breunig2000lof_fig5; sabokrou2018adversarially_alocc. In a series of experiments, each category of digits is individually taken as inliers. Whereas, randomly sampled images of the other categories with a proportion of 10% to 50% are taken as outliers.
USCD Ped2. This dataset chan2008ucsd comprises of 2,550 frames in 16 training and 2,010 frames in 12 test videos. Each frame is of pixels resolution. Pedestrians dominate most of the frames whereas anomalies include skateboards, vehicles, bicycles, etc. Similar to tudor2017unmasking_novelty; luo2017revisit_novelty; nguyen2019hybrid; hinami2017joint_novelty; liu2018classifier_novelty; ravanbakhsh2017abnormal_novelty; luo2017remembering; Gong_2019_ICCV; ravanbakhsh2018plug_novelty; sun2017online; hasan2016learning_novelty; liu2018future_novelty; xu2015learning_denoise; Nguyen_2019_ICCV; zhang2016video_novelty; ionescu2019object; zhao2017spatio; sabokrou2018adversarially_alocc, frame-level AUC and EER metrics are adopted to evaluate performance on this dataset.
2 Outlier Detection in Images
One of the significant applications of a one-class learning algorithm is outlier detection. In this problem, objects belonging to known classes are treated as inliers based on which the model is trained. Other objects that do not belong to these classes are treated as outliers, which the model is supposed to detect based on its training. Results of the experiments conducted using Caltech-256 griffin2007caltech and MNIST mnist datasets are reported and comparisons with state-of-the-art outlier detection models kim2009observe; you2017provable_novelty; xia2015learning_novelty_fig5; sabokrou2018adversarially_alocc; tsakiris2015dual; lerman2015robust; xu2010robust; rahmani2017coherence; liu2010robust are provided.
Results on Caltech-256. Figure 4b shows outlier examples reconstructed using . It is interesting to observe that although the generated images are of reasonably good quality, our model still depicts superior results in terms of score and area under the curve (AUC), as listed in Table 1, which demonstrates that our model is robust to the over-training of .
Results on MNIST. As it is a well-studied dataset, various outlier detection related works use MNIST as a stepping-stone to evaluate their approaches. Following xia2015learning_novelty_fig5; breunig2000lof_fig5; sabokrou2018adversarially_alocc, we also report score as an evaluation metric of our method on this dataset. A comparison provided in Figure 5 shows that our approach performs robustly to detect outliers even when the percentage of outliers is increased. An insight of the performance improvement by our approach is shown in Figure 6. It can be observed that as the phase two training continues, score distribution of inliers and outliers output by our network smoothly distributes to a wider range.
3 Anomaly Detection in Videos
One-class classifiers are finding their best applications in the domain of anomaly detection for surveillance purposes sultani2018real_novelty; tudor2017unmasking_novelty; dutta2015online; zhang2016video_novelty; ravanbakhsh2019training; ravanbakhsh2017abnormal_novelty. However, this task is more complicated than the outlier detection because of the involvement of moving objects, which cause variations in appearance.
Experimental setup. Each frame of the Ped2 dataset is divided into grayscale patches of size pixels. Normal videos, which only contain scenes of walking pedestrians, are used to extract training patches. Test patches are extracted from abnormal videos which contain abnormal as well as normal scenes. In order to remove unnecessary inference of the patches, a motion detection criteria based on frame difference is set to discard patches without motion. A maximum of all patch-level anomaly scores is declared as the frame-level anomaly score of that particular frame as:
Performance evaluation. Frame-level AUC and EER are the two evaluation metrics used to compare our approach with a series of existing works tudor2017unmasking_novelty; luo2017revisit_novelty; nguyen2019hybrid; hinami2017joint_novelty; liu2018classifier_novelty; ravanbakhsh2017abnormal_novelty; luo2017remembering; Gong_2019_ICCV; ravanbakhsh2018plug_novelty; sun2017online; hasan2016learning_novelty; liu2018future_novelty; xu2015learning_denoise; Nguyen_2019_ICCV; zhang2016video_novelty; ionescu2019object; zhao2017spatio; sabokrou2018adversarially_alocc published within last 5 years. The corresponding results provided in Table 2 and Table 3 show that our method outperforms recent state-of-the-art methodologies in the task of anomaly detection. Comparing with the baseline, our approach achieves an absolute gain of 5.2% in terms of AUC. Examples of the reconstructed patches are provided in Figure 4. As shown in Figure 4b, although generates noticeably good reconstructions of anomalous inputs, due to the presence of our proposed pseudo-anomaly module, gets to learn the underlying patterns of reconstructed anomalous images. This is why, in contrast to the baseline, our framework provides consistent performance across a wide range of training epochs (Figure 1).
4 Discussion
When to stop phase one training? The convergence of our framework is not strictly dependent on phase one training. Figure 7 shows the AUC performance of phase two training applied after various epochs of phase one on Ped2 dataset chan2008ucsd. Values plotted at iterations = 0, representing the performance of the baseline, show a high variance. Interestingly, it can be seen that after few iterations into phase two training of our proposed approach, the model starts to converge better. Irrespective of the initial epoch in phase one training, models converged successfully showing consistent AUC performances.
When to stop phase two training? As seen in Figure 7 and Figure 8, it can be observed that once a specific model is converged, further iterations do not deteriorate its performance. Hence, a model can be trained for any number of iterations as deemed necessary.
Which low epoch generator is better? For the selection of , as mentioned earlier, the generator after the epoch of training was arbitrarily chosen in our experiments. This selection is intuitive and mostly based on the fact that the generator has seen all dataset once. In addition, we visually observed that after first epoch, although the generator was capable of reconstructing its input, the quality was not ‘good enough’, which is a suitable property for in our model. However, this way of selection is not a generalized solution across various datasets. Hence, to investigate the matter further, we evaluate a range of low epoch numbers as candidates for . The baseline epoch of is kept fixed throughout this experiment. Results in Figure 8a show that irrespective of the low epoch number chosen as , the model converges and achieves state-of-the-art or comparable AUC. In pursuit of another more systematic way to obtain , we also explored the possibility of using average parameters of all previous models. Hence, for each given epoch of the baseline that we pick as , a is formulated by taking an average of all previous models until that point. The results plotted in Figure 8b show that such also depicts comparable performances. Note that this formulation completely eradicates the need of handpicking a specific epoch number for , thus making our formulation generic towards the size of a training dataset.
5 Ablation
Ablation results of our framework on UCSD Ped2 dataset chan2008ucsd are summarized in Table 4. As shown, while each input component of our training model (i.e. real images , high quality reconstructions , low quality reconstructions , and pseudo anomaly reconstructions ) contributes towards a robust training, removing any of these at a time still shows better performance than the baseline. One interesting observation can be seen in the fourth column of the phase two training results. In this case, the performance is measured after we remove the last step of pseudo-anomaly module, which is responsible for providing regenerated pseudo-anomaly () through , as in Equation 4. Hence, by removing this part, the fake anomalies () obtained using Equation 3 are channeled directly to the discriminator as one of the two sets of bad reconstruction examples. With this configuration, the performance deteriorates significantly (i.e. 9.6% drop in the AUC). The model shows even worse performance than the baseline after phase one training. This shows the significance of our proposed pseudo-anomaly module. Once pseudo-anomalies are created within the module, it is necessary to obtain a regeneration result of these by inferring . This helps to learn the underlying patterns of reconstructed anomalous images, which results in a more robust anomaly detection model.
Conclusion
This paper presents an adversarially learned approach in which both the generator () and the discriminator () are utilized to perform a stable and robust anomaly detection. A unified and model employed towards such problems often produces unstable results due to the adversary. However, we attempted to tweak the basic role of the discriminator from distinguishing between real and fake to discriminating between good and bad quality reconstructions, a formulation that aligns well with the philosophy of conventional anomaly detection using generative networks. We also propose a pseudo-anomaly module which is employed to create fake anomaly examples from normal training data. These fake anomaly examples help to learn about the behavior of in the case of unusual input data.
Our extensive experimentation shows that the approach not only generates stable results across a wide range of training epochs but also outperforms a series of state-of-the-art methods tudor2017unmasking_novelty; luo2017revisit_novelty; nguyen2019hybrid; hinami2017joint_novelty; liu2018classifier_novelty; ravanbakhsh2017abnormal_novelty; luo2017remembering; Gong_2019_ICCV; ravanbakhsh2018plug_novelty; sun2017online; hasan2016learning_novelty; liu2018future_novelty; xu2015learning_denoise; Nguyen_2019_ICCV; zhang2016video_novelty; ionescu2019object; zhao2017spatio; sabokrou2018adversarially_alocc for outliers and anomaly detection.
Acknowledgment
This work was supported by the ICT R&D program of MSIP/IITP. [2017-0-00306, Development of Multimodal Sensor-based Intelligent Systems for Outdoor Surveillance Robots]. Also, we thank HoChul Shin, Ki-In Na, Hamza Saleem, Ayesha Zaheer, Arif Mahmood, and Shah Nawaz for the discussions and support in improving our work.