Understanding Anomaly Detection with Deep Invertible Networks through Hierarchies of Distributions and Features

Robin Tibor Schirrmeister, Yuxuan Zhou, Tonio Ball, Dan Zhang

Introduction

One line of work for anomaly detection - to detect if a given input is from the same distribution as the training data - uses the likelihoods provided by generative models. Through likelihood maximization, they are trained to yield high likelihoods on the in-distribution inputs (a.k.a. inliers). We note that likelihood is a function of the model given the data. As the model parameters are trained to model the data density, we may abuse likelihood and density in the paper for simplicity. After training, one may expect out-of-distribution inputs (a.k.a. outliers) to have lower likelihoods than the inliers. However, this is often not the case. For example, Nalisnick et al. showed that generative models trained on CIFAR10 assign higher likelihoods to SVHN than to CIFAR10 images.

Several works have investigated a potential reason for this failure: The image likelihoods of deep generative networks can be well-predicted from simple factors. For example, the deep generative networks’ image likelihoods highly correlate with: the image encoding sizes from a lossless compressor such as PNG ; background statistics, e.g., the number of zeros in Fashion-MNIST/MNIST images ; smoothness and size of the background . These factors do not directly correspond to the type of object, hence the type of object does not affect the likelihood much.

In this work, we first synthesize these findings into the following hypothesis: A convolutional deep generative network trained on any image dataset learns low-level local feature relationships common to all images - such as smooth local patches - and these local features, forming the domain prior, dominate the likelihood. One can therefore expect a smoother dataset like SVHN to have higher likelihoods than a less smooth one like CIFAR10, irrespective of the image dataset the network was trained on. Following prior works, we take Glow networks as the baseline model for our study.

Next, we report several new findings to support the hypothesis: (1) Using a fully-connected instead of a convolutional Glow network, likelihood-based anomaly detection works much better for Fashion-MNIST vs. MNIST, indicating a convolutional model bias. (2) Image likelihoods of Glow models trained on more general datasets, e.g., 80 Million Tiny Images (Tiny), have the highest average correlation with image likelihoods of models trained on other datasets, indicating a hierarchy of distributions from more general distributions (better for learning domain prior) to more specific distributions. (3) The likelihood contributions of the final scale of the Glow network correlate less between different Glow networks than the likelihood contributions of the earlier scales, while the overall likelihood is dominated by the earlier scales. This indicates a hierarchy of features inside the Glow network scales, from more generic low-level features that dominate the likelihood to more distribution-specific high-level features that are more informative about object categories.

Finally, leveraging the two novel views of a hierarchy of distributions and a hierarchy of features, we propose two likelihood-based anomaly detection methods. From the hierarchy-of-distributions view, we use likelihood ratios of two identical generative architectures (e.g., Glow), one trained on the in-distribution data (e.g., CIFAR10) and the other one on a more general distribution (e.g., 80 Million Tiny Images), refining previous likelihood-ratio-based methods. To further improve the performance, we additionally train our in-distribution model on samples from the general distribution using a novel outlier loss. From the hierarchy-of-features view, we show that using the likelihood contribution of the final scale of a multi-scale Glow model performs remarkably well for anomaly detection.

Our manuscript advances the understanding of the anomaly detection behavior of deep generative neural networks by synthesizing previous findings into two novel viewpoints accounting for the hierarchical nature of natural images. Based on this concept, we propose two new anomaly detection methods which reach strong performance, especially in the unsupervised setting. Our experiments are more extensive than previous likelihood-ratio-based methods on images, especially in the unsupervised setting, and therefore also fill an important empirical gap in the literature.

Common Low-Level Features Dominate the Model Likelihood

The following hypotheses synthesized from the findings of prior work motivate our methods to detect if an image is from a different object recognition dataset:

Distributions of low-level features form a domain prior valid for all natural images (see Fig. 1A).

These low-level features contribute the most to the overall likelihood assigned by a deep generative network to any image (see Fig. 1B and C).

How strongly which type of features contributes to the likelihood is influenced by the model bias (e.g., for convolutional networks, local features dominate) (see Fig. 1C).

For the first hypothesis, we start from defining low-level features. They are features that can be extracted in few computational steps (these will be local features for convolutional models like Glow). As an example, we use the difference of a pixel value to the mean of its 3×33\times 3 neighbouring pixels. As natural images are smooth, the distributions over such per-pixel difference of SVHN, CIFAR10 and 80 Million Tiny Image (Tiny) depicted in Fig. 1A are all zero centered. Smoother images will produce smaller differences among neighbouring pixels, therefore the density of SVHN has the highest peak around zero. A significant overlapping high-density regions for SVHN, CIFAR10 and Tiny show that these low-level features are common to natural images, and thus not useful for anomaly detection.

Next, we examine the second hypothesis by showing that low-level per-image likelihoods highly correlate with Glow network likelihoods. On the pixel level, we compute the pixel difference and estimate its density using a histogram with 100100 equally distanced bins. Using the estimated density, we can get the conditional (on 3×33\times 3 neighbours) per-pixel likelihoods of each image. Summing the per-pixel likelihoods over the entire image, we obtain pixel-level per-image pseudo-likelihood, which is not a correct likelihood of image, but a proxy measure of low-level feature contributions to the image-level likelihood. The low-level pseudo-likelihoods of SVHN and CIFAR10 images have Spearman correlationsAll results in this section are qualitatively the same with Pearson correlations. >0.83>0.83 with likelihoods of Glow networks trained on CIFAR10 (see Fig 1B), SVHN or Tiny. We also trained small modified Glow networks on 8×88\times 8 patches cropped from the original image. These local Glow networks’ likelihoods correlate even more with the full Glow networks’ likelihoods (>0.95>0.95), further suggesting low-level local features dominate the total likelihood (see Fig 1C). To validate that the low-level features dominating the likelihoods are independent of the semantic content of the image, we mix two images in Fourier space by combining the amplitudes of one image’s Fourier transform with the phases of the other image’s Fourier transform, and then evaluate the likelihood of the resulting image using the same pretrained Glow model (see supp. material S3 for details). The mixed images are semantically much more coherent with the images that provide the phase information, yet their Glow likelihoods correlate more strongly with the Glow likelihoods of images sharing the same amplitudes (>0.8>0.8 vs. <0.05<0.05).

Lastly, we show that what features are extracted and how much each feature contributes to the likelihood depend on the type of model. When training a modified Glow network that uses fully connected instead of convolutional blocks (see supp. material S2.2) on Fashion-MNIST and MNIST, the image likelihoods among them do not correlate (Spearman correlation −0.2-0.2). The fully-connected Fashion-MNIST network achieves worse likelihoods (4.84.8 vs. 2.92.9 bpd), but is much better at anomaly detection (81%81\% AUROC for Fashion-MNIST vs. MNIST; 15%15\% AUROC for convolutional Glow).

Consistent with non-distribution-specific low-level features dominating the likelihoods, we find full Glow networks trained independently on CIFAR10, CIFAR100, SVHN or Tiny produce highly correlated image likelihoods (Spearman correlation >0.96>0.96 for all pairs of models, see Fig. 1 C). The same is true to a lesser degree for Fashion-MNIST and MNIST (Spearman correlation 0.850.85).

Taken together, the evidence suggests convolutional generative networks have model biases that guide them to learn low-level feature distributions of natural images (domain prior) well, at the expense of anomaly detection performance. Based on this understanding, next we propose two methods to remove this influence of model bias and domain prior on likelihood-based anomaly detection.

Hierarchy of Distributions

The models trained on Tiny have the highest average likelihood across all evaluated datasets. This inspired us to use a hierarchy of distributions: CIFAR10 and SVHN are subdistributions of natural images, CIFAR10-planes are a subdistribution of CIFAR10, etc. (see Fig. 2).

We use this hierarchy of distributions to derive a log-likelihood-ratio-based anomaly detection method:

Train a generative network on a general image distribution like 80 Million Tiny Images

Train another generative network on images drawn from the in-distribution, e.g., CIFAR10

Use their likelihood ratio for anomaly detection

Formally, given the general-distribution-network likelihood pgp_{g} and the specific-in-distribution-network likelihood pinp_{in}, our anomaly detection score (low scores indicate outliers) is:

where σ\sigma is the sigmoid function and λ\lambda is a weighting factor.

2 Extension to the Supervised Setting

Hierarchy of Features

The image likelihood correlations between models trained on different datasets reduce substantially when evaluating the likelihood contributions of the final scale of the Glow network (see Fig. 3). Here, the adopted Glow network has three scales. At the first two scales, i.e., i=1i=1 and 22, the layer output is split into two parts hih_{i} and ziz_{i}, where hih_{i} is passed onto the next scale and ziz_{i} is output as one part of the latent code zz. The output at the last scale is z3z_{3}, which together with z1z_{1} and z2z_{2} makes up the complete latent code zz. In terms of y1=(h1,z1)y_{1}=(h_{1},z_{1}), y2=(h2,z2)y_{2}=(h_{2},z_{2}), y3=z3y_{3}=z_{3} and h0=xh_{0}=x, the logarithm of the learned density p(x)p(x) can be decomposed into per-scale likelihood contributions ci(x)c_{i}(x) as

The log-likelihood contributions c3(x)c_{3}(x) of the final third scale of Glow networks correlate substantially less than the full likelihoods for Glow networks trained on the different datasets (0.260.26 mean correlation vs. 0.990.99 mean correlation). This is consistent with the observation that last-scale dimensions encode more global object-specific features (see Fig. 3 and ). Therefore, we use c3(x)c_{3}(x) as our anomaly detection score (low scores indicate outliers).

Note that, here we do not condition z1z_{1} on z2z_{2} or z2z_{2} on z3z_{3}, whereas other implementations often make z1z_{1} dependent of z2z_{2} such as z1∼N(f(z2),g(z2)2)z_{1}\sim N(f(z_{2}),g(z_{2})^{2}) with f,gf,g being small neural networks. Such dependency is removable by transforming to z1′=(z1−f(z2))g(z2))z_{1}^{\prime}=\frac{(z_{1}-f(z_{2}))}{g(z_{2}))}, with z1′z_{1}^{\prime} now independent of z2z_{2} as z1′∼N(0,1)z_{1}^{\prime}\sim N(0,1), and this type of transformation can already be learned by an affine coupling layer applied to z1z_{1} and z2z_{2}, hence the explicit conditioning of other implementations does not fundamentally change network expressiveness. We do not use it here and do not observe bits/dim differences between our implementation and those that use it (see supp. material S7.1 for details).

Experiments

For the main experiments, we use SVHN , CIFAR10 , CIFAR100 as inlier datasets and use the same and LSUN as outlier datasets. Results for further outlier datasets can be found in the supp. material S7.4. 80 Million Tiny Images serve as our general distribution dataset in the log likelihood-ratio based anomaly detection experiments and is also used in the outlier loss as given in Eq. 2 when training two generative models, i.e., Glow and PixelCNN++ (PCNN) (see supp. material S4 for their training details). In-distribution Glow and PixelCNN++ models are finetuned from the models pre-trained on Tiny for more rapid training, see supp. material S7.2 for details and ablation studies. Our reported results are averaged over 3 random seeds.

In Tab. 1, we compare the raw log-likelihood based anomaly detection (i.e., diff to: None) with the log-likelihood ratio based ones (i.e., diff to: PNG, Tiny-Glow, Tiny-PCNN). The raw log-likelihood based scheme underperforms the log-likelihood ratio based ones that use Tiny-Glow and Tiny-PCNN. However, when using PNG to remove the domain prior as proposed in The original results of are not comparable since they used the training set of the in-distribution for evaluation. We provide supplementary code that uses a publicly available CIFAR10 Glow-network with pretrained weights that roughly matches the results reported here., it sometimes performs worse than the raw log-likelihood based scheme, e.g., SVHN as the in-distribution vs. the other three OODs. This relates to the remaining model bias, as Glow trained on the in-distribution encodes the domain prior differently to PNG. Also note that using PCNN as the general-distribution model for Glow does not work well for CIFAR10/100. This is because Tiny-PCNN has very large bpd gains over Glow for the CIFAR-datasets and less large gains for SVHN. On average across datasets, it works best to use matching general and in-distribution models (Glow for Glow, PCNN for PCNN), validating our idea of a model bias. Also note that our SVHN vs. CIFAR10 results already outperform the likelihood-ratio-based results of Ren et al. slightly (93.9% vs. 93.1% AUROC), and we observe further improvements with outlier loss in Section 5.3. Ren et al. used a noised version of the in-distribution as the general distribution and only tested it on SVHN vs. CIFAR10. Comparing to Tiny, it is less representative as a domain prior, and thus its performance on more complicated datasets requires further assessment.

To further validate our log-likelihood ratio approach on a different domain, we setup an experiment on the medical BRATS Magnetic Resonance Imaging (MRI) dataset. We use one MRI modality as in-distribution and the other three as OOD. The raw likelihood, the log-likelihood ratio to Tiny-Glow and to BRATS-Glow (trained on all modalities) yield AUROC 53.3%53.3\%, 68.3%68.3\% and 78.3%78.3\%, respectively. So, Tiny also serves as a general distribution for the very different medical images, and a distribution from the more specific domain further improves the performance.

The log-likelihood ratio approach can likely be applied to more than images. In the above, we have already shown the application to typical image datasets and, without adaptation, to medical MRI images. In the text/NLP domain, it may be used with Wikitext-2 as the general dataset, since Wikitext-2 already worked well as an outlier dataset in . In the audio domain, the domain prior may come from strong dependencies of the signal values on short timescales, similar to the smoothness of natural images. If a suitable general dataset needs to be created, it does not require labels and may even profit from noisy/unclean data. Therefore there is no principal obstacle preventing collection of such data, including concatenating existing datasets.

2 Anomaly Detection based on Last-scale Log-likelihood Contribution

As an alternative to remove the domain prior by using, e.g., Tiny-Glow, our hierarchy-of-features view suggests to use the log-likelihood contributed by the high-level features attained at the last scale of the Glow model. It is orthogonal to the log-likelihood ratio based scheme, and can be used when the general distribution is unavailable. As shown in Tab. 3, using the raw log-likelihood on the last scale (4×44\times 4), consistently outperforms the conventional log-likelihood comparison on the full scale, but performs slightly worse than using the log-likelihood ratio in the full scale. Note that we don’t expect the log-likelihood ratio on the last scale to be the top performer, as the domain prior is mainly reflected by the earlier two scales, see Fig. 3.

3 Outlier Losses

When training the in-distribution model, we can use the images from Tiny as the outliers to improve the training. Tab. 3 shows that our outlier loss as formulated in Eq. 2 consistently outperforms the margin loss when combining with three different types of log-likelihood based schemes, i.e., raw log-likelihood, raw log-likelihood at the last scale 4×44\times 4 and log-likelihood ratio. We note that as the margin loss leads to substantially less stable training than our loss, see the supp. material S7.5.

We also experiment on adding the outlier loss to the training loss of Tiny-Glow, i.e., using the in-distribution samples as outliers. This further improves the anomaly detection performance, see Diff† of Tab. 4, while Diff only uses the outlier loss for training the in-distribution Glow-network.

4 Unsupervised vs. Supervised Setting

From the unsupervised to the supervised setting, Tab. 4 further reports the numbers achieved by using the class-conditional in-distribution Glow-network and treating inputs from other classes as outliers. We observe further improved anomaly detection performance.

Overall, our approach only slightly underperforms the approach MSP-OE with inlier class labels (Supervised), while being substantially better without inlier class labels (Unsupervised), see Tab. 4. In contrast to observations by Hendrycks et al. for their unsupervised setup, we do not experience a severe degradation of the anomaly detection performance from the lack of class labels.

Related Work

We present an overview over anomaly detection approaches with a focus on recent work closely related to the ideas of a hierarchy of distributions and a hierarchy of features.

Classifier-based Methods Multi-class classifiers trained to discriminate in-distribution classes have been used for anomaly detection. Hendrycks and Gimpel used maximum softmax response as the score of normality. Different data augmentation schemes further enforced its performance. Lee et al. alternatively modeled the class-conditional features attained by the hidden layers of the classifier as multivariate Gaussians, and then used the Mahalanobis distance of Gaussians for anomaly detection. Another recent work used the gradient norm of the log-sum-exp of the class logits over the input for anomaly detection. In the context of self-supervised learning, class becomes the type of transformations . Self-supervised contrastive training improved the anomaly detection performance of multi-class classifiers .

Instead of exploiting multi-class classifiers, a different approach is to train a one-class classifier to directly discriminate inliers and outliers. One-class support vector machines are trained to return positive values only in a small region containing the inliers and negative values elsewhere . This approach has also been used in forming the latent space of deep autoencoders . Ruff et al. also used samples from a general distribution as outliers. Steinwart et al. drawn outliers from uniform distribution. In the supp. material S9, we also report results using a in-distribution-vs-general-distribution classifier.

Reconstruction-based Methods Another line of work is to learn the features and generation of inliers by reconstructing the training samples either in their input space or latent space, e.g., . At test time, an outlier is then detected if reconstruction is poor. However, owing to large capacity of deep neural networks, reconstruction loss alone may not be a reliable metric for anomaly detection. Huang et al. proposed to additionally use the joint likelihood of latent variables, which is obtained by using a neural rendering model to invert multi-class CNN-based classifier.

Input Likelihood-based Methods Generative modeling through maximum likelihood estimation tries to enforce high likelihoods on inliers. Under the normalization constraint, the likelihoods of outliers are expected to be low (ideally zero). However, their anomaly detection performance is often unsatisfactory . Outliers may attain even higher likelihoods than inliers. Recent work related the poor performance to sampling in a high-dimensional space, namely, inliers being mapped to the typical set of the latent code rather than the high likelihood area. They proposed to address this issue by batch-wise anomaly detection, whose application is more limited than instance-wise anomaly detection. A different approach combined input likelihoods with inlier classifiers. Che et al. trained a class-conditional generative model with an auxiliary adversarial loss to disentangle the class information from the rest latent representation. The achieved performance is better than ours in the supervised setting, while our methods mainly target and work in the unsupervised setting. It can be interesting to exploit their way of incorporating the label information into our model training.

Our work is also input likelihood-based. Our analysis in Section 2 showed that convolutional networks trained on one natural image dataset will learn low-level feature distribution that is common to the whole domain, and such domain prior dominates the likelihood. The concurrent work found results consistent with ours. We further exploited hierarchies of distributions or hierarchies of features as explained in Section 3 and 4 to improve the anomaly detection performance.

Hierarchy of Distributions: Our hierarchy-of-distributions likelihood-ratio method relates to prior and concurrent methods as follows. Ren et al. used a noised version of the in-distribution as the general distribution. Their method always requires training of two models for each in-distribution, our method has the option of only require one training of a general distribution model which we can reuse for any new in-distribution as long as it is as a subdistribution of the general distribution. In our method, the challenge is to find a suitable general distribution, while in their method the challenge is to find a suitable noise model. In the only rgb-image-setting they evaluated in , we show improved results over theirs, see Section 5.1. Serrà et al. used a generic lossless compressor such as PNG as their general distribution model and unfortunately only reported results on the training set of the in-distribution, making their results incomparable with any other works. We show improved performance over a reimplementation (see code repository) in Section 5.1. Note that both and did not evaluate the cases where raw likelihoods work well, even though using the likelihood ratio may decrease performance in that case as seen in Section 5.1.

Outlier Loss: Hendrycks et al. used a margin loss on images from a known outlier distribution (Tiny without CIFAR-images). We develop an outlier loss to improve the performance of likelihood models. It can be viewed as a combination of their idea of an outlier loss with our view of a hierarchy of distributions, and achieved an improved performance in the unsupervised setting in Section 5.3.

Hierarchy of Features: Regarding hierarchy of features, the closest related work investigated deep variational autoencoders that use a hierarchy of stochastic variables and found that later stochastic variables perform better at anomaly detection . Furthermore, Krusinga et al. developed a method to approximate probability densities from generative adversarial networks. The log densities are the sum of a change-of-volume determinant and a latent prior probability density. The latent log densities better reflected semantic similarity to the in-distribution than the full log density, however, no comparable anomaly detection results were reported. Nalisnick et al. also found the same for latent vs. full log densities of Glow networks, but also did not report anomaly detection results and did not look at individual scales of Glow networks.

Discussion and Conclusion

In this work, we proposed two log-likelihood based metrics for anomaly detection, outperforming state of the art methods in the unsupervised setting and only slightly underperforming classifier-based methods in the supervised setting. For good anomaly detection performance from raw likelihoods, an additional loss (such as our outlier loss) that forces the model to assign low likelihoods for images with OOD-high-level features, e.g. wrong objects, is particularly beneficial according to empirical results. Our analysis points to a potential reason, namely that without an outlier loss, the likelihoods are almost fully determined by low-level features such as smoothness or the dominant color in an image. As such low level features are common to natural image datasets, they form a strong domain prior, presenting a difficult task to detect high-level differences between inliers and outliers, e.g., object classes.

An interesting future direction is how to best combine the hierarchy-of-distributions and hierarchy-of-features views into a single approach. Preliminary experiments freezing the first two scales of the Tiny-Glow model and only finetuning the last scale on the in-distribution have shown promise, awaiting further evaluation.

In summary, our approach shows strong anomaly detection performance particularly in the more challenging unsupervised setting and also allows a better understanding of generative-model-based anomaly detection by leveraging hierarchical views of distributions and features.

Broader Impact

A better understanding of deep generative networks with regards to anomaly detection can help the machine learning research community in multiple ways. It allows to estimate which tasks deep generative networks may be suitable or unsuitable for when trained via maximum likelihood. With regards to that, our work helps more precisely understand the outcomes of maximum likelihood training. This more precise understanding can also help guide the design of training regimes that combine maximum likelihood training with other objectives depending on the task, if the task is unlikely to be solved by maximum likelihood training alone.

Anomaly detection in general itself has positive uses. For example, detecting anomalies in medical data can detect existing and developing medical problems earlier. Safety of machine learning systems in healthcare, autonomous driving, etc., can be improved by detecting if they are processing data that is unlike their training distribution.

Negative uses and consequences of anomaly detection can be that it allows tighter control of people by those with access to large computer and data, as they can more easily find unusual patterns deviating from the norm. For example, it may also allow health insurance companies to detect unusual behavioral patterns and associate them with higher insurance costs. Similarly repressive governments may detect unusual behavioral patterns to target tighter surveillance.

These developments may be steered in a better direction by a better public understanding and regulation for what purpose anomaly-detection machine-learning systems are developed and used.

Acknowledgments and Disclosure of Funding

This work was partially done during an internship of Robin Tibor Schirrmeister at the Bosch Center for Artificial Intelligence. A part of this work was supported by the german Federal Ministry of Education and Research (BMBF, grant RenormalizedFlows 01IS19077C).

Yuxuan Zhou wants to thank Zhongyu Lou and Duc Tam Nguyen for the helpful discussions about the preliminary experiment results.

Robin Tibor Schirrmeister wants to thank Jan Hendrik Metzen, Polina Kirichenko, Pavel Izmailov, Manuel Watter and Dengfeng Huang for discussions and support.

References

Supplementary outline

This document completes the presentation of the main paper with the following:

Details about Glow and PixelCNN++ architectures, and the likelihood decomposition equation Eq. (3) of Sec. 4 in the main paper;

Details about the modified (local/fully connected) Glow architectures for the analysis in Sec. 2 of the main paper;

Fourier-based analysis of influence of amplitude and phase on likelihoods;

Details about training and evaluation of Glow and PixelCNN++, including hyperparameter choices and computing infrastructure;

Details on the used datasets and dataset splits;

Reasons why the results of Serrà et al. are not comparable as is, and details of our reimplementation;

Further quantitative results, including maximum-likelihood performance (S7.1), finetuning vs. from-scratch training (S7.2), variance over seeds (S7.3), further outlier datasets (S7.4) and different outlier losses (S7.5);

Qualitative analysis of the different anomaly detection metrics;

Generative vs. discriminative approach for anomaly detection.

Please also note the attached supplementary codes.

Appendix S1 Glow and PixelCNN++ architectures

Our implementation of the Glow network is based on a publicly available Glow implementation https://github.com/y0ast/Glow-PyTorch/ with one modification explained in S1.3. The multi-scale Glow network consists of three sequential scales processing representations of size 12×16×1612\times 16\times 16, 24×8×824\times 8\times 8 and 48×4×448\times 4\times 4 (channel ×\times width ×\times height). Each scale consists of a repeating sequence of activation normalization, invertible 1×11\times 1 convolution and affine coupling blocks, see the original paper for details. Our Glow network, consistent with aforementioned public implementation, uses 3232 actnorm-11 conv-affine sequences per scale.

S1.2 PixelCNN++ architecture

We use a publicly available PixelCNN++ implementation https://github.com/pclucas14/pixel-cnn-pp/tree/16c8b2fb8f53e838d705105751e3c56536f3968a with only a single change. We reduce the number of filters used across the model from 160160 to 120120 for fast single-GPU training.

For a multi-scale model like Glow, the overall likelihood consists of the contributions from different scales, see Eq. (3) in the main paper. In contrast to other implementations, we do not condition z1z_{1} on z2z_{2} or z2z_{2} on z3z_{3} in our Glow model as described in Sec. S1.1. Recall that Glow splits the complete latent code into per-scales latent codes z1z_{1}, z2z_{2}, z3z_{3}. Here, z1z_{1} is the half of the output of the first scale that is not processed further. Many implementations make z1z_{1} dependent of z2z_{2} (and same for z2z_{2} and z3z_{3}) as z1∼N(f(z2),g(z2)2)z_{1}\sim N(f(z_{2}),g(z_{2})^{2}) with f,gf,g being small neural networks. For ease of implementation, we do not do that, instead we directly evaluate z1z_{1} under a standard-normal gaussian, so z1∼N(0,1)z_{1}\sim N(0,1).

Note that this does not fundamentally alter network expressiveness. An affine coupling layer can already implement the same computation achieved by z1∼N(f(z2),g(z2)2)z_{1}\sim N(f(z_{2}),g(z_{2})^{2}). Imagine zz is split for the affine coupling layer into z1z_{1} and z2z_{2}, with a coefficient network on z2z_{2} used to compute the affine scale and translation coefficients s,ts,t to transform z1′=z1⊙s(z2)+t(z2)z_{1}^{\prime}=z_{1}\odot s(z_{2})+t(z_{2}). Then if s(z2)=−f(z2)s(z_{2})=-f(z_{2}) and t(z2)=1g(z2)t(z_{2})=\frac{1}{g(z_{2})} and z1∼N(f(z2),g(z2)2)z_{1}\sim N(f(z_{2}),g(z_{2})^{2}), it follows that z1′∼N(0,1)z_{1}^{\prime}\sim N(0,1). In other words, computing the mean and standard deviation for z1z_{1} from z2z_{2} is the same as normalizing z1z_{1} by subtracting the mean and dividing by the standard deviation computed from z2z_{2}, which a regular affine coupling block can already learn.

In practice, there could still be differences due to the additional parameters, different kind of blocks used to implement ff and gg, and the difference of computing the (log)stdstd or its inverse. However, we observe no appreciable bits/dim differences between our implementation and those using the explicit conditioning step, see Section S7.1.

Appendix S2 Local and Fully Connected Architectures

We designed our local patches experiment to train compact Glow-like models that can only process information from 8×88\times 8 patches in the original datasets. Full-sized Glow networks process the full image using three scales as written in Section S1.1. Our local Glow network instead processes local 8×88\times 8 patches using a single scale. The 32×3232\times 32 input image is first cropped into 1616 non-overlapping 8×88\times 8 patches. These 8×88\times 8 patches are then processed independently by a local Glow network corresponding to a single scale of the full Glow network. In other words, we treat the image as if it consists of independent 8×88\times 8 patches. Evaluating the likelihoods of these patches, their sum is the likelihood of the image assigned by the local Glow model. Note that we aimed to create a network restricted to learn a general local domain prior and not one with the best maximum-likelihood performance.

S2.2 Fully Connected

We designed our fully-connected experiment to train fully-connected Glow networks that have a different model bias to regular convolutional Glow networks. We kept the three-scale architecture of Glow including the invertible subsampling steps at the beginning of each scale. Within each scale, the fully-connected Glow first flattens the representation, e.g. from a 12×16×1612\times 16\times 16 tensor per rgb-image to a 30723072-sized vector in the scale 11. This vector is then processed by the usual sequence of actnorm-1×11\times 1-affine. We next detail the processing at the scale 11, whereas the other scales follow the same design pattern. Activation normalization now processes 30723072 dimensions, so has substantially more parameters. The 1×11\times 1 is now an invertible linear projection keeping the dimensionality, so a projection from 30723072 dimensions to 30723072 dimensions. To ensure training stability, we did not train the 1×11\times 1-projections, but kept their parameters in the randomly initialized starting state. The fully-connected affine coupling block uses a sequence of linear layer (1536×5121536\times 512) - ReLU - linear layer (512×512512\times 512) - ReLU - linear layer (512×3072512\times 3072) modules to compute the 15361536 translation and 15361536 scale coefficients. To allow fast single-GPU training, we reduced the number of actnorm-1×11\times 1-affine sequences from 3232 to 88 per scale. Similar to Section S2.1, the fully-connected network was designed to highlight the influence different model biases and not to reach the best maximum-likelihood performance.

Appendix S3 Fourier-based Amplitude/Phase Analysis

To validate that low-level features dominate the likelihoods independent of the semantic content of the image, we create mixed images in Fourier space. Concretely, we:

Compute the amplitudes and phases of a batch Fourier transformed images;

Mix up images in their frequency domain by using one image’ amplitudes and phases of the other image

Apply the inverse Fourier transformation to invert these mixed images to the input domain

We show examples in Figure S1. Note that the mixed images are semantically much more similar to the image the phases were extracted from. We then compare the CIFAR10-Glow likelihoods on the original images and the mixed images for SVHN and CIFAR10 images. The likelihoods of the mixed images correlate much more with the amplitude-image likelihoods (Spearman correlation >0.8>0.8) than with the phase-image likelihoods (Spearman correlation <0.05<0.05).

Appendix S4 Training and Evaluation

We stayed close to the training setting of a publicly available Glow repositoryhttps://github.com/y0ast/Glow-PyTorch/blob/master/train.py (we do not use warmstart). Namely, we use Adamax as the optimizer (learning rate 5⋅10−45\cdot 10^{-4}, weight decay 5⋅10−55\cdot 10^{-5}) and 250250 training epochs. The training setting also includes data augmentation (translations and horizontal flipping for CIFAR10/100, only translations for SVHN, Fashion-MNIST and MNIST). These settings are the same for all experiments (from-scratch training, finetuning, with and without outlier loss, unsupervised/supervised). On 80 Million Tiny Images, we use substantially less training epochs, so that the number of batch updates is identical between Tiny and the other experiments. All datasets were preprocessed to be in the range [−0.5,0.5−1256][-0.5,0.5-\frac{1}{256}] as is standard practice for Glow-network training, see also supplementary code.

S4.2 PixelCNN++ Training

We stayed close to the training setting of the public PixelCNN++ repositoryhttps://github.com/pclucas14/pixel-cnn-pp/tree/16c8b2fb8f53e838d705105751e3c56536f3968a. Namely, we use Adam as the optimizer (learning rate 2⋅10−42\cdot 10^{-4}, no weight decay, negligible learning rate decay 5⋅10−65\cdot 10^{-6} every epoch). We substantially reduced the number of epochs from 50005000 to 120120 to save GPU resources, and since our aim here was to show general applicability of our methods to another type of model and not to reach maximum possible performance. Consistent with the public implementation, no data augmentation is performed.

S4.3 Numerical stabilization of the training

When training networks with outlier loss, negative infinities can appear for numerical reasons. This is actually expected as true outliers should ideally have likelihood zero and therefore log likelihood equal to negative infinity. These do not contribute to the loss in theory, but cause numerical issues. To ensure numerical stability during training, we remove examples that get assigned negative infinite likelihoods from the current minibatch. In case this would remove more than 75%75\% of the minibatch, we skip the entire minibatch. These methods are only meant to ensure numerical stability, no specific training stability methods like gradient norm clipping are used.

S4.4 Outlier Loss Hyperparameters

As described in the main manuscript, our outlier loss is:

where σ\sigma is the sigmoid function, TT a temperature and λ\lambda is a weighting factor. Based on a brief manual search on CIFAR10, we use T=1000T=1000 λ=6000\lambda=6000 as they had the highest train set anomaly detection performance while retaining stable training. PixelCNN++ training works well with the same exact values, validating the choice.

S4.5 Evaluation Details

We use the likelihoods computed on noise-free inputs for anomaly detection with Glow networks. In practice, this means not adding dequantization noise and instead adding a constant, namely half of the dequantization interval. We found this to yield slightly better anomaly detection performance in preliminary experiments. Noise-free inputs are only used during evaluation for the anomaly detection performance. Training is done adding the standard dequantization noise introduced as in . The BPD numbers reported in Section S7.1 and Table S1 also are obtained using single samples with the standard dequantization noise.

We clip log likelihoods from below by a very small number. When computing the log likelihoods from networks trained with outlier loss, negative infinities are to be expected, see Section S4.3. To include these inputs in the AUC computation, we set non-finite log-likelihoods to a very small constant (−3000000-3000000) before computing any log-likelihood difference.

S4.6 Computing Infrastructure

All experiments were computed on single GPUs. Runtimes vary between 22 to 88 days on Nvidia Geforce RTX 2080 depending on the experiment setting (outlier loss or not, supervised/unsupervised).

Appendix S5 Datasets

We use the pre-defined train/test folds on CIFAR10, CIFAR100, SVHN, Fashion-MNIST and MNIST and only train on the training fold. All results are reported on the test folds. For CelebA, we only use the first 6000060000 images for faster computations. We use a 11 million random subset of 80 Million Tiny Images in all of our experiments. For the Fashion-MNIST/MNIST experiments, we create a greyscaled Tiny dataset from the rgb data as x=r⋅0.2989+g⋅0.5870+b⋅0.1140x=r\cdot 0.2989+g\cdot 0.5870+b\cdot 0.1140.

S5.2 MRI Dataset

For the MRI BRATS dataset, we also use the official train/test split. The dataset was introduced to verify the log likelihood difference on a data from a slightly different domain (medical imaging). Since our outliers defined as different modalities are not defined by the object type, we did not expect last-scale likelihood contributions c3(x)c_{3}(x) to perform as good as on the object recognition datasets. However, they still outperform raw likelihoods, by 59.2% to 53.3%.

Appendix S6 Replication of [32]

The anomaly detection results from were obtained using the training folds of the in-distribution datasets, preventing a fair comparison to our results. In written communication with Serrà et al. , they explained to us that the AUROC-results reported in their paper compare in-distribution training-fold examples with out-of-distribution test-fold examples. This makes a fair comparison to our results and other works impossible. In contrast and in line with standard practice, our results were obtained using the test folds of the in-distribution datasets. Unfortunately, Serrà et al. are unable to provide their training code and models at the current time, so we cannot recompute their anomaly detection performance for the test fold of the in-distribution datasets. We also confirmed that depending on the training setting, the anomaly-detection AUROC values can differ substantially between the training and test fold of the in-distribution dataset.

In any case, we provide supplementary code to reproduce the method of to the best of our understanding. Using a publicly available pretrained Glow-modelhttps://github.com/y0ast/Glow-PyTorch, we find anomaly detection performance results similar to the ones we report for PNG as a general-distribution model.

Appendix S7 Further Quantitative Results

The maximum-likelihood-performance of our finetuned Glow networks are similar to the performance reported for from-scratch training in the original Glow paper . We show the bits-per-dimension values obtained using single dequantization samples in Table S1. Note that the Glow model trained on 80 Million Tiny Images already reaches bits per dimensions on CIFAR10 and CIFAR100 close to the Glow models trained on the actual dataset (CIFAR10/CIFAR100), in line with our view that the bits per dimension are dominated by the domain prior (results also do not substantially change when including or excluding CIFAR-images from Tiny).

Our Glow-model architecture was chosen from a reimplementation of Nalisnick et al. (see Section S1.1), in order to facilitate comparison of our results to other anomaly detection works. In future work, evaluating anomaly detection performance of our method with newer types of normalizing flows could be interesting.

S7.2 Finetuning

Training Glow networks on an in-distribution dataset by finetuning a Glow network trained on Tiny substantially speeds up the training progress over training from scratch. As can be seen in Figures S3 and S2, the Glow networks reach better results after less training epochs for both maximum-likelihood performance and anomaly-detection performance. The improvements are strongest for CIFAR100 and weakest for SVHN, in line with CIFAR100 being the most diverse dataset and most similar to 80 Million Tiny Images.

There are are still gains on log-likelihood ratio based anomaly detection for CIFAR100 when comparing finetuning a Glow a Glow network trained on Tiny to finetuning a Glow network already trained on CIFAR100 (or in other words, training the Glow Network from scratch on CIFAR100 for twice the number of epochs), see Figure S4. This is likely because the exact Tiny-model be used as the general distribution model later on, validating more similar models better cancel the model bias.

S7.3 Result Variance across Seeds

We present the original results including standard deviation in Table S2 and provide a graphical overview over our per-seed anomaly detection results in Figure S5. Results are relatively stable across seeds.

S7.4 Additional OOD Datasets

We report results on CelebA and Tiny Imagenet https://tiny-imagenet.herokuapp.com as additional out-of-distribution (OOD) datasets in Table S3.

S7.5 Margin Loss vs. Outlier Loss

In our experiments, the margin-based loss introduced in is less stable than our outlier loss for longer training runs, see Figure S6. Note that our results for the margin-based loss already substantially outperform the results reported for PixelCNN with a margin-based loss in .

Appendix S8 Qualitative Analyses

Our different metrics (raw likelihoods, log-likelihood ratios and last-scale likelihood contributions) result in qualitatively different highest-scoring images on 80 Million Tiny Images (see Fig. S7 and S8). We take a random 120000-images subset of 80 Million Tiny Images and use our Glow network trained on CIFAR10 either with or without outlier loss to compute the metrics. Looking at the top 12 images per metric shows that using the log-likelihood ratio results in more reasonable images (closer to the inliers) than the raw likelihood, albeit mostly still simple images (see Fig. S7) and that using the Glow network trained with outlier loss results in more fitting images for all metrics (see Fig. S8).

Appendix S9 Pure Discriminative Approach

As an additional baseline, we also evaluated using a purely discriminative approach. We trained a Wide-ResNet classifier to distinguish between the in-distribution and 80 Million Tiny-Images, without using any in-distribution labels. Concretely, we trained the classifier using samples of the in-distribution as the positive class and samples from 80 Million Tiny Images as the negative class in a normal supervised training setting. We use the training settings and architecture from a publicly available Wide-ResNet repository https://github.com/meliketoy/wide-resnet.pytorch, with depth=28 and widen-factor=10. After training, we use the prediction pResNet(yindist∣x)p_{ResNet}(y_{indist}|x) as our anomaly metric.

For CIFAR10/100 in-distribution, results show this baseline performs better for OOD dataset CIFAR100/10 (89% and 70% vs. 87% and 63% AUROC), similar for OOD dataset LSUN (93% and 89% vs. 96% and 86%) and worse for OOD dataset SVHN (93% and 73% vs. 99% and 85%) compared to our unsupervised generative methods (compare Table S4 to unsupervised in S2). Future work may further show what properties, advantages and disadvantages these different approaches have.