Full-Spectrum Out-of-Distribution Detection
Jingkang Yang, Kaiyang Zhou, Ziwei Liu
Introduction
State-of-the-art deep neural networks are notorious for their overconfident predictions on out-of-distribution (OOD) data yang2021generalized , defined as those not belonging to in-distribution (ID) classes. Such a behavior makes real-world deployments of neural network models untrustworthy and could endanger users involved in the systems. To solve the problem, various OOD detection methods have been proposed in the past few years baseline ; odin ; energyood ; mahalanobis ; likelihood ; duq20icml ; gram20icml . The main idea for an OOD detection algorithm is to assign to each test image a score that can represent the likelihood of whether the image comes from in- or out-of-distribution. Images whose scores fail to pass a threshold are rejected, and the decision-making process should be transferred to humans for better handling.
A critical problem in existing research of OOD detection is that only semantic shift is considered in the detection benchmarks while covariate shift—a type of distribution shift that is mainly concerned with changes in appearances like image contrast, lighting or viewpoint—is either excluded from the evaluation stage or simply treated as a sign of OOD yang2021generalized , which contradicts with the primary goal in machine learning, i.e., to generalize beyond the training distribution zhou2021domain .
In this paper, we introduce a more challenging yet realistic problem setting called full-spectrum out-of-distribution detection, or FS-OOD detection. The new setting takes into account both the detection of semantic shift and the ability to recognize covariate-shifted data as ID. To this end, we design three benchmarks, namely DIGITS, OBJECTS and COVID, each targeting a specific visual recognition task and together constituting a comprehensive testbed. We also provide a more fine-grained categorization of distributions for the purpose of thoroughly evaluating an algorithm. Specifically, we divide distributions into four groups: training ID, covariate-shifted ID, near-OOD, and far-OOD (the latter two are inspired by a recent study ming2021impact ). Figure 1-a shows example images from the DIGITS benchmark: the covariate-shifted images contain the same semantics as the training images, i.e., digits from 0 to 9, and should be classified as ID, whereas the two OOD groups clearly differ in semantics but represent two different levels of covariate shift.
Ideally, an OOD detection system is expected to produce high scores for samples from the training ID and covariate-shifted ID groups, while assign low scores to samples from the two OOD groups. However, when applying a state-of-the-art OOD detection method, e.g. the energy-based EBO energyood , to the proposed benchmarks like DIGITS (see Figure 1-b), we observe that the resulting scores completely fail to distinguish between ID and OOD. As shown in Figure 1-b, all data are classified as ID including both near-OOD and far-OOD samples.
To address the more challenging but realistic FS-OOD detection problem, we propose SEM, a simple feature-based semantics score function. Unlike existing score functions that are based on either marginal distribution energyood or predictive confidence baseline , SEM leverages features from both top and shallow layers to deduce a single score that is only relevant to semantics, hence more suitable for identifying semantic shift while ensuring robustness under covariate shift. Specifically, SEM is mainly composed of two probability measures: one is based on high-level features containing both semantic and non-semantic information, while the other is based on low-level feature statistics only capturing non-semantic image styles. With a simple combination, the non-semantic part is cancelled out, which leaves only semantic information in SEM. Figure 1-c illustrates that SEM’s scores are much clearer to distinguish between ID and OOD.
We summarize the contributions of this paper as follows. 1) For the first time, we introduce the full-spectrum OOD detection problem, which represents a more realistic scenario considering both semantic and covariate shift in the evaluation pipeline. 2) Three benchmark datasets are designed for research of FS-OOD detection. They cover a diverse set of recognition tasks and have a detailed categorization over distributions. 3) A simple yet effective OOD detection score function called SEM is proposed. Through extensive experiments on the three new benchmarks, we demonstrate that SEM significantly outperforms current state-of-the-art methods in FS-OOD detection. The source code and new datasets are open-sourced in https://github.com/Jingkang50/OpenOOD.
Related Work
The key idea in out-of-distribution (OOD) detection is to design a metric, known as score function, to assess whether a test sample comes from in- or out-of-distribution. The most commonly used metric is based on the conditional probability . An early OOD detection method is maximum softmax probability (MSP) baseline , which is motivated by the observation that deep neural networks tend to give lower confidence to mis-classified or OOD data. A follow-up work ODIN odin applies a temperature scaling parameter to soften the probability distribution, and further improves the performance by injecting adversarial perturbations to the input. Model ensembling has also been found effective in enhancing robustness in OOD detection waic ; eloc .
Another direction is to design the metric in a way that it reflects the marginal probability . Liu et al. energyood connect their OOD score to the marginal distribution using an energy-based formulation, which essentially sums up the prediction logits over all classes. Lee et al. mahalanobis assume the source data follow a normal distribution and learn a Mahalanobis distance to compute the discrepancy between test images and the estimated distribution parameters. Generative modeling has also been investigated to estimate a likelihood ratio for scoring test images waic ; likelihood ; S .
Some methods exploit external OOD datasets. For example, Hendrycks et al. hendrycks18oe extend MSP by training the model to produce uniform distributions on external OOD data. Later works introduce re-sampling strategy backgroundsample and cluster-based methodology yang2021scood to better leverage the background data. However, this work do not use external OOD datasets for model design.
Different from all existing methods, our approach aims to address a more challenging scenario, i.e., FS-OOD detection, which has not been investigated in the literature but is critical to real-world applications. The experiments show that current state-of-the-art methods mostly fail in the new setting while our approach gains significant improvements.
Methodology
Key to detect out-of-distribution (OOD) data lies in the design of a score function, which is used as a quantitative measure to distinguish between in- and out-of-distribution data. Our idea is to design the function in such a way that the degree of semantic shift is effectively captured, i.e., the designed score to be only sensitive to semantic shift while being robust to covariate shift. For data belonging to the in-distribution classes, the score is high, and vice versa.
Our score function, called SEM, has the following design:
where denotes image features learned by a neural network; and denotes features that only capture the semantics. The probability can be computed by a probabilistic model, such as a Gaussian mixture model.
The straightforward way to model is to learn a neural network for image recognition and hope that the output features only contain semantic information, i.e., . If so, the score can be simply computed by . However, numerous studies have suggested that the output features often contain both semantic and non-semantic information while decoupling them is still an open research problem zhou2021domain ; lin2021domain ; peng2019domain . Let denote non-semantic features, we assume that semantic features and non-semantic features are generated independently, namely
We propose a simple method to model the score function so that it becomes only relevant to the semantics of an image. This is achieved by leveraging low-level feature statistics, i.e., means and standard deviations, learned in a CNN, which have been shown effective in capturing image styles that are essentially irrelevant to semantics mixstyle . Specifically, the score function in Eq. 1 is rewritten as
where is computed using the output features while is based on low-level feature statistics.
Below we first discuss how to compute feature statistics and then detail the approach of how to model the distributions for and .
Feature Statistics Computation
As shown in Zhou et al. mixstyle , the feature statistics in shallow CNN layers are strongly correlated with domain information (i.e., image style) while those in higher layers pick up more semantics. Therefore, we choose to extract feature statistics in the first CNN layer and represent by concatenating the means and standard deviations, i.e., .
Distribution Modeling
For simplicity, we model and in Eq. 3 using the same approach, which consists of two steps: dimension reduction and distribution modeling. Below we only discuss for clarity.
Motivated by the manifold assumption in Bengio et al. bengio2013representation that suggests data typically lie in a manifold of much lower dimension than the input space, we transform features to a new low-dimensional space, with a hope that the structure makes it easier to distinguish between in- and out-of-distribution. To this end, we propose a variant of the principal component analysis (PCA) approach. Specifically, rather than maximizing the variance for the entire population, we maximize the sum of variances computed within each class with respect to the transformation matrix. In doing so, we can identify a space that is less correlated with classes.
Given a training dataset, we build a Gaussian mixture model (GMM) to capture . Formally, is defined as
where denotes the number of mixture components, the mixture weight s.t. , and and the means and variances of a normal distribution. A GMM model can be efficiently trained by the expectation-maximization (EM) algorithm.
2 Source-Awareness Enhancement
While feature statistics exhibit a higher correlation with source distributions mixstyle , the boundary between in- and out-of-distribution in complicated real-world data is not guaranteed to be clear enough for differentiation. Inspired by Liu et al. energyood who fine-tune a pretrained model to increase the energy values assigned to OOD data and lower down those for ID data, we propose a fine-tuning scheme to enhance source-awareness in feature statistics. An overview of the fine-tuning scheme is illustrated in Figure 2-b.
The motivation behind our fine-tuning scheme is to obtain a better estimate of non-semantic score, in hope that it will help SEM better capture the semantics with the combination in Eq. 3. This can be achieved by explicitly training feature statistics of ID data to become more compact, while pushing OOD data’s feature statistics away from the ID support areas. A straightforward way is to collect auxiliary OOD data like Liu et al. energyood for building a contrastive objective. In this work, we propose a more efficient way by using negative data augmentation sinha2021negative to synthesize OOD samples. The key idea is to choose data augmentation methods to easily generate samples with covariate shift. One example augmentation is Mixup mixup .
Learning Objectives
Given a source dataset ,With a slight abuse of notation, we use here to denote an image. we employ negative data augmentation methods to synthesize an OOD dataset where . For fine-tuning, we combine a classification loss with a source-awareness enhancement loss . These two losses are formally defined as
where the marginal probability is computed based on a GMM model described previously. Note that the GMM model is updated every epoch to adapt to the changing features.
After fine-tuning, we learn a new GMM model using the original source dataset. This model is then used to estimate the marginal probability at test time.
FS-OOD Benchmarks
To evaluate full-spectrum out-of-distribution (FS-OOD) detection algorithms, we design three benchmarks: DIGITS, OBJECTS, and COVID. Examples for DIGITS are shown in Figure 1 and the other two are shown in Figure 3.
We construct the DIGITS benchmark based on the popular digit datasets: MNIST mnist , which contains 60,000 images for training. During testing, the model will be exposed to 10,000 MNIST test images, with 26,032 covariate-shifted ID images from SVHN svhn and another 9,298 from USPS usps . The near-OOD datasets are notMNIST notmnist and FashionMNIST fashionmnist , which share a similar background style with MNIST. The far-OOD datasets consist of a textural dataset (Texture texture ), two object datasets (CIFAR-10 cifar & Tiny-ImageNet imagenet ), and one scene dataset (Places365 places365 ). The CIFAR-10 and Tiny-ImageNet test sets have 10,000 images for each. The Places365 test set contains 36,500 scene images.
Benchmark-2: OBJECTS
The OBJECTS benchmark is built on top of CIFAR-10 cifar , which contains 50,000 images for training. During testing, the model will be exposed to 10,000 CIFAR-10 test images, and another 10,000 images selected from ImageNet-22K imagenet with the same categories as CIFAR-10 (so it is called ImageNet-10). For ImageNet-10, we choose five ImageNet-22K classes corresponding to one CIFAR-10 class, with each class selecting 1,000 training images and 200 testing images. Details of the selected classes are shown in Table 1. In addition to ImageNet, CIFAR-10-C is used as a covariate-shifted ID dataset, which is essentially a corrupted version of CIFAR-10. For near-OOD, we choose CIFAR-100 and Tiny-ImageNet. For far-OOD, we choose MNIST, FashionMNIST, Texture and CIFAR-100-C.
Benchmark-3: COVID
We construct a real-world benchmark to show the practical value of FS-OOD. We simulate the scenario where an AI-assisted diagnostic system is trained to identify COVID-19 infection from chest x-ray images. The training data come from a single source (e.g., a hospital) while the covariate-shifted ID test data are from other hospitals or machines, to which the system needs to be robust and produce reliable predictions. Specifically, we refer to the COVID-19 chest X-ray dataset review santa2021public , and use the large-scale image collection from Valencian Region Medical ImageBank vaya2020bimcv (referred to as BIMCV) as training ID images (randomly sampled 2443 positive cases and 2501 negative cases with necessary cleaning). Images from two other sources, i.e., ACTUALMED Wang2020 (referred to as ActMed with 132 positive images), and Hannover winther12275009covid (from Hannover Medical School with 243 positive images), are considered as the covariate-shifted ID group. OOD images are from completely different classes. Near-OOD images are obtained from other medical datasets, i.e., the RSNA Bone Age dataset with 200 bone X-ray images boneage and 544 COVID CT images covidct . Far-OOD samples are defined as those with drastic visual and concept differences than the ID images. We use MNIST, CIFAR-10, Texture and Tiny-ImageNet.
Evaluation Metrics
In the FS-OOD setting, different datasets belonging to one OOD type (i.e., near-OOD or far-OOD) are grouped together. We also report the performance on contrasting covariate-shifted ID with training ID, although covariate-shifted ID are not OOD samples. We use three metrics to evaluate the OOD detection performance, which are detailed as follows: 1) FPR95 stands for false positive rate measured when true positive rate (TPR) sits at 95%. Intuitively, FPR95 measures the portion of samples that are falsely recognized as in-distribution data when most true in-distribution samples are recalled. 2) AUROC refers to the Area Under the Receiver Operating Characteristic curve, which is concerned with both FPR and TPR. 3) AUPR means the Area Under the Precision-Recall curve, which considers both precision and recall. For FPR95, the lower the value, the better the model. For AUROC and AUPR, the higher the value, the better the model.
Experiments
We conduct experiments on the three proposed FS-OOD benchmarks, i.e., DIGITS, OBJECTS, and COVID. In terms of architectures, we use LeNet-5 lenet for DIGITS and ResNet-18 resnet for both OBJECTS and COVID. All models are trained by the SGD optimizer with a weight decay of and a momentum of 0.9. For DIGITS and OBJECTS, we set the initial learning rate to 0.1, which is decayed by the cosine annealing rule, and the total epochs to 100. For COVID benchmark, the initial learning rate is set to 0.001 and the model is trained for 200 epochs. When fine-tuning for source-awareness enhancement, the learning rate is set to 0.005 and the total number of epochs is 10. The batch size is set to 128 for all benchmarks.
Notice that the baseline implementations of ODIN odin and MDS mahalanobis require validation set for hyperparameter tuning, we spare a certain portion of near-OOD for validation. More specifically, we use 1,000 notMNIST images for the DIGITS benchmark, 1,000 CIFAR-100 images for the OBJECTS benchmark, and 54 images from CT-SCAN dataset for the COVID benchmark. The proposed method SEM relies on the hyperparameter of for low-layer and number of classes for high-layer in Gaussian mixture model. For output features with dimensions over 50, PCA is performed to reduce the dimensions to 50.
1 Results on FS-OOD Setting
We first discuss the results on near- and far-OOD datasets. Table 2 summarizes the results where the proposed SEM is compared with current state-of-the-art methods including MSP baseline , ODIN odin , Mahalanobis distance score (MDS), and Energy-based OOD energyood .
For the DIGITS benchmark, SEM gains significant improvements in all metrics (FPR95, AUROC, and AUPR). A huge gain is observed on notMNIST, which is a challenging dataset due to its closeness in background to the training ID MNIST. While none of the previous softmax/logits-based methods (e.g., MSP, ODIN, and EBO) are capable to solve the notMNIST problem, the proposed SEM largely reduces the FPR95 metric from 99% to 10.93%, and the AUROC is increased from around 30% to beyond 95%. One explanation of the clear advantage is that, the previous output-based OOD detection methods largely depend on the covariate shift to detect OOD samples, while the feature-based MDS (partly rely on top-layer semantic-aware features) and the proposed SEM uses more semantic information, which is critical to distinguish MNIST and notMNIST. In other words, in the MNIST/notMNIST scenario where ID and OOD have high visual similarity, large dependency on covariate shift while ignorance on the semantic information will lead to the failure of OOD separation. Similar advantages are also achieved with the other near-OOD dataset.
OBJECTS Benchmark
Similar to DIGITS benchmark, the proposed SEM surpasses the previous state-of-the-art methods on the near-OOD scenario of the OBJECTS benchmark, especially on the more robust metrics of AUROC and AUPR. However, the performance gap is not as large as DIGITS. One explanation is that images in OBJECTS benchmark are more complex than DIGITS, leading the neural networks to be more semantics-orientated. Therefore, more semantic information is encoded in the previous output-based methods. Nevertheless, the proposed SEM method still outperforms others on most of the metrics. We also notice that SEM score does not reach the best performance on MNIST and FashionMNIST. One explanation is that two black-and-white images in these two datasets inherently contain significant covariate shifts comparing to both training ID and covariate-shifted ID, so that the scores that efficient on covariate shift detection (e.g., ODIN) can also achieve good results on these datasets. However, these methods fail in near-OOD scenario, as they might believe CIFAR-10-C should be more likely to be OOD than CIFAR-100.
COVID Benchmark
In this new and real-world application of OOD detection, the proposed SEM score achieves an extraordinary performance on all metrics, which surpasses the previous state-of-the-art methods by a large margin in both near and far-OOD scenarios. The result also indicates that previous output-based methods generally breaks down on this setting, e.g., their FPR@95 scores are generally beyond 90% in near-OOD setting which means ID and OOD are totally mixed. However, the proposed SEM achieves around 10% in near-OOD setting. On far-OOD samples, the output-based methods are still unable to be sensitive to the ID/OOD discrepancy. The phenomenon matches the performance in DIGITS dataset, where the training data is simple and the logits might learn much non-semantic knowledge to be cancelled out.
Observation Summary
We summarize the following two take-away messages from the experiments on all three FS-OOD benchmarks: 1) SEM score performs consistently well on near-OOD, which classic output-based methods (e.g., MSP, ODIN, EBO) majorly fail on. The reason can be that output-based methods use too much covariate shift information for OOD detection, which by nature cannot distinguish between covariate-shifted ID and near-OOD. The proposed SEM score also outperforms the similar feature-based baseline MDS. 2) SEM score sometimes underperforms on far-OOD, with a similar reason that classic OOD detectors use covariate shift to distinguish ID and OOD, which is sometimes sufficient to detect far-OOD samples. Nevertheless, SEM reaches more balanced good results on near-OOD and far-OOD.
2 Results on Classic OOD Detection Setting
Table 3 shows the performance on the classic OOD detection benchmark. The result shows that without the introduction of covariate-shifted ID data, the previous methods reach a near-perfect performance on the classic benchmark, which matches the reported results in their origin papers. However, by comparing with Table 2, their performance significantly breakdown when covariate-shifted ID is introduced, showing the fragility of previous methods, and therefore we advocate the more realistic FS-OOD benchmark. Furthermore, we also report the results that by using the value of , the score from low-layer feature statistics for detecting covariate shift is shown surprisingly effective on classic OOD benchmark, which exceeds all the previous methods and achieve a near-perfect result across all the metrics. This phenomenon shows that only taking covariate shift score can completely solve the classic OOD detection benchmark with MNIST, which, in fact, contradicts the goal of OOD detection. It also advocates the significance of the proposed FS-OOD benchmark.
3 Ablation Study
In this section, we validate the effectiveness of the main components that contribute to the proposed SEM score, and also analyze the effects of fine-tuning scheme for source-awareness enhancement. All the experiments in this part are conducted on the DIGITS benchmark.
According to Equation 2 in the Section 3, SEM score can be decomposed by the estimations of and . While our final SEM score uses output flattened features of the CNN model for estimation and low-layer feature statistics for , there are actually several options for the estimation, which is discussed in Table 4. In this analysis, we set top flattened features as the default usage for and only explore , which is the key part of SEM score.
Exp#1 shows the result that only uses as the final score, which can be interpreted as a simple method using GMM to estimate ID likelihood on the final-layer features. Compared to the MDS result in Table 2, this simple method already obtains a better performance on near-OOD. Notice that we use LeNet-5 on DIGITS, the final-layer features are identical to their feature statistics (ref. Exp#2). Therefore, everything is cancelled out if is top-layer feature statistics (ref. Exp#3).
Exp#4 and Exp#6 shows comparison between using low-layer flattened features (L-FF) and low-layer feature statistics (L-FS) only. The performance on detecting covariate-shifted ID shows that both L-FF and L-FS have significant sensitivity to covariate shifts, but with a poor performance on FS-OOD detection. The result indicates that with only the usage of low-level features, the score has a strong correlation to covariate shift but barely contains semantic information, and the feature statistics show the stronger characteristics compared to flattened feature. This observation indicates our selection of low-level feature statistics for estimating , which is further supported by the results of Exp#5 and Exp#7, and visually illustrated by Figure 4.
Fine-Tuning Scheme
Here we evaluate the designed fine-tuning scheme of SEM. As elaborated in Section 3.2, this learning procedure is designed to enhance the source-aware compactness. Specifically, a source-awareness enhancement loss is proposed to aggregate the ID training data and separate from the generated negative augmented images at the same time. Table 5 demonstrates the effectiveness of the fine-tuning scheme. When combining both in-distribution training and negative augmented data training, our framework achieves the best performance.
Hyperparameter of M𝑀M
Table 6 shows the analysis of hyperparameter . In the DIGITS dataset, leads to a slightly better performance comparing to other choices. Nevertheless, the overall difference among various is not obvious on near-OOD, showing that the model is robust to the hyperparameter.
Discussion and Conclusion
Existing OOD detection literature has shown mostly relied on covariate shift even though they are intended to detect semantic shift. This is very effective when test OOD data only come from the far-OOD group—where the covariate shift is large and is further exacerbated by semantic shift, so using covariate shift as a measure to detect OOD fares well. However, when it comes to near-OOD data, especially with covariate-shifted ID (i.e., data experiencing covariate shift but still belonging to the same in-distribution data), current state-of-the-art methods would suffer a significant drop in performance, as shown in the experiments.
We find the gap is caused by a shortcoming in existing evaluation benchmarks: they either exclude covariate-shifted data during testing or treat them as OOD, which is conceptually contradictory with the primary goal that a machine learning model should generalize beyond the training distribution. To fill the gap, we introduce a new problem setting that better matches the design principles of machine learning models: they should be robust in terms of good generalization to covariate-shifted datasets, and trustworthy as they also need to be capable of detecting abnormal semantic shift.
The empirical results suggest that current state-of-the-art methods rely too heavily on covariate shift and hence could easily mis-classify covariate-shifted ID data as OOD data. In contrast, our SEM score function, despite having a simple design, provides a more reliable measure for solving full-spectrum OOD detection.
In fact, to detecting samples with covariate shift, we find that a simple probabilistic model using low-level feature statistics can reach a near-perfect result.
As the OOD detection community getting common awareness of the saturated performance problem of classic OOD benchmarks, several works have taken one-step further towards the more realistic setting and proposed large-scale benchmarks wang2022vim ; srivastava2022out . However, this paper shows that even under the classic MNIST/CIFAR-scale OOD benchmarks, current OOD methods in fact cannot achieve satisfactory results when the generalization ability is required. We hope that the future OOD detection works could also consider the generalization capability on covariate-shifted ID data, in parallel to exploring larger-scale models and datasets.
Broader Impacts
Our research aims to improve the robustness of machine learning systems in terms of the capability to safely handle abnormal data to avoid catastrophic failures. This could have positive impacts on a number of applications, ranging from consumer (e.g., AI-powered mobile phones) to transportation (e.g., autonomous driving) to medical care (e.g., abnormality detection). The new problem setting introduced in the paper includes an important but largely missing element in existing research, namely data experiencing covariate shift but belonging to the same in-distribution classes. We hope the new setting, along with the simple approach based on SEM and the findings presented in the paper, can pave the way for future research for more reliable and practical OOD detection.