Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks

Shiyu Liang, Yixuan Li, R. Srikant

Introduction

Modern neural networks are known to generalize well when the training and testing data are sampled from the same distribution (Krizhevsky et al., 2012; Simonyan & Zisserman, 2015; He et al., 2016; Cho et al., 2014; Zhang et al., 2017). However, when deploying neural networks in real-world applications, there is often very little control over the testing data distribution. Recent works have shown that neural networks tend to make high confidence predictions even for completely unrecognizable (Nguyen et al., 2015) or irrelevant inputs (Hendrycks & Gimpel, 2017; Szegedy et al., 2014; Moosavi-Dezfooli et al., 2017). It has been well documented (Amodei et al., 2016) that it is important for classifiers to be aware of uncertainty when shown new kinds of inputs, i.e., out-of-distribution examples. Therefore, being able to accurately detect out-of-distribution examples can be practically important for visual recognition tasks (Krizhevsky et al., 2012; Farabet et al., 2013; Ji et al., 2013).

A seemingly straightforward approach of detecting out-of-distribution images is to enlarge the training set of both in- and out-of-distribution examples. However, the number of out-of-distribution examples can be infinitely many, making the re-training approach computationally expensive and intractable. Moreover, to ensure that a neural network accurately classifies in-distribution samples into correct classes while correctly detecting out-of-distribution samples, one might need to employ exceedingly large neural network architectures, which further complicates the training process.

Hendrycks & Gimpel proposed a baseline method to detect out-of-distribution examples without further re-training networks. The method is based on an observation that a well-trained neural network tends to assign higher softmax scores to in-distribution examples than out-of-distribution examples. In this paper, we go further. We observe that after using temperature scaling in the softmax function (Hinton et al., 2015; Pereyra et al., 2017) and adding small controlled perturbations to inputs, the softmax score gap between in - and out-of-distribution examples is further enlarged. We show that the combination of these two techniques (temperature scaling and input perturbation) can lead to better detection performance. For example, provided with a pre-trained DenseNet (Huang et al., 2016) on CIFAR-10 dataset (positive samples), we test against images from TinyImageNet dataset (negative samples). Our method reduces the False Positive Rate (FPR), i.e., the fraction of misclassified out-of-distribution samples, from 34.7%34.7\% to 4.3%4.3\%, when 95%95\% of in-distribution images are correctly classified. We summarize the main contributions of this paper as the following:

We propose a simple and effective method, ODIN (Out-of-DIstribution detector for Neural networks), for detecting out-of-distribution examples in neural networks. Our method does not require re-training the neural network and is easily implementable on any modern neural architecture.

We test ODIN on state-of-the-art network architectures (e.g., DenseNet (Huang et al., 2016) and Wide ResNet (Zagoruyko & Komodakis, 2016)) under a diverse set of in- and out-distribution dataset pairs. We show ODIN can significantly improve the detection performance, and consistently outperforms the baseline method (Hendrycks & Gimpel, 2017) by a large margin.

We empirically analyze how parameter settings affect the performance, and further provide simple analysis that provides some intuition behind our method.

The outline of this paper is as follows. In Section 2, we present the necessary definitions and the problem statement. In Section 3, we introduce ODIN and present performance results in Section 4. We experimentally analyze the proposed method and provide some justification for our method in Section 5. We summarize the related works and future directions in Section 6 and conclude the paper in Section 7.

Problem Statement

In this paper, we focus on detecting out-of-distribution images. However, it is equally important to correctly classify an image into the right class if it is an in-distribution image. But this can be easily done: once it has been detected that an image is in-distribution, we can simply use the original image and run it through the neural network to classify it. Thus, we do not change the predictions of the neural network for in-distribution images and only focus on improving the detection performance for out-of-distribution images.

ODIN: Out-of-distribution Detector

In this section, we present our method, ODIN, for detecting out-of-distribution samples. The detector is built on two components: temperature scaling and input preprocessing. We describe the details of both components below.

Temperature Scaling. Assume that the neural network f=(f1,...,fN)\bm{f}=(f_{1},...,f_{N}) is trained to classify NN classes. For each input x\bm{x}, the neural network assigns a label y^(x)=arg⁡max⁡iSi(x;T)\hat{y}(\bm{x})=\arg\max_{i}S_{i}(\bm{x};T) by computing the softmax output for each class. Specifically,

Input Preprocessing. In addition to temperature scaling, we preprocess the input by adding small perturbations:

where the parameter ε\varepsilon is the perturbation magnitude. The method is inspired by the idea of adversarial examples (Goodfellow et al., 2015), where small perturbations are added to decrease the softmax score for the true label and force the neural network to make a wrong prediction. Here, our goal and setting are the opposite: we aim to increase the softmax score of any given input, without the need for a class label at all. As we shall see later, the perturbation can have stronger effect on the in- distribution images than that on out-of-distribution images, making them more separable. Note that the perturbations can be easily computed by back-propagating the gradient of the cross-entropy loss w.r.t the input.

The parameters T,εT,\varepsilon and δ\delta are chosen so that the true positive rate (i.e., the fraction of in-distribution images correctly classified as in-distribution images) is 95%.95\%.

Experiments

In this section, we demonstrate the effectiveness of ODIN on several computer vision benchmark datasets. We run all experiments with PyTorchhttp://pytorch.org and we release the code to reproduce all experimental resultshttps://github.com/facebookresearch/odin.

Architectures and training configurations. We adopt two state-of-the-art neural network architectures, including DenseNet (Huang et al., 2016) and Wide ResNet (Zagoruyko & Komodakis, 2016). For DenseNet, our model follows the same setup as in (Huang et al., 2016), with depth L=100L=100, growth rate k=12k=12 (Dense-BC) and dropout rate 0. In addition, we evaluate the method on a Wide ResNet, with depth 2828, width 10 (WRN-28-10) and dropout rate 0. The hyper-parameters of neural networks are set identical to the original Wide ResNet (Zagoruyko & Komodakis, 2016) and DenseNet (Huang et al., 2016) implementations. All neural networks are trained with stochastic gradient descent with Nesterov momentum (Duchi et al., 2011; Kingma & Ba, 2014). Specifically, we train Dense-BC for 300 epochs with batch size 64 and momentum 0.9; and Wide ResNet for 200 epochs with batch size 128 and momentum 0.9. The learning rate starts at 0.1, and is dropped by a factor of 10 at 50%50\% and 75%75\% of the training progress, respectively.

Accuracy of pre-trained networks. Each neural network architecture is trained on CIFAR-10 (C-10) and CIFAR-100 (C-100) datasets (Krizhevsky & Hinton, 2009), respectively. CIFAR-10 and CIFAR-100 images are drawn from 10 and 100 classes, respectively. Both datasets consist of 50,000 training images and 10,000 test images. The test error on CIFAR datasets are given in Table 1.

2 Out-of-distribution Datasets

At test time, the test images from CIFAR-10 (CIFAR-100) datasets can be viewed as the in-distribution (positive) examples. For out-of-distribution (negative) examples, we follow the setting in (Hendrycks & Gimpel, 2017) and test on several different natural image datasets and synthetic noise datasets. We consider the following out-of-distribution test datasets.

TinyImageNet. The Tiny ImageNet datasethttps://tiny-imagenet.herokuapp.com consists of a subset of ImageNet images (Deng et al., 2009). It contains 10,000 test images from 200 different classes. We construct two datasets, TinyImageNet (crop) and TinyImageNet (resize), by either randomly cropping image patches of size 32×3232\times 32 or downsampling each image to size 32×3232\times 32.

LSUN. The Large-scale Scene UNderstanding dataset (LSUN) has a testing set of 10,000 images of 10 different scenes categories such as bedroom, kitchen room, living room, etc. (Yu et al., 2015). Similar to TinyImageNet, we construct two datasets, LSUN (crop) and LSUN (resize), by randomly cropping and downsampling the LSUN testing set, respectively.

Gaussian Noise. The synthetic Gaussian noise dataset consists of 10,000 random 2D Gaussian noise images, where each RGB value of every pixel is sampled from an i.i.d Gaussian distribution with mean 0.5 and unit variance. We further clip each pixel value into the range $$.

Uniform Noise. The synthetic uniform noise dataset consists of 10,000 images where each RGB value of every pixel is independently and identically sampled from a uniform distribution on $$.

For hyperparameter tuning, we use a separate validation dataset iSUN (Xu et al., 2015), which is independent from the OOD test datasets. iSUN (Xu et al., 2015) consists of natural scene images. We include the entire collection of 8925 images in iSUN and downsample each image to size 3232 by 3232.

3 Evaluation metrics

We adopt the following four different metrics to measure the effectiveness of a neural network in distinguishing in- and out-of-distribution images.

FPR at 95%95\% TPR can be interpreted as the probability that a negative (out-of-distribution) example is misclassified as positive (in-distribution) when the true positive rate (TPR) is as high as 95%95\%.

Detection Error, i.e., PeP_{e} measures the misclassification probability when TPR is 95%. The definition of PeP_{e} is given by Pe=0.5(1−TPR)+0.5FPRP_{e}=0.5(1-\text{TPR})+0.5\text{FPR}, where we assume that both positive and negative examples have the equal probability of appearing in the test set.

AUROC is the Area Under the Receiver Operating Characteristic curve, which is also a threshold-independent metric (Davis & Goadrich, 2006). The ROC curve depicts the relationship between TPR and FPR. The AUROC can be interpreted as the probability that a positive example is assigned a higher detection score than a negative example (Fawcett, 2006). A perfect detector corresponds to an AUROC score of 100%100\%.

AUPR is the Area under the Precision-Recall curve, which is another threshold independent metric (Manning et al., 1999; Saito & Rehmsmeier, 2015). The PR curve is a graph showing the precision=TP/(TP+FP) and recall=TP/(TP+FN) against each other. The metric AUPR-In and AUPR-Out in Table 2 denote the area under the precision-recall curve where in-distribution and out-of-distribution images are specified as positives, respectively.

4 Experimental Results

Comparison with baseline. In Figure 1, we show the ROC curves when DenseNet-BC-100 is evaluated on CIFAR-10 (positive) images against TinyImageNet (negative) test examples. The red curve corresponds to the ROC curve when using baseline method (Hendrycks & Gimpel, 2017), whereas the blue curve corresponds to ODIN. We observe a strikingly large gap between the blue and red ROC curves. For example, when TPR=95%=95\%, the FPR can be reduced from 34%34\% to 4.2%4.2\% by using our approach.

Hyperparameters. We use a separate OOD validation dataset for hyperparameter selection, which is independent from the OOD test datasets. For temperature TT, we select among 1, 2, 5, 10, 20, 50, 100, 200, 500, 1000; and for perturbation magnitude ε\varepsilon we choose from 21 evenly spaced numbers starting from 0 and ending at 0.004. The optimal parameters are chosen to minimize the FPR at TPR 95% on the validation OOD dataset.

Main results. The main results are summarized in Table 2, where we use iSUN (Xu et al., 2015) as validation set. We use T=1000T=1000 for all settings. For DenseNet, we use ε=0.0014\varepsilon=0.0014 for CIFAR-10 and ε=0.002\varepsilon=0.002 for CIFAR-100. We provide additional details on the effect of parameters in Section 5. For each in- and out-of-distribution dataset pair, we report both the performance of the baseline (Hendrycks & Gimpel, 2017) and ODIN. In Table 2, we observe significant performance improvement across all dataset pairs.

Parameter transferability. In Table 3, we show how the parameters tuned on one validation set can generalize across datasets. Specifically, we tune the parameters using one validation dataset and then evaluated on the remaining OOD test datasets. The results are very similar across different validation sets, which suggests the insensitivity of our method w.r.t the tuning set.

Data distributional distance vs. detection performance. To measure the statistical distance between in- and out-of-distribution datasets, we adopt a commonly used metric, maximum mean discrepancy (MMD) with Gaussian RBF kernel (Sriperumbudur et al., 2010; Gretton et al., 2012; Sutherland et al., 2016). Specifically, given two image sets, V={v1,...,vm}V=\{v_{1},...,v_{m}\} and W={w1,...,wm}W=\{w_{1},...,w_{m}\}, the maximum mean discrepancy between VV and QQ is defined as

where k(⋅,⋅)k(\cdot,\cdot) is the Gaussian RBF kernel, i.e., k(x,x′)=exp⁡(−∥x−x′∥222σ2)k(x,x^{\prime})=\exp\left(-\frac{\|x-x^{\prime}\|_{2}^{2}}{2\sigma^{2}}\right). We use the same method used by Sutherland et al. (2016) to choose σ\sigma, where 2σ22\sigma^{2} is set to the median of all Euclidean distances between all images in the aggregate set V∪WV\cup W.

In Figure 2 (a)(b), we show how the performance of ODIN varies against the MMD distances between in- and out-of-distribution datasets. The datasets (on x-axis) are ranked in the descending order of MMD distances with CIFAR-100. There are two interesting observations can be drawn from these figures. First, we find that the MMD distances between the cropped datasets and CIFAR-100 tend to be larger. This is likely due to the fact that cropped images only contain local image context and are therefore more distinct from CIFAR-100 images, while resized images contain global patterns and are thus similar to images in CIFAR-100. Second, we observe that the MMD distance tends to be negatively correlated with the detection performance. This suggests that the detection task becomes harder as in- and out-of-distribution images are more similar to each other.

Discussions

In this subsection, we analyze the effectiveness of the temperature scaling method. As shown in Figure 3 (a) and (b), we observe that a sufficiently large temperature yields better detection performance although the effects diminish when TT is too large. To gain insight, we can use the Taylor expansion of the softmax score (details provided in Appendix B). When TT is sufficiently large, we have

by omitting the third and higher orders. For simplicity of notation, we define

Interpretations of U1U_{1} and U2U_{2}. By definition, U1U_{1} measures the extent to which the largest unnormalized output of the neural network deviates from the remaining outputs; while U2U_{2} measures the extent to which the remaining smaller outputs deviate from each other. We provide formal mathematical derivations in Appendix D. In Figure 5(a), we show the distribution of U1U_{1} for each out-of-distribution dataset vs. the in-distribution dataset (in red). We observe that the largest outputs of the neural network on in-distribution images deviate more from the remaining outputs. This is likely due to the fact that neural networks tend to make more confident predictions on in-distribution images.

Further, we show in Figure 5(b) the expectation of U2U_{2} conditioned on U1U_{1}, i.e., E[U2∣U1]E[U_{2}|U_{1}], for each dataset. The red curve (in-distribution images) has overall higher expectation. This indicates that, when two images have similar values on U1U_{1}, the in-distribution image tends to have a much higher value of U2U_{2} than the out-of-distribution image. In other words, for in-distribution images, the remaining outputs (excluding the largest output) tend to be more separated from each other compared to out-of-distribution datasets. This may happen when some classes in the in-distribution dataset share common features while others differ significantly. To illustrate this, in Figure 5 (f)(g), we show the outputs of each class using a DenseNet (trained on CIFAR-10) on a dog image from CIFAR-10, and another image from TinyImageNet (crop). For the image of dog, we can observe that the largest output for the label dog is close to the output for the label cat but is quite separated from the outputs for the label car and truck. This is likely due to the fact that, in CIFAR-10, images of dogs are very similar to the images of cats but are quite distinct from images of car and truck. For the image from TinyImageNet (crop), despite having one large output, the remaining outputs are close to each other and thus have a smaller deviation.

The effects of TT. To see the usefulness of adopting a large TT, we can first rewrite the softmax score function in Equation (3) as S∝(U1−U2/2T)/T{S\propto{(U_{1}-U_{2}/2T)/T}}. Hence the softmax score is largely determined by U1U_{1} and U2/2TU_{2}/2T. As noted earlier, U1U_{1} makes in-distribution images produce larger softmax scores than out-of-distribution images since S∝U1S\propto U_{1}, while U2U_{2} has the exact opposite effect since S∝−U2S\propto-U_{2}. Therefore, by choosing a sufficiently large temperature, we can compensate the negative impacts of U2/2TU_{2}/2T on the detection performance, making the softmax scores between in- and out-of-distribution images more separable. Eventually, when TT is sufficiently large, the distribution of softmax score is almost dominated by the distribution of U1U_{1} and thus increasing the temperature further is no longer effective. This explains why we see in Figure 3 (a)(b) that the performance does not change when TT is too large (e.g., T>100T>100). In Appendix C, we provide a formal proof showing that the detection error eventually converges to a constant number when TT goes to infinity.

2 Analysis on Input Preprocessing

As noted previously, using the temperature scaling method by itself can be effective in improving the detection performance. However, the effectiveness quickly diminishes as TT becomes very large. In order to make further improvement, we complement temperature scaling with input preprocessing. This has already been seen in Figure 4, where the detection performance is improved by a large margin on most datasets when T=1000T=1000, provided with an appropriate perturbation magnitude ε\varepsilon is chosen. In this subsection, we provide some intuition behind this.

The effects of gradient. In Figure 5 (c), we present the distribution of ∥∇xlog⁡S(x;T)∥1\|\nabla_{\bm{x}}\log S(\bm{x};T)\|_{1} — the 1-norm of gradient of log-softmax with respect to the input x\bm{x} — for all datasets. A salient observation is that CIFAR-10 images (in-distribution) tend to have larger values on the norm of gradient than most out-of-distribution images. To further see the effects of the norm of gradient on the softmax score, we provide in Figures 5 (d) the conditional expectation E[∥∇xlog⁡S(x;T)∥1∣S]E[\|\nabla_{\bm{x}}\log S(\bm{x};T)\|_{1}|S]. We can observe that, when an in-distribution image and an out-of-distribution image have the same softmax score, the value of ∥∇xlog⁡S(x;T)∥1\left\|\nabla_{\bm{x}}\log S(\bm{x};T)\right\|_{1} for in-distribution image tends to be larger.

We illustrate the effects of the norm of gradient in Figure 6. Suppose that an in-distribution image x1\bm{x}_{1} (blue) and an out-of-distribution image x2\bm{x}_{2} (red) have similar softmax scores, i.e., S(x1)≈S(x2)S(\bm{x}_{1})\approx S(\bm{x}_{2}). After input processing, the in-distribution image can have a much larger softmax score than the out-of-distribution image x2\bm{x}_{2} since x1\bm{x}_{1} results in a much larger value on the norm of softmax gradient than that of x2\bm{x}_{2}. Therefore, in- and out-of-distribution images are more separable from each other after input preprocessingSimilar observation can be seen when T=1T=1, where we present the conditional expectation of the norm of softmax gradient in Figure 5 (e)..

Related Works and Future Directions

The problem of detecting out-of-distribution examples in low-dimensional space has been well-studied in various contexts (see the survey by Pimentel et al. (2014)). Conventional methods such as density estimation, nearest neighbor and clustering analysis are widely used in detecting low-dimensional out-of-distribution examples (Chow, 1970; Vincent & Bengio, 2003; Ghoting et al., 2008; Devroye et al., 2013), . The density estimation approach uses probabilistic models to estimate the in-distribution density and declares a test example to be out-of-distribution if it locates in the low-density areas. The clustering method is based on the statistical distance, and declares an example to be out-of-distribution if it locates far from its neighborhood. Despite various applications in low-dimensional spaces, unfortunately, these methods are known to be unreliable in high-dimensional space such as image space (Wasserman, 2006; Theis et al., 2015). In recent years, out-of-distribution detectors based on deep models have been proposed. Schlegl et al. (2017) train a generative adversarial networks to detect out-of-distribution examples in clinical scenario. Sabokrou et al. (2016) train a convolutional network to detect anomaly in scenes. Andrews et al. (2016) adopt transfer representation-learning for anomaly detection. All these works require enlarging or modifying the neural networks. In a more recent work, Hendrycks & Gimpel (2017) found that pre-trained neural networks can be overconfident to out-of-distribution example, limiting the effectiveness of detection. Our paper aims to improve the performance of detecting out-of-distribution examples, without requiring any change to an existing well-trained model.

Our approach leverages the following two interesting observations to help better distinguish between in- and out-of-distribution examples: (1) On in-distribution images, modern neural networks tend to produce outputs with larger variance across class labels, and (2) neural networks have larger norm of gradient of log-softmax scores when applied on in-distribution images. We believe that having a better understanding of these phenomenon can lead to further insights into this problem.

Conclusions

In this paper, we propose a simple and effective method to detect out-of-distribution data samples in neural networks. Our method does not require retraining the neural network and significantly improves on the baseline method Hendrycks & Gimpel (2017) on different neural architectures across various in and out-distribution dataset pairs. We empirically analyze the method under different parameter settings, and provide some insights behind the approach. Future work involves exploring our method in other applications such as speech recognition and natural language processing.

Acknowledgments

The research reported here was supported by NSF Grant CPS ECCS 1739189.

References

Appendix A Supplementary Results in Section 5.1 and 5.2

Appendix B Taylor Expansion

In this section, we present the Taylor expansion of the soft-max score function:

Appendix C Proposition 1

The following proposition 1 shows that the detection error Pe(T,0)≈cP_{e}(T,0)\approx c if TT is sufficiently large. Thus, increasing the temperature further can only slightly improve the detection performance.

There exists a constant cc only depending on function U1U_{1}, in-distribution PXP_{\bm{X}} and out-of-distribution QXQ_{\bm{X}} such that lim⁡T→∞Pe(T,ε)=c\lim_{T\rightarrow\infty}P_{e}(T,\varepsilon)=c, when ε=0\varepsilon=0 (i.e., no input preprocessing).

as T→∞T\rightarrow\infty. This means that for a specific α>0\alpha>0, choosing the threshold δT=1/(N−α/T)\delta_{T}=1/(N-\alpha/T), then the false positive rate

Choosing α∗\alpha^{*} such that PX((N−1)U1(X)>α∗)=0.95P_{\bm{X}}\left((N-1)U_{1}(\bm{X})>\alpha^{*}\right)=0.95, then TPR(T)→0.95(T)\rightarrow 0.95 as T→∞T\rightarrow\infty and at the same time FPR(T)→QX((N−1)U1(X)>α∗)(T)\rightarrow Q_{\bm{X}}\left((N-1)U_{1}(\bm{X})>\alpha^{*}\right) as T→∞T\rightarrow\infty. There exists a constant cc depending on U1,PX,QXU_{1},P_{\bm{X}},Q_{\bm{X}} and PZP_{Z}, such that

Appendix D Analysis of Temperature

For simplicity of the notations, let Δi=fy^−fi\Delta_{i}=f_{\hat{y}}-f_{i} and thus Δ={Δi}i≠y^\Delta=\{\Delta_{i}\}_{i\neq\hat{y}}. Besides, let Δˉ\bar{\Delta} denote the mean of the set Δ\Delta. Therefore,

Appendix E Additional Results on Distance Measurement

Apart from the Maximum Mean Discrepancy, we also calculate the Energy distance between in- and out-of-distribution datasets. Let PP and QQ denote two different distributions. Then the energy distance between distributions PP and QQ is defined as

Therefore, the energy distance between two datasets V={V1,...,Vm}∼iidPV=\{V_{1},...,V_{m}\}\stackrel{{\scriptstyle iid}}{{\sim}}P and W={W1,...,Wm}∼iidQW=\{W_{1},...,W_{m}\}\stackrel{{\scriptstyle iid}}{{\sim}}Q is defined as

In the experiment, we use the 2-norm ∥⋅∥2\|\cdot\|_{2}.

Appendix F Additional Discussions

In this section, we present additional discussion on the proposed method. We first empirically show how the threshold δ\delta affects the detection performance. We next show how the proposed method performs when the parameters are tuned on a certain out-of-distribution dataset and are evaluated on other out-of-distribution datasets.

Effects of the threshold. We analyze how the threshold affects the following metrics: (1) FPR, i.e., the fraction of out-of-distribution images misclassified as in-distribution images; (2) TPR, i.e, the fraction of in-distribution images correctly classified as in-distribution images. In Figure 10, we show how the thresholds affect FPR and TPR when the temperature and perturbation magnitude are chosen optimally (i.e., T=1,000T=1,000, ε=0.0014\varepsilon=0.0014). From the figure, we can observe that the threshold corresponding to 95% TPR can produce small FPRs on all out-of-distribution datasets.

Difficult-to-classify images and difficult-to-detect images. We analyze the correlation between the images that tend to be out-of-distribution and images on which the neural network tend to make incorrect predictions. To understand the correlation, we devise the following experiment. For the fixed temperature TT and perturbation magnitude ε\varepsilon, we first set δ\delta to the softmax score threshold corresponding to a certain true positive rate. Next, we calculate the test accuracy on the images with softmax scores above δ\delta and the test accuracy on the images with softmax score below δ\delta, respectively. We report the results in Figure 11(a) and (b). From these two figures, we can observe that the images that are difficult to detect are more likely to be the images that are difficult to classify. For example, the DenseNet can achieve up to 98.5% test accuracy on the images having softmax scores above the threshold corresponding to 80% TPR, but can only achieve around 82% test accuracy on the images having softmax scores below the threshold corresponding to 80% TPR.