Estimating Example Difficulty Using Variance of Gradients

Chirag Agarwal, Daniel D'souza, Sara Hooker

Introduction

Over the past decade, machine learning models are increasingly deployed to high-stake decision applications such as healthcare , self-driving cars and finance . For gaining trust from stakeholders and model practitioners, it is important for deep neural networks (DNNs) to make decisions that are interpretable to both researchers and end-users. To this end, for sensitive domains, there is an urgent need for auditing tools which are scalable and help domain experts audit models.

Reasoning about model behavior is often easier when presented with a subset of data points that are relatively more difficult for a model to learn. Besides aiding interpretability through case-based reasoning , it can also be used to surface a tractable subset of atypical examples for further human auditing , for active learning to inform model improvements, and to choose not to classify some instances when the model is uncertain . One of the biggest bottlenecks for human auditing is the large scale of modern datasets and the cost of annotating individual features . Methods which automatically surface a subset of relatively more challenging examples for human inspection help prioritize limited human annotation and auditing time. Despite the urgency of this use-case, ranking examples by difficulty has had limited treatment in the context of deep neural networks due to the computational cost of ranking a high dimensional feature space.

Present work. A popular interpretability tool is saliency maps, where each of the features of the input data are scored based on their contribution to the final output . However, these explanations are typically for a single prediction and generated after the model is trained. Our goal is to leverage these explanations to automatically surface a subset of relatively more challenging examples for human inspection to help prioritize limited human annotation and auditing time. To this end, we propose a ranking method across all examples that instead measures the per-example change in explanations over training. Examples that are difficult for a model to learn will exhibit higher variance in gradient updates throughout training. On the other hand, the backpropagated gradients of the samples that are relatively easier will exhibit lower variance because the loss from these examples does not consistently dominate the model training.

We term this class normalized ranking mechanism Variance of Gradients (VoG) and demonstrate that VoG is a meaningful way for ranking data by difficulty and surfacing a tractable subset of the most challenging examples for human-in-the-loop auditing across a variety of large-scale datasets. VoG assigns higher scores to test set examples that are more challenging for the model to classify and proves to be an efficient tool for detecting out-of-distribution (OoD) samples. VoG is model and domain-agnostic as all that is required is the backpropagated gradients from the model.

Contributions. We demonstrate consistent results across two architectures and three datasets – Cifar-10, Cifar-100 and ImageNet . Our contributions can be enumerated as follows:

We present Variance of Gradients (VoG) – a class-normalized gradient variance score for determining the relative ease of learning data samples within a given class (Sec. 2). VoG identifies clusters of images with clearly distinct semantic properties, where images with low VoG scores feature far less cluttered backgrounds and more prototypical vantage points of the object (Fig. 4). In contrast, images with high VoG scores over-index on images with cluttered backgrounds and atypical vantage points of the object of interest.

VoG effectively surfaces memorized examples, i.e. it allocates higher scores to images that require memorization (Sec. 4). Further, VoG aids in understanding the model behavior at different training stages and provides insight into the learning cycle of the model.

We show the reliability of VoG as an OoD detection technique and compare its performance to 9 existing OoD methods, where it outperforms several methods, such as PCA and KDE . VoG presents an overall improvement of 9.26%9.26\% in precision compared to all other methods.

VoG Framework

We consider a supervised classification problem where a DNN is trained to approximate the function F\mathcal{F} that maps an input variable X\mathbf{X} to an output variable Y\mathbf{Y}, formally F:X↦Y\mathcal{F}:\mathbf{X}\mapsto\mathbf{Y}, where Y\mathbf{Y} is a discrete label vector associated with each input X\mathbf{X} and y∈Yy\in\mathbf{Y} corresponds to one of CC categories or classes in the dataset.

A given input image X\mathbf{X} can be decomposed into a set of pixels xix_{i}, where i={1,…,N}i=\{1,\dots,N\} and NN is the total number of pixels in the image. For a given image, we compute the gradient of the activation AplA_{p}^{l} with respect to each pixel xix_{i}, where ll designates the pre-softmax layer of the network and pp is the index of either the true or predicted class probability. We would like to note that the pre-softmax layer is responsible for connecting activations from previous layers in the network to individual class scores. Hence, computing the gradients w.r.t. this class indexed score measures the contribution of features to the final class prediction .

Note our goal is to rank examples, so for each example, we compute the pre-softmax activation gradient indexed at predicted/true label with respect to the input. This is far more computationally efficient than computing the full Jacobian matrix with individual layers.

Let S\mathbf{S} be a matrix that represents the gradient of AplA_{p}^{l} with respect to individual pixels xix_{i}, i.e. for an image of size 3×32×323{\times}32{\times}32, the gradient matrix S\mathbf{S} will be of dimensions 3×32×323{\times}32{\times}32.

This formulation may feel familiar as it is often computed based upon the weights of a trained model and visualized as a image heatmap for interpretability purposes . In contrast to saliency maps which are inherently local explanation tools, we are leveraging relative changes in gradients across training to rank all examples globally.

We average the pixel-wise variance of gradients to compute a scalar VoG score for the given input image:

where NN is the total number of pixels in a given image. First calculating the pixel-wise variance (Eqn. 3) and then average over the pixels (Eqn. 4) is consistent with previous XAI works where the gradients of an input image are computed independently for each pixel in an image .

In order to account for inherent differences in variance between classes, we normalize the absolute VoG score by class-level VoG mean and standard deviation. This amounts to asking: What is the variance of gradients for a given image with respect to all other exemplars of this class category?

In Fig. 1(a), we illustrate the principle and effectiveness of VoG in a controlled toy example setting. The data was generated using two separate isotropic Gaussian clusters. In such a simple low dimensional problem, the most challenging examples for the model to classify can be quantified by distance to the decision boundary. In Fig. 1(a), we visualize the trained decision boundary of a multiple layer perceptron (MLP) with a single hidden layer trained for 1515 epochs. We compute VoG for each training data point and plot final VoG score for each point against the distance to the trained boundary. In Fig. 1(b), we can see that VoG successfully ranks highest the examples closest to the decision boundary. The most challenging examples exhibit the greatest variance in gradient updates over the course of the training process. In the following sections, we will scale this toy problem and show consistent results across multiple architectures and datasets.

2 Experimental Setup

Datasets. We evaluate our methodology on Cifar-10 and Cifar-100 , and ImageNet datasets. For all datasets, we compute VoG for both training and test sets.

Cifar Training. We use a ResNet-18 network for both Cifar-10 and Cifar-100. For each dataset, we train the model for 350350 epochs using stochastic gradient descent (SGD) and compute the input gradients for each sample every 1010 epochs. We implemented standard data augmentation by applying cropping and horizontal flips of input images. We use a base learning rate schedule of 0.10.1 and adaptively change to 0.010.01 at 150th150^{\text{th}} and 0.0010.001 at 250th250^{\text{th}} training epochs. The top-1 test set accuracy for Cifar-10 and Cifar-100 were 89.57%89.57\% and 66.86%66.86\% respectively.

ImageNet Training. We use a ResNet-50 model for training on ImageNet. The network was trained with batch normalization , weight decay, decreasing learning rate schedules, and augmented training data. We train for 32,00032,000 steps (approximately 9090 epochs) on ImageNet with a batch size of 10241024. We store 3232 checkpoints over the course of training, but in practice observe that VoG ranking is very stable computed with as few as 33 checkpoints. Our model achieves a top-1 accuracy of 76.68%76.68\% and top-5 accuracy of 93.29%93.29\%.

Number of checkpoints. The number of checkpoints used to compute VoG balances efficiency for practitioners to use with the robustness of ranking. This can be set by the practitioner, and we note that in practice the last 3 checkpoints are sufficient for a robust VoG ranking (minimal difference when restricting to the last 3 in Figs. 5b,8b,11b vs. evaluating on all checkpoints in Fig. 4). For all experiments, VoG(early-stage) is computed using checkpoints from the first 3 epochs and VoG(late-stage) is computed using checkpoints from the last 3 epochs. The test set accuracy at the early-stage is 44.65%44.65\%, 14.16%14.16\%, and 51.87%51.87\% for Cifar-10, Cifar-100, and ImageNet, respectively. In the late-stage it is 89.57%89.57\%, 66.86%66.86\%, and 76.68%76.68\% for Cifar-10, Cifar-100, and ImageNet, respectively.

Utility of VoG as an Auditing Tool

In this section, we evaluate the merits of VoG as an auditing tool. Specifically, we (1) present the qualitative properties of images at both ends of the VoG spectrum, (2) measure how discriminative VoG is at separating easy examples from difficult, (3) quantify the stability of the VoG ranking, (4) use VoG as an auditing tool for test dataset, and (5) leverage VoG to understand the training dynamics of a DNN.

1) Qualitative inspection of ranking. A qualitative inspection of examples with high and low VoG scores shows that there are distinct semantic properties to the images at either end of the ranking. We visualize 2525 images ranked lowest and highest according to VoG for both the entire dataset (visualized for ImageNet in Fig. 7) and for specific classes (visualized for ImageNet in Fig. 3 and for Cifar-10 and Cifar-100 in Fig. 2). Images with low VoG score tend to have uncluttered and often white backgrounds with the object of interest centered clearly in the frame. Images with the high VoG scores have cluttered backgrounds and the object of interest is not easily distinguishable from the background. We also note that images with high VoG scores tend to feature atypical vantage points of the objects such as highly zoomed frames, side profiles of the object or shots taken from above. Often, the object of interest is partially occluded or there are image corruptions present such as heavy blur.

2) Test set error and VoG. A valuable property of an auditing tool is to effectively discriminate between easy and challenging examples. In Fig. 4, we plot the test set error of examples bucketed by VoG decile. Note that we plot error, so lower is better. We show that examples at the lowest percentiles of VoG have low error rates, and misclassification increases with an increase in VoG scores. Our results are consistent across all datasets, yet the trend is more pronounced for more complex datasets such as Cifar-100 and ImageNet. We ascribe this to differences in underlying model complexity. Furthermore, in Fig. 10, we observe that test set error on the lowest VoG scored images are lower than the baseline test set performance.

3) Stability of VoG ranking. To build trust with an end-user, a key desirable property of any auditing tool is consistency in performance. We would expect a consistent method to produce a ranking with a closely bounded distribution of scores across independently trained runs for a given model and dataset. To measure the consistency of the VoG ranking, we train five Cifar-10 networks from random initialization following the training methodology described in Sec. 2.2. Empirically, Fig. 6 shows that VoG rankings evidence a consistent distribution of test-error at each percentile given the same model and dataset. For completeness, we also measure instance-wise VoG stability by computing the standard deviation of VoG scores for 50k Cifar-10 samples across 10 independent initializations. The standard deviation of the VoG scores is negligible with a mean deviation of 3.81e−9{3.81e^{-9}} across all samples. In addition, we find similar results for Cifar-100 dataset where the output VoG scores are stable (mean std of 9.6e−69.6e{-}6) across different model initializations. Finally, we extend our stability experiments to understand the effect of different training hyperparameter settings (e.g., batch size) on the VoG scores. Here, we train 5 Cifar-10 models using different batch sizes, i.e., {128, 256, 384, 512, 640}, and find that the mean VoG standard deviation across 50k Cifar-10 samples was 1.9e−51.9e{-}5.

4) VoG as an unsupervised auditing tool. Many auditing tools used to evaluate and understand possible model bias require the presence of labels for protected attributes and underlying variables. However, this is highly infeasible in real-world settings . For image and language datasets, the high dimensionality of the problem makes it hard to identify a priori what underlying variables one needs to be aware of. Even acquiring the labels for a limited number of attributes protected by law (gender, race) is expensive and/or may be perceived as intrusive, leading to noisy or incomplete labels . This means that ranking techniques which do not require labels at test time are very valuable.

One key advantage of VoG is that we show it continues to produce a reliable ranking even when the gradients are computed w.r.t. the predicted label. In Fig. 7, we include the top and bottom 25 VoG ImageNet test images using predicted labels from the model. Finally, we also computed the mean test-error for the predicted VoG distribution, and find that it also effectively discriminates between top-10 and bottom-10 examples, respectively (Fig. 12(a)).

5) VoG understands early and late training dynamics. Recent works have shown that there are distinct stages to training in deep neural networks . To this end, we investigate whether VoG rankings are sensitive to the stage of the training process. We compute VoG separately for two different stages of the training process: (i) the Early-stage (first three epochs) and (ii) the Late-stage (last three epochs). We plot VoG scores against the test set error at each decile in early- and late-stage and find a flipping behavior across all datasets and networks (Fig. 5 for ImageNet, Fig. 8 for Cifar-100, and Fig. 11 for Cifar-10). In the early training stage, samples having higher VoG scores have a lower average error rate as the gradient updates hinge on easy examples. This phenomenon reverses during the late-stage of the training, where, across all datasets, high VoG scores in the late-stage have the highest error rates as updates to the challenging examples dominate the computation of variance. Further, we note a noticeable visual difference between the image ranking computed for early- and late-stages of training. As seen in Fig. 2, for some classes such as apple, it appears that VoG scores also capture the network’s color bias during the early training stage, where images with the lowest VoG scores over-index on red-colored apples.

Relationship between VoG Scores and Memorized/OoD Examples

Recent works have highlighted that DNNs produce uncalibrated output probabilities that cannot be interpreted as a measure of certainty . To this end, we argue that if VoG is a reliable auditing tool, it should capture model uncertainty even when it’s not reflected in the output probabilities. We consider VoG rankings on a task where the network produces highly confident predictions for incorrect/out-of-distribution inputs and evaluate VoG on two separate tasks: (1) identifying examples memorized by the model and (2) detecting out-of-distribution examples.

Overparameterized networks have been shown to achieve zero training error by memorizing examples . We explore whether VoG can distinguish between examples that require memorization and the rest of the dataset. To do this, we replicate the general experiment setup of Zhang et al. and replace 20%20\% of all labels in the training set with randomly shuffled labels. We re-train the model from random initialization and compute VoG scores across training for all examples in the training set. Our network achieves 0%0\% training error which would only be possible given successful memorization of the noisy examples with shuffled labels. We now answer the question: Is VoG able to discriminate between these memorized examples and the rest of the dataset?

We perform a two-sample tt-test with unequal variances and show that this difference is statistically significant at a pp-value of 0.0010.001, i.e. shuffled labels have a different VoG distribution than the non-shuffled dataset. Intuitively, the two-sample tt-test produces a pp-value that can be used to decide whether there is evidence of a significant difference between the two distributions of VoG scores. The pp-value represents the probability that the difference between the sample means is large, i.e. the smaller the pp-value, the stronger is the evidence that the two populations have different means. For both Cifar-10 and Cifar-100, we find a statistically significant difference in VoG scores for each population (pp-value is <0.001<0.001), which shows that VoG is discriminative at distinguishing between memorized and non-memorized examples. We include more details about the statistical testing in Sec. A.3.

2 Out-of-Distribution detection

We have already established that VoG is very effective at distinguishing between easy and challenging examples (Fig. 10). Here, we ask whether this makes VoG an effective out of distribution (OoD) detection tool. It also gives us a setting in which to compare VoG as a ranking mechanism to other methods

Ruff et al. benchmark a variety of OoD detection techniques on MNIST-C . For completeness, we replicate this precise setup by using a trained LeNet model and evaluate VoG on MNIST-C against 9 other methods.

Evaluation metrics. We evaluate OoD detection performance using the following metrics:

i) AUROC. The Area Under the Receiver Operator Characteristic (AUROC) curve can be interpreted as the probability that a positive example is assigned a higher detection score than a negative example .

ii) AUPR (In). The Area Under the Precision Recall (AUPR) curve computes the precision-recall pairs for different probability thresholds by considering the in-distribution examples as the positive class.

iii) AUPR (Out). AUPR (Out) is AUPR as described above, but calculated considering the OoD examples as the positive class. We treat this outlier class as positive by multiplying the VoG scores by −1-1 and labelling them positive when calculating AUPR (Out).

Findings. In Table 1, we observe that VoG outperforms all methods except AutoEncoders (AE) and AutoEncoder GAN (AEGAN). In stark contrast to VoG, AE and AEGAN require complex training of auxiliary models and do not feasibly scale beyond small-scale datasets like MNIST. Given these limitations, VoG remains a valuable and scalable OoD detection method as it can be used for large-scale datasets (e.g. ImageNet) and networks (e.g. ResNet-50). Unlike generative models, VoG does not require an uncorrupted training dataset for learning image distributions. Further, VoG only leverages data from training itself, is computed from checkpoints already stored over the course of training, and does not require the true label to rank.

Related Work

Our work proposes a method to rank training and testing data by estimating example difficulty. Given the size of current datasets, this can be a powerful interpretability tool to isolate a tractable subset of examples for human-in-the-loop auditing and aid in curriculum learning or distinguishing between sources of uncertainty . While prior works have proposed different notions of what subset merits surfacing, introduced the concept of prototypes and quintessential examples in the dataset, but did not focus on large-scale deep neural networks models .

Unlike previous works, we propose a measure that can be extended to rank the entire dataset by estimating example difficulty (rather than surfacing a prototypical subset). In addition, VoG is far more efficient than other global rankings like and .

VoG also does not require modifying the architecture or making any assumptions about the statistics of the input distribution. In particular, works such as require assumptions about the statistics of the input distribution and requires modifying the architecture to prefix an autoencoder to surface a set of prototypes, leverages pruning of the model to identify difficult examples and requires the addition of an auxiliary k-nn model after each layer.

Our work is complementary to recent works by that proposes a c-score to rank examples by aligning them with training instances, that classifies examples as outliers according to sensitivity to varying model capacity, and that considers different measures to isolate prototypes for ranking the entire dataset. We note that the c-score method proposed by is considerably more computationally intensive to compute than VoG as it requires training up to 20,000 network replications per dataset. Several of the prototype methods considered by require training ensembles of models, as does the compression sensitivity measure proposed by . Finally, our proposed VoG is both different in the formulation and can be computed using a small number of existing checkpoints saved over the course of training.

Conclusion and Future Work

In this work, we proposed VoG as a valuable and efficient way to rank data by difficulty and surface a tractable subset of the most challenging examples for human-in-the-loop auditing. High VoG samples are challenging to classify for algorithm and surfaces clusters of images with distinct visual properties. Moreover, VoG is domain agnostic as it uses only the vanilla gradient explanation from the model, and can be used to rank both training and test examples. We show that it is also a useful unsupervised protocol, as it can effectively rank examples using the predicted label.

References

Appendix A Appendix

We generate the clusters for classification using scikit-learn and use a 90-10% split for dividing the dataset into train and test set Code and datasets available at https://github.com/chirag126/VOG.git. We train a linear Multiple Layer Perceptron network with a hidden layer of 10 neurons using Stochastic Gradient Descent optimizer for 15 epochs. We divided the training process into three epoch stages: (1) EarlyEarly [0, 5), (2) MiddleMiddle [5, 10), and (3) LateLate stage [10, 15). The trained model achieves a 0% test set error using a linear boundary (Fig. 1(a)).

A.2 Class Level Error Metrics and VoG

Here, we explore whether VoG is able to capture class level differences in difficulty. We compute VoG scores for each image in the test set of Cifar-10 and Cifar-100 (both test sets have 10,00010,000 images). In Fig. 9, we plot the average absolute VoG score for each class against the false negative rate for each class. We find that there is a positive, albeit weak, correlation between the two, classes with higher VoG scores have higher mis-classification error rate. The correlation between these metrics is 0.650.65 and 0.590.59 for Cifar-10 and Cifar-100 respectively. Given that VoG is computed on a per-example level, we find it interesting that the aggregate average of VoG is able to capture class level differences in difficulty.

A.3 Statistical Significance of Memorization Experiments

The two-sample tt-test produces a pp-value that can be used to decide whether there is evidence of a significant difference between the two distributions of VoG scores. The pp-value represents the probability that the difference between the sample means is large, i.e. smaller the pp-value, stronger is the evidence that the two populations have different means.

Null Hypothesis: μ1=μ2\mu_{1}=\mu_{2} Alternative Hypothesis: μ1≠μ2\mu_{1}\neq\mu_{2}

If the pp-value is less than your significance level (α=0.05\alpha=0.05 in this experiment), you can reject the null hypothesis, i.e. the difference between the two means is statistically significant. The details for the individual tt-tests for Cifar-10 and Cifar-100 are given below:

Cifar-10: The statistics for the samples in the correct and shuffled labels are: Corrected labels: μ1=0.62\mu_{1}=0.62; σ1=0.54\sigma_{1}=0.54; N1=40000N_{1}=40000 Shuffled labels: μ2=0.85\mu_{2}=0.85; σ2=0.75\sigma_{2}=0.75; N2=10000N_{2}=10000 Result: pp-value is <0.001<0.001 — Reject Null Hypothesis (the two populations have different VoG means)

Cifar-100: The statistics for the samples in the correct and shuffled labels are: Corrected labels: μ1=0.54\mu_{1}=0.54; σ1=0.46\sigma_{1}=0.46; N1=40000N_{1}=40000 Shuffled labels: μ2=0.82\mu_{2}=0.82; σ2=0.71\sigma_{2}=0.71; N2=10000N_{2}=10000 Result: pp-value is <0.001<0.001 — Reject Null Hypothesis (the two populations have different VoG means)

A.4 Early training dynamics of Deep Neural Networks

Following Sec. 8, we plot the relationship between VoG and error rate of the testing dataset for Cifar-10 and Cifar-100. As in ImageNet, we observe a flipping trend between the early and late stages for both datasets (Figs. 8,11). We find that for easier datasets like Cifar, this point is only seen on using a lower learning rate (1e-3 in our experiments) for the early training stages.

A.5 Detection of Distribution Shifts

We consider ImageNet-O , an open source curated out-of-distribution (OoD) dataset designed to fool classifiers. ImageNet-O consists of images that are not included in the original 10001000 ImageNet classes. These images were selected with the goal of producing high confidence incorrect ImageNet-1K predictions of labels from within the training distribution. We are interested in understanding if VoG can correctly rank ImageNet-O examples as being atypical or OoD and expect to observe that ImageNet-O examples would be over-represented in top percentiles of VoG scores. In Fig. 12(b), we observe that the percentage of ImageNet-O images are relatively over-represented at high levels of VoG, with 30% of all images in the top-25th percentile vs 24% in the bottom 25th percentile.

A.6 Out-of-Distribution Detection (OoD) Datasets and Model Architectures

Here, we carry out additional experiments to measure the effectiveness of VoG to detect OoD data. We run experiments using three DNN architectures: ResNet-18 , DenseNet and WideResNet , and benchmark against Maximum Softmax Probability (MSP) , which is widely considered a strong baseline in OoD detection We follow the setup in by setting all test set examples in CIFAR-10 as in-distribution (positive). For OoD examples (negative), we benchmark across four datasets: CIFAR-100, iSUN , TinyImageNet (Resize) , LSUN (Resize) , and Gaussian Noise. The Gaussian dataset was generated as described in , with N(0.5,1)\mathcal{N}(0.5,1). For the various ablations, the size of the OoD dataset can be seen in Table 2.

Findings. From Table 3, we observe that VoG is a valuable ranking for OoD detection and improves upon state-of-the-art uncertainty measures for many different tasks. On average, VoG outperforms MSP by large margins with a mean gain of 2.62% in AUROC, 2.33% in AUPR/In, and 2.47% in AUPR/Out across all three architectures and five datasets.