Can You Trust This Prediction? Auditing Pointwise Reliability After Learning
Peter Schulam, Suchi Saria
Introduction
Machine learning is quickly becoming an important component of everyday software systems. This is due, in large part, to advances in tools that make it easy to express and train models. Using automatic differentiation software, a programmer can describe the “forward pass” of a predictive model and then train the model using gradient-based methods without much additional effort (e.g. Maclaurin et al. 2015; Abadi et al. 2016). As the barriers to building machine learning systems become lower, there is rising excitement around the idea of applying the technology in high-impact domains (e.g. Soleimani et al. 2018; Lipton et al. 2018).
Tools for quickly building machine learning systems, however, have generally outpaced the growth and adoption of tools to understand whether they are reliable. The average error on heldout training data is a common measure of model performance, but is inadequate for ensuring model reliability. For instance, low average error on heldout training data does not give assurances under data set shift (see e.g. Quiñonero Candela et al. 2009). A model with low average error can also still make large pointwise errors on individual predictions for some types of inputs (see e.g. Nushi et al. 2018).
Reliability is commonly addressed during learning. In this paper, we focus on methods that improve reliability by preventing errors at test time (there are other ways to make a system reliable; Saria and Subbaswamy 2019). For instance, we can proactively account for data set shift by using learning algorithms that adjust for possible causes of the difference between train and test distributions (e.g. Schulam and Saria 2017; Subbaswamy and Saria 2018; Subbaswamy et al. 2019). To detect pointwise errors, we can use specialized learning algorithms that produce models equipped with uncertainty estimates of predictions (e.g. Gal and Ghahramani 2016; Lakshminarayanan et al. 2017). More generally, we can build confidence in a model by criticizing and refining it (e.g. Blei 2014; Kim et al. 2016).
In this paper, we consider the problem of improving model reliability after learning with tools that audit the predictions of a fixed model. To perform the audit, we introduce the Resampling Uncertainty Estimation (RUE) algorithm to calculate an uncertainty score for each prediction at test time. RUE makes minimal assumptions about the structure of the model and how it was learned. This makes RUE easy to use alongside the rich ecosystem of tools for training models (like automatic differentiation).
The most common approach in machine learning to assessing uncertainty is using Bayesian methods (e.g. Graves 2011; Blundell et al. 2015; Gal and Ghahramani 2016). This approach is tightly coupled to the model structure and the learning algorithm. As a result, Bayesian methods cannot be applied after learning as an auditing tool. Many non-Bayesian methods like the bootstrap (Efron and Tibshirani 1986) or deep ensembles (Lakshminarayanan et al. 2017) also depend on specific train-time procedures, and so cannot be applied after learning either.
There are several existing tools for improving model reliability after learning. For instance, Courrieu 1994 computes the convex hull of inputs in the training data and uses it to define a region wherein outcomes can be safely interpolated by a neural network. Bishop 1994 generalizes this idea by using density estimation to characterize regions where there are sufficient training samples to make reliable predictions. This approach can be further generalized using ideas from novelty detection where the aim is to characterize the support of a distribution without an explicit generative model (see e.g. Schölkopf et al. 2000). Another approach is to fit a second model to errors made by the first, which is related to stacked generalization (Wolpert 1992).
One limitation of many existing tools for improving model reliability after learning is that they rely on Euclidean distance, which often ignores model-specific structure. For example, the predictive model may use an internal representation wherein similar points (according to the internal representation) are in fact far away in Euclidean space. This is an issue especially for deep learning, where the empirical successes are thought to be due to rich internal representations of the data that are learned to reflect task-specific invariances (Bengio et al. 2013; Farabet et al. 2013). Characterizing similarity among inputs for these models is difficult because it is not clear how internal representations can be extracted (see e.g. Hinton et al. 2015; Mahendran and Vedaldi 2015).
In Section 3.2, we show that RUE’s uncertainty score reflects two important criteria to judge whether a prediction is reliable: the density criterion and local fit criterion (Leonard et al. 1992). The density criterion states that a prediction at input is reliable if there are samples in the training data that are similar to . The local fit criterion states that a prediction at is reliable if the model has small error on samples in the training data similar to . These criteria hinge upon a definition of similarity, and we show that RUE implements these criteria using a model-dependent inner product to measure similarity.
Contributions. We describe Resampling Uncertainty Estimation (RUE), an algorithm used to audit the pointwise reliability of predictions by computing an uncertainty score. To quantify uncertainty, the RUE algorithm uses the gradients and Hessian of the model’s loss on training data (Line 2 in Algorithm 1) and bootstrap samples (Line 7) to produce an ensemble of predictions for a test input (Lines 9-11). Intuitively, this estimates the amount that the prediction would change had the model been fit on different data sets drawn from the same distribution as the training data. This intuition has connections to the classical bootstrap (Efron and Tibshirani 1986), and in Section 3.1 we show that RUE can be derived as an approximation to this procedure that does not require fitting multiple models. We also show that RUE implements the density and local fit criteria using a model-dependent inner product as a measure of similarity. Experimentally, we show that RUE more effectively detects inaccurate predictions than existing methods for improving reliability after learning. Moreover, we show that RUE’s uncertainty score can create predictive distributions that are well-calibrated and competitive with state-of-the-art integrated methods (i.e. methods like Hernández-Lobato and Adams 2015; Gal and Ghahramani 2016; and Lakshminarayanan et al. 2017) that estimate uncertainty during learning.
Related Work
The dominant approach to modeling uncertainty in machine learning is to use Bayesian inference. Gaussian processes are a flexible class of Bayesian nonparametric priors over functions and provide a natural framework for reasoning about posterior predictive distributions at test time given training data (Rasmussen and Williams 2006). Bayesian neural networks are another important class of regression models, but have historically been much more challenging to work with given the nonlinearities in the likelihood function. Early important work on this subject was done by Buntine and Weigend 1991 and by MacKay 1992, who proposed using a Laplace approximation to the posterior distribution in order to obtain posterior predictive distributions and estimates of the evidence (marginal probability of the data) for model selection. A key challenge with this technique is that it requires computing and inverting the Hessian with respect to the network parameters, which can be prohibitively expensive for larger neural networks. There has, however, been recent work in scalable approximations to the Hessian for second-order optimization (e.g. Martens and Grosse 2015; Botev et al. 2017), which could make the Laplace approximation to the posterior feasible (Ritter et al. 2018).
Efforts to scale Bayesian inference to modern neural networks have revolved around mean field variational approximations to the posterior distribution (e.g. Graves 2011; Blundell et al. 2015). There has also been work to apply expectation propagation to Bayesian neural networks (Hernández-Lobato and Adams 2015; Hernández-Lobato et al. 2016). Dropout, a non-Bayesian technique for regularizing neural networks (Srivastava et al. 2014), was recently shown to approximate a variational approximation of the posterior (Gal and Ghahramani 2016), and the insight led to an approach for computing posterior predictive distributions using a simple simulation algorithm that leverages existing software for fitting neural networks that implements dropout layers. There has also been recent work to scale MCMC for posterior inference by designing approximate transition operators that trade off between fidelity to the desired stationary distribution and computational cost (e.g. Welling and Teh 2011; Ahn et al. 2012). Although the Bayesian approach is conceptually appealing, it is difficult to place meaningful priors on neural networks (see e.g. Buntine and Weigend 1991).
Many authors have explored alternatives to Bayesian methods for modeling uncertainty in predictive models. Early work by Leonard et al. 1992 explored heuristic scores that reflect the density and local fit principles. Bishop 1994 proposes a technique for identifying “novel” inputs as a way to detect when a model is likely to be unreliable. Hooker 2004 proposes a tree-based density estimator for detecting novel inputs. More recently, Hendrycks and Gimpel 2016 proposed a simple heuristic using softmax outputs to detect misclassifications and out-of-distribution samples for classification problems. They also propose an “abnormality module” for detecting novel inputs that is similar in spirit to the validity index network proposed by Leonard et al. 1992. Guo et al. 2017 investigate whether softmax outputs are calibrated estimates of the conditional distribution . They find that the softmax probabilities are not calibrated in large neural networks and propose an approach related to Platt scaling (Platt 1999) to address the issue. Ensembles can also be used to express uncertainties (Breiman 1996; Lakshminarayanan et al. 2017).
The bootstrap (Efron and Tibshirani 1986) is a classic approach to estimating uncertainty that uses Monte Carlo simulation to estimate the sampling distribution of arbitrary functions of a data set. It cannot be used after learning because we must fit many models at train time rather than just one (this also makes it computationally expensive). We show that RUE is an approximation of the boostrap by building upon sensitivity analysis ideas pioneered by Cook 1986 that were recently extended to modern machine learning problems by Koh and Liang 2017 in order to shed light on a model’s provenance and improve interpretability.
There are several other threads of research related to improving model reliability. One issue that in classification problems is when the current set of labels that a model assigns to inputs is not exhaustive. If this occurs, a model that automatically identifies and learns to classify new categories can make the system more robust (e.g. Bendale and Boult 2016; Liu et al. 2018). It is also useful for a machine learning system to implement a policy for rejecting an input due to uncertainty. Such a policy can be easy to implement if the model outputs well-calibrated probability distributions (e.g. Chow 1970), but models can also be trained to directly minimize a risk function incorporating the reject option (e.g. Herbei and Wegkamp 2006; Bartlett and Wegkamp 2008; Grandvalet et al. 2009).
Resampling Uncertainty Estimation
In this section, we connect RUE to Efron’s bootstrap procedure. The bootstrap approximates the sampling distribution of an estimate by repeatedly simulating bootstrap samples, which are new data sets of size created by sampling with replacement from the uniform distribution over the original data set. To bootstrap a supervised learning algorithm, one would need to sample bootstrap datasets and run the learning procedure from scratch each time. When learning on the original sample, let denote the learned parameters that satisfy
The bootstrap is equivalent to choosing new weights that count the number of times a particular example is included in the bootstrap sample and then refitting the objective. Considering our objective function as a function of both the model parameters and sample weights , the following equality holds at the solution :
If we assume that is smooth in (this amounts to assuming that the loss and regularizer are twice continuously differentiable), we can use the implicit function theorem to conclude that there exists a local function such that . In other words, there is a map from the sample weights to learned model parameters. If we knew this map, then the bootstrap procedure would be straightforward: sample new weights and map them to the parameters . Our approach approximates this function using a first-order Taylor expansion around , which means we can estimate the parameters we would obtain with a bootstrap sample as
where the equality above also follows from the implicit function theorem (recall that is the Hessian of the objective evaluated at ). The Jacobian of can be computed once (this is the matrix in Algorithm 1), then sampling only requires drawing new weights and a matrix multiplication. To approximate the sampling distribution of the model’s prediction at a new point , we compute the predictions made by a collection of models with parameters .
2 Density and Local Fit Criteria
where is the base measure of the exponential family, is the natural sufficient statistic, the predicted value is the natural parameter, and is the log normalizing constant. Our argument rests upon the following claim.
As the number of training samples grows, we can approximate as
Let denote a parameter vector resampled according to lines 7-8 in Algorithm 1, then
Definine the kernel as in 8 and write the matrix expression as a sum of scalars to complete the argument. ∎
Note that the approximation of resembles a kernel smoothing estimate of the exponential family residual at a test input . Interpeting the kernel function as a measure of similarity between the test point and training point , we can see that reflects the local fit criterion. In other words, the uncertainty score will be smaller (larger) if the model’s predictions have small (large) errors on the training samples that are similar to .
Finally, we note that the similarity metric depends on both the structure of the model and the learned parameter values through the gradient and Hessian. We therefore see that RUE uses a measure of similarity that is model-dependent for both the local fit and density criteria, which is a distinguishing feature among tools for improving reliability after learning.
2.2 Related Uncertainty Scores
RUE has connections to robust statistics. When the learning objective is the log likelihood of independent samples, the negative Hessian of the objective is known as the Fisher information matrix, and it is often used to approximate the covariance matrix under the sampling distribution of a maximum likelihood estimator (Tibshirani 1996). In likelihood theory, an alternative estimator of the covariance matrix of the parameters that is robust to model misspecification is the sandwich covariance estimator (Kent 1982), which is . Recall that we arrived at a similar expression in Equation 7 by letting . This draws a connection between RUE and robust statistics (Huber and Ronchetti 2009). In our experiments, we find that RUE outperforms the Laplace approximation and these connections to robust statistics may help to develop theory that explains this observation.
2.3 Scalability
To address storage and computation issues related to the Hessian, we note that we do not need to explicitly store the Hessian so long as we have access to the model code and can run backwards-mode automatic differentiation, which allows us to efficiently compute Hessian-vector products (Pearlmutter 1994). Moreover, we can perform the required inversion approximately using conjugate gradients (which only requires Hessian-vector products), or can use more sophisticated techniques discussed by Koh and Liang 2017. Alternatively, we could use recently developed approximations to the Hessian such as K-FAC (Martens and Grosse 2015). We leave exploration and evaluation of these approximations for future work.
Experiments
Overview. We first evaluate RUE on an auditing task where the goal is to detect when model predictions will be wrong. We compare to existing methods for improving reliability after learning, and show that RUE generally outperforms them. Next, we show that RUE can create predictive distributions over outputs that are competitive with state-of-the-art integrated methods that rely on particular algorithms at train time (e.g. Bayesian inference).
We run our experiments on eight common benchmark regression data sets (e.g. Hernández-Lobato and Adams 2015; Gal and Ghahramani 2016; Lakshminarayanan et al. 2017). The data sets we use are: Boston housing, concrete compression strength, energy efficiency, robot arm kinematics, naval propulsion, combined cycle power plant, red wine quality, and yacht hydrodynamics. These data sets are all available from https://archive.ics.uci.edu/ml/index.php For each data set, we sample twenty train-test splits. For the smaller data sets (housing, concrete, energy, and yacht), we sample of the points for training. The remaining data sets are larger and so we use smaller percentages of the data set for each training split in order to simulate a small data regime, where uncertainty scores are most useful. For the larger data sets we sample approximately training points per split and test on the remaining points.
To control for model structure, learning objective, and optimization algorithm across all data sets and uncertainty estimation algorithms, we fix a single feedforward neural network architecture with a single hidden layer comprised of 50 hidden units, softplus nonlinearities, and loss. We regularize with weight decay and set the penalty . Moreover, we use a single set of learning hyperparameters: we learn for a total of epochs with samples per minibatch, and use the default settings () of Adam (Kingma and Ba 2015) with a fixed learning rate of . We fit a single model for each split of each data set, and compute uncertainty scores for the predictions made by that model on the test set using the four alternative uncertainty score methods (RUE, Laplace, KDE, and Bootstrap SGD).
RUE is better at detecting errors than Laplace, KDE, and Bootstrap SGD for most data sets and most error tolerances. Note that the range of AUC scores varies widely across data sets. For instance, the RUE scores are able to discriminate well across all thresholds in the energy data set. The power data set, however, is challenging; all AUCs are .
2 RUE Predictive Distributions
Existing methods for quantifying uncertainty output probabilities. Although the methods we compare to in the first experiments above do not output probability distributions, there is a natural way to create one for the methods that use ensembles of model predictions (RUE, Laplace, and Bootstrap SGD). When the observation noise is additive—that is, —the posterior variance of a prediction at input for a Bayesian model is the posterior variance of plus the variance of the observation noise . For the ensemble methods, the uncertainty scores can be seen as approximations of the variance of , which we denote using . We can also estimate the variance of using the residuals on the training data, which we denote using . A natural predictive distribution is a Gaussian centered at the prediction under the learned model parameters with variance . We use this predictive distribution to compute negative log likelihoods (NLL) and compare the methods used after learning to state-of-the-art alternatives that estimate uncertainty during training using specialized objectives (Gal and Ghahramani 2016; Hernández-Lobato and Adams 2015; Lakshminarayanan et al. 2017), which we refer to as integrated methods.
Table 1 shows the results of this experiment. The negative log likelihoods for the integrated methods are the results reported in Table 1 of Lakshminarayanan et al. 2017. We find that RUE has better NLL than Laplace and Bootstrap SGD on all but two data sets. The improvements over Bootstrap SGD, however, are small, but recall that RUE generally outperformed Bootstrap SGD in the detection experiments shown in Figure 2. We also compare RUE to three integrated methods: Probabilistic Backpropagation (Hernández-Lobato and Adams 2015), Monte Carlo Dropout (Gal and Ghahramani 2016), and Deep Ensembles (Lakshminarayanan et al. 2017). We find that RUE is competitive with these methods (e.g., on Housing, Power and Wine), but does not outperform them. We are not surprised by this result because integrated methods jointly fit predictions and uncertainty using a specialized objective. Furthermore, we suspect that RUE’s negative log likelihood would be further improved if the learning rates and regularization weights were tuned using cross validation for each data set (these choices were fixed upfront at a single setting for all data sets in order to control for effects of learning algorithms).
Conclusion
We introduced RUE, an algorithm that audits a model by estimating the uncertainty of predictions without changing the learning algorithm. This makes it easy to use RUE alongside the rich ecosystem of existing tools for training machine learning models (e.g. automatic differentiation libraries). This is in contrast to methods like Bayesian inference that depend on specific learning algorithms to estimate uncertainty. By making it easier to estimate uncertainty, we suspect that after-learning auditing methods like RUE will help to increase the adoption of machine learning in high-stakes domains such as medicine. We evaluated RUE on an error detection task, and showed that it outperforms existing tools for improving model reliability after learning. Moreover, we showed that RUE generates predictive distributions that are competitive with state-of-the-art integrated methods that use specialized training algorithms to quantify uncertainty (e.g. Bayesian inference). We also showed that RUE’s uncertainty score reflects the density and local fit criteria using a model-dependent similarity measure. In future work we can define new audits by leveraging different principles. For instance, methods for interpreting model predictions (e.g. Ribeiro et al. 2016) or performing sensitivity analyses (e.g. Koh and Liang 2017) after learning may inspire new criteria for reliability that could be integrated into different auditing tools.