Semi-supervised Deep Kernel Learning: Regression with Unlabeled Data by Minimizing Predictive Variance

Neal Jean, Sang Michael Xie, Stefano Ermon

Introduction

The prevailing trend in machine learning is to automatically discover good feature representations through end-to-end optimization of neural networks. However, most success stories have been enabled by vast quantities of labeled data . This need for supervision poses a major challenge when we encounter critical scientific and societal problems where fine-grained labels are difficult to obtain. Accurately measuring the outcomes that we care about—e.g., childhood mortality, environmental damage, or extreme poverty—can be prohibitively expensive . Although these problems have limited data, they often contain underlying structure that can be used for learning; for example, poverty and other socioeconomic outcomes are strongly correlated over both space and time.

Semi-supervised learning approaches offer promise when few labels are available by allowing models to supplement their training with unlabeled data . Mostly focusing on classification tasks, these methods often rely on strong assumptions about the structure of the data (e.g., cluster assumptions, low data density at decision boundaries ) that generally do not apply to regression .

In this paper, we present semi-supervised deep kernel learning, which addresses the challenge of semi-supervised regression by building on previous work combining the feature learning capabilities of deep neural networks with the ability of Gaussian processes to capture uncertainty . SSDKL incorporates unlabeled training data by minimizing predictive variance in the posterior regularization framework, a flexible way of encoding prior knowledge in Bayesian models .

Our main contributions are the following:

We introduce semi-supervised deep kernel learning (SSDKL) for the largely unexplored domain of deep semi-supervised regression. SSDKL is a regression model that combines the strengths of heavily parameterized deep neural networks and nonparametric Gaussian processes. While the deep Gaussian process kernel induces structure in an embedding space, the model also allows a priori knowledge of structure (i.e., spatial or temporal) in the input features to be naturally incorporated through kernel composition.

By formalizing the semi-supervised variance minimization objective in the posterior regularization framework, we unify previous semi-supervised approaches such as minimum entropy and minimum variance regularization under a common framework. To our knowledge, this is the first paper connecting semi-supervised methods to posterior regularization.

We demonstrate that SSDKL can use unlabeled data to learn more generalizable features and improve performance on a range of regression tasks, outperforming the supervised deep kernel learning method and semi-supervised methods such as virtual adversarial training (VAT) and mean teacher . In a challenging real-world task of predicting poverty from satellite images, SSDKL outperforms the state-of-the-art by 15.5%15.5\%—by incorporating prior knowledge of spatial structure, the median improvement increases to 17.9%17.9\%.

Preliminaries

There are two major paradigms in semi-supervised learning, inductive and transductive. In inductive semi-supervised learning, the labeled data (XL,yL)(X_{L},\mathbf{y}_{L}) and unlabeled data XUX_{U} are used to learn a function f:X↦Yf:\mathcal{X}\mapsto\mathcal{Y} that generalizes well and is a good predictor on unseen test examples XTX_{T} . In transductive semi-supervised learning, the unlabeled examples are exactly the test data that we would like to predict, i.e., XT=XUX_{T}=X_{U} . A transductive learning approach tries to find a function f:Xn+m↦Yn+mf:\mathcal{X}^{n+m}\mapsto\mathcal{Y}^{n+m}, with no requirement of generalizing to additional test examples. Although the theoretical development of SSDKL is general to both the inductive and transductive regimes, we only test SSDKL in the inductive setting in our experiments for direct comparison against supervised learning methods.

with mean function μ(x)\mu(\mathbf{x}) and covariance kernel function kϕ(xi,xj)k_{\phi}(\mathbf{x}_{i},\mathbf{x}_{j}) parameterized by ϕ\phi, then any collection of function values is jointly Gaussian,

with mean vector and covariance matrix defined by the GP, s.t. μi=μ(xi)\bm{\mu}_{i}=\mu(\mathbf{x}_{i}) and (KX,X)ij=kϕ(xi,xj)(K_{X,X})_{ij}=k_{\phi}(\mathbf{x}_{i},\mathbf{x}_{j}). In practice, we often assume that observations include i.i.d. Gaussian noise, i.e., y(x)=f(x)+ϵ(x)y(\mathbf{x})=f(\mathbf{x})+\epsilon(\mathbf{x}) where ϵ∼N(0,ϕn2)\epsilon\sim\mathcal{N}(0,\phi_{n}^{2}), and the covariance function becomes

Deep kernel learning

for some mean function μ(⋅)\mu(\cdot) and base kernel function kϕ(⋅,⋅)k_{\phi}(\cdot,\cdot) with parameters ϕ\phi. Parameters θ=(w,ϕ)\theta=(w,\phi) of the deep kernel are learned jointly by minimizing the negative log likelihood of the labeled data :

For Gaussian distributions, the marginal likelihood is a closed-form, differentiable expression, allowing DKL models to be trained via backpropagation.

Posterior regularization

In probabilistic models, domain knowledge is generally imposed through the specification of priors. These priors, along with the observed data, determine the posterior distribution through the application of Bayes’ rule. However, it can be difficult to encode our knowledge in a Bayesian prior. Posterior regularization offers a more direct and flexible mechanism for controlling the posterior distribution.

Let D=(XL,yL)\mathcal{D}=(X_{L},\mathbf{y}_{L}) be a collection of observed data. present a regularized optimization formulation called regularized Bayesian inference, or RegBayes. In this framework, the regularized posterior is the solution of the following optimization problem:

where L(q(M∣D))\mathcal{L}(q(M|\mathcal{D})) is defined as the KL-divergence between the desired post-data posterior q(M∣D)q(M|\mathcal{D}) over models MM and the standard Bayesian posterior p(M∣D)p(M|\mathcal{D}) and Ω(q(M∣D))\Omega(q(M|\mathcal{D})) is a posterior regularizer. The goal is to learn a posterior distribution that is not too far from the standard Bayesian posterior while also fulfilling some requirements imposed by the regularization.

Semi-supervised deep kernel learning

We introduce semi-supervised deep kernel learning (SSDKL) for problems where labeled data is limited but unlabeled data is plentiful. To learn from unlabeled data, we observe that a Bayesian approach provides us with a predictive posterior distribution—i.e., we are able to quantify predictive uncertainty. Thus, we regularize the posterior by adding an unsupervised loss term that minimizes the predictive variance at unlabeled data points:

where nn and mm are the numbers of labeled and unlabeled training examples, α\alpha is a hyperparameter controlling the trade-off between supervised and unsupervised components, and θ\theta represents the model parameters.

Optimizing LsemisupL_{semisup} is equivalent to computing a regularized posterior through solving a specific instance of the RegBayes optimization problem (2), where our choice of regularizer corresponds to variance minimization.

Let θˉ\bar{\theta} be a specific instance of the model parameters. Instead of maximizing the marginal likelihood of the labeled training data in a purely supervised approach, we train our model in a semi-supervised fashion by minimizing the compound objective

where the variance is with respect to p(f∣θˉ,D)p(f|\bar{\theta},\mathcal{D}), the Bayesian posterior given θˉ\bar{\theta} and D\mathcal{D}.

Let observed data D\mathcal{D}, a suitable space of functions F\mathcal{F}, and a bounded parameter space Θ\Theta be given. As in , we assume that F\mathcal{F} is a complete separable metric space and Π\Pi is an absolutely continuous probability measure (with respect to background measure η\eta) on (F,B(F))(\mathcal{F},\mathcal{B}(\mathcal{F})), where B(F)\mathcal{B}(\mathcal{F}) is the Borel σ\sigma-algebra, such that a density π\pi exists where dΠ=πdηd\Pi=\pi d\eta and we have prior density π(f,θ)\pi(f,\theta) and likelihood density p(D∣f,θ)p(\mathcal{D}|f,\theta). Assume we choose the prior such that π(θ)\pi(\theta) is uniform on Θ\Theta. Then the semi-supervised variance minimization problem (5)

is equivalent to the RegBayes optimization problem (2)

where α′=αnm\alpha^{\prime}=\frac{\alpha n}{m}, and Pprob={q:q(f,θ∣D)=q(f∣θ,D)δθˉ(θ∣D),θˉ∈Θ}\mathcal{P}_{prob}=\{q:q(f,\theta|\mathcal{D})=q(f|\theta,\mathcal{D})\delta_{\bar{\theta}}(\theta|\mathcal{D}),\bar{\theta}\in\Theta\} is a variational family of distributions where q(θ∣D)q(\theta|\mathcal{D}) is restricted to be a Dirac delta centered on θˉ∈Θ\bar{\theta}\in\Theta.

We include a formal derivation in Appendix A.1 and give a brief outline here. It can be shown that solving the variational optimization objective

is equivalent to minimizing the unconstrained form of the first term L(q(f,θ∣D))\mathcal{L}(q(f,\theta|\mathcal{D})) of the RegBayes objective in Theorem 1, and the minimizer is precisely the Bayesian posterior p(f,θ∣D)p(f,\theta|\mathcal{D}). When we restrict the optimization to q∈Pprobq\in\mathcal{P}_{prob} the solution is of the form q∗(f,θ∣D)=p(f∣θ,D)δθˉ(θ∣D)q^{*}(f,\theta|\mathcal{D})=p(f|\theta,\mathcal{D})\delta_{\bar{\theta}}(\theta|\mathcal{D}) for some θˉ\bar{\theta}. This allows us to show that (6) is also equivalent to minimizing the first term of Lsemisup(θˉ)L_{semisup}(\bar{\theta}). Finally, noting that the regularization function Ω\Omega only depends on θˉ\bar{\theta} (through q(θ∣D)=δθˉ(θ)q(\theta|\mathcal{D})=\delta_{\bar{\theta}}(\theta)), the form of q∗(f,θ∣D)q^{*}(f,\theta|\mathcal{D}) is unchanged after adding Ω\Omega. Therefore the choice of Ω\Omega reduces to minimizing the predictive variance with respect to q∗(f∣θ,D)=p(f∣θˉ,D)q^{*}(f|\theta,\mathcal{D})=p(f|\bar{\theta},\mathcal{D}).

By minimizing LsemisupL_{semisup}, we trade off maximizing the likelihood of our observations with minimizing the posterior variance on unlabeled data that we wish to predict. The posterior variance acts as a proxy for distance with respect to the kernel function in the deep feature space, and the regularizer is an inductive bias on the structure of the feature space. Since the deep kernel parameters are jointly learned, the neural net is encouraged to learn a feature representation in which the unlabeled examples are closer to the labeled examples, thereby reducing the variance on our predictions. If we imagine the labeled data as “supports” for the surface representing the posterior mean, we are optimizing for embeddings where unlabeled data tend to cluster around these labeled supports. In contrast, the variance regularizer would not benefit conventional GP learning since fixed kernels would not allow for adapting the relative distances between data points.

Another interpretation is that the semi-supervised objective is a regularizer that reduces overfitting to labeled data. The model is discouraged from learning features from labeled data that are not also useful for making low-variance predictions at unlabeled data points. In settings where unlabeled data provide additional variation beyond labeled examples, this can improve model generalization.

Training and inference

Semi-supervised deep kernel learning scales well with large amounts of unlabeled data since the unsupervised objective LvarianceL_{variance} naturally decomposes into a sum over conditionally independent terms. This allows for mini-batch training on unlabeled data with stochastic gradient descent. Since all of the labeled examples are interdependent, computing exact gradients for labeled examples requires full batch gradient descent on the labeled data. Therefore, assuming a constant batch size, each iteration of training requires O(n3)O(n^{3}) computations for a Cholesky decomposition, where nn is the number of labeled training examples. Performing the GP inference requires O(n3)O(n^{3}) one-time cost in the labeled points. However, existing approximation methods based on kernel interpolation and structured matrices used in DKL can be directly incorporated in SSDKL and would reduce the training complexity to close to linear in labeled dataset size and inference to constant time per test point . While DKL is designed for the supervised setting where scaling to large labeled datasets is a very practical concern, our focus is on semi-supervised settings where labels are limited but unlabeled data is abundant.

Experiments and results

We apply SSDKL to a variety of real-world regression tasks in the inductive semi-supervised learning setting, beginning with eight datasets from the UCI repository . We also explore the challenging task of predicting local poverty measures from high-resolution satellite imagery . In our reported results, we use the squared exponential or radial basis function kernel. We also experimented with polynomial kernels, but saw generally worse performance. Our SSDKL model is implemented in TensorFlow . Additional training details are provided in Appendix A.3, and code and data for reproducing experimental results can be found on GitHub.https://github.com/ermongroup/ssdkl

We first compare SSDKL to the purely supervised DKL, showing the contribution of unlabeled data. In addition to the supervised DKL method, we compare against semi-supervised methods including co-training, consistency regularization, generative modeling, and label propagation. Many of these methods were originally developed for semi-supervised classification, so we adapt them here for regression. All models, including SSDKL, were trained from random initializations.

Coreg, or Co-training Regressors, uses two kk-nearest neighbor (kkNN) regressors, each of which generates labels for the other during the learning process . Unlike traditional co-training, which requires splitting features into sufficient and redundant views, Coreg achieves regressor diversity by using different distance metrics for its two regressors .

Label propagation defines a graph structure over the data with edges that define the probability for a categorical label to propagate from one data point to another . If we encode this graph in a transition matrix TT and let the current class probabilities be yy, then the algorithm iteratively propagates y←Tyy\leftarrow Ty, row-normalizes yy, clamps the labeled data to their known values, and repeats until convergence. We make the extension to regression by letting yy be real-valued labels and normalizing TT. As in , we use a fully-connected graph and the radial-basis kernel for edge weights. The kernel scale hyperparameter is chosen using a validation set.

Generative models such as the variational autoencoder (VAE) have shown promise in semi-supervised classification especially for visual and sequential tasks . We compare against a semi-supervised VAE by first learning an unsupervised embedding of the data and then using the embeddings as input to a supervised multilayer perceptron.

2 UCI regression experiments

We evaluate SSDKL on eight regression datasets from the UCI repository. For each dataset, we train on n={50,100,200,300,400,500}n=\{50,100,200,300,400,500\} labeled examples, retain 10001000 examples as the hold out test set, and treat the remaining data as unlabeled examples. Following , the labeled data is randomly split 90-10 into training and validation samples, giving a realistically small validation set. For example, for n=100n=100 labeled examples, we use 90 random examples for training and the remaining 10 for validation in every random split. We report test RMSE averaged over 10 trials of random splits to combat the small data sizes. All kernel hyperparameters are optimized directly through LsemisupL_{semisup}, and we use the validation set for early stopping to prevent overfitting and for selecting α∈{0.1,1,10}\alpha\in\{0.1,1,10\}. We did not use approximate GP procedures in our SSDKL or DKL experiments, so the only difference is the addition of the variance regularizer. For all combinations of input feature dimensions and labeled data sizes in the UCI experiments, each SSDKL trial (including all training and testing) ran on the order of minutes.

Following , we choose a neural network with a similar [dd-100-50-50-2] architecture and two-dimensional embedding. Following , we use this same base model for all deep models, including SSDKL, DKL, VAT, mean teacher, and the VAE encoder, in order to make results comparable across methods. Since label propagation creates a kernel matrix of all data points, we limit the number of unlabeled examples for label propagation to a maximum of 20000 due to memory constraints. We initialize labels in label propagation with a kNN regressor with k=5k=5 to speed up convergence.

Table 1 displays the results for n=100n=100 and n=300n=300; full results are included in Appendix A.3. SSDKL gives a 4.20%4.20\% and 5.81%5.81\% median RMSE improvement over the supervised DKL in the n=100,300n=100,300 cases respectively, superior to other semi-supervised methods adapted for regression. A Wilcoxon signed-rank test versus DKL shows significance at the p=0.05p=0.05 level for at least one labeled training set size for all 8 datasets.

The same learning rates and initializations are used across all UCI datasets for SSDKL. We use learning rates of 1\times10−31\text{\times}{10}^{-3} and 0.10.1 for the neural network and GP parameters respectively and initialize all GP parameters to 11. In Fig. 2 (right), we study the effect of varying α\alpha to trade off between maximizing the likelihood of labeled data and minimizing the variance of unlabeled data. A large α\alpha emphasizes minimization of the predictive variance while a small α\alpha focuses on fitting labeled data. SSDKL improves on DKL for values of α\alpha between 0.10.1 and 10.010.0, indicating that performance is not overly reliant on the choice of this hyperparameter. Fig. 2 (left) compares SSDKL to purely supervised DKL, Coreg, and VAT as we vary the labeled training set size. For the Elevators dataset, DKL is able to close the gap on SSDKL as it gains access to more labeled data. Relative to the other methods, which require more data to fit neural network parameters, Coreg performs well in the low-data regime.

Surprisingly, Coreg outperformed SSDKL on the Blog, CTslice, and Buzz datasets. We found that these datasets happen to be better-suited for nearest neighbors-based methods such as Coreg. A kNN regressor using only the labeled data outperformed DKL on two of three datasets for n=100n=100, beat SSDKL on all three for n=100n=100, beat DKL on two of three for n=300n=300, and beat SSDKL on one of three for n=300n=300. Thus, the kNN regressor is often already outperforming SSDKL with only labeled data—it is unsurprising that SSDKL is unable to close the gap on a semi-supervised nearest neighbors method like Coreg.

To gain some intuition about how the unlabeled data helps in the learning process, we visualize the neural network embeddings learned by the DKL and SSDKL models on the Skillcraft dataset. In Fig. 3 (left), we first train DKL on n=100n=100 labeled training examples and plot the two-dimensional neural network embedding that is learned. In Fig. 3 (right), we train SSDKL on n=100n=100 labeled training examples along with m=1000m=1000 additional unlabeled data points and plot the resulting embedding. In the left panel, DKL learns a poor embedding—different colors representing different output magnitudes are intermingled. In the right panel, SSDKL is able to use the unlabeled data for regularization, and learns a better representation of the dataset.

3 Poverty prediction

High-resolution satellite imagery offers the potential for cheap, scalable, and accurate tracking of changing socioeconomic indicators. In this task, we predict local poverty measures from satellite images using limited amounts of poverty labels. As described in , the dataset consists of 3,0663,066 villages across five Africa countries: Nigeria, Tanzania, Uganda, Malawi, and Rwanda. These include some of the poorest countries in the world (Malawi and Rwanda) as well as some that are relatively better off (Nigeria), making for a challenging and realistically diverse problem.

In this experiment, we use n=300n=300 labeled satellite images for training. With such a small dataset, we can not expect to train a deep convolutional neural network (CNN) from scratch. Instead we take a transfer learning approach as in , extracting 4096-dimensional visual features and using these as input. More details are provided in Appendix A.5.

In order to highlight the usefulness of kernel composition, we explore extending SSDKL with a spatial kernel. Spatial SSDKL composes two kernels by summing an image feature kernel and a separate location kernel that operates on location coordinates (lat/lon). By treating them separately, it explicitly encodes the knowledge that location coordinates are spatially structured and distinct from image features.

As shown in Table 2, all models outperform the baseline state-of-the-art ridge regression model from . Spatial SSDKL significantly outperforms the DKL and SSDKL models that use only image features. Spatial SSDKL outperforms the other models by directly modeling location coordinates as spatial features, showing that kernel composition can effectively incorporate prior knowledge of structure.

Related work

introduced deep Gaussian processes, which stack GPs in a hierarchy by modeling the outputs of one layer with a Gaussian process in the next layer. Despite the suggestive name, these models do not integrate deep neural networks and Gaussian processes.

proposed deep kernel learning, combining neural networks with the non-parametric flexibility of GPs and training end-to-end in a fully supervised setting. Extensions have explored approximate inference, stochastic gradient training, and recurrent deep kernels for sequential data .

Our method draws inspiration from transductive experimental design, which chooses the most informative points (experiments) to measure by seeking data points that are both hard to predict and informative for the unexplored test data . Similar prediction uncertainty approaches have been explored in semi-supervised classification models, such as minimum entropy and minimum variance regularization, which can now also be understood in the posterior regularization framework .

Recent work in generative adversarial networks (GANs) , variational autoencoders (VAEs) , and other generative models have achieved promising results on various semi-supervised classification tasks . However, we find that these models are not as well-suited for generic regression tasks such as in the UCI repository as for audio-visual tasks.

Consistency regularization posits that the model’s output should be invariant to reasonable perturbations of the input . Combining adversarial training with consistency regularization, virtual adversarial training uses a label-free regularization term that allows for semi-supervised training . Mean teacher adds a regularization term that penalizes deviation from a exponential weighted average of the parameters over SGD iterations . For semi-supervised classification, found that VAT and mean teacher were the best methods across a series of fair evaluations.

Label propagation defines a graph structure over the data points and propagates labels from labeled data over the graph. The method must assume a graph structure and edge distances on the input feature space without the ability to adapt the space to the assumptions. Label propagation is also subject to memory constraints since it forms a kernel matrix over all data points, requiring quadratic space in general, although sparser graph structures can reduce this to a linear scaling.

Co-training regressors trains two kNN regressors with different distance metrics that label each others’ unlabeled data. This works when neighbors in the given input space have similar target distributions, but unlike kernel learning approaches, the features are fixed. Thus, Coreg cannot adapt the space to a misspecified distance measure. In addition, as a fully nonparametric method, inference requires retaining the full dataset.

Much of the previous work in semi-supervised learning is in classification and the assumptions do not generally translate to regression. Our experiments show that SSDKL outperforms other adapted semi-supervised methods in a battery of regression tasks.

Conclusions

Many important problems are challenging because of the limited availability of training data, making the ability to learn from unlabeled data critical. In experiments with UCI datasets and a real-world poverty prediction task, we find that minimizing posterior variance can be an effective way to incorporate unlabeled data when labeled training data is scarce. SSDKL models are naturally suited for many real-world problems, as spatial and temporal structure can be explicitly modeled through the composition of kernel functions. While our focus is on regression problems, we believe the SSDKL framework is equally applicable to classification problems—we leave this to future work.

This research was supported by NSF (#1651565, #1522054, #1733686), ONR, Sony, and FLI. NJ was supported by the Department of Defense (DoD) through the National Defense Science & Engineering Graduate Fellowship (NDSEG) Program. We are thankful to Volodymyr Kuleshov and Aditya Grover for helpful discussions.

References

Appendix A Appendix

Let D=(XL,yL,XU)\mathcal{D}=(X_{L},\mathbf{y}_{L},X_{U}) be a collection of observed data. Let X=(XL,XU)X=(X_{L},X_{U}) be the observed input data points. As in , we assume that F\mathcal{F} is a complete separable metric space and Π\Pi is an absolutely continuous probability measure (with respect to background measure η\eta) on (F,B(F))(\mathcal{F},\mathcal{B}(\mathcal{F})), where B(F)\mathcal{B}(\mathcal{F}) is the the Borel σ\sigma-algebra, such that a density π\pi exists where dΠ=πdηd\Pi=\pi d\eta. Let Θ\Theta be a space of parameters to the model, where we treat θ∈Θ\theta\in\Theta as random variables. With regards to the notation in the RegBayes framework, the model is the pair M=(f,θ)M=(f,\theta). We assume as in that the likelihood function P(⋅∣M)P(\cdot|M) is the likelihood distribution which is dominated by a σ\sigma-finite measure λ\lambda for all MM with positive density, such that a density p(⋅∣M)p(\cdot|M) exists where dP(⋅∣M)=p(⋅∣M)dλdP(\cdot|M)=p(\cdot|M)d\lambda.

We would like to compute the posterior distribution

We claim that the solution of the following optimization problem is precisely the Bayesian posterior p(f,θ∣D)p(f,\theta|\mathcal{D}):

By adding the constant log⁡p(D)\log p(\mathcal{D}) to the objective,

where by definition p(f,θ,D)=p(D∣f,θ)π(f,θ)p(f,\theta,\mathcal{D})=p(\mathcal{D}|f,\theta)\pi(f,\theta) and we see that the objective is minimized when q(f,θ∣D)=p(f,θ∣D)q(f,\theta|\mathcal{D})=p(f,\theta|\mathcal{D}) as claimed. We note that the objective is equivalent to the first term of the RegBayes objective (Section 2.3).

We introduce a variational approximation q∈Pprobq\in\mathcal{P}_{prob} which approximates p(f,θ∣D)p(f,\theta|\mathcal{D}), where Pprob={q:q(f,θ∣D)=q(f∣θ,D)δθˉ(θ∣D)}\mathcal{P}_{prob}=\{q:q(f,\theta|\mathcal{D})=q(f|\theta,\mathcal{D})\delta_{\bar{\theta}}(\theta|\mathcal{D})\} is the family of approximating distributions such that q(θ∣D)q(\theta|\mathcal{D}) is restricted to be a Dirac delta centered on θˉ\bar{\theta}. When we restrict q∈Pprobq\in\mathcal{P}_{prob},

where in equation (13) we use the form from equation (9), and in equation (15) we note that ∫θδθˉ(θ)log⁡δθˉ(θ)\int_{\theta}\delta_{\bar{\theta}}(\theta)\log\delta_{\bar{\theta}}(\theta) does not vary with θˉ\bar{\theta} or qq, and can be removed from the optimization. For every θˉ\bar{\theta}, the optimizing distribution is q∗(f∣θ,D)=p(f∣θˉ,D)q^{*}(f|\theta,\mathcal{D})=p(f|\bar{\theta},\mathcal{D}), which is the Bayesian posterior given the model parameters. Substituting this optimal value into (15),

using that DKL(q∗(f∣θ,D)∥p(f∣θ,D))=0D_{KL}(q^{*}(f|\theta,\mathcal{D})\|p(f|\theta,\mathcal{D}))=0 in (17) and ∫θδθˉ(θ)log⁡p(yL∣X)\int_{\theta}\delta_{\bar{\theta}}(\theta)\log p(\mathbf{y}_{L}|X) is a constant. We also treat log⁡p(θˉ∣X)\log p(\bar{\theta}|X) as a constant since θˉ\bar{\theta} is independent of XX without any observations yLy_{L}, and we choose a uniform prior over θˉ\bar{\theta}. The optimization problem over θˉ\bar{\theta} reflects maximizing the likelihood of the data.

We defined the regularization function as

Note that the regularization function only depends on θˉ\bar{\theta}, through q(θ∣D)=δθˉ(θ)q(\theta\mid\mathcal{D})=\delta_{\bar{\theta}}(\theta). Therefore the optimal post-data posterior qq in the regularized objective is still in the form q∗(f,θ∣D)=p(f∣θ,D)δθˉ(θ)q^{*}(f,\theta|\mathcal{D})=p(f|\theta,\mathcal{D})\delta_{\bar{\theta}}(\theta), and qq is modified by the regularization function only through δθˉ(θ)\delta_{\bar{\theta}}(\theta).

Thus, using the optimal post-data posterior q∗(f,θ∣D)=p(f∣θ,D)δθˉ(θ)q^{*}(f,\theta|\mathcal{D})=p(f|\theta,\mathcal{D})\delta_{\bar{\theta}}(\theta), the RegBayes problem is equivalent to the objective optimized by SSDKL:

A.2 Virtual Adversarial Training

Virtual adversarial training (VAT) is a general training mechanism which enforces local distributional smoothness (LDS) by optimizing the model to be less sensitive to adversarial perturbations of the input . The VAT objective is to augment the marginal likelihood with an LDS objective:

where Hi=∇∇rDKL(p(y∣Xi,θˉ)∣r=0H_{i}=\nabla\nabla_{r}D_{KL}(p(\mathbf{y}|X_{i},\bar{\theta})|_{r=0}. Note that the first derivative is zero since DKL(p(y∣Xi,θˉ)D_{KL}(p(\mathbf{y}|X_{i},\bar{\theta}) is minimized at r=0r=0. Therefore rv-adv(i)r_{\text{v-adv}(i)} can be approximated with the first dominant eigenvector of the HiH_{i} scaled to norm ϵ\epsilon. The eigenvector calculation is done via a finite-difference approximation to the power iteration method. As in , one step of the finite-difference approximation was used in all of our experiments.

A.3 Training details

In our reported results, we use the standard squared exponential or radial basis function (RBF) kernel,

A.4 UCI results

In section 4.2 of the main text, we include results on the UCI datasets for n=100n=100 and n=300n=300. Here we include the rest of the experimental results, for n∈{50,200,400,500}n\in\{50,200,400,500\}. For n=50n=50, we note that both Coreg and label propagation perform quite well — we expect that this is true because these methods do not require learning neural network parameters from data.

A.5 Poverty prediction

High-resolution satellite imagery offers the potential for cheap, scalable, and accurate tracking of changing socioeconomic indicators. The United Nations has set 17 Sustainable Development Goals (SDGs) for the year 2030—the first of these is the worldwide elimination of extreme poverty, but a lack of reliable data makes it difficult to distribute aid and target interventions effectively. Traditional data collection methods such as large-scale household surveys or censuses are slow and expensive, requiring years to complete and costing billions of dollars . Because data on the outputs that we care about are scarce, it is difficult to train models on satellite imagery using traditional supervised methods. In this task, we attempt to predict local poverty measures from satellite images using limited amounts of poverty labels. As described in , the dataset consists of 3,0663,066 villages across five Africa countries: Nigeria, Tanzania, Uganda, Malawi, and Rwanda. These countries include some of the poorest in the world (Malawi, Rwanda) as well as regions of Africa that are relatively better off (Nigeria), making for a challenging and realistically diverse problem. The raw satellite inputs consist of 400×400400\times 400 pixel RGB satellite images downloaded from Google Static Maps at zoom level 16, corresponding to 2.42.4 m ground resolution. The target variable that we attempt to predict is a wealth index provided in the publicly available Demographic and Health Surveys (DHS) .