Deep Gaussian Processes with Convolutional Kernels

Vinayak Kumar, Vaibhav Singh, P. K. Srijith, Andreas Damianou

Introduction

Deep learning models have made tremendous progress in computer vision problems through their ability to learn complex functions and representations . They learn complex functions mapping some input x to output y through composition of linear and non-linear functions. However, popular deep learning models based on convolutional and recurrent neural networks have significant limitations. The parametric form of the functions lead them to have millions of parameters to estimate which is less suitable for problems where the data are scarce. Deep learning models though probabilistic in nature, do not provide any uncertainty estimates on its predictions. Knowledge of uncertainty helps in better decision making and is crucial in high risk applications such as disease diagnosis and autonomous driving . Another major limitation with the existing deep learning networks is model selection. Developing an appropriate deep learning model to solve a problem is time consuming and computationally expensive. Deep Gaussian processes (DGPs) constitute a deep Bayesian non-parametric approach based on Gaussian processes (GPs) and have the potential to overcome the aforementioned limitations.

The original DGP model was introduced by inspired by the hierarchical GP-LVM structure and variations have emerged in recent years, mainly differing in the employed inference procedure. While employs a mean field variational posterior over the latent layers, extends this formulation with amortized inference, considers a nested variational inference approach, uses an approximate Expectation Propagation procedure. Further, achieves scalability through random Fourier features while the approach of considers the variational posterior to be conditioned over the previous layer, preserving correlations across the layers, and uses a doubly stochastic variational inference approach.

All the DGP models use kernels such as radial basis function (RBF) which is inadequate for problems in computer vision, such as object detection. They fail to capture wide variability of objects in images due to pose, illumination and complex backgrounds. RBF captures similarity between images on a global scale and is not invariant to unwanted variations in the image. On the other hand, convolutional neural networks (CNN) learn image representations from raw pixel data which are invariant to such perturbations in the image. They learn features important for the object detection task by successively convolving the representations by filters, applying non-linearity and performing feature pooling. We propose to use convolutional kernels in DGPs to learn salient features from the images which are invariant to transformations. This is different from recent works which combine CNNs and GPs in hybrid mode, such as . In particular, replaces the fully connected layers of a CNN with GPs, aiming at obtaining well-calibrated probabilities. While, in deep kernel learning , the kernel in GPs are computed using deep neural networks. In contrast, our approach brings the convolutional structure inside the deep GP model, through kernels, and remains fully non-parametric.

Convolutional kernels could effectively learn rich representations of the data. The similarity between structured objects such as images are computed by considering the similarity of the sub-structures in the object which makes them invariant to transformations in the image. They have been used to compute similarities between structured objects such as graphs and trees . Recently, they were used as a covariance function in GPs and were found to be very effective for object recognition tasks . Here, the kernel computations between images are done by summing the base kernel acting over different patches of the images.

We introduce convolutional kernels in the DGP framework in order to extract discriminative features from images for object classification. Our work builds on the convolutional GP and extends it for the deep learning case, allowing the resulting model to additionally perform hierarchical feature learning. We consider various DGP architectures obtained by stacking together convolutional and RBF kernels in various combinations. Further, we consider variants of the convolutional kernel such as weighted convolutional kernels which provide more discriminative features, and combination of RBF kernels as the base kernel. Convolutional kernels are computationally expensive as they require performing summation over all patches of the image. We propose an approach to improve the computational efficiency by random sub-sampling of the patches. We demonstrate the effectiveness of the proposed approaches for image classification on benchmark data sets such as MNIST, Rectangles-image, CIFAR10 and Caltech101. The experiments show that DGP models typically achieve better generalization performance by using convolutional kernels compared to state-of-the-art shallow GP models.

Background

We consider the image classification problem with CC classes and NN training data points, X={xi}i=1NX=\{\bf{x_{i}}\}_{i=1}^{N} and the corresponding labels y={yi}i=1N\bf{y}=\{y_{i}\}_{i=1}^{N}, where xi∈RW×H\bf{x_{i}}\in\mathcal{R}^{W\times H} and yi∈Y={1,2,…C}y_{i}\in\mathcal{Y}=\{1,2,\ldots C\}. Assume there exists a latent function f:RW×H→Yf:\mathcal{R}^{W\times H}\rightarrow\mathcal{Y} mapping the training data to outputs. In a Bayesian setting, we strive to learn a posterior distribution over this function, so that we can use it to compute the predictive distribution over the test labels. It helps one to make sound predictions about the test data labels, taking into account the uncertainty about them. Gaussian processes provide a Bayesian non-parametric approach to perform classification. In this section, we summarize Gaussian process classification and Deep Gaussian Processes (DGP) that will lay the groundwork for our model.

A GP is defined as a collection of random variables such that any finite subset of which is Gaussian distributed . It allows one to specify a prior distribution over real valued functions ff, represented as f(x)∼GP(m(x),k(x,x′))f(\mathbf{x})\sim\mathcal{GP}(m(\mathbf{x}),k(\mathbf{x},\mathbf{x}^{\prime})) where m(x)m(\mathbf{x}) is the mean function and k(x,x′)k(\mathbf{x},\mathbf{x}^{\prime}) provides the covariance across the function values at two data points x\mathbf{x} and x′\mathbf{x}^{\prime}.

The kernel function determines various properties of the function such as stationarity, smoothness etc. A popular kernel function is the radial basis function (RBF) (squared exponential kernel), as it can model any smooth function. It is given by σf2exp⁡(−12κ∣∣x−x′∣∣2)\sigma_{f}^{2}\exp(-\frac{1}{2\kappa}||\mathbf{x}-\mathbf{x}^{\prime}||^{2}) where the length scale κ\kappa determines the variations in function values across the inputs.

For multi-class classification problems, we associate a separate function fcf_{c} with each class cc. An independent GP prior is placed over each of these functions, fc(x)∼GP(mc(x),k(x,x′))f_{c}(\mathbf{x})\sim{\mathcal{GP}}(m_{c}(\mathbf{x}),k(\mathbf{x},\mathbf{x}^{\prime})). Let fc=[fc(x1),fc(x2),⋯ ,fc(xN)]{\bf f_{c}}=[f_{c}(\mathbf{x}_{1}),f_{c}(\mathbf{x}_{2}),\cdots,f_{c}(\mathbf{x}_{N})] be a column vector indicating function values at the input data points for a class cc. Further, let FF be the matrix formed by stacking all column vectors {fc}c=1C\{\bf f_{c}\}_{c=1}^{C} , with Fn,cF_{n,c} representing the latent function value of nthn^{th} sample belonging to class cc and FnF_{n} representing the vector of latent function values over classes for the nthn^{th} sample. The GP prior over FF takes the following form : p(F)=∏c=1CN(fc;mc(X),KXX)p(F)=\prod\limits_{c=1}^{C}\mathcal{N}(\mathbf{f}_{c};m_{c}(X),K_{XX}), where KXXK_{XX} is the N×NN\times N covariance matrix formed by evaluating kernel over all pairs of training data points. For a data point nn, the likelihood of it belonging to class cc, p(yn=c∣Fn)p(y_{n}=c|F_{n}), is obtained by considering a soft-max link function. The posterior distribution over FF is obtained by combining the prior and the likelihood using Bayes theorem:

In GP multi-class classification, the posterior distribution cannot be computed in closed form due to the non-conjugacy between likelihood and prior. Learning in GPs involves learning the kernel hyper-parameters by maximizing the evidence p(y)=∫∏n=1Np(yn∣Fn)p(F)dFp(\mathbf{y})=\int\prod_{n=1}^{N}p(y_{n}|F_{n})p(F)dF, which also cannot be computed in closed form. The posterior distribution can be approximated as a Gaussian using approximate inference techniques such as Laplace approximation and variational inference . The Gaussian approximated posterior is then used to make predictions on the test data points. Variational inference has received a lot interest recently as it does not suffer from convergence problems unlike Markov chain Monte Carlo techniques and it provides a posterior approximation quickly by solving an optimization problem. It is scalable to large data sets and amenable to distributed processing. It also provides a lower bound on the marginal likelihood which can be used to perform model selection. The variational inference approach learns an approximate posterior distribution q(F)q(F) by minimizing the KL divergence between q(F)q(F) and p(F∣y)p(F|{\bf y}). Choosing a mean field family of variational distributions, q(F)q(F) factorizes across dimensions(or columns), i.e q(F)=∏q(fc)q(F)=\prod q({\mathbf{f}_{c}}). Each variational factor q(fc)q(\mathbf{f}_{c}) is assumed to be a Gaussian with variational parameters, mean vector μc\bf\mu_{c} and covariance Σc\Sigma_{c}. In the variational inference framework, minimizing the KL divergence with respect to the variational parameters is equivalent to maximizing the so-called variational Evidence Lower BOund(ELBO) which is given by

The variational parameters {μc,Σc}c=1C\{{\bf\mu_{c}},\Sigma_{c}\}_{c=1}^{C} and the kernel hyperparameters {σf2,l}\{\sigma_{f}^{2},l\} are learnt by jointly maximizing the variational lower bound in eq. (1) using any gradient based approach.

The KL divergence term in eq. (1) involves inversion of the covariance matrix KXXK_{XX} which scales as O(N3){\mathcal{O}}(N^{3}) computationally. Therefore, we opt for the variational sparse Gaussian process approximation which reduces the computational complexity to O(NM2){\mathcal{O}}(NM^{2}) , where M≪NM\ll N represents the number of inducing points. Specifically, the variational sparse approximation expands the latent function space with MM inducing variables u∈RM\mathbf{u}\in\mathcal{R}^{M} which are latent function values at inducing points Z={zi}i=1MZ=\{\mathbf{z}_{i}\}_{i=1}^{M}. Within the context of GP multi-class classification, we additionally have the inducing variable outputs uc\mathbf{u}_{c} for each class cc which are stacked together to form the matrix U∈RM×CU\in\mathcal{R}^{M\times C}. The joint GP prior over {f,u}\{f,u\} is then

where KXZK_{XZ} is the N×MN\times M covariance matrix over training inputs XX and inducing inputs ZZ and KZZK_{ZZ} is the M×MM\times M covariance matrix over inducing points ZZ. The conditional distribution of fc\mathbf{f}_{c} given uc\mathbf{u}_{c} is given by

and the marginal distribution over uc\mathbf{u}_{c} is p(uc)=N(uc;mc(Z),KZZ)p(\mathbf{u}_{c})=\mathcal{N}(\mathbf{u}_{c};m_{c}(Z),K_{ZZ}). The variational sparse approximation of considers a joint variational posterior over {fc,uc}\{\mathbf{f}_{c},\mathbf{u}_{c}\} in factorized form and is written as q(fc,uc)=p(fc∣uc,X)q(uc)q(\mathbf{f}_{c},\mathbf{u}_{c})=p(\mathbf{f}_{c}|\mathbf{u}_{c},X)q(\mathbf{u}_{c}). Assuming Gaussian variational factors for inducing points q(uc)=N(uc;mc,Sc)q(\mathbf{u}_{c})=\mathcal{N}(\mathbf{u}_{c};{\bf m_{c}},S_{c}), the variational lower bound (ELBO) can be derived as

2 Deep Gaussian Process

Deep Gaussian processes (DGPs) learn complex functions by stacking GPs one over the other resulting in a deep architecture of GPs. The function mapping one hidden layer to the next in DGPs is more expressive and data dependent compared to the pre-fixed sigmoid non-linear function used in standard parametric deep learning approaches. In addition, it is devoid of large number of parameters but only a few kernel hyper-parameters and few variational parameters (few due to sparse GP approach). Deep GPs do not typically overfit on small data due to Bayesian model averaging, and the stochasticity inherent in GPs naturally allows them to handle uncertainty in the data. Furthermore, by using a specific kernel which enables automatic relevance determination, one can automatically learn the dimensionality of hidden layers (number of neurons) . This overcomes the model selection problem in deep learning to a great extent.

DGPs consider the function mapping input to output to be represented as a composition of functions, f(x)=fL∘(fL−1…∘(f1(x)))f(\mathbf{x})=f^{L}\circ(f^{L-1}\ldots\circ(f^{1}(\mathbf{x}))), assuming there are LL layers. The lthl^{th} layer consists of DlD^{l} functions fl={fjl}j=1Dlf^{l}=\{f^{l}_{j}\}_{j=1}^{D^{l}} mapping representations in layer l−1l-1 to obtain DlD^{l} representation for layer ll. Independent GP priors are placed over the function fjlf^{l}_{j} producing jthj^{th} representation in layer ll, fjl(⋅)∼GP(mjl(⋅),kl(⋅,⋅))f_{j}^{l}(\cdot)\sim{\mathcal{GP}}(\it m_{j}^{l}(\cdot),k^{l}(\cdot,\cdot)). The jthj^{th} function in layer ll, fj1f^{1}_{j}, acts on the input data point xi\mathbf{x}_{i} to produce the representation Fi,j1=fj1(xi)F^{1}_{i,j}=f^{1}_{j}(\mathbf{x}_{i}). In general, the jthj^{th} function in layer ll, fjl(⋅)f^{l}_{j}(\cdot) acts on the representation of the data point xi\mathbf{x}_{i} at layer l−1l-1, Fil−1F^{l-1}_{i} to produce the representation Fi,jl=fjl(Fil−1)F^{l}_{i,j}=f^{l}_{j}(F^{l-1}_{i}). Let fjl\mathbf{f}^{l}_{j} denote the jthj^{th} representation at layer ll computed over all inputs. The final layer LL will have CC functions corresponding to the classes and these functions values are squashed through a soft-max function to produce the class probabilities.

We follow the DGP variant presented in where the noise between layers is absorbed into the kernel. The kernel function associated with a GP in layer ll is defined as kl(Fil,Fjl)=σfl2exp⁡(−12κl∣∣Fil−Fjl∣∣2)+σnl2δijk^{l}(F^{l}_{i},F^{l}_{j})={\sigma^{l}_{f}}^{2}\exp(\frac{-1}{2\kappa^{l}}||F^{l}_{i}-F^{l}_{j}||^{2})+{\sigma^{l}_{n}}^{2}\delta_{ij}. Following the variational sparse Gaussian process approximation as explained in the section 2.1, each layer ll is associated with inducing variables {Ul}\{U^{l}\} which are function values over MM inducing points ZlZ^{l} associated with layer ll, Zl={zil}i=1MZ^{l}=\{{\bf z}^{l}_{i}\}_{i=1}^{M}. Let ujl\mathbf{u}^{l}_{j} represent the inducing variables associated with the jthj^{th} representation at layer ll. The number of inducing points are kept fixed for all layers (only for convenience) as MM and a joint GP prior is considered over latent function values and inducing points. The joint distribution p(y,F,U)p(\mathbf{y},F,U) is given by

where a deep GP prior is put recursively over the entire latent space with F0=XF^{0}=X and a soft-max likelihood is used for classification. The conditional above is:

The posterior distribution p(F,U∣y)p(F,U|\mathbf{y}) and marginal likelihood p(y)p(\mathbf{y}) cannot be computed in closed form due to the intractability in obtaining the marginal prior over {Fl}l=2L\{F^{l}\}_{l=2}^{L}. This involves integrating out the previous layer, which is present in a non linear manner inside the covariance matrices (KFl−1Fl−1lK^{l}_{F^{l-1}F^{l-1}}) appearing in (5). Along with non-conjugate likelihood, this brings in additional difficulty to the DGP model. Multiple approaches have been suggested in the literature for achieving tractability in DGPs, such as variational inference , amortized inference , expectation propagation and random Fourier features . Here we follow the variational inference approach, and we assume the variational posterior to be having form q(F,U)=∏l=1L∏j=1Dlp(fjl∣ujl,Fl−1,Zl)q(ujl)q(F,U)=\prod\limits_{l=1}^{L}\prod\limits_{j=1}^{D^{l}}p(\mathbf{f}_{j}^{l}|\mathbf{u}_{j}^{l},F^{l-1},Z^{l})q(\mathbf{u}_{j}^{l}), where q(ujl)=N(ujl;mjl,Sjl)q(\mathbf{u}^{l}_{j})=\mathcal{N}(\mathbf{u}^{l}_{j};{\bf m}^{l}_{j},S^{l}_{j}) . Let ml{\bf m}^{l} be a vector formed by concatenating the vectors mjl{\bf m}^{l}_{j} and SlS^{l} be the block diagonal covariance matrix formed from SjlS^{l}_{j}. We can formulate the ELBO by extending the methodology described in Section 2.1 to multiple layers as follows:

where, the marginal distribution of the functions values for layer LL over all the data points is obtained as

and the conditional distribution in (7) is computed as

The marginal distribution in (7) is intractable, due to presence of stochastic term {Fl−1}l=2L\{F^{l-1}\}_{l=2}^{L} inside the conditional distributions {q(Fl∣Fl−1,Zl,ml,Sl)}l=2L−1\{q(F^{l}|F^{l-1},Z^{l},{\bf m}^{l},S^{l})\}_{l=2}^{L-1} in a non-linear manner. This intractability results in the expected log likelihood in (6) to be intractable even for Gaussian likelihood. We approximate it via Monte Carlo sampling as done in .

The lower bound can be written as sum over data points and the parameters can be updated based gradients computed on a mini-batch of data. This enables one to use stochastic gradient techniques for maximizing the variational lower bound. This stochasticity in gradient computation combined with the stochasticity introduced by the Monte Carlo sampling in variational lower bound computation results in the doubly stochastic variational inference method for deep GPs.

Convolutional Deep Gaussian Processes

We combine the convolutional GP kernels with deep Gaussian processes in order to obtain the convolutional deep Gaussian process (CDGP). A CDGP can capture salient features which are invariant to variations in the image through the convolutional structures and is simultaneously performing strong function learning through out its depth, all within a Bayesian framework. This results in a powerful well-calibrated model for tasks like image classification.

Our starting point is the recently introduced convolutional Gaussian processes (CGP) where the function evaluation on an image is considered as sum of functions over the patches of the input image. Assuming there are PP patches in x\mathbf{x} with each patch x[p]\mathbf{x}^{[p]} to be w×hw\times h dimensional, CGP considers f(x)=∑p=1Pg(x[p])f(\mathbf{x})=\sum_{p=1}^{P}g(\mathbf{x}^{[p]}). Placing a zero mean GP{\mathcal{GP}} prior over the function g(x[p])g(\mathbf{x}^{[p]}), g(x[p])∼GP(0,kg(xi[p],xj[p]))g(\mathbf{x}^{[p]})\sim{\mathcal{GP}}(0,k_{g}(\mathbf{x}_{i}^{[p]},\mathbf{x}_{j}^{[p]})), induces a zero mean GP{\mathcal{GP}} prior over the function f(x)f(\mathbf{x}) with a convolutional kernel (Conv kernel) kfk_{f},

We refer to kgk_{g} as the base kernel. Considering a convolutional kernel in computing the similarities between the images is useful in capturing non-local similarities among the images. The convolutional kernel compares one region in the image xi\mathbf{x}_{i} with another region in the image xj\mathbf{x}_{j}, and could provide a high similarity even under transformations in the image. The kernel computation over patches (xi[p],xj[p′])(\mathbf{x}_{i}^{[p]},\mathbf{x}_{j}^{[p^{\prime}]}) considers similarity in a spatial neighborhood, whereas with other kernels (such as RBF kernel) only global similarity across images can be computed and fails to capture similarity in images due to transformations.

Convolutional Neural Networks(CNNs) convolve image with multiple kernels (filters), apply a non-linear operation and then feature pooling (average, max) multiple times to learn discriminative features useful for the object detection task. Similar to CNN, the function f(x)f(\mathbf{x}) could be seen to perform average pooling of the non-linear feature maps produced by the patch response functions g(x[p])g(\mathbf{x}^{[p]}). This pooling operation results in convolution operation in kernel space. The convolutional kernel computation between two images xi\mathbf{x}_{i} and xj\mathbf{x}_{j} is expanded as

The convolution operation between pthp^{th} patch of image xi\mathbf{x}_{i} (which now acts as a filter) and the image xj\mathbf{x}_{j} results in a convolution signal, where signal value at any point p′p^{\prime} is obtained by computing the dot product between the filter xi[p]\mathbf{x}_{i}^{[p]} and patch xj[p′]\mathbf{x}_{j}^{[p^{\prime}]}. This dot product is performed by the base kernel which transforms these patches into feature vectors in a high dimensional space and computes the dot product between them in that space. Any pthp^{th} summand is the sum of the convolution signal values obtained at all the points.

2 Deep Gaussian processes with convolutional kernels

Convolutional DGP considers multiple functions from a GP prior with convolutional kernels to form a representation of the image in the first layer. The function corresponding to otho^{th} representation for layer 11 is obtained as

Each output in layer 11 captures different features of the image. The feature representations of the image obtained in the first layer are then mapped using a GP with convolutional or RBF kernel to obtain further representations. In general, the function corresponding to otho^{th} representation for layer ll is considered as

The kernel matrices involved in the computation of the conditional distribution in eq. (8) such as KFl−1Fl−1lK^{l}_{F^{l-1}F^{l-1}}, KFl−1ZllK^{l}_{F^{l-1}Z^{l}} and KZlZll{K^{l}_{Z^{l}Z^{l}}} use the convolutional kernel defined in (3.2). As before, ZlZ^{l} represents the inducing points associated with layer ll and has the same dimension as Fl−1F^{l-1}. The variational lower bound expression and “reparameterization trick” remains the same as has been derived for deep GPs in Section 2.2.

We also consider variants of the convolutional kernel such as weighted convolutional kernels (Wconv kernels) . It associates a weight with each patch which allows the kernel to provide differential weightage to the patches which is useful for object detection. The function f(x)f(\mathbf{x}) in general for any layer is considered as

3 Reducing computational complexity through patch subsets

Convolutional kernels provide an effective way to capture the similarity across images, but are computationally expensive. Computing the similarity between two images involves O(P2)\mathcal{O}(P^{2}) computational cost, where PP is the number of patches in the input image or the feature representation. For the input image of size W×HW\times H, it is of the order of O(WH)\mathcal{O}(WH) when stride length and patch sizes are small. This is costly even for image data sets such as MNIST and rectangles which contain images of size (28×2828\times 28). This makes the computations impractical on higher dimensional data such as Caltech101 (250×250250\times 250). This can be addressed to some extent using the idea of treating the inducing points in the patch space , where Zjl∈Rw×hZ^{l}_{j}\in\mathcal{R}^{w\times h} rather than in the input space RW×H\mathcal{R}^{W\times H}. In this case, computation of the entries in the matrix KFl−1ZllK^{l}_{F^{l-1}Z^{l}} can be performed in O(P)\mathcal{O}(P) time, and that of KZlZllK^{l}_{Z^{l}Z^{l}} can be performed in constant time.

However, computation of the entries in the matrix KFl−1Fl−1lK^{l}_{F^{l-1}F^{l-1}} matrix which appears in the conditional distribution in (8) still requires O(P2)\mathcal{O}(P^{2}) computations for the first layer making it a costly operation. This makes the approach practically inapplicable to high dimensional data sets such as Caltech101 even with a reduced image size. Moreover in these images, a lot of information will be shared by overlapping patches and will be redundant for the computation of the similarity across images. We propose to use random sub-sampling of the patches in computing the convolutional kernel for the entries in the matrix KFl−1Fl−1lK^{l}_{F^{l-1}F^{l-1}} and KFl−1ZllK^{l}_{F^{l-1}Z^{l}}. Let S,S′⊂{1,2,…,P}{S,S^{\prime}}\subset\{1,2,\ldots,P\} represent the random subsets. For the otho^{th} representation of layer 11 (F0=XF^{0}=X), we consider the covariance functions to be as follows

This reduces the cost of computing the matrix KFl−1Fl−1lK^{l}_{F^{l-1}F^{l-1}} for layer 11 to O(∣S∣∣S′∣)\mathcal{O}(|S||S^{\prime}|) where the size of the subsets ∣S∣,∣S′∣≪P|S|,|S^{\prime}|\ll P. Computational speedup achieved through random sub-sampling of patches is testified in our experiments on Caltech101.

Experiments

We evaluate the generalization performance of the proposed model, convolutional deep Gaussian processes (CDGP), on various image classification data sets, namely MNIST, Rectangle-Images, CIFAR10 and Caltech101. We consider different kernel architectures of the proposed CDGP model and compare it with sparse GPs (SGP) Results as reported in ., deep GP (DGP) models with RBF kernel and with convolutional GPs (CGP) with different convolutional kernels. The convolutional deep GP uses the same inference procedure as in deep GP (“re-parameterization trick”) and uses an a priori fixed inducing input points by considering centroids of the clustered images . The inducing points and the linear mean function for each of the inner layers is obtained using the singular value decomposition approach mentioned in . The number of inducing points is taken to be 100. We follow the same approach for convolutional GPs also to maintain a fair playground. The kernel parameters are kept the same across various outputs in a layer while it is different across the layers. The number of outputs in the latent layers is taken to be 3030 for MNIST and 5050 for other datasets (except for the final layer which will be equal to the number of classes). For the models considering convolutional kernels, the patch size is taken to be 3×33\times 3 with a stride length of 11 for the rectangles data while a patch size of 5×55\times 5 is considered for the rest of the data sets. We consider the RBF kernel as the base kernel kgk_{g} for all our experiments. The approaches are compared in terms of their accuracy in making predictions on the test data and the negative log predictive probability (NLPP) on test data which considers uncertainty in predictions. The code has been developed on top of GPflow framework with ADAM optimizer to learn the kernel and variational parameters by maximizing the variational lower bound. The variational mean parameters are initialized to , variance parameters to 1e−51e^{-5} and length-scales are initialized to 22 for MNIST and 1010 for other datasets.

We performed experiments with MNIST dataset with 10 classes corresponding to the digits 0−90-9. We consider the standard train/test split with 60K training and 10K test images. We considered CDGP and DGP models with various architectures as described in Table 1. Parameters of the model are learned by running the ADAM optimizer for 400400 epochs with 0.010.01 step size and a mini-batch of 1000. Experimental comparison indicates that the proposed CDGP models with 2 layers, first layer with a weighted convolutional kernel and the second layer with an RBF kernel gave the best performance, an accuracy of 98.6698.66 and an NLPP score of 0.04630.0463. Second best performance was given again by a CDGP model with 4 layers, 2 weighted convolutional kernels followed by 2 RBF layers. We could observe that all the CDGP models performed better than the DGP and CGP models in the MNIST data. We also conducted experiments with the combinations of two RBF kernels with length scales initialized to 2 different values 0.010.01 and 1010, as the base kernel in a convolutional kernel. The approach gave an accuracy of 98.4698.46. We found that the learned length scales are also quite far apart which shows that one RBF kernel is trying to capture long distance correlations while the other one captures short distant correlations. This did not result in better results as MNIST is quite simple dataset for which capturing such information might not be necessary.

2 Rectangles-Image

We consider the rectangles-image data set used in , where a rectangle of varying height and width is placed inside images. The patches in the border and inside of the rectangle and the background patches are sampled to make the rectangle hard to detect Rectangles-image data is different from the simpler rectangles data used in , where a random size rectangle is placed in black background with the pixels corresponding to the border of the rectangle in white, while that of inside in black. . The task is to classify if a rectangle in an image has a larger height or width. The data set consists of 12K training images and 50K test images, and is known to require deep architectures for correct classification. We consider two different architectures of CDGP, and compare it against sparse GPs, deep GPs with 2 and 3 layers and convolutional GPs. Parameters of the models are learnt by running the ADAM optimizer for 200200 epochs with 0.010.01 step size and a mini-batch of 1000. Experimental comparison across different approaches is provided in Table 2. We could observe that the proposed CDGP model with 2 layers, first layer using a weighted convolutional kernel and the second layer using an RBF, provided the best performance beating DGP, CGP and SGP models by a large margin. To the best of our knowledge, this is the highest accuracy reported by a GP model on the rectangles-image data. This indicates the usefulness of the representation learning capability of CDGP model for complex image classification.

3 CIFAR-10

The CIFAR-10 dataset consists of total 60K images out of which 50K are used as training images while the rest 10K images are being used for testing. The dataset contains colored images of objects like airplane, automobile, etc. There are 10 classes in total having 6K images per class. The dimensionality of each image is 32×32×332\times 32\times 3 (3 is for channels). We compare the performance of CDGP, DGP and CGP models in Table 3. Parameters of the models are learned by running the ADAM optimizer for 200200 epochs with a mini-batch size of 40Learning took around 11 hours on Nvidia GTX 1080 Ti GPU, while the best results reported in is obtained after running the optimization for 40 hours..

We observe that DGP models gave a relatively low performance on the CIFAR10 datasets. Equipping DGP models with convolutional kernels have boosted the performance by 10%10\% showing the effectiveness of convolutional kernels for image classification. However, CDGP models were not able to obtain a performance close to CGP. This could be an indication that, for this particular dataset, the properties of a single-layer CDGP i.e, CGP is enough to learn a good classifier. In fact, the previous experiments have shown that 2-layer CDGPs typically result in the best accuracy (in comparison with deeper models), implying that a CGP has already very large capacity for classification and therefore the addition of one layer is usually enough to improve on the results.

4 Experiments with Random Sub-sampling of Patches on Caltech-101 Dataset

Computation of convolutional kernels becomes prohibitive on data sets such as Caltech101 with very high dimensionality. It consists of 101 classes with 20 images per class for training and 10 images per class for testing. The size varies slightly for each image in the actual dataset but is roughly around 300 ×\times 200 pixels per channel. The images are colored so each image has 3 channels. The experiments are conducted on images resized to 50×50×350\times 50\times 3. Instead of taking all the patches of the image for computing the convolutional kernel, we randomly picked up one-tenth of the total number of image patches for computing the kernel. This resulted in a very significant speed-up in learning time without much loss in accuracy, as can be seen from table 4 We ran the experiments on Nvidia GTX 1080 Ti GPU.. The test accuracy obtained with CDGP1 is 20.39%20.39\% and time taken for training is 11 hrs 18min. On the other hand, for the same model considering random subset of image patches, training time drops to only 1 hr 15 min with an accuracy drop of only 0.39%0.39\% making it around 10 times faster. Similar phenomenon has been observed in case of CDGP2 considering random subset of patches, where training time improved from 12 hrs 2 min to 1 hrs 19 min with an accuracy drop of just 0.69%0.69\%, providing a speedup of around 10 hours. The classification accuracies of the models presented in Table 4 are low due to resizing of the original image to size 50×5050\times 50, resulting in loss of information. As a future work, we will conduct experiments by keeping the original image size, and study the effectiveness of random sub-sampling of patches and generalization performance of the proposed approach.

Conclusion

Deep GP models provide a lot of advantages in terms of capacity control and predictive uncertainty, but they are less effective in computer vision tasks. Commonly used RBF kernels in the DGP models fail to capture variations in image data and are not invariant to translations. In this paper we proposed a DGP model which captures convolutional structure in image data using convolutional kernels. Our model extends the convolutional GPs with the ability to learn hierarchical latent representations making it a useful model for image classification. We incorporated different types of convolutional kernels in the DGP models and demonstrated their usefulness for image classification in benchmark data sets such as MNIST, Rectangles-Image and CIFAR10. In the future, we plan to develop methods to further reduce the cost of convolutional kernel computation and memory requirements of the CDGP model for high dimensional datasets. This will allow us to consider a higher mini-batch size, leading to reduced stochastic gradient variance and faster convergence of the optimization routine. We found that increasing the number of layers in CDGP did not bring much improvements in performance contrary to what we expected. We hope that our future research on faster and more effective variational inference techniques will address these limitations with convolutional DGPs.

References