Understanding image representations by measuring their equivariance and equivalence

Karel Lenc, Andrea Vedaldi

Introduction

Image representations have been a key focus of the research in computer vision for at least two decades. Notable examples include textons , histogram of oriented gradients (SIFT and HOG ), bag of visual words , sparse and local coding , super vector coding , VLAD , Fisher Vectors , and the latest generation of deep convolutional networks . However, despite their popularity, our theoretical understanding of representations remains limited. It is generally believed that a good representation should combine invariance and discriminability, but this characterisation is rather vague; for example, it is often unclear what invariances are contained in a representation and how they are obtained.

In this work, we propose a new approach to study image representations. We look at a representation ϕ\phi as an abstract function mapping an image x\mathbf{x} to a vector ϕ(x)∈d\phi(\mathbf{x})\in^{d} and we empirically establish key mathematical properties of this function. We focus in particular on three such properties (Sect. 2). The first one is equivariance, which looks at how the representation changes upon transformations of the input image. We demonstrate that most representations, including HOG and most of the layers in deep neural networks, change in a easily predictable manner with the input (Fig. 1). We show that such equivariant transformations can be learned empirically from data (Sect. 2.1) and that, importantly, they amount to simple linear transformations of the representation output (Sect. 3.1 and 3.2). In the case of convolutional networks, we obtain this by introducing and learning a new transformation layer. By analysing the learned equivariant transformations we are also able to find and characterise the invariances of the representation, our second property. This allows us to quantify invariance and show how it builds up with depth in deep models.

The third property, equivalence, looks at whether the information captured by heterogeneous representations is in fact the same. CNN models, in particular, contain millions of redundant parameters that, due to non-convex optimisation in learning, may differ even when retrained on the same data. The question then is whether the resulting differences are genuine or just apparent. To answer this question we learn stitching layers that allow swapping parts of different networks. Equivalence is then obtained if the resulting “Franken-CNNs” perform as well as the original ones (Sect. 3.3).

The rest of the paper is organised as follows. Sect. 2 discussed methods to learn empirically representation equivariance, invariance, and equivalence. Sect. 3.1 and 3.2 present experiments on shallow and deep representation equivariance respectively, and Sect. 3.3 on representation equivalence. Sect. 3.4 demonstrates a practical application of equivariant representations to structured-output regression. Finally, Sect. 4 summarises our findings.

The problem of designing invariant or equivariant features has been widely explored in computer vision. For example, a popular strategy is to extract invariant local descriptors on top of equivariant (also called co-variant) detectors . Various authors have also looked at incorporating equivariance explicitly in the representations . Deep CNNs, including the one of Krizhevsky et al. and related state-of-the-art architecutres, are deemed to build an increasing amount of invariance layer after layer. This is even more explicit in the scattering transform of Sifre and Mallat .

In all these examples, invariance is a design aim that may or may not be achieved by a given architecture. By contrast, our aim is not to propose yet another mechanism to learn invariances, but rather a method to systematically tease out invariance, equivariance, and other properties that a given representation may have. To the best of our knowledge, there is very limited work in conducting this type of analysis. Perhaps the contributions that come closer study only invariances of neural networks to specific image transformations . However, we believe to be the first to functionally characterise and quantify these properties in a systematic manner, as well as being the first to investigate the equivalence of different representations.

Notable properties of representations

Image representations such as HOG, SIFT, or CNNs can be thought of as functions ϕ\phi mapping an image x∈X\mathbf{x}\in\mathcal{X} to a vector ϕ(x)∈d\phi(\mathbf{x})\in^{d}. This section describes three notable properties of representations — equivariance, invariance, and equivalence — and gives algorithms to establish them empirically.

A representation ϕ\phi is equivariant with a transformation gg of the input image if the transformation can be transferred to the representation output. Formally, equivariance with gg is obtained when there exists a map Mg:d→dM_{g}:^{d}\rightarrow^{d} such that:

A sufficient condition for the existence of MgM_{g} is that the representation ϕ\phi is invertible, because in this case Mg=ϕ∘g∘ϕ−1M_{g}=\phi\circ g\circ\phi^{-1}. It is known that representations such as HOG are at least approximately invertible . Hence it is not just the existence, but also the structure of the mapping MgM_{g} that is of interest. In particular, MgM_{g} should be simple, for example a linear function. This is important because the representation is often used in simple predictors such as linear classifiers, or in the case of CNNs, is further processed by linear filters. Furthermore, by requiring the same mapping MgM_{g} to work for any input image, intrinsic geometric properties of the representations are captured.

The nature of the transformation gg is in principle arbitrary; in practice, in this paper we will focus on geometric transformations such as affine warps and flips of the image.

Invariance.

Invariance is a special case of equivariance obtained when MgM_{g} (or a subset of MgM_{g}) acts as the simplest possible transformation, i.e. the identity map. Invariance is often regarded as a key property of representations since one of the goals of computer vision is to establish invariant properties of images. For example, the category of the objects contained in an image is invariant to viewpoint changes. By studying invariance systematically, it is possible to clarify if and where the representation achieves it.

Equivalence.

While equi/invariance look at how a representation is affected by transformations of the image, equivalence studies the relationship between different representations. Two heterogenous representations ϕ\phi and ϕ′\phi^{\prime} are equivalent if there exist a map Eϕ→ϕ′E_{\phi\rightarrow\phi^{\prime}} such that

If ϕ\phi is invertible, then Eϕ→ϕ′=ϕ′∘ϕ−1E_{\phi\rightarrow\phi^{\prime}}=\phi^{\prime}\circ\phi^{-1} satisfies this condition; hence, as for the mapping MgM_{g} before, the interest is not just in the existence but also in the structure of the mapping Eϕ→ϕ′E_{\phi\rightarrow\phi^{\prime}}.

Example: equivariant HOG transformations.

Let ϕ\phi denote the HOG feature extractor. In this case ϕ(x)\phi(\mathbf{x}) can be interpreted as a H×WH\times W vector field of of DD-dimensional feature vectors or cells. If gg denotes image flipping around the vertical axis, then ϕ(x)\phi(\mathbf{x}) and ϕ(gx)\phi(g\mathbf{x}) are related by a well defined permutation of the feature components. This permutation swaps the HOG cells in the horizontal direction and, within each HOG cell, swaps the components corresponding to symmetric orientations of the gradient. Hence the mapping MgM_{g} is a permutation and one has exactly ϕ(gx)=Mgϕ(x)\phi(g\mathbf{x})=M_{g}\phi(\mathbf{x}). The same is true for horizontal flips and 180°180\degree rotations, and, approximately,Most HOG implementations use 9 orientation bins, breaking rotational symmetry. for 90°90\degree rotations. HOG implementations do in fact explicitly provide such permutations.

Example: translation equivariance in convolutional representations.

HOG, densely-computed SIFT (DSIFT), and convolutional networks are examples of convolutional representations in the sense that they are obtained from local and translation invariant operators. Barring boundary and sampling effects, any convolutional representation is equivariant to translations of the input image as this result in a translation of the feature field.

1 Learning properties with structured sparsity

When studying equivariance and equivalence, the transformation MgM_{g} and Eϕ→ϕ′E_{\phi\rightarrow\phi^{\prime}} are usually not available in closed form and must be estimated from data. This section discusses a number of algorithms to do so. The discussion focuses on equivariant transformations MgM_{g}, but dealing with equivalence transformations Eϕ→ϕ′E_{\phi\rightarrow\phi^{\prime}} is similar.

Given a representation ϕ\phi and a transformation gg, the goal is to find a mapping MgM_{g} satisfying (1). In the simplest case Mg=(Ag,bg),M_{g}=(A_{g},\mathbf{b}_{g}), Ag∈d×d,A_{g}\in^{d\times d}, bg∈d\mathbf{b}_{g}\in^{d} is an affine transformation ϕ(gx)≈Agϕ(x)+bg\phi(g\mathbf{x})\approx A_{g}\phi(\mathbf{x})+\mathbf{b}_{g}. This choice is not as restrictive as it may initially seem: in the examples above MgM_{g} is a permutation, and hence can be implemented by a corresponding permutation matrix AgA_{g}.

Estimating (Ag,bg)(A_{g},\mathbf{b}_{g}) is naturally formulated as an empirical risk minimisation problem. Given data x\mathbf{x} sampled from a set of natural images, learning amounts to optimising the regularised reconstruction error

The choice of regulariser is particularly important as Ag∈d×dA_{g}\in^{d\times d} has a Ω(d2)\Omega(d^{2}) parameters. Since dd can be quite large (for example, in HOG one has d=DWHd=DWH), regularisation is essential. The standard l2l^{2} regulariser ∥Ag∥F2\|A_{g}\|_{F}^{2} was found to be inadequate; instead, sparsity-inducting priors work much better for this problem as they encourage AgA_{g} to be similar to a permutation matrix.

We consider two such sparsity-inducing regularisers. The first regulariser allows AgA_{g} to contain a fixed number kk of non-zero entries for each row:

Regularising rows independently reflects the fact that each row is a predictor of a particular component of ϕ(gx)\phi(g\mathbf{x}).

The second sparsity-inducing regulariser is similar, but exploits the convolutional structure of many representations. Convolutional features are obtained from translation invariant and local operators (non-linear filters), such that the representation ϕ(x)\phi(\mathbf{x}) can be interpreted as a feature field with spatial indexes (u,v)(u,v) and channel index tt. Due to the locality of the representation, the component (u,v,t)(u,v,t) of ϕ(gx)\phi(g\mathbf{x}) should be predictable from a corresponding neighbourhood Ωg,m(u,v)\Omega_{g,m}(u,v) of features in the feature field ϕ(x)\phi(\mathbf{x}) (Fig. 2). This results in a particular sparsity structure for AgA_{g} that can be imposed by the regulariser

where mm denotes the neighbour size and indexes of AA have been identified with triplets (u,v,t)(u,v,t). The neighbourhood itself is defined as the m×mm\times m input feature sites closer to the back-projection of the output feature (u,v)(u,v).Formally, denote by (x,y)(x,y) the coordinates of a pixel in the input image x\mathbf{x} and by p:(u,v)↦(x,y)p:(u,v)\mapsto(x,y) the affine function mapping the feature index (u,v)(u,v) to the centre (x,y)(x,y) of the corresponding receptive field (measurement region) in the input image. Denote by Nk(u,v)\mathcal{N}_{k}(u,v) the kk feature sites (u′,v′)(u^{\prime},v^{\prime}) that are closer to (u,v)(u,v) (the latter can have fractional coordinates) and use this to define the neighbourhood of the back-transformed site (u,v)(u,v) as Ωg,k(u,v)=Nk(p−1∘g−1∘p(u,v))\Omega_{g,k}(u,v)=\mathcal{N}_{k}(p^{-1}\circ g^{-1}\circ p(u,v)). In practice (3) and (4) will be combined in order to limit the number of regression coefficients activated in each neighbourhood.

Loss.

2 Equivariance in CNNs: transformation layers

The method of Sect. 2.1 can be substantially refined for the case of convolutional representations and certain transformation classes. The structured sparsity regulariser (4) encourages AgA_{g} to match the convolutional structure of the representation. If gg is an affine transformation more can be said: up to sampling artefacts, the equivariant transformation MgM_{g} is local and translation invariant, i.e. convolutional. The reason is that an affine gg acts uniformly on the image domainThis means that g(x+u,y+v)=g(x,y)+(u′,v′)g(x+u,y+v)=g(x,y)+(u^{\prime},v^{\prime}). so that the same is true for MgM_{g}. This has two key advantages: it reduces dramatically the number of parameters to learn and it can be implemented efficiently as an additional layer of a CNN. Such a transformation layer consists of a permutation layer that maps input feature sites (u,v,t)(u,v,t) to output feature sites (g(u,v),t)(g(u,v),t) followed by a bank of DD linear filters, each of dimension m×m×Dm\times m\times D. Here mm corresponds to the size of the neighbourhood Ωg,m(u,v)\Omega_{g,m}(u,v) in Sect. 2.1. Intuitively, the main purpose of these filters is to permute and interpolate feature channels.

Note that g(u,v)g(u,v) does not, in general, fall at integer coordinates. In our case, the permutation layer assigns g(u,v)g(u,v) to the closest lattice site by rounding but it can be also distributed to the nearest 2×22\times 2 sites by using bilinear interpolation.Better accuracy could be obtained by using image warping techniques. For example, sub-pixel accuracy can be obtained by upsampling in the permutation layer and then allowing the transformation filter to be translation variant (or, equivalently, by introducing a suitable non-linear mapping between the permutation layer and transformation filters).

3 Equivalence in CNNs: stitching layers

The previous section looked at how equivariance can be studied more efficiently in CNNs; this section does the same for equivalence. Following the task-oriented loss formulation of Sect. 2.1, consider two representations ϕ1\phi_{1} and ϕ1′\phi_{1}^{\prime} and a predictor ϕ2′\phi_{2}^{\prime} learned to solve a reference task using the representation ϕ1′\phi_{1}^{\prime}. For example, these could be obtained by decomposing two CNNs ϕ=ϕ2∘ϕ1\phi=\phi_{2}\circ\phi_{1} and ϕ′=ϕ2′∘ϕ1′\phi^{\prime}=\phi_{2}^{\prime}\circ\phi_{1}^{\prime} trained on the ImageNet ILSVCR data (but ϕ1\phi_{1} could also be learned on a different problem or be handcrafted).

Experiments

The experiments begin in Sect. 3.1 by studying the problem of learning equivariant mappings for shallow representations. Sect. 3.2 and 3.3 move on to deep convolutional representations, examining equivariance and equivalence respectively. In Sect. 3.4 equivariant mappings are applied to structure-output regression.

This section applies the methods of Sect. 2.1 to learn equivariant maps for shallow representations, and HOG features in particular. The first method to be evaluated is sparse regression, followed by structured sparsity. Finally, the learned equivariant maps are validated in example recognition tasks.

The first experiment (Fig. 3) explores variants of the sparse regression formulation (2). The goal is to learn a mapping Mg=(Ag,g)M_{g}=(A_{g},_{g}) that predicts the effect of selected image transformations gg on the HOG features of an image. For each transformation, the mapping MgM_{g} is learned from 1,000 training images by minimising the regularised empirical risk (5). The performance is measured as the average Hellinger’s distance ∥ϕ(gx)−Mgϕ(x)∥Hell.\|\phi(g\mathbf{x})-M_{g}\phi(\mathbf{x})\|_{\text{Hell.}} on a test set of further 1,000 images.The Hellinger’s distance (∑i(xi−yi)2)1/2(\sum_{i}(\sqrt{x_{i}}-\sqrt{y_{i}})^{2})^{1/2} is preferred to the Euclidean distance as the HOG features are histograms. Images are randomly sampled from the ILSVRC12 train and validation datasets respectively.

This experiment focuses on predicting a small array of 5×55\times 5 of HOG cells, which allows to train full regression matrices even with naive baseline regression algorithms. Furthermore, the 5×55\times 5 array is predicted from a larger 9×99\times 9 input array to avoid boundary issues when images are rotated or rescaled. Both these restrictions will be relaxed later. Fig. 3 compares the following methods to learn MgM_{g}: choosing the identity transformation Mg=1M_{g}=\mathbf{1}, learning MgM_{g} by optimising the objective (2) without regularisation (Least Square – LS), with the Frobenius norm regulariser for different values of λ\lambda (Ridge Regression – RR), and with the sparsity-inducing regulariser (3) (Forward-Selection – FS, using ) for a different number kk of regression coefficients per output dimension.

As can be seen in Fig. 3a, 3b, LS overfits badly, which is not surprising given that MgM_{g} contains 1M parameters even for these small HOG arrays. RR performs significantly better, but it is easily outperformed by FS, confirming the very sparse nature of the solution (e.g. for k=5k=5 just 0.2% of the 1M coefficients are non-zero). The best result is obtained by FS with k=5k=5. As expected, the prediction error of FS is zero for a 180°180\degree rotation as this transformation is exact (Sect. 2), but note that LS and RR fail to recover it. As one might expect, errors are smaller for transformations close to identity, although in the case of FS the error remains small throughout the range.

Structured sparse regression.

The conclusion of the previous experiments is that sparsity is essential to achieve good generalisation. However, learning MgM_{g} directly, e.g. by forward-selection or by l1l^{1} regularisation, can be quite expensive even if the solution is ultimately sparse. Next, we evaluate using the structured sparsity regulariser (4), where each output feature is predicted from a prespecified neighbourhood of input features dependent on the image transformation gg. Fig. 3c repeats the experiment of Fig. 3a for a 45°45\degree rotation, but this time limited to neighbourhoods of m×mm\times m input HOG cells. To be able to span larger intervals of mm, an array of 15×1515\times 15 HOG cells is used. Since spatial sparsity is now imposed a-priori, LS, RR, and FS perform nearly equivalently for m≤3m\leq 3, with the best result achieved by FS with k=5k=5 and a small neighbourhood of m=3m=3 cells. There is also a significant computational advantage in structured sparsity (Tab. 1) as it limits the effective size of the regression problems to be solved. We conclude that structured sparsity is highly preferable over generic sparsity.

Regression quality.

So far results have been given in term of the reconstruction error of the features; this paragraph relates this measure to the practical performance of the learned mappings. The first experiment is qualitative and uses the HOGgle technique to visualise the transformed features. As shown in Fig. 5, the visualisations of ϕ(gx)\phi(g\mathbf{x}) and Mgϕ(x)M_{g}\phi(\mathbf{x}) are indeed nearly identical, validating the mapping MgM_{g}. The second experiment (Fig. 4) evaluates instead the performance of transformed HOG features quantitatively, in a classification problem. To this end, an SVM classifier ⟨w,ϕ(x)⟩\langle\mathbf{w},\phi(\mathbf{x})\rangle is trained to discriminate between dog and cat faces using the data of (using 15×1515\times 15 HOG templates, 400 training and 1,000 testing images evenly split among cats and dogs). Then a progressively larger rotation or scaling g−1g^{-1} is applied to the input image and the effect compensated by MgM_{g}, computing the SVM score as ⟨w,Mgϕ(g−1x)⟩\langle\mathbf{w},M_{g}\phi(g^{-1}\mathbf{x})\rangle (equivalently the model is transformed by Mg⊤M_{g}^{\top}). The performance of the compensated classifier is nearly identical to the original classifier for all angles and scales, whereas the uncompensated classifier ⟨w,ϕ(g−1x)⟩\langle\mathbf{w},\phi(g^{-1}\mathbf{x})\rangle rapidly fails, particularly for rotation. We conclude that equivariant transformations encode visual information effectively.

2 Equivariance in deep representations

The previous section validated learning equivariant transformations in shallow representations such as HOG. This section extends these results to deep representations, using the Alexn CNN as a reference state-of-the-art deep feature extractor using the MatConvNet framework . Alexn is the composition of twenty functions, grouped into five convolutional layers (comprising filtering, max-pooling, normalisation and ReLU) and three fully-connected layers (filtering and ReLU). The experiments look at the convolutional layers Conv1 to Conv5 right after the linear filters (learning the linear transformation layers after the ReLU was found to be harder due to the non-negativity of the features).

The first experiment (Fig. 6a) compares different methods to learn equivariant mappings MgM_{g} in a CNN. The first method is FS, computed for different neighbourhood sizes mm and sparsity kk. The second is the task oriented formulation of Sect. 2.1 using a transformation layer. Both the l2l^{2} reconstruction error of the features and the classification error (task-oriented loss) are reported. As in Sect. 2.2, the latter is the classification error of the compensated network ϕ2∘Mg∘ϕ1(g−1x)\phi_{2}\circ M_{g}\circ\phi_{1}(g^{-1}\mathbf{x}) in the ImageNet ILSVCR data (the reported error is measured on the validation data, but optimised on the training data). The figure reports the evolution of the loss as more training samples are used. For the purpose of this experiment, gg is set to vertical image flip. Fig. 7 repeats the experiments for the task-oriented objective and rotations gg from 0 to 90 degrees (the fact that intermediate rotations are slightly harder to reconstruct suggests that a better MgM_{g} could be learned by addressing more carefully interpolation and boundary effects).

Several observations can be made. First, all methods perform substantially better than doing nothing (∼75%\sim 75\% top-1 error), recovering most if not all the performance of the original classifier (43%43\%). This demonstrates that linear equivariant mappings MgM_{g} can be learned successfully for CNNs too. Second, for the shallower features up to Conv2, FS is better: it requires less training samples and it has a smaller reconstruction error and comparable classification error than the task-oriented loss. Compared to Sect. 3.1, however, the best setting m=3m=3, k=25k=25 is substantially less sparse. However, from Conv3 onwards, the task-oriented loss is better, converging to a much lower classification error than FS. FS still achieves a significantly smaller reconstruction error, showing that feature reconstruction is not always predictive of classification performance. Third, the classification error increases somewhat with depth, matching the intuition that deeper layers contain more specialised information: as such, perfectly transforming these layers for transformations not experienced during training (such as vertical flips) may not be possible.

Testing transformations.

Next, we investigate which geometric transformations can be represented by different layers of a CNN (Tab. 4), considering in particular horizontal and vertical flips, rescaling by half, and rotation of 90°90\degree. First, for transformations such as horizontal flips and scaling, learning equivariant mappings is not better than leaving the features unchanged: the reason is that the CNN is implicitly learned to be invariant to such factors. For vertical flips and rotations, however, the learned equivariant mapping substantially reduce the error. In particular, the first few layers are easily transformable, confirming their generic nature.

Quantifying invariance.

One use of the mapping MgM_{g} is the identification of invariant features in the representation. These are the ones that are best predicted by themselves after a transformation. In practice, a transformation layer in a CNN (Sect. 2.2) identifies invariant feature channels since the same transformation filters are applied uniformly at all spatial locations. In practice, invariance is almost never achieved exactly; instead, the degree of invariance of a feature channel is scored as the ratio of the Euclidean norm of the corresponding row of MgM_{g} with the same after suppressing the “diagonal” component of that row. Then, the pp rows of MgM_{g} with the highest invariance score are replaced by (scaled) rows of the identity matrix. Finally, the performance of the modified transformation Mˉg\bar{M}_{g} is evaluated and accepted if the classification performance does not deteriorate by more than 5%5\% relative to MgM_{g}. The corresponding feature channels for the largest possible pp are then be considered approximately invariant.

Table 4 reports the result of this analysis for horizontal and vertical flips, rescaling, and 90°90\degree rotation in the Alexn CNN. There are several notable observations. First, for transformations for which the network is overall invariant such as horizontal flips and rescaling, invariance is obtained largely in Conv3 or Conv4. Second, invariance is not always increasing with depth, as for example Conv1 tends to be more invariant than Conv2. This is possible because, even if the feature channels in a layer are invariant, the spatial pooling in the subsequent layer may not be. Third, the number of invariant features is significantly smaller for unexpected transformations such as vertical flips and 90°90\degree rotations, further validating the approach.

3 Equivalence of deep representations

While the previous two sections studied the equivariance of representations, this section looks at their equivalence. The goal is to clarify whether heterogeneous representations may in fact capture the same visual information by replacing part of a representation with another using the methods of Sect. 2 and Sect. 2.3.

To validate this idea, the first several layers ϕ1′\phi_{1}^{\prime} of the Alexn CNN ϕ′=ϕ2′∘ϕ1′\phi^{\prime}=\phi_{2}^{\prime}\circ\phi_{1}^{\prime} are swapped with layers ϕ1\phi_{1} from Imnet, also trained on the ILSVRC12 data, Plcs , trained on the MIT Places data, and Plcs-h, trained on a mixture of MIT Places and ILSVRC12 images. These representations have a similar, but not identical, structure and entirely different parametrisations.

Table 4 shows the top-1 performance of hybrid models ϕ2′∘Eϕ1→ϕ1′∘ϕ1\phi_{2}^{\prime}\circ E_{\phi_{1}\rightarrow\phi_{1}^{\prime}}\circ\phi_{1}, where the equivalence map Eϕ1→ϕ1′E_{\phi_{1}\rightarrow\phi_{1}^{\prime}} is learned as a stitching layer (Sect. 2.3) from ILSVRC12 training images. There are a number of notable facts. First, setting Eϕ→ϕ′=1E_{\phi\rightarrow\phi^{\prime}}=\mathbf{1} to the identity map has a top-1 error >99%>99\% (not shown in the table), matching the intuition that different parametrisations make feature channels not directly compatible. Second, a very good level of equivalence can be established up to Conv4 between Alexn and Imnet, and a slightly less good one between Alexn and Plcs-h; however, in Plcs deeper layers are substantially less compatible. Specifically, Conv1 and Conv2 are interchangeable in all cases, whereas Conv5 is not fully interchangeable, particularly for Plcs. This corroborates the intuition that Conv1 and Conv2 are generic image codes, whereas Conv5 is more task-specific. Note however that, even in the worst case, performance is dramatically better than chance, demonstrating that all such features are compatible to an extent.

4 Application to structured-output regression

As a complement of the theoretical investigation so far, this section shows a direct practical application of the learned equivariant mappings of Sect. 2 to structured-output regression . In structured regression an input image x\mathbf{x} is mapped to a label y\mathbf{y} by the function y^(x)=argmax⁡y,z⟨ϕ(x,y,z),w⟩\hat{\mathbf{y}}(\mathbf{x})=\operatorname{argmax}_{\mathbf{y},\mathbf{z}}\langle\phi(\mathbf{x},\mathbf{y},\mathbf{z}),\mathbf{w}\rangle (direct regression) where z\mathbf{z} is an optional latent variable and ϕ\phi a joint feature map. If y\mathbf{y} and/or z\mathbf{z} include geometric parameters, the joint feature can be partially of fully rewritten as ϕ(x,y,z)=My,zϕ(x)\phi(\mathbf{x},\mathbf{y},\mathbf{z})=M_{\mathbf{y},\mathbf{z}}\phi(\mathbf{x}), reducing inference to the maximisation of ⟨My,z⊤w,ϕ(x)⟩\langle M_{\mathbf{y},\mathbf{z}}^{\top}\mathbf{w},\phi(\mathbf{x})\rangle (equivariant regression). There are two computational advantages: (i) the representation ϕ(x)\phi(\mathbf{x}) needs to be computed just once and (ii) the vectors My,z⊤wM_{\mathbf{y},\mathbf{z}}^{\top}\mathbf{w} can be precomputed.

This idea is demonstrated on the task of pose estimation, where y=g\mathbf{y}=g is a geometric transformation in a class g−1∈Gg^{-1}\in G of possible poses of an object. As an example, consider estimating the pose of cat faces in the PASCAL VOC 2007 (VOC07) data using for GG either (i) rotations or (ii) affine transformations (Fig. 9). The rotations in GG are sampled uniformly every 10 degrees and the ground-truth rotation of a face is defined by the line connecting the nose to the midpoints between the eyes. These keypoints are obtained as the center of gravity of the corresponding regions in the VOC07 part annotations . The affine transformations in GG are obtained instead by clustering the vectors [cl⊤,cr⊤,cn⊤]⊤[\mathbf{c}_{l}^{\top},\mathbf{c}_{r}^{\top},\mathbf{c}_{n}^{\top}]^{\top} containing the location of eyes and nose of 300 example faces in the VOC07 data. The clusters are obtained using GMM-EM on the training data and used to map the test data to the same pose classes for evaluation. GG then contains the set of affine transformations mapping the keypoints [cˉl⊤,cˉr⊤,cˉn⊤]⊤[\bar{\mathbf{c}}_{l}^{\top},\bar{\mathbf{c}}_{r}^{\top},\bar{\mathbf{c}}_{n}^{\top}]^{\top} in a canonical frame to each cluster center.

The matrices MgM_{g} are pre-learned (from generic images, not containing cats) using FS with k=5k=5 and m=3m=3 as in Sect. 2. Since cat faces in VOC07 data are usually upright, a second more challenging version of the data (denoted by the symbol ↺\circlearrowleft) augmented with random image rotations is considered as well. The direct ⟨w,ϕ(gx)⟩\langle\mathbf{w},\phi(g\mathbf{x})\rangle and equivariant ⟨w,Mgϕ(x)⟩\langle\mathbf{w},M_{g}\phi(\mathbf{x})\rangle scoring functions are learned using 300 training samples and evaluated on 300 test ones.

Table 5 reports the accuracy and speed obtained for HOG and CNN Conv3, Conv4, and Conv5 features for direct and equivariant regression. The latter is generally as good or nearly as good as direct regression, but up to 22 times faster validating once more the mappings MgM_{g}. Fig. 8 shows the cumulative error curves for the different regressors.

Summary

This paper introduced the idea of studying representations by learning their equivariant and equivalence properties. It was shown that shallow representations and the first several layers of deep state-of-the-art CNNs transform in an easily predictable manner with image warps and that they are interchangeable, and hence equivalent, in different architectures. Deeper layers share some of these properties but to a lesser degree, being more task-specific. In addition to the use as analytical tools, these methods have practical applications such as accelerating structured-output regressors classifier in a simple and elegant manner.

Karel Lenc was partially supported by an Oxford Engineering Science DTA.

References