Warped Convolutions: Efficient Invariance to Spatial Transformations
João F. Henriques, Andrea Vedaldi
Introduction
A crucial aspect of current deep learning architectures is the encoding of invariances. This fact is epitomized in the success of convolutional neural networks (CNN), where equivariance to image translation is key: translating the input results in a translated output. When invariances are present in the data, encoding them explicitly in an architecture provides an important source of regularization, which reduces the amount of training data required for learning.
Invariances may also be used to improve the efficiency of implementations. For instance, a convolutional layer requires orders of magnitude less memory (by reusing filters across space) and less computation (due to their limited support) compared to a fully-connected layer. Its local and predictable memory access pattern also makes better use of modern hardware’s caching mechanisms.
The success of CNNs indicates that translation invariance is an important property of images. However, this does not explain why translation equivariant operators work well for image understanding. The common interpretation is that such operators are matched to the statistics of natural images, which are well known to be translation invariant (Hyvärinen et al., 2009). However, natural image statistics are also (largely) invariant to other transformations such as isotropic scaling and rotation, which suggests that alternative neural network designs may also work well with images. Furthermore, in specific applications, invariances other than translation may be more appropriate.
Therefore, it is natural to consider generalizing convolutional architectures to other image transformations, and this has been the subject of extensive study (Kanazawa et al., 2014; Bruna et al., 2013; Cohen & Welling, 2016). Unfortunately these approaches do not possess the same memory and speed benefits that CNNs enjoy. The reason is that, ultimately, they have to transform (warp) an image or filter several times (Kanazawa et al., 2014; Marcos et al., 2016; Dieleman et al., 2015), incurring a high computational burden. Another approach is to consider a basis of filters (analogous to eigen-images) encoding the desired invariance (Cohen & Welling, 2014; Bruna et al., 2013; Cohen & Welling, 2016). Although they are able to handle transformations with many pose parameters, in practice most recent proposals are limited to very coarsely discretized transformations, such as horizontal/vertical flips and 90∘ rotations (Dieleman et al., 2015; Cohen & Welling, 2014).
In this work we consider generalizations of CNNs that overcome these disadvantages. Well known constructions in group theory enable the extension of convolution to general transformation groups (Folland, 1995). However, this generality usually comes at an increased computational cost or complexity. Here we show that, by making appropriate assumptions, we can design convolution operators that are equivariant to a large class of two-parameter transformations while reducing to a standard convolution in a warped image space. The fixed image warp can be implemented using bilinear resampling, a simple and fast operation that has been popularized by spatial transformer networks (Jaderberg et al., 2015; Heckbert, 1989) and is part of most deep learning toolboxes. Unlike previous proposals, the proposed warped convolutions can handle continuous transformations, such as fine rotation and scaling.
This makes generalized convolution easily implementable in neural networks, reusing fast convolution algorithms on GPU hardware, such as Winograd (Lavin, 2015) or the Fast Fourier Transform (Lyons, 2010).
Generalizing convolution
In the neural network literature, this is often written using the cross-correlation convention (Goodfellow et al., 2016), by considering the reflected filter :
To handle continuous deformations of the input, it is more natural to express eq. 2 as an integral over continuous rather than discrete inputs:
2 Convolution on groups
The standard convolution operator of eq. 4 can be interpreted as applying the filter to translated versions of the image. We wish to replace translations with other image transformations , belonging to a group . In the context of machine learning models for images, this generalized (group) convolution can be understood to exhaustively search for a pattern at various poses (e.g. rotation angles or scale factors) (Dieleman et al., 2015; Kanazawa et al., 2014).
The (left) Haar measure is the most natural choice for ; it is the only measure (up to scaling factors) that is (left) invariant to group translation. In other words, satisfies the equation:
From the viewpoint of statistical learning, a key property of convolution is equivariance. Consider the (left) translation operator
(Folland, 1995) Convolution is equivariant with group translations, in the sense that commutes with :
3 From groups to images
We can then update eq. 5 to express group convolution as a function of images on the real plane:
4 Standard convolutions with exponential maps
and group convolution reduces to the standard notion of convolution on :
We refer to this standard convolution in warped space (eq. 8) as warped convolution.
Note that the result is an image whose dimensionality is that of the vector space ; in the following, we mainly work in the case . By far, the strongest requirement is that the map is additive: this is the same as requiring the transformation group to be Abelian, in the sense that transformations commute (). In section 5 we will show a variety of useful image transformations that respect this property. The advantage of introducing this restriction is that calculations simplify tremendously, ultimately enabling a simple and efficient implementation of the operator as discussed below.
Warped convolutions
Our main contribution is to note that certain group convolutions can be implemented efficiently by a standard convolution, by pre-warping the input image and filter appropriately. The warp is the same for any image, depending solely on the nature of the relevant transformations, and can be written in closed form. This result allows one to implement very efficient group convolutions using simple computational blocks, as shown in section 3.2.
We can now reinterpret these results in terms of a new neural network convolutional layer. The input of the layer is an image and the learnable parameters are the coefficients of the filter . The output is a new “image” defined on the vector space . This image is obtained by first warping using the exponential map and then by convolving the result with in the standard sense:
The most important property of this layer is equivariance: if we warp the image by the transformation , then the convolution result translates by :
Note that the output action equivalent to warping the input is simply to translate the result, as for standard convolution. The second most important property is that this operator can be implemented efficiently as the combination of warping and standard convolution.
2 Implementation and intuition
The warp (exponential map) that is applied to the input image eq. 9 can be implemented as follows. We start with an arbitrary pivot point in the image, and then sample all possible transformations of that point, . For discrete images, will be implemented as a 2D grid of discrete transformations (e.g. rotations and scales at regular intervals), and will be a 2D grid as well, referred to as the warp grid. Finally, sampling the input image at the points (for example, by bilinear interpolation) yields the warped image.
An illustration is given in fig. 1, for various transformations (each one is discussed in more detail in section 5). The red dot shows the pivot point , and the two arrows pointing away from it show the original axes of the sampled grid of parameters. The grids were generated by sampling the transformation parameters at regular intervals. Note that the warp grids are independent of the image contents – they can be computed once offline and then applied to any image.
The steps for implementing a warped convolution block are outlined in algorithm 1. The main advantage of implementing group convolution as warped convolution is that it replaces a large number of warping operations (one per group element) with a single warp.
Discussion
Next, let and note that . Hence
Let the filter be the reflectionThis is well defined because is invertible. In fact, if , then and of around the pivot point by the group , i.e. It follows that:
We can now use the change of variable to write
Thus we see that the group convolution amounts to applying a certain filter to the warped image . The filter elements are reweighed by the determinant of the Jacobian of , which accounts for the stretching and shrinking of space due to the non-linear map. In practice, both the reflection and Jacobian can be absorbed into a learned filter, making such calculations unnecessary. Nevertheless, they offer a complementary view of warped convolutions.
2 Efficiency vs. generality
By reducing to standard convolution, warped convolution allows one to take full advantage of modern convolution implementations (Lavin, 2015; Lyons, 2010), including those with lower computational complexity (e.g. FFT (Lyons, 2010)). However, while warped convolution works with an important class of transformations (including the ones considered in previous works Kanazawa et al. (2014); Cohen & Welling (2014); Marcos et al. (2016)), non-trivial restrictions are imposed on the transformation group: it must be Abelian and have only two parameters.
By contrast, the group-theoretic convolution operator of eq. 5 does not make (almost) any restriction on the transformation group. Unfortunately, it is in general significantly more difficult to compute efficiently than the special case we consider here. To understand some of the implementation challenges, consider specializing eq. 7 to a discrete group such as a discrete set of planar rotations. In this case the Haar measure is trivial and equal to 1, and one has:
Direct computation of this equation has complexity where is the cardinality of the discrete group. Assuming that is in the order of where is the resolution of the input image (as it would be for standard convolution), one would obtain a complexity of . In practice, since usually the support of a filter is much smaller than the image, this complexity might reduce to , which is the complexity for the spatial domain implementation of convolution; however, compared to the standard case, this has two major disadvantages. First, the image is sampled in a spatially-varying manner, using bilinear or other interpolation, which foregoes the benefit of the regular, predictable, and local pattern of computations in standard convolution. This makes high-performance implementation of the naive algorithm difficult, particularly on GPUs. Secondly, it precludes the use of faster convolution routines such as Winograd’s algorithm (Lavin, 2015) or the Fast Fourier Transform (Lyons, 2010), the later having lower computational complexity than exhaustive search (). The development of analogues of the FFT for other general groups remains the subject of active research (Tygert, 2010; Li & Yang, 2016), which we sidestep by reusing highly optimized standard convolutions.
In practice, most recent works focus on very coarse transformations that do not change the filter support and can be implemented strictly via permutations, like horizontal/vertical flips and 90∘ rotations (Dieleman et al., 2015; Cohen & Welling, 2014). Such difficulties may explain why group convolutions are not as widespread as CNNs.
Examples of spatial transformations
We now give some concrete examples of two-parameter spatial transformations that obey the conditions of section 2.4, and can be useful in practice.
Visual object detection tasks require predicting the extent of an object as a bounding box. While the location can be found accurately by a standard CNN, which is equivariant to translation, the size prediction could similarly benefit from equivariance to horizontal and vertical scale (equivalently, scale and aspect ratio).
Such a spatial transformation, from which a warp can be constructed, is given by:
2 Scale and rotation (log-polar warp)
Planar scale and rotation are perhaps the most obvious spatial transformations in images, and are a natural test case for works on spatial transformations (Kanazawa et al., 2014; Marcos et al., 2016). Rotating a point by radians and scaling it by , around the origin, can be performed with
The resulting warp grid can be visualized in fig. 1-c. It is interesting to observe that it corresponds exactly to the log-polar domain, which is used in the signal processing literature to perform correlation across scale and rotation (Tzimiropoulos et al., 2010; Reddy & Chatterji, 1996). In fact, it was the source of inspiration for this work, which can be seen as a generalization of the log-polar domain to other spatial transformations.
3 3D sphere rotation under perspective
We will now tackle a more difficult spatial transformation, in an attempt to demonstrate the generality of our result. The transformations we will consider are yaw and pitch rotations in 3D space, as seen by a perspective camera. In the experiments (section 6) we will show how to apply it to face pose estimation.
In order to maintain additivity, the rotated 3D points must remain on the surface of a sphere. We consider a simplified camera and world model, whose only hyperparameters are a focal length , the radius of a sphere , and its distance from the camera center . The equations for the spatial transformation corresponding to yaw and pitch rotation under this model are in appendix A.
The corresponding warp grid can be seen in fig. 1-d. It can be observed that the grid corresponds to what we would expect of a 3D rendering of a sphere with a discrete mesh. An intuitive picture of the effect of the warp grid in such cases is that it wraps the 2D image around the surface of the 3D object, so that translation in the warped space corresponds to moving between vertexes of the 3D geometry.
Experiments
As mentioned in section 2.2, group convolution can perform an exhaustive search for patterns across spatial transformations, by varying pose parameters. For tasks where invariance to that transformation is important, it is usual to pool the detection responses across all poses (Marcos et al., 2016; Kanazawa et al., 2014).
In the experiments, however, we will test the framework in pose prediction tasks. As such, we do not want to pool the detection responses (e.g. with a max operation) but rather find the pose with the strongest response (i.e., an argmax operation). To perform this operation in a differentiable manner, we implement a soft argmax operation, defined as follows:
Our base architecture then consists of the following blocks, outlined in fig. 2. First, the input image is warped with a pre-generated grid, according to algorithm 1. The warped image is then processed by a standard CNN, which is now equivariant to the spatial transformation that was used to generate the warp grid. A soft argmax (eq. 14) then finds the maximum over pose-space. To ensure the pose prediction is well registered to the reference coordinate system, a learnable scale and bias are applied to the outputs; multiple predictions can be linearly combined into a single one at this stage. Training proceeds by minimizing the loss between the predicted pose and ground truth pose.
2 Google Earth
For the first task in our experiments, we will consider aerial photos of vehicles, which have been used in several works that deal with rotation invariance (Liu et al., 2014; Schmidt & Roth, 2012; Henriques et al., 2014).
The Google Earth dataset (Heitz & Koller, 2008) contains bounding box annotations, supplemented with angle annotations from (Henriques et al., 2014), for 697 vehicles in 15 large images. We use the first 10 for training and the rest for validation. Going beyond these previous works, we focus on the estimation of both rotation and scale parameters. The object scale is taken to be the diagonal length of the bounding box. Due to its small size, we augment the dataset during training, by randomly rotating the images uniformly over 360∘ and scaling them by up to 20%.
Implementation.
A image is cropped around each vehicle, and then fed to a network for pose prediction. The proposed method, Warped CNN, follows the architecture of section 6.1 (visualized in fig. 2). The CNN block contains 3 convolutional layers with filters, with 50, 20 and 50 output channels respectively. We use dilation factors of 2, 4 and 8 respectively (à trous convolution (Chen et al., 2015)), which increases the receptive field and resolution without adding paramenters. There is a batch normalization and ReLU layer after each convolution, and a max-pooling operator (stride 2) after the second one. The CNN block outputs response maps over 2D pose-space, which in this case consists of rotation and scale. All networks are trained for 40 epochs using the ADAM solver (Kingma & Ba, 2015) and implemented in MatConvNet (Vedaldi & Lenc, 2015). Angular error is taken modulo 180∘ due to annotation ambiguity. We report the average validation error over 3 runs.
Baselines and results.
The results of the experiments are presented in table 1, which shows angular and scale error in the validation set. To verify whether the proposed warped convolution is indeed responsible for a boost in performance, rather than other architectural details, we compare it against a number of baselines with different components removed. The first baseline, CNN+softargmax, consists of the same architecture but without the warp (section 5.2). This is a standard CNN, with the soft argmax at the end. Since CNNs are equivariant to translation, rather than scale and rotation, we observe a drop in performance. For the second baseline, CNN+FC, we replace the soft argmax with a fully-connected (FC) layer, to allow a prediction that is not equivariant with translation. Due to the larger number of parameters, there is overfitting and a large drop in performance. We also compare against the method of (Dieleman et al., 2015), which applies the same CNN to 90∘ rotations and flips of the image, combining the result with a FC layer. Just like with CNN+FC, there is overfitting on this small dataset, which requires very fine angular predictions. The proposed Warped CNN achieves the best results, except for scale prediction where (Dieleman et al., 2015) performs better. Our method has the same runtime performance as the CNN baselines, since the cost of a single warp is negligible, however (Dieleman et al., 2015) is 16 slower, since it applies the same CNN to multiple transformed images.
3 Faces
We now turn to face pose estimation in unconstrained photos, which requires handling more complex 3D rotations under perspective.
For this task we use the Annotated Facial Landmarks in the Wild (AFLW) dataset (Koestinger et al., 2011). It contains about 25K faces found in Flickr photos, and includes yaw (left-right) and pitch (up-down) annotations. We removed 933 faces with yaw larger than 90 degrees (i.e., facing away from the camera), resulting in a set of 24,384 samples. 20% of the faces were set aside for validation.
Implementation.
The region in each face’s bounding box is resized to a image, which is then processed by the network. Recall that our simplified 3D model of yaw and pitch rotation (section 5.3) assumes a spherical geometry. Although a person’s head roughly follows a spherical shape, the sample images are centered around the face, not the head. As such, we use an affine Spatial Transformer Network (STN) (Jaderberg et al., 2015) as a first step, to center the image correctly. Similarly, because the optimal camera parameters (, and ) are difficult to set by hand, we let the network learn them, by computing their derivatives numerically (which has a low overhead, since they are scalars). The rest of the network follows the same diagram as before (fig. 2). The main CNN has 4 convolutional layers, the first two with filters, the others being . The numbers of output channels are 20, 50, 20 and 50, respectively. A max-pooling with a stride of 2 is performed after the first layer, and there are ReLU non-linearities between the others. As for the STN, it has 3 convolutional layers (), with 20, 50 and 6 output channels respectively, and max-pooling (stride 2) between them. The remaining experimental settings are as in section 6.2.
Baselines and results.
The angular error of the proposed equivariant pose estimation, Warped CNN, is shown in table 2, along with a number of baselines. The goal of these experiments is to demonstrate that it is possible to achieve equivariance to complex 3D rotations. To compare with non-equivariant models, we test two baselines with the same CNN architecture, where the softargmax is replaced with a fully-connected (FC) layer. We include the Spatial Transformer Network (Jaderberg et al., 2015) and CNN+FC, which is a standard CNN of equivalent (slightly larger) capacity. We observe that neither the FC or the STN components account for the performance of the warped convolution, which better exploits the natural 3D rotation equivariance of the data.
Conclusions
In this work we show that it is possible to reuse highly optimized convolutional blocks, which are equivariant to image translation, and coax them to exhibit equivariance to other operators, including 3D transformations. This is achieved by a simple warp of the input image, implemented with off-the-shelf components of deep networks, and can be used for image recognition tasks involving a large range of image transformations. Compared to other works, warped convolutions are simpler, relying on highly optimized convolution routines, and can flexibly handle many types of continuous transformations. Studying generalizations that support more than two parameters seems like a fruitful direction for future work. In addition to the practical aspects, our analysis offers some insights into the fundamental relationships between arbitrary image transformations and convolutional architectures.
Appendix A Spatial transformation for 3D sphere rotation under perspective
Our simplified model consists of a perspective camera with focal length and all other camera parameters equal to identity, at a distance from a centered sphere of radius (see fig. 1-d).
A 2D point in image-space corresponds to the 3D point
Raycasting it along the axis, it will intersect the sphere surface at the 3D point
If the argument of the square-root is negative, the ray does not intersect the sphere and so the point transformation is undefined. This means that the domain of the image should be restricted to the sphere region. In practice, in such cases we simply leave the point unmodified. Then, the yaw and pitch coordinates of the point on the surface of the sphere are
These polar coordinates are now rotated by the spatial transformation parameters, . Converting them back to a 3D point
Finally, projection of into image-space yields
This research was funded by ERC StG 638009-IDIU.