Scale equivariance in CNNs with vector fields
Diego Marcos, Benjamin Kellenberger, Sylvain Lobry, Devis Tuia
Introduction
Equivariance to a predefined set of transformations can be a very desirable property of computer vision models. For instance, translation equivariance has a big role in the success of Convolutional Neural Networks (CNN) when applied to images, since in many applications a pattern on the image should be detected independently from its location. This can be clearly seen in problems such as semantic segmentation or optical flow estimation, where a translation of the input image should results in a translation of the output.
In this work, we propose a method for enforcing scale equivariance within CNNs. Our proposition has three main steps:
1) A convolutional filter is applied at multiple scales and only information about the maximally activating scale at each location is passed to the output.
2) The output is composed of the magnitude of the maximally activating scale and of the corresponding scale value itself. These two values are combined into a 2D vector.
3) Convolutions on this vector field use a similarity metric that takes into account both magnitudes and scale factors.
We propose an end-to-end learnable framework based on these principles that can be applied to any kind of convolutional architecture and allow encoding both invariance and equivariance to scale, depending on the nature of the task. In this paper, we validate the idea on a scale-invariant image classification problem, where digits from the MNIST dataset are randomly scaled, and on a scale equivariant regression problem, where we aim at regressing the scaling factor applied to the digits. In both cases, we outperform state of the art scale-invariant models.
Related Work
Our model builds on the work on local scale invariance by (Kanazawa et al., 2014), where filters are applied at multiple scales and only the maximum activation value per location is kept, as well as on an application to scale of the vector field representation of (Marcos et al., 2017), that was proposed to disentangle rotation and content information.
The modified convolution operator presented by (Kanazawa et al., 2014) is locally scale invariant and consists of all the elements shown in Fig. 1, except that the output consists only of the magnitude map. If we apply it to an image and look at the value at a single location of the output, small variations in scale of the underlying local object will have no effect on this value (see Fig. 2 middle). This is precisely the sought behavior. However, if these scale variations contained information useful for the task at hand but were discarded, we would be hampering the model’s performance. For instance, in histology and cytology, the sizes of cells and organelles can provide hints about their functionality, while in scene understanding, apparent object sizes implicitly encode distance from the camera and can be valuable information for scene segmentation (Zhang et al., 2017).
Other methods in the literature can be used to obtain global (but not local) scale invariance, and consist of applying multiple Siamese networks with either differently scaled filters (Xu et al., 2014) or differently scaled input images (Laptev et al., 2016). As noted by the authors of (Laptev et al., 2016), enforcing scale invariance can lead to a loss in performance. This might happen when the relative sizes of certain features on the image are important for the task: suppose we want a model that detects whether an image contains a duck family. A scale-invariant duck detector with a single appearance model will simply detect multiple ducks, aggravating the task of distinguishing a duck family from a flock of similarly sized ducks. Making equivariance explicit in the model can disentangle appearance and scale, allowing to keep a single appearance model. This maintains the advantages of scale-invariant models by still allowing a single appearance model, while retaining the information about relative sizes of the detected elements. Such a model will be able to distinguish the presence of a larger duck accompanied by a number of ducklings.
Equivariance to other transformations, chiefly rotation, through weight sharing has gathered a lot of attention in recent literature (Cohen & Welling, 2016; Dieleman et al., 2016; Ravanbakhsh et al., 2017; Weiler et al., 2017; Zhou et al., 2017). Some works propose to use enriched representations (beyond scalar fields) encoding equivariance while maintaining more compact intermediate representations, based on vector, tensor or complex number fields (Marcos et al., 2017; Thomas et al., 2018; Worrall et al., 2017). In this contribution, we propose to use a vector field representation similar to the one in (Marcos et al., 2017) to obtain a locally scale-equivariant deep CNN.
Scale equivariance
In our case, let correspond to the scaling of an object or pattern and to a linear operator that implements this scaling, such as through some interpolation method. We say that is scale invariant if:
while we say that is scale equivariant if we can find some representation of on , other than the identity, such that:
where the phase changes proportionally to the scaling factor.
Scale-Equivariant convolution with vector fields
We propose to build a scale equivariant CNN by using convolutions that apply each learned filter at different scales and then pool across scales, as shown in Fig. 1. Similarly to (Kanazawa et al., 2014), we obtain multiple copies of the input image or tensor using bilinear interpolation. We then apply a convolution with the same filter to this set of inputs, interpolate the outputs back to the original size and perform a max-pooling operation across the different scales. After such scale-pooling, we enrich the output of the modified convolution (the largest magnitude) with the argmax across scales, i.e. information about the scale that most activated the filter at that location. Fig. 2 (right) exemplifies the equivariance to local scaling provided by this operator. A convolution filter applied on this output also has to be formed by elements containing magnitude and scale. It is straightforward to see that the naïve use of the dot product as a similarity metric would not provide the desired results: we want the similarity to be high when the scales in both the input and the filter are close to each other, and not just when both values are high. Although it would be worth investigating the use of distance based metrics, in this work we follow an approach comparable to (Marcos et al., 2017) and encode the two values as a 2D vector in Cartesian coordinates where the length is the magnitude of the max-pooling and the angle is proportional to the argmax. The vector field convolution operator we use is the same as described in (Marcos et al., 2017) and is based on applying two standard convolutions, corresponding to each orthogonal component in the vector fields, and summing both results.
The total range of angles from the minimum to the maximum scale determines the interaction between different scales in the input and the filter. If the angle range is smaller than , there will be a certain level of positive interaction between the smallest and the largest scales, while a range larger than means that the smallest and the largest scales interact negatively in terms of similarity. The range should not surpass . A larger range would mean that the smallest and the largest scales would not be the most dissimilar.
Experiments
In order to showcase the potential of the proposed model, we apply it to the task of simultaneous classification and scale factor regression.
MNIST-scale is a variation of the MNIST digit classification benchmark introduced by (Sohn & Lee, 2012). It is built by rescaling each image in the original dataset by a factor randomly sampled from a uniform distribution between 0.3 and 1, followed by zero-padding to a size of pixels. We randomly rescale all the images in the original dataset and randomly select 10k samples for training, 2k for validation and 50k for testing. All the results shown are averages over six realizations of the modified dataset.
2 Architecture
In all our experiments, we use an architecture with three convolutional layers, all with filters. The scale invariant and equivariant architectures use 12, 32 and 48 filters in each layer, while the standard CNN was found to provide the best results with three times as many filters. In all models, the last convolutional block is followed by a fully connected layer with 256 hidden units that is then mapped to predict the class, as well as the scale factor in the standard and the invariant models. The regression is based on an MSE loss. In the scale equivariant model, only the magnitudes of the 48-dimensional representation are used for the fully connected layer, which is then solely used to predict the class. The scale factor is directly predicted by linearly combining the angles in the 48-dimensional representation, without an additional hidden layer. We used a total of eight scales, three larger and four smaller than the original, using a scale factor of 1.25 between them. The vector representation was empirically chosen to have an angle range of .
Results and discussion
MNIST-scale classification. Table 1 shows the results of the three tested models, compared to previously published results on the same dataset. (Kanazawa et al., 2014) observed a improvement by making a two-layer CNN locally invariant to scale. Similarly, we obtain a improvement by making a three-layers CNN locally scale-invariant. Interestingly, an additional can be obtained by switching from local invariance to local equivariance, i.e. moving from our model using only magnitude to the full model using 2D vector fields. This means that, even for a scale invariant task like MNIST-scale classification, keeping the information about the selected scales, thus allowing the model to learn the interactions between different relative scales, can add a substantial boost in performance.
MNIST-scale scale factor regression. The results on the scale factor regression are shown in Table 2. In this case, we observe no improvement at all from injecting scale invariance into the model, and even a slight decrease in accuracy. This is to be expected, since the scale-invariant model explicitly removes information on scale, potentially hampering the regression task. On the other hand, there is a substantial improvement in the scale factor prediction by using the scale-equivariant model, since the orientation of the vectors in the vector field layers is built to be linearly dependent on the scale of the features found in the input image.
Conclusion
We presented a method for injecting equivariance to local scale variations in images. This is done by applying each convolutional filter at multiple scales and outputting, for each location, the magnitude and the scale value of the maximally activated scale. These two values are represented as a 2D vector whose length encodes the magnitude and whose orientation encodes the scale. A vector field convolution can be applied on this output to build a deep architecture.
The results on MNIST-scale show that a locally scale equivariant model provides a substantial improvement over a scale-invariant one even in a task that is itself scale-invariant by nature and a 3-fold reduction in the number of learnable filters with respect to the best standard CNN model. This highlights the importance of allowing the model to learn the relationships between different relative scales. Less surprisingly, we also found that scale equivariance helps in the task of predicting the scale factor that has been applied to an object in the input image.