Learning Steerable Filters for Rotation Equivariant CNNs

Maurice Weiler, Fred A. Hamprecht, Martin Storath

Introduction

Convolutional neural networks are extremely successful predictive models when the input data has spatial structure. One principal reason is that the convolution operation exhibits translational equivariance so that feature extraction is independent of the spatial position. For many types of images it is desirable to make feature extraction orientation independent as well. Typical examples are biomedical microscopy images or astronomical data which do not show a prevailing global orientation. Consequently, the output of a network processing such data should be equivariant w.r.t. the orientation of its input – if the input is rotated, the output should transform accordingly. Even when there is a predominant direction in an image as a whole, the low level features in the first layers such as edges usually appear in all orientations; see e.g. the filterbanks visualized in . In both cases, conventional CNNs are compelled to learn rotated versions of the same filter, introducing redundant degrees of freedom and increasing the risk of overfitting.

We propose a rotation-equivariant CNN architecture which shares weights over filter orientations to improve generalization and to reduce sample complexity. A key property of our network is that its filters are learned such that they are steerable. This approach avoids interpolation artifacts which can be severe at the small length scale of typical filter kernels. We accomplish the steerability of the learned filters by representing them as linear combinations of a fixed system of atomic steerable filters.

In all intermediate layers of the network, we utilize group convolutions to ensure an equivariant mapping of feature maps. Group-convolutional networks were proposed by Cohen and Welling who considered four filter orientations. An advantage of our construction is that we can achieve an arbitrary angular resolution w.r.t. the sampled filter orientations. Indeed, our experiments show that results improve significantly when using more than four orientations.

An important practical aspect of CNNs is a proper weight initialization. Since the weights to be learned serve as expansion coefficients for the steerable function space, common weight initialization schemes need to be adapted. Here, we generalize the results found by Glorot and Bengio and He et al. to networks which learn filters as a composition of (not necessarily steerable) atomic filters.

Our network achieves state-of-the-art results on two important rotation-equivariant/invariant recognition tasks: (i) The proposed approach is the first to obtain an accuracy higher than 99% on the rotated MNIST dataset, which is the standard benchmark for rotation-invariant classification. (ii) A processing pipeline based on the proposed SFCNN layers ranks among the top three entries in the ISBI 2012 electron microscopy segmentation challenge .

Figure 1 gives an overview over the key concepts utilized in Steerable Filter CNNs.

Equivariance properties of CNNs

Equivariance is the property of a function to commute with the actions of a symmetry group acting on its domain and codomain. Formally, given a transformation group G,G, a function f:X→Yf:X\to Y is said to be equivariant if

In many machine learning tasks a set of transformations is known a-priori under which the prediction should transform in an equivariant way. Including such knowledge directly into the model can greatly facilitate learning by freeing up model capacity for other factors of variation. As an example consider a segmentation problem where the goal is to learn a mapping from an image space I\mathcal{I} to label images in L\mathcal{L}, which we formalize by a ground truth segmentation map S:I→L.\mathcal{S}:\mathcal{I}\to\mathcal{L}. The learning process involves fitting a model M:I→L\mathcal{M}:\mathcal{I}\to\mathcal{L} to approximate the ground truth. For segmentation tasks, however, translations of the input image I∈II\in\mathcal{I} should typically lead to a translated segmentation map. Specifically, one has

CNN layers, which transform feature maps ζ\zeta by convolving them with filters Ψ,\Psi, are by construction equivariant under translations, that is, (Tdζ)∗Ψ=Td(ζ∗Ψ).\left(\mathcal{T}_{d}\zeta\right)\ast\Psi=\mathcal{T}_{d}\left(\zeta\ast\Psi\right). Therefore, their hypothesis space is restricted to M~.\widetilde{M}. In practice one often uses strided pooling layers which make the prediction more robust to local deformations but reduce the equivariance to a subgroup determined by the stride. As consequence, patterns learned at one specific location evoke the same response at each other location which leads to reduced sample complexity and enhanced generalization.

Besides translations, there are often further transformations like rotations, mirroring or dilations under which the model should be equivariant. Enforcing equivariance under an extended transformation group GG leads to an enhanced generalization over larger orbits G.IG.I and reduces the hypothesis space further to M~:I/G→L/G.\widetilde{\mathcal{M}}:\mathcal{I}/G\to\mathcal{L}/G.

Steerable Filter CNNs

Here, we develop Steerable Filter CNNs (SFCNNs) which achieve equivariance under joint translations and discrete rotations. The key concept leading to translation equivariance of CNNs is translational weight sharing. We extend the transformation group under which our networks’ layers are equivariant by additionally sharing weights over filter orientations. This implies to perform convolutions with several rotated versions of each filter. The rotational weight sharing leads to an improved sample complexity and to an enhanced generalization over orbits consisting of images connected by translations and discrete rotations.

At the heart of convolutional neural networks lies the concept of learning filter kernels. Our construction demands for filters whose responses can be computed accurately and economically for several filter orientations. Simultaneously the filters should not be restricted in their expressive power, i.e. in the patterns to be learned. All of these requirements are met by learning linear combinations of a system of steerable filters. Here we describe a suitable construction of steerable filters for learning in CNNs.

for all angles θ∈(−π,π]\theta\in(-\pi,\pi] and for angular expansion coefficient functions κq\kappa_{q}. Here ρθ\rho_{\theta} denotes both the rotation operator defined by ρθΨ(x)=Ψ(ρ−θx)\rho_{\theta}\Psi(x)=\Psi(\rho_{-\theta}x) when acting on a function as well as a counterclockwise rotation by the angle θ\theta when acting on a coordinate vector. As pointed out by Freeman and Adelson , the rotation by steerability is analytic and exact even for signals sampled on a grid. In contrast to rotations by interpolation the approach does not suffer from interpolation artifacts. An important practical consequence of steerability is that the response of each orientation can be synthesized from the atomic responses f∗ψqf\ast\psi_{q}; that is, (f∗ρθΨ)(x)=∑q=1Qκq(θ)(f∗ψq)(x).\left(f\ast\rho_{\theta}\Psi\right)(x)=\sum\nolimits_{q=1}^{Q}\kappa_{q}(\theta)\left(f\ast\psi_{q}\right)(x).

The learned filters are then defined as linear combinations of the elementary filters, that is,

We select a single orientation by taking their real part

2 Equivariant network architecture

The basic building blocks of the proposed SFCNN are three equivariant layer types which we introduce in this section.

where the filters are rotated by in total Λ\Lambda equidistant orientations θ∈Θ={0,…,2πΛ−1Λ}\theta\in\Theta=\{0,\ldots,2\pi\frac{\Lambda-1}{\Lambda}\}. In this setting the rotational weight sharing is reflected by the phase manipulation of the weights wc^cjkw_{\hat{c}cjk} which themselves are independent of the angle θ\theta. A higher resolution in orientations can be achieved by simply expanding the tensor containing the phase-factors.

As usual, after the convolution step a bias βc^(1)\beta^{(1)}_{\hat{c}} is added and a nonlinearity σ\sigma is applied, so that we end up with the first layer’s feature map given by

Note that the resulting representation ζc^(1)\zeta^{(1)}_{\hat{c}} depends on a spatial location xx and an orientation angle θ,\theta, i.e. on the transformation group applied to the filters.

Here the multiplication with the inverse group element, (u,ϕ)−1(x,θ)=(ρ−ϕ(x−u),θ−ϕ)(u,\phi)^{-1}(x,\theta)=(\rho_{-\phi}(x-u),\theta-\phi), was evaluated by switching to a representation of the group. We further introduced the action Rϕ\mathcal{R}_{\phi} defined by

which transforms functions on the group by rotating them spatially and shifting their orientation components cyclically. The above equation reveals that the group convolution can be decomposed into a spatial convolution, rotation and linear combination. In analogy to the first layer we make use of the steerable filters which on the group are defined by Ψc^c(l)(x,θ)=∑j=1J∑k=0Kjwc^cjkθψjk(x).\Psi_{\hat{c}c}^{(l)}(x,\theta)=\sum\nolimits_{j=1}^{J}\sum\nolimits_{k=0}^{K_{j}}w_{\hat{c}cjk\theta}\psi_{jk}(x). Note that the additional orientation dimension is reflected by an additional index of the weight tensor. Inserting the steerable filters in (9) we obtain the pre-nonlinearity feature maps of the group-convolutional layers

As before, a bias βc^(l)\beta^{(l)}_{\hat{c}} is added and the activation function σ\sigma is applied, ζc^(l)(x,θ)=σ(yc^(l)(x,θ)+βc^(l)).\zeta^{(l)}_{\hat{c}}(x,\theta)=\sigma(y^{(l)}_{\hat{c}}(x,\theta)+\beta^{(l)}_{\hat{c}}).

By the linearity of the steerability and the convolution, one can implement the layers either by a direct convolution with linearly combined filters, or by linearly combining the responses of the atomic filters. We implemented both approaches and found that in typical operation regimes the first option is faster since the kernels to be linearly combined have a smaller spatial extent than the atomic responses of the second option.

Output layer: After the last group-convolutional layer we extract the information of interest for the specific task. For rotation-invariant classification we pool globally over both the orientation dimension and the remaining spatial resolution. A pooling over orientations is also done for rotation-equivariant segmentation where spatial dimensions remain and the output rotates according to the rotation of the network’s input. If the orientation itself is of interest it could be kept as extra feature.

where dd is the number of group-convolutional layers. The layers’ equivariance is formally proven in appendix A.

The top part of Figure 3 visualizes the building blocks of a typical SFCNN for rotation-equivariant segmentation. An overview over the transformation behavior of the feature maps under rotation of the input is given in the bottom part. The spatial rotation and cyclic shift over orientation channels Rϕ\mathcal{R}_{\phi} of the feature maps on the group can be understood intuitively when paying attention to the relative orientation of each layer’s input and filters.

Compared to a conventional CNN which independently learns filters in Λ\Lambda orientations in a rotation-invariant recognition task, a corresponding SFCNN consumes Λ\Lambda times less parameters to extract the same representation.

SFCNN incur a small computational overhead for building the filter kernels from the circular harmonics basis which we found to be negligible. The computational cost of SFCNNs is therefore equivalent to that of a conventional CNN when the effective number of channels coincide, i.e. when ICNN=ΛISFCNN.I_{\text{CNN}}=\Lambda I_{\text{SFCNN}}.

3 Generalizing He’s weight initialization scheme

An important practical aspect of training deep networks is an appropriate initialization of their weights. When the weights’ variances are chosen too high or low, the signals propagating through the network are amplified or suppressed exponentially with depth. Glorot and Bengio and He et al. investigated this issue and came up with initialization schemes which are accepted as a standard for random weight initialization. In contrast to and our filters are not parameterized in a pixel basis but as a linear combination of a system of atomic filters with weights serving as expansion coefficients. To be specific, we consider filters Ψc^cx=∑q=1Qwc^cqψqx\Psi_{\hat{c}cx}=\sum_{q=1}^{Q}w_{\hat{c}cq}\psi_{qx} which are built from QQ, not necessarily steerable, real valued atomic filters which map CC input channels to C^\hat{C} output channels. This assumption is more general than that of the aforementioned works since they only consider the pixel basis ψqxDirac=δq,x\psi_{qx}^{\text{Dirac}}=\delta_{q,x}, i.e. atomic filters which are zero everywhere but at one pixel.

Most of the further assumptions are identical to those in : We assume the activations and gradients to be i.i.d. and to be independent from the weights. Further, the weights themselves are initialized to be mutually independent and have zero mean. An important difference is that we do not restrict the weights to be identically distributed because of the inherent asymmetry of the different atomic filters. All biases are initialized to be zero and the nonlinearities are chosen to be ReLUs. These assumptions lead to the initialization conditions

for the forward or backward pass, respectively. A detailed derivation is given in appendix B.

As discussed in , the difference between both initializations cancels out for intermediate layers. Note that our results include those of He et al. , that is, Var⁡[wq]=2nin\operatorname{Var}\left[w_{q}\right]=\frac{2}{n_{\text{in}}} or Var⁡[wq]=2nout,\operatorname{Var}\left[w_{q}\right]=\frac{2}{n_{\text{out}}}, for ψqxDirac=δq,x\psi_{qx}^{\text{Dirac}}=\delta_{q,x} with ∥ψqxDirac∥22=1.\left\lVert\psi_{qx}^{\text{Dirac}}\right\rVert_{2}^{2}=1. We further want to point out that the learned filters are combined of products wqψqw_{q}\psi_{q} which implies that the factors ∥ψq∥22\left\lVert\psi_{q}\right\rVert_{2}^{2} counterbalance different energies of the basis filters. A convenient way to initialize the network is hence to normalize all filters to unit norm and subsequently initialize the weights uniformly by Var⁡[wq]=2CQ\operatorname{Var}\left[w_{q}\right]=\frac{2}{CQ} or Var⁡[wq]=2C^Q.\operatorname{Var}\left[w_{q}\right]=\frac{2}{\hat{C}Q}. In our group-convolutional layers the filters additionally comprise orientation channels. From the perspective of weight initialization these have the same effect as conventional channels, therefore we propose to normalize their weights variance with an additional factor of Λ.\Lambda. We emphasize that using normalization layers like batch normalization does not obviate the need for a proper weight initialization. This is because such layers scale activations as a whole while our initialization conditions indicates that the relative scale of the summands contributing to each activation needs to be adapted. Further details, in particular on initializing weights of complex-valued filters, are given in the appendix.

Prior and related work

A priori knowledge about transformation-invariance of images can be exploited in manifold ways. A commonly utilized technique is data augmentation, see e.g. . The basic idea is to enrich the training set by transformed samples. Augmenting datasets allows to train larger models and is easily applicable without modifying the network architectures. When the augmenting transformations form a group GG the additional images I⊆II\subseteq\mathcal{I} lie on the orbit G.IG.I. In contrast to equivariant models the hypothesis space is not restricted to the quotient space I/G\mathcal{I}/G under the utilized symmetry group but the equivariance needs to be learned explicitly by the network. This demands for a high learning capacity which makes the network prone to overfitting.

Recent work focuses on incorporating equivariance to various transformations directly into the network’s architecture. Invariance to specific transformations can be achieved by applying them to the input and subsequently pooling their responses . In the regions in symmetry space to pool over are learned to become invariant only to nuisance deformations. Another approach is to resample the input and apply standard convolutions. Henriques and Vedaldi achieve equivariance w.r.t. Abelian symmetry groups by fixing a sampling grid according to the symmetry while in the network itself estimates the grid. In transformations are dealt with by convolving with filters which are steered by a subnetwork.

In particular, there has recently been a considerable interest in rotation-equivariant CNNs. The work introduces four operations which are easily included into existing networks and enrich both the batch- and feature dimension with transformed versions of their content. In , the feature maps resulting from transformed filters are treated as functions of the corresponding symmetry-group which allows to use group-convolutional layers. As their computational cost is coupled to the size of the group, Cohen and Welling propose to alternatively use steerable representations as composition of elementary feature types. Besides translations and rotations, the aforementioned works also incorporate reflections, i.e. they operate on the dihedral group. Their current limitation is the restriction to rotations by the angle π2,\frac{\pi}{2}, thus to four orientations. In , several rotated versions of the same image are sent through a conventional CNN. The resulting features are subsequently pooled over the orientation dimension. The approach can be easily extended to other transformations. On the downside, the equivariance is only w.r.t. global transformations. Marcos et al. perform convolutions with rotated versions of a each filter in a shallow network followed by a global pooling over orientations. These ideas were extended to networks which additionally propagate the orientation of the maximum response . In both approaches the filter rotation is based on bicubic interpolation, allowing for fine resolutions with respect to the orientation but causing interpolation artifacts. Worrall et al. achieve continuous resolution in orientations by working with complex valued steerable filters and feature maps. However, this requires the angular frequencies of the feature maps to be kept disentangled. Rotation-equivariant feature extraction can also be achieved by using group-convolutional scattering transforms . A fundamental difference to our work is that the filter banks are fixed rather than learned.

Experimental results

We evaluate the proposed SFCNNs on two datasets exhibiting rotational symmetries. On the rotated MNIST dataset we first investigate specific network properties like the accuracy’s dependence on the number of sampled orientations and the generalization of learned patterns over orientations. With the insights gained in these experiments we benchmark the model and the proposed initialization scheme. To evaluate the segmentation capabilities of SFCNNs on real world data we run a further experiment on the ISBI 2012 EM segmentation challenge.

In our first experiments we investigate the equivariance properties of the proposed network architecture on the rotated MNIST dataset (mnist-rot) which is the standard benchmark for rotation-equivariant models. The dataset contains the handwritten digits of the classical MNIST dataset, rotated to random orientations in [0,2π)[0,2\pi). It is split in 1200012000 training and 5000050000 test images; model selection is done by training on 1000010000 images and validating on the 20002000 remaining samples in the training set.

For our initial experiments we utilize the classification-SFCNN given in Table 2 in the appendix as baseline. It consists of one steerable input layer which maps the input images to the group, five following group-convolutional layers and three fully connected layers. After every two steerable filter layers we perform a spatial 2×22\times 2 max-pooling. The orientation dimension and the remaining spatial dimensions are pooled out globally after the last convolutional layer. Details on the further training setup are given in appendix C.

Sampled orientations: The number of sampled orientations Λ\Lambda is a parameter specific to our network, so we first explore its influence on the test accuracy. We are further interested in the network’s sample complexity, i.e. the dependence on the size of the training set. The accuracies resulting when varying these parameters are reported in Figure 4 (left). As expected, the test error and its standard deviation decrease with the size of the training data set. We observe that the accuracy improves significantly when increasing the number of orientations until it saturates at around 1212 to 1616 angles. Up to this point, the gain of adding more sampled orientations is considerable. For example, in almost all cases, increasing the angular resolution from 22 to 44 sampled orientations provides a higher gain in accuracy than sticking with 22 orientations and doubling the number of training samples. We want to emphasize that the possibility of SFCNNs to go beyond the four sampled orientations of leads to a significant gain in accuracy. Note that the case Λ=1\Lambda=1 corresponds, up to the different filter parameterization, to conventional CNNs.

Rotational generalization: In order to test how well the networks generalize learned patterns over orientations we conduct an experiment where we train them on unrotated digits and record their accuracy over the orientation of rotated digits. Specifically, we take the the first 1200012000 samples of the conventional MNIST dataset to train a SFCNN with Λ=16\Lambda=16 as well as a conventional CNN of comparable size using either no augmentation, augmentation by rotations which are multiples of either π4\frac{\pi}{4} or π2\frac{\pi}{2} or augmentation by rotations which are densely sampled from [0,2π)[0,2\pi). As test set we take the remaining 5800058000 samples and record the test errors’ dependence on the orientation of this dataset. To obtain a fair comparison between the networks we experiment with conventional CNNs with the same number of parameters or the same number of channels like the SFCNN. Since both show the same behavior we only report the accuracies of the network with the same number of channels which performs slightly better. The results are plotted in Figure 4 (right). One can see that, lacking rotational equivariance, the conventional CNN does not generalize well over orientations. When using rotational augmentation the error reduces considerably on average, it grows, however, for small angles in a neighborhood of zero. This is the case because the network needs to learn to detect the augmented samples additionally which demands an increased learning capacity. The SFCNN on the other hand generalizes quite well over orientations even without augmentation. In continuous space we would expect the test error curve to be 2πΛ\frac{2\pi}{\Lambda}-periodic because of the rotational equivariance. The deviations from this behavior can be attributed to the sampling effects of using digitized images. As to be expected for Λ=16\Lambda=16 orientations, the accuracy is not influenced by augmentation with π2\frac{\pi}{2}-rotations since the additional samples lie on the group orbit on which the network is invariant. In contrast to conventional CNNs, SFCNNs do not show an increased error for small angles in a neighborhood of zero when using augmentation. This indicates that the cost of learning rotated versions of each digit is negligible thanks to the approximate rotation equivariance. An augmentation by rotations which are multiples of π4\frac{\pi}{4} or by continuous rotations give very similar results. Both seem to act as a regularization preventing the filters to overfit on the pixel grid.

We conclude that SFCNNs outperform the rotational generalization of CNNs for all levels of augmentation.

Benchmarking: Based on the insights from the above experiments we fix the number of sampled orientations to Λ=16\Lambda=16 and tune the network further to the slightly larger architecture given in Table 3 in the appendix. The results are reported in Table 1. Using the SFCNN with He’s weight initialization and no data augmentation, we obtain a test error of 0.957%0.957\% which already exceeds the previous state-of-the-art. The proposed initialization scheme, adapted to filter coefficients, significantly improves the test error to 0.880%.0.880\%. When additionally augmenting the dataset with continuous rotations during training time the error decreases further to 0.714%.0.714\%. To summarize, our approach reduces the best previously published error by a factor of 29%.29\%.

2 ISBI 2012 2D EM segmentation challenge

In a second experiment we evaluate the performance of our model on the ISBI 2012 electron microscopy segmentation challenge . The goal of the challenge is to predict the locations of the cell boundaries in the Drosophila ventral nerve cord from EM images which is a key step for investigating the connectome of the brain. The dataset consists of 3030 train and test slices of size 512×512512\times 512 px with a binary segmentation ground truth provided for the training set. Figure 5 shows an exemplary raw EM image with the corresponding ground truth segmentation mask and our network’s prediction. An important property of the dataset is that the images have no preferred orientation which makes it suitable for evaluating rotation-equivariant networks.

We build on an established pipeline introduced in where a crucial step is the boundary prediction via a conventional CNN. In the present experiment, we replaced their network by a SFCNN with a U-net design . The network architecture is visualized in Figure 6 in the appendix. As loss function we chose a pixel wise binary cross entropy loss. The dataset was augmented by random elastic deformations, flips and rotations by multiples of π2\frac{\pi}{2} during train time. In the experiment on rotational generalization we found that augmenting samples by transformations in a subgroup under which the network is equivariant does not have any effect. We therefore sampled Λ=17\Lambda=17 orientations which is mutually prime with the 44 augmented orientations. This way the augmented images do not fall into a subgroup w.r.t. which the network is invariant.

Segmentation predictions are evaluated by the challenge hosters and ranked w.r.t. the foreground-restricted Rand score VRandV^{\text{Rand}} and the information score VInfoV^{\text{Info}}; for an explanation of these metrics see . The current leaderboard in Figure 5 (right) shows that our approach yields top-tier results. In particular, it improves upon the results of .

Conclusion

We have developed a rotation-equivariant CNN whose filters are learned such that they are steerable. Layerwise equivariance is obtained by using group convolutions. He’s weight initialization scheme is extended to general filter bases which empirically leads to an increased accuracy. Our network allows sampling an arbitrary number of filter orientations which improves the performance until a saturation is reached. We confirmed experimentally that SFCNNs generalize learned patterns over orientations and therefore achieve a lower sampling complexity than CNNs in rotation-equivariant recognition tasks. The proposed SFCNNs achieve state-of-the-art results on rotated MNIST and the ISBI 2012 2D EM segmentation challenge.

We would like to thank T. Beier, C. Pape, N. Rahaman and I. Arganda-Carreras for their technical support and U. Köthe and T. Cohen for valuable discussions. This work was partially supported by the German Research Foundation (DFG grant STO1126/2-1).

Appendix A Equivariance properties

The mutual transformation behavior is visualized in the following commutative diagram:

ρα<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>∗</mo><msub><mi>ρ</mi><mi>θ</mi></msub><mimathvariant="normal">Ψ</mi></mrow><annotationencoding="application/x−tex">∗ρθΨ</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8778em;vertical−align:−0.1944em;"></span><spanclass="mord">∗</span><spanclass="mord"><spanclass="mordmathnormal">ρ</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3361em;"><spanstyle="top:−2.55em;margin−left:0em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmathnormalmtight"style="margin−right:0.0278em;">θ</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span><spanclass="mord">Ψ</span></span></span></span></span>Rα\rho_{\alpha}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>∗</mo><msub><mi>ρ</mi><mi>θ</mi></msub><mi mathvariant="normal">Ψ</mi></mrow><annotation encoding="application/x-tex">\ast\rho_{\theta}\Psi</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8778em;vertical-align:-0.1944em;"></span><span class="mord">∗</span><span class="mord"><span class="mord mathnormal">ρ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">θ</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord">Ψ</span></span></span></span></span>\mathcal{R}_{\alpha}∗ρθΨ\ast\rho_{\theta}\Psi Adding a bias β\beta to each feature map channel and applying a nonlinearity σ\sigma does not interfere with translational- or rotational equivariance since both operations do neither depend on the spatial position nor orientation channel:

A.2 Group-convolutional layers

Given feature maps ζ(l)(x,θ)\zeta^{(l)}(x,\theta), the group-convolutional layers perform an equivariant mapping of Rαζ(l)(x,θ)\mathcal{R}_{\alpha}\zeta^{(l)}(x,\theta) to Rαζ(l+1)(x,θ)\mathcal{R}_{\alpha}\zeta^{(l+1)}(x,\theta) under the group action R.\mathcal{R}. The step of adding the bias and applying the activation function is equivariant by the same argument as in the first layer. What is left to show is the equivariance (Rαζ(l)⊛Ψ)(x,θ)=Rα(ζ(l)⊛Ψ)(x,θ)=Rαy(l)(x,θ)\left(\mathcal{R}_{\alpha}\zeta^{(l)}\circledast\Psi\right)(x,\theta)=\mathcal{R}_{\alpha}\left(\zeta^{(l)}\circledast\Psi\right)(x,\theta)=\mathcal{R}_{\alpha}y^{(l)}(x,\theta) of the group convolution. Inserting a transformed feature map and writing the group convolution out explicitly yields:

This proves the equivariance of the intermediate layers. Again, the relations are illustrated in a commutative diagram:

A.3 Orientation max-pooling layer

For rotation-invariant segmentation or classification we max-pool over orientations after the last group-convolutional layer. The pooling step is itself equivariant and results in a rotated version of its output:

The rotation operator commutes with the maximum over orientation channels because it acts on spatial coordinates only. We again visualize the transformation behavior by a commutative diagram:

Rα<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><munder><mrow><mi>max</mi><mo>⁡</mo></mrow><mi>θ</mi></munder></mrow><annotationencoding="application/x−tex">max⁡θ</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:1.1827em;vertical−align:−0.7521em;"></span><spanclass="mopop−limits"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.4306em;"><spanstyle="top:−2.3479em;margin−left:0em;"><spanclass="pstrut"style="height:3em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmathnormalmtight"style="margin−right:0.0278em;">θ</span></span></span></span><spanstyle="top:−3em;"><spanclass="pstrut"style="height:3em;"></span><span><spanclass="mop">max</span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.7521em;"><span></span></span></span></span></span></span></span></span></span>ρα\mathcal{R}_{\alpha}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><munder><mrow><mi>max</mi><mo>⁡</mo></mrow><mi>θ</mi></munder></mrow><annotation encoding="application/x-tex">\max\limits_{\theta}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1827em;vertical-align:-0.7521em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.4306em;"><span style="top:-2.3479em;margin-left:0em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">θ</span></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span><span class="mop">max</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.7521em;"><span></span></span></span></span></span></span></span></span></span>\rho_{\alpha}max⁡θ\max\limits_{\theta} In the case of classification the remaining spatial structure is pooled out such that the output is invariant under transformations of the input.

Instead of the maximum pooling which we applied in our experiments, one could also utilize average pooling layers. The equivariance of average pooling can be derived in analogy to the derivation for maximum pooling.

Appendix B Derivation of the generalized He weight initialization scheme

In this section we give the derivation of the generalized weight initialization scheme whose results are stated in the main paper. For completeness we recall the assumptions going into the following calculations. We consider the activation of a single neuron in layer l,l,

where rectified linear units were chosen as nonlinearities. The pre-nonlinearity activations are given by the convolution with filters Ψ\Psi and summing over the input channels:

For convenience we shifted the addition of the bias to the pre-nonlinearity activations. The filters are defined by

that is, they are built from QQ real valued atomic filters ψq\psi_{q}. We keep the discussion general by not restricting the atomic filters to be steerable. In analogy to Glorot and Bengio and He et al. we assume the activations and gradients to be i.i.d. and to be independent from the weights. We let the weights themselves be mutually independent and have zero mean but do not restrict them to be identically distributed because of the inherent asymmetry coming from the different atomic filters. Furthermore we initialize all biases to be zero.

In order to prevent vanishing or exploding gradients of the loss E\mathcal{E} due to inappropriate initialization we demand their variance Var⁡[∂E∂ζ(l)]\operatorname{Var}\left[\frac{\partial\mathcal{E}}{\partial\zeta^{(l)}}\right] to be constant across all layers. It follows from (11) and (12) that the gradient with respect to the activation ζc0x0(l)\zeta_{c_{0}x_{0}}^{(l)} of a particular neuron in layer ll is given by

The factor 12\frac{1}{2} in the last line originates from the symmetric distribution of y(l)y^{(l)} in conjunction with the indicator function. Using the fact that the weights’ variances are initialized to only depend on qq and the assumption of identically distributed gradients, both can be pulled out of the sums:

It seems reasonable to assign the contribution to the overall variance equally to the QQ summands. Demanding the gradients’ variances to be constant over layers then leads to the initialization condition

B.2 Forward pass

where δ\delta denotes the delta distribution and Θ\Theta is the Heaviside step function. As before, we drop all indices which the random variables are independent from to compute the sums. This leads to

which in turn suggests a weight initialization according to

to ensure that the activations’ variances are not amplified.

B.3 Normalization of complex atomic filters

The results derived above suggest to initialize the weights of each layer uniformly by

Appendix C Details on the experimental setup

Here we give further details on the network architectures and the training setup of our experiments.

For our initial experiments on the dependence on sampled orientations and the networks’ rotational generalization capabilities we utilize the architecture given in Table 2 as baseline. Based on the results of these experiment we fix the number of sampled orientations to Λ=16\Lambda=16 and tune the network architecture further. We achieve the best benchmark results using the slightly larger network given in Table 3. In particular, we found that increasing the size of the filter masks improved the results. Both architectures consist of one steerable input layer which maps the input images to the group, five following group convolutional layers and three fully connected layers. After every two steerable filter layers we perform a spatial 2×22\times 2 max-pooling. The orientation dimension and the remaining spatial dimensions are pooled out globally after the last convolutional layer. We normalize the activations by adding batch normalization layers after each convolutional and fully connected layer. The batch normalization on the group does not interfere with the equivariance when the responses are normalized by averaging over both spatial and orientation dimensions.

The number of feature channels stated in the tables refers to the number of learned filters C^\hat{C} of the corresponding layer. As these filters are themselves applied with respect to Λ\Lambda orientations we end up with C^Λ\hat{C}\Lambda responses; e.g. 24⋅16=38424\cdot 16=384 effective responses in the first layer of the smaller network. Note that the extraction of this comparatively large number of responses without overfitting is possible because the rotational weight sharing leads to an increased parameter utilization (in the sense of Cohen and Welling ) by a factor Λ\Lambda.

All networks are trained for 4040 epochs using the Adam optimizer with standard parameters. The initial learning rate is set to 0.0150.015 and is decayed exponentially with a rate of 0.80.8 per epoch starting from epoch 15.15. We regularize the weights with an elastic net penalty with hyperparameters λL1=λL2\lambda_{L1}=\lambda_{L2} which are set to 10−710^{-7} and 10−810^{-8} for the convolutional and fully connected layers respectively. Dropout is used only in the fully connected layers with a dropping probability of p=0.3p=0.3.

C.2 ISBI 2012 EM segmentation challenge

The network architecture used to segment the membranes from raw EM images of neural tissue for the ISBI EM segmentation challenge is visualized in Figure 6. Inspired by the U-Net it is build as a symmetric encoder-decoder network with additional skip-connections between stages of the same resolution. This allows to extract semantic information from a large field of view while at the same time preserving precise spatial localization. Further, we adopt two modifications from : we do not concatenate the skipped feature maps but add it to the decoder features upsampled from the previous stage, and we use intermediate residual blocks (here of depth 1). On the highest resolution level we learn C^=12\hat{C}=12 filters, applied in Λ=17\Lambda=17 orientations which corresponds to C^Λ=204\hat{C}\Lambda=204 effective channels. The number of filters is doubled when going to the second and third level and is afterwards kept constant since we did not observe further gains in performance when adding more channels. All group-convolutional layers utilize kernels of size 7×77\times 7 pixels while the input layer applies 11×1111\times 11 pixel kernels.

As input, we feed the network cropped regions of 256×256256\times 256 pixels which are padded to 320×320320\times 320 pixels by reflecting a region of 3232 pixels around the borders to alleviate boundary artifacts. The padded regions are augmented by random elastic deformations, reflections and rotations by multiples of π2\frac{\pi}{2}. After the decoder we max-pool over orientations to obtain locally invariant features and crop out 256×256256\times 256 pixels centrally. Two subsequent 1×11\times 1 convolution layers map these features pixel-wise to the desired probability map.

The network is optimized by minimizing the spatially averaged binary cross-entropy loss between predictions and the ground truth segmentation masks using the ADAM optimizer. As on the rotated MNIST dataset we regularize the convolutional weights with an elastic net penalty with hyperparameters λL1=λL2\lambda_{L1}=\lambda_{L2} set to 10−710^{-7} and 10−810^{-8} for the steerable and 1×11\times 1 convolution layers respectively. Here we chose a dropout probability of p=0.4p=0.4 both in the steerable as well as in the 1×11\times 1 convolution layers. The learning rate is decayed exponentially by a factor of 0.850.85 per epoch starting from an initial rate of 5⋅10−25\cdot 10^{-2}.

References