Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks

Caglar Gulcehre, Kyunghyun Cho, Razvan Pascanu, Yoshua Bengio

Introduction

The importance of well-designed nonlinear activation functions when building a deep neural network has become more apparent recently. Novel nonlinear activation functions that are unbounded and often piecewise linear but not continuous such as rectified linear units (ReLU) , or rectifier, and maxout units have been found to be particularly well suited for deep neural networks on many object recognition tasks.

A pooling operator, an idea which dates back to the work in , has been adopted in many object recognizers. Convolutional neural networks which often employ max pooling have achieved state-of-the-art recognition performances on various benchmark datasets . Also, biologically inspired models such as HMAX have employed max pooling . A pooling operator, in this context, is understood as a way to summarize a high-dimensional collection of neural responses and produce features that are invariant to some variations in the input (across the filter outputs that are being pooled).

Recently, the authors of proposed to understand a pooling operator itself as a nonlinear activation function. The proposed maxout unit pools a group of linear responses, or outputs, of neurons, which overall acts as a piecewise linear activation function. This approach has achieved many state-of-the-art results on various benchmark datasets.

In this paper, we attempt to generalize this approach by noticing that most pooling operators including max pooling as well as maxout units can be understood as special cases of computing a normalized LpL_{p} norm over the outputs of a set of filter outputs. Unlike those conventional pooling operators, however, we claim here that it is beneficial to estimate the order pp of the LpL_{p} norm instead of fixing it to a certain predefined value such as ∞\infty, as in max pooling.

The benefit of learning the order pp, and thereby a neural network with LpL_{p} units of different orders, can be understood from geometrical perspective. As each LpL_{p} unit defines a spherical shape in a non-Euclidean space whose metric is defined by the LpL_{p} norm, the combination of multiple such units leads to a non-trivial separating boundary in the input space. In particular, an MLP may learn a highly curved boundary efficiently by taking advantage of different values of pp. In contrast, using a more conventional nonlinear activation function, such as the rectifier, results in boundaries that are piece-wise linear. Approximating a curved separation of classes would be more expensive in this case, in terms of the number of hidden units or piece-wise linear segments.

In Sec. 2 a basic description of a multi-layer perceptron (MLP) is given followed by an explanation of how a pooling operator may be considered a nonlinear activation function in an MLP. We propose a novel LpL_{p} unit for an MLP by generalizing pooling operators as LpL_{p} norms in Sec. 3. In Sec. 4 the proposed LpL_{p} unit is further analyzed from the geometrical perspective. We describe how the proposed LpL_{p} unit may be used by recurrent neural networks in Sec. 5. Sec. 6 provides empirical evaluation of the LpL_{p} unit on a number of object recognition tasks.

Background

A multi-layer perceptron (MLP) is a feedforward neural network consisting of multiple layers of nonlinear neurons . Each neuron uju_{j} of an MLP typically receives a weighted sum of the incoming signals {a1,…,aN}\left\{a_{1},\dots,a_{N}\right\} and applies a nonlinear activation function ϕ\phi to generate a scalar output such that

With this definition of each neuron We omit a bias to make equations less cluttered., we define the output of an MLP having LL hidden layers and qq output neurons given an input x\mathbf{x} by

where W[l]\mathbf{W}_{\left[l\right]} and ϕ[l]\phi_{\left[l\right]} are the weights and the nonlinear activation function of the ll-th hidden layer, and W[1]\mathbf{W}_{\left[1\right]} and U\mathbf{U} are the weights associated with the input and output, respectively.

2 Pooling as a Nonlinear Unit in MLP

Pooling operators have been widely used in convolutional neural networks (CNN) to reduce the dimensionality of a high-dimensional output of a convolutional layer. When used to group spatially neighboring neurons, this operator which summarizes a group of neurons in a lower layer is able to achieve the property of (local) translation invariance. Various types of pooling operator have been proposed and used successfully, such as average pooling, root-of-mean-squared (RMS) pooling and max pooling .

A pooling operator may be viewed instead as a nonlinear activation function. It receives input signal from the layer below, and it returns a scalar value. The output is the result of applying some nonlinear function such as max⁡\max (max pooling). The difference from traditional nonlinearities is that the pooling operator is not applied element-wise on the lower layer, but rather on groups of hidden units. A maxout nonlinear activation function proposed recently in is a representative example of max pooling in this respect.

LpL_{p} Unit

The recent success of maxout has motivated us to consider a more general nonlinear activation function that is rooted in a pooling operator. In this section, we propose and discuss a new nonlinear activation function called an LpL_{p} unit which replaces the max operator in a maxout unit by an LpL_{p} norm.

Given a finite vector/set of input signals [a1,…,aN]\left[a_{1},\dots,a_{N}\right] a normalized LpL_{p} norm is defined as

where pjp_{j} indicates that the order of the norm may differ for each neuron. It should be noticed that when 0<pj<1{0<p_{j}<1} this definition is not a norm anymore due to the violation of triangle inequality. In practice, we re-parameterize pjp_{j} by 1+log⁡(1+eρj)1+\log\left(1+e^{\rho_{j}}\right) to satisfy this constraint.

The input signals (also called filter outputs) aia_{i} are defined by

where x\mathbf{x} is a vector of activations from the lower layer. cic_{i} is a center, or bias, of the ii-th input signal aia_{i}. Both pjp_{j} and cic_{i} are model parameters that are learned.

We call a neuron with this nonlinear activation function an LpL_{p} unit. An illustration of a single LpL_{p} unit is presented in Fig. 1 (a).

Each LpL_{p} unit in a single layer receives input signal from a subset of linear projections of the activations of the layer immediately below. In other words, we project the activations of the layer immediately below linearly to A={a1,…,aN}{A=\left\{a_{1},\dots,a_{N}\right\}}. We then divide AA into equal-sized, non-overlapping groups of which each is fed into a single LpL_{p} unit. Equivalently, each LpL_{p} unit has its private set of filters.

The parameters of an MLP having one or more layers of LpL_{p} units can be estimated by using backpropagation , and in particular we adapt the order pp of the norm. The activation function is continuous everywhere except a finite set of points, namely when ai−cia_{i}-c_{i} is 0 and the absolute value function becomes discontinuous. We ignore these discontinuities, as it is done, for instance, in maxouts and rectifiers. In our experiments, we use Theano to compute these partial derivatives and update the orders pjp_{j} (through the parametrization of pjp_{j} in terms of ρj\rho_{j}), as usual with any other parameters.

2 Related Approaches

Thanks to the definition of the proposed LpL_{p} unit based on the LpL_{p} norm, it is straightforward to see that many previously described nonlinear activation functions or pooling operators are closely related to or special cases of the LpL_{p} unit. Here we discuss some of them.

If we further assume ai≥0a_{i}\geq 0, for instance, by using a logistic sigmoid activation function on the projection of the lower layer, the activation is reduced to computing the average of these projections. This is a form of average pooling, where the non-linear projections represent the pooled layer. With a single filter, this is equivalent to the absolute value rectification proposed in . If pjp_{j} is 22 instead of 11, the root-of-mean-squared pooling from is recovered.

As pjp_{j} grows and ultimately approaches ∞\infty, the LpL_{p} norm becomes

When N=2N=2, this is a generalization of a rectified linear unit (ReLU) as well as the absolute value unit . If each aia_{i} is constrained to be non-negative, this corresponds exactly to the maxout unit.

In short, the proposed LpL_{p} unit interpolates among different pooling operators by the choice of its order pjp_{j}. This was noticed earlier in as well as . However, both of them stopped at analyzing the LpL_{p} norm as a pooling operator with a fixed order and comparing those conventional pooling operators against each other. The authors of investigated a similar nonlinear activation function that was inspired by the cells in the primary visual cortex. In , the possibility of learning pp has been investigated in a probabilistic setting in computer vision.

On the other hand, in this paper, we claim that the order pjp_{j} needs to, and can be learned, just like all other parameters of a deep neural network. Furthermore, we conjecture that (1) an optimal distribution of the orders of LpL_{p} units differs from one dataset to another, and (2) each LpL_{p} unit in a MLP requires a different order from the other LpL_{p}. These properties also distinguish the proposed LpL_{p} unit from the conventional radial-basis function network (see, e.g., )

Geometrical Interpretation

We analyze the proposed LpL_{p} unit from a geometrical perspective in order to motivate our conjecture regarding the order of the LpL_{p} units. Let the value of an LpL_{p} unit uu be given by:

The space onto which x\mathbf{x} is projected may be spanned by linearly dependent vectors wi\mathbf{w}_{i}’s. Due to the possible lack of the linear independence among these vectors, they span a subspace S\mathcal{S} of dimensionality k≤Nk\leq N. The subspace S\mathcal{S} has its origin at c=[c1,…,cN]\mathbf{c}=\left[c_{1},\dots,c_{N}\right].

We impose a non-Euclidean geometry on this subspace by defining a norm in the space to be LpL_{p} with pp potentially not 22, as in Eq. (4). The geometrical object to which a particular value of the LpL_{p} unit corresponds forms a superellipse when projected back into the original input space. Since k≤Nk\leq N, the superellipse may be degenerate in the sense that in some of the N−kN-k axes the width may become infinitely large. However, as this does not invalidate our further argument, we continue to refer this kind of (degenerate) superellipse simply by an superellipse. The superellipse is centered at the inverse projection of c\mathbf{c} in the Euclidean input space. Its shape varies according to the order pp of the LpL_{p} unit and due to the potentially linearly-dependent bases. As long as p≥1p\geq 1 the shape remains convex. Fig. 1 (b) draws some of the superellipses one can get with different orders of pp, as a function of a1a_{1} (with a single filter).

In this way each LpL_{p} unit partitions the input space into two regions – inside and outside the superellipse. Each LpL_{p} unit uses a curved boundary of learned curvature to divide the space. This is in contrast to, for instance, a maxout unit which uses piecewise linear hyperplanes and might require more linear pieces to approximate the same curved segment.

When the dimensionality of the input space is 22 and each LpL_{p} receives 22 input signals, we can visualize the partitions of the input space obtained using LpL_{p} units as well as conventional nonlinear activation functions. Here, we examine some artificially generated cases in a 22-D space.

Fig. 2 shows a case of having two classes (∙\bullet and ∙\bullet) of which each corresponds to a Gaussian distribution. We trained MLPs having a single hidden neuron. When the MLPs had an LpL_{p} unit, we fixed pp to either 22 or ∞\infty. We can see in Fig. 2 (a) that the MLP with the L2L_{2} unit divides the input space into two regions -- inside and outside a rotated superellipse. Even though we use p=2p=2, which means an Euclidean space, we get a superellipse instead of a circle because of the linearly-dependent bases {w1,…,wN}.\left\{\mathbf{w}_{1},\dots,\mathbf{w}_{N}\right\}. The superellipse correctly identified one of the classes (red).

In the case of p=∞p=\infty, what we see is a degenerate rectangle which is an extreme form of a superellipse. The superellipse again spotted one of the classes and appropriately draws a separating curve between the two classes.

In the case of rectifier units it could find a correct separating curve, but it is clear that a single rectifier unit can only partition the input space linearly unlike LpL_{p} units. A combination of several rectifier units can result in a nonlinear boundary, specifically a piecewise-linear one, though our claim is that you need more such rectifier units to get an arbitrarily shaped curve whose curvature changes in a highly nonlinear way.

Three Classes, Two LpL_{p} Units

Similarly to the previous experiment, we trained two MLPs having two LpL_{p} units on data generated from a mixture of three Gaussian distribution. Again, each mixture component corresponds to each class.

For one MLP we fixed the orders of the two LpL_{p} units to 22. In this case, see Fig. 3 (a), the separating curves are constructed by combining two translated superellipses represented by the LpL_{p} units. These units were able to locate the two classes, which is sufficient for classifying the three classes (∙\bullet, ∙\bullet and ∙\bullet).

The other MLP had two LpL_{p} units with pp fixed to 22 and ∞\infty, respectively. The L2L_{2} unit defines, as usual, a superellipse, while the L∞L_{\infty} unit defines a rectangle. The separating curves are constructed as a combination of the translated superellipse and rectangle and may have more non-trivial curvature as in Fig. 3 (b).

Furthermore, it is clear from the two plots in Fig. 3 that the curvature of the separating curves may change over the input space. It will be easier to model this non-stationary curvature using multiple LpL_{p} units with different pp’s.

Decision Boundary with Non-Stationary Curvature: Representational Efficiency

In order to test the potential efficiency of the proposed LpL_{p} unit from its ability to learn the order pp, we have designed a binary classification task that has a decision boundary with a non-stationary curvature. We use 5000 data points of which a subset is shown in Fig. 4 (a), where two classes are marked with blue dots (∙\bullet) and red crosses (++), respectively.

On this dataset, we have trained MLPs with either LpL_{p} units, L2L_{2} units (LpL_{p} units with fixed p=2p=2), maxout units, rectifiers or logistic sigmoid units. We varied the number of parameters, which correspond to the number of units in the case of rectifiers and logistic sigmoid units and to the number of inputs signals to the hidden layer in the case of LpL_{p} units, L2L_{2} units and maxout units, from 22 to 1616. For each setting, we trained ten randomly initialized MLPs. In order to reduce effects due to optimization difficulties, we used in all cases natural conjugate gradient .

From Fig. 4 (c), it is clear that the MLPs with LpL_{p} units outperform all others in terms of representing this specific curve. They were able to achieve the zero training error with only three units (i.e., 6 filters) on all ten random runs and achieved the lowest average training error even with less units. Importantly, the comparison to the performance of the MLPs with L2L_{2} units shows that it is beneficial to learn the orders pp of LpL_{p} units. For example, with only two L2L_{2} units none of the ten random runs succeed while at least one succeeds with two LpL_{p} units. All the other MLPs, especially ones with rectifiers and maxout units which can only model the decision boundary with piecewise linear functions, were not able to achieve the similar efficiency of the MLPs with LpL_{p} units (see Fig 4 (b)).

Fig. 4 (a) also shows the decision boundary found by the MLP with two LpL_{p} units after training. As can be observed from the shapes of the LpL_{p} units (purple and cyan dashed curves), each LpL_{p} unit learned an appropriate order pp that enables them to model the non-stationary decision boundary. Fig. 4 (b) shows the boundary obtained by a rectifier model with four units. We can see that it has to use linear segments to compose the boundary, resulting in not perfectly solving the task. The rectifier model represented here has 64 mistakes, versus 0 obtained by the LpL_{p} model.

Although this is a low-dimensional, artificially generated example, it demonstrates that the proposed LpL_{p} units are efficient at representing decision boundaries which have non-stationary curvatures.

Application to Recurrent Neural Networks

A conventional recurrent neural network (RNN) mostly uses saturating nonlinear activation functions such as tanh⁡\tanh to compute the hidden state at each time step. This prevents the possible explosion of the activations of hidden states over time and in general results in more stable learning dynamics. However, at the same time, this does not allow us to build an RNN with recently proposed non-saturating activation functions such as rectifiers and maxout as well as the proposed LpL_{p} units.

The authors of recently proposed three ways to extend the conventional, shallow RNN into a deep RNN. Among those three proposals, we notice that it is possible to use non-saturating activations functions for a deep RNN with deep transition without causing the instability of the model, because a saturating non-linearity (tanh⁡\tanh) is applied in sandwich between the LpL_{p} MLP associated with each step.

The deep transition RNN (DT-RNN) has one or more intermediate layers between a pair of consecutive hidden states. The transition from a hidden state ht−1\mathbf{h}_{t-1} at time t−1t-1 to the next hidden state ht\mathbf{h}_{t} is

When a usual saturating nonlinear activation function is used for gg, the activations of the hidden state ht\mathbf{h}_{t} are bounded. This allows us to use any, potentially non-saturating nonlinear function for ff. We can simply use a layer of the proposed LpL_{p} unit in the place of ff.

As argued in , if the procedure of constructing a new summary which corresponds to the new hidden state ht\mathbf{h}_{t} from the combination of the current input xt\mathbf{x}_{t} and the previous summary ht−1\mathbf{h}_{t-1} is highly nonlinear, any benefit of the proposed LpL_{p} unit over the existing, conventional activation functions in feedforward neural networks should naturally translate to these deep RNNs as well. We show this effect empirically later by training a deep output, deep transition RNN (DOT-RNN) with the proposed LpL_{p} units.

Experiments

In this section, we provide empirical evidences showing the advantages of utilizing the LpL_{p} units. In order to clearly distinguish the effect of employing LpL_{p} units from introducing data-specific model architectures, all the experiments in this section are performed by neural networks having densely connected hidden layers.

Let us first list our claims about the proposed LpL_{p} units that need to be verified through a set of experiments. We expect the following from adopting LpL_{p} units in an MLP:

The optimal orders of LpL_{p} units vary across datasets

An optimal distribution of the orders of LpL_{p} units is not close to a (shifted) Dirac delta distribution

The first claim states that there is no universally optimal order pjp_{j}. We train MLPs on a number of benchmark datasets to see the resulting distribution of pjp_{j}’s. If the distributions had been similar between tasks, claim 1 would be rejected.

This naturally connects to the second claim. As the orders are estimated via learning, it is unlikely that the orders of all LpL_{p} units will convergence to a single value such as ∞\infty (maxout or max pooling), 11 (average pooling) or 22 (RMS pooling). We expect that the response of each LpL_{p} unit will specialize by using a distinct order. The inspection of the trained MLPs to confirm the first claim will validate this claim as well.

On top of these claims, we expect that an MLP having LpL_{p} units, when the parameters including the orders of the LpL_{p} units are well estimated, will achieve highly competitive classification performance. In addition to classification tasks using feedforward neural networks, we anticipate that a recurrent neural network benefits from having LpL_{p} units in the intermediate layer between the consecutive hidden states, as well.

2 Datasets

For feedforward neural networks or MLPs, we have used four datasets; MNIST , Pentomino , the Toronto Face Database (TFD) and Forest Covertype We use the first 16 principal components only. (data split DS2-581) . MNIST, TFD and Forest Covertype are three representative benchmark datasets, and Pentomino is a relatively recently proposed dataset that is known to induce a difficult optimization challenge for a deep neural network. We have used three music datasets from for evaluating the effect of LpL_{p} units on deep recurrent neural networks.

3 Distributions of the Orders of LpL_{p} Units

To understand how the estimated orders pp of the proposed LpL_{p} unit are distributed we trained MLPs with a single LpL_{p} layer on MNIST, TFD and Pentomino. We measured validation error to search for good hyperparameters, including the number of LpL_{p} units and number of filters (input signals) per LpL_{p} unit. However, for Pentomino, we simply fixed the size of the LpL_{p} layer to 400400, and each LpL_{p} unit received signals from six hidden units below.

In Table 1, the averages and standard deviations of the estimated orders of the LpL_{p} units in the single-layer MLPs are listed for MNIST, TFD and Pentomino. It is clear that the distribution of the orders depend heavily on the dataset, which confirms our first claim described earlier. From Fig. 5 we can clearly see that even in a single model the estimated orders vary quite a lot, which confirms our second claim. Interestingly, in the case of Pentomino, the distribution of the orders consists of two distinct modes.

The plots in Fig. 5 clearly show that the orders of the LpL_{p} units change significantly from their initial values over training. Although we initialized the orders of the LpL_{p} units around 33 for all the datasets, the resulting distributions of the orders are significantly different among those three datasets. This further confirms both of our claims. As a simple empirical confirmation we tried the same experiment with the fixed p=2p=2 on TFD and achieved a worse test error of 0.210.21.

4 Generalization Performance

The ultimate goal of any novel nonlinear activation function for an MLP is to achieve better generalization performance. We conjectured that by learning the orders of LpL_{p} units an MLP with LpL_{p} layers will achieve highly competitive classification performance.

For MNIST we trained an MLP having two LpL_{p} layers followed by a softmax output layer. We used a recently introduced regularization technique called dropout . With this MLP we were able to achieve 99.0399.03% accuracy on the test set, which is comparable to the state-of-the-art accuracy of 99.0699.06% obtained by the MLP with maxout units .

On TFD we used the same MLP from the previous experiment to evaluate generalization performance. We achieved a recognition rate of 79.2579.25%. Although we use neither pretraining nor unlabeled samples, our result is close to the current state-of-the-art rate of 82.482.4% on the permutation-invariant version of the task reported by who pretrained their models with a large amount of unlabeled samples.

As we have used the five-fold cross validation to find the optimal hyperparameters, we were able to use this to investigate the variance of the estimations of the pp values. Table 3 shows the averages and standard deviations of the estimated orders for MLPs trained on the five folds using the best hyperparameters. It is clear that in all the cases the orders ended up in a similar region near two without too much difference in the variance.

Similarly, we have trained five randomly initialized MLPs on MNIST and observed the similar phenomenon of all the resulting MLPs having similar distributions of the orders. The standard deviation of the averages of the learned orders was only 0.0280.028, while its mean is 2.162.16.

The MLP having a single LpL_{p} layer was able to classify the test samples of Pentomino with 31.3831.38% error rate. This is the best result reported so far on Pentomino dataset without using any kind of prior information about the task (the best previous result was 44.6% error).

On Forest Covertype an MLP having three LpL_{p} layers was trained. The MLP was able to classify the test samples with only 2.832.83% error. The improvement is large compared to the previous state-of-the-art rate of 3.133.13% achieved by the manifold tangent classifier having four hidden layers of logistic sigmoid units . The result obtained with the LpL_{p} is comparable to that obtained with the MLP having maxout units.

These results as well as previous best results for all datasets are summarized in Table 2.

In all experiments, we optimized hyperparameters such as an initial learning rate and its scheduling to minimize validation error, using random search , which is generally more efficient than grid search when the number of hyperparameters is not tiny. Each MLP was trained by stochastic gradient descent. All the experiments in this paper were done using the Pylearn2 library .

5 Deep Recurrent Neural Networks

We tried the polyphonic music prediction tasks with three music datasets; Nottingam, JSB and MuseData . The DOT-RNNs we trained had deep transition with LpL_{p} units and tanh⁡\tanh units and deep output function with maxout in the intermediate layer (see Fig. 6 for the illustration). We coarsely optimized the size of the models and the initial leaning rate as well as its schedule to maximize the performance on validation sets. Also, we chose whether to threshold the norm of the gradient based on the validation performance . All the models were trained with dropout .

As shown in Table 4, we were able to achieve the state-of-the-art results (RNN-only case) on all the three datasets. These results are much better than those achieved by the same DOT-RNNs using logistic sigmoid units in both deep transition and deep output, which suggests the superiority of the proposed LpL_{p} units over the conventional saturating activation functions. This suggests that the proposed LpL_{p} units are well suited not only to feedforward neural networks, but also to recurrent neural networks. However, we acknowledge that more investigation into applying LpL_{p} units is needed in the future to draw more concrete conclusion on the benefits of the LpL_{p} units in recurrent neural networks.

Conclusion

In this paper, we have proposed a novel nonlinear activation function based on the generalization of widely used pooling operators. The proposed nonlinear activation function computes the LpL_{p} norm of several projections of the lower layer. Max-, average- and root-of-mean-squared pooling operators are special cases of the proposed activation function, and naturally the recently proposed maxout unit is closely related under an assumption of non-negative input signals.

An important difference of the LpL_{p} unit from conventional pooling operators is that the order of the unit is learned rather than pre-defined. We claimed that this estimation of the orders is important and that the optimal model should have LpL_{p} units with various orders.

Our analysis has shown that an LpL_{p} unit defines a non-Euclidean subspace whose metric is defined by the LpL_{p} norm. When projected back into the input space, the LpL_{p} unit defines an ellipsoidal boundary. We conjectured and showed in a small scale experiment that the combination of these curved boundaries may more efficiently model separating curves of data with non-stationary curvature.

These claims were empirically verified via training both deep feedforward neural networks and deep recurrent neural networks. We tested the feedforward neural network on on four benchmark datasets; MNIST, Toronto Face Database, Pentomino and Forest Covertype, and tested the recurrent neural networks on the task of polyphonic music prediction. The experiments revealed that the distribution of the estimated orders of LpL_{p} units indeed depends highly on dataset and is far away from a Dirac delta distribution. Additionally, our conjecture that deep neural networks with LpL_{p} units will be able to achieve competitive generalization performance was empirically confirmed.

Acknowledgments

We would like to thank the developers of Pylearn2 and Theano . We would also like to thank CIFAR, and Canada Research Chairs for funding, and Compute Canada, and Calcul Québec for providing computational resources.

References