Deep SimNets

Nadav Cohen, Or Sharir, Amnon Shashua

Introduction

Deep neural networks, and convolutional neural networks (ConvNets) in particular, have had a dramatic impact in advancing the state of the art in computer vision, speech analysis, and many other domains (cf. ). It has been demonstrated time and time again, that when ConvNets are trained in an end-to-end manner, they deliver significantly better results than systems relying on manually engineered features.

The goal of this paper is to introduce a generalization of ConvNets we call Similarity Networks (SimNets), that preserves the simplicity and effectiveness of ConvNets, yet has a higher abstraction level. In a nutshell, the inner-product operator, which lies at the core of the ConvNet architecture, is replaced by an inner-product in “feature space”. The feature spaces are controlled by a family of kernel functions which include in particular the conventional (linear) inner-product as a special case.

We argue that the incentive for designing deep networks with a higher abstraction level than ConvNets, arises from the need for small networks that could fit into mobile platforms in terms of space and run-time. With small networks the approximation error becomes a limiting factor, which could be ameliorated through network architectures that are based on a higher level of abstraction.

The SimNet architecture is based on two operators. The first is analogous to, and generalizes, the inner-product operator of neural networks. The second, as special cases, plays the role of non-linear activation and pooling, but has additional capabilities that take SimNets far beyond ConvNets. In a detailed set of experiments, the SimNet architecture achieves state of the art accuracy using networks with complexity comparable to that of top performing ConvNets. However, when network complexity is limited, SimNets deliver a significant boost in accuracy.

Recently, the task of reducing run-time complexity of ConvNets is receiving increased attention. For example, a method named FitNets (), based on the knowledge distillation principle (), has been suggested in order to assist in compressing deep networks. In , a form of gating inspired by Long Short-Term Memory recurrent networks is introduced, allowing training of very deep and narrow networks. Another line of work considers imposing structural constrains on network weights, such as sparsity, in order to improve run-time efficiency (). Alternatively, network weights may be factorized using matrix or tensor decompositions, reducing storage and computational complexity, at the expense of marginal deterioration in accuracy (). All of these approaches consider ConvNets (or neural networks) as a baseline, and use supplementary techniques to reduce run-time complexity. In this work, we propose the alternative (generalized) SimNet architecture, and argue that it is inherently more efficient than ConvNets. The techniques listed here for reducing run-time complexity of ConvNets could just as well be applied to SimNets, thereby resulting in even more computationally efficient models.

The SimNet architecture

For the second SimNet operator we define MEX – a log-mean-exp function:

Note that unlike a conventional MLP unit which has a bias scalar, a MEX unit has a vector of biases. We may choose to omit part or all of the biases as part of a network design. For example, when all biases are dropped the MEX operator implements a soft trade-off between maximum and average.

SimNet MLP

Denoting by ψθ\psi_{\theta} a feature mapping associated with KθK_{\theta}, we get:

Deep SimNets for processing images

When used to classify images, the prediction rule associated with SimNet MLPConv is given by: y^(input)=argmax⁡rMEXβ2{MEXβ1{ul⊤ϕ(xij,zl)+br,l}l}i,j\hat{y}(input)=\operatorname*{argmax}_{r}MEX_{\beta_{2}}\left\{MEX_{\beta_{1}}\left\{{\mathbf{u}}_{l}^{\top}\phi({\mathbf{x}}_{ij},{\mathbf{z}}_{l})+b_{r,l}\right\}_{l}\right\}_{i,j}. Setting β1=β2=β\beta_{1}=\beta_{2}=\beta, and using the collapsing property of MEX, we get a “patch-based” version of SimNet MLP’s classification:

It can be shown () that all results put forth in sec. 3 for relating SimNet MLP to kernel machines apply to SimNet MLPConv as well, but with the underlying kernels being based on “patch-representations”. In other words, SimNet MLPConv – a “patch-based” extension of SimNet MLP, maintains all kernel relations of the latter, with a “patch-based” extension of the underlying kernels.

2 Whitening with convolutional layer

3 Going deep with SimNet MLPConv

Pre-training

The log probability density of a vector drawn from this distribution being equal to y{\mathbf{y}} and originating from component ll is: log⁡P(y∧comp. l)=−∑t=1dαl,t−β∣yt−μl,t∣β+cl\log P({\mathbf{y}}\land\text{comp.}~{}l)=-\sum_{t=1}^{d}\alpha_{l,t}^{-\beta}\lvert y_{t}-\mu_{l,t}\rvert^{\beta}+c_{l}, where cl:=log⁡{λl∏t=1dβ2αl,tΓ(1/β)}c_{l}:=\log\left\{\lambda_{l}\prod_{t=1}^{d}\frac{\beta}{2\alpha_{l,t}\Gamma(1/\beta)}\right\} is a constant that does not depend on y{\mathbf{y}}. This implies that if we model whitened patches yij{\mathbf{y}}_{ij} with a Generalized Gaussian mixture as above, initializing the similarity templates via zl,t=μl,tz_{l,t}=\mu_{l,t}, the weights via ul,t=αl,t−βu_{l,t}=\alpha_{l,t}^{-\beta} and the order via p=βp=\beta would give:

In words, similarity channel ll would hold, up to a constant, the probabilistic heat map of component ll and the whitened patches yij{\mathbf{y}}_{ij}. This observation suggests estimating the parameters of the mixture (shape β\beta, scales αl,t\alpha_{l,t} and means μl,t\mu_{l,t}) based on whitened patches (via EM, cf. ), and initializing the similarity parameters accordingly. We note in passing that it is possible to append additive biases blb_{l} to the similarity (through offsets of the succeeding MEX operator), in which case initializing these via bl=clb_{l}=c_{l} would make the probabilistic heat maps exact (not up to a constant).

Experiments

The datasets used in our experiments are CIFAR-10 and CIFAR-100 (), as well as SVHN (). These three datasets together form an image recognition benchmark that is diverse and challenging on one hand, yet simple enough to enable granular controlled experiments such as those needed to evaluate a new architecture. All datasets consist of 32x32 color images. SVHN (Street View House Numbers) represents a rather simple classification benchmark, where various methods are known to produce near-human accuracies. It contains approximately 600K images for training and 26K images for testing, partitioned into 10 categories that correspond to the digits 0 through 9. CIFAR-100 contains 50K images for training and 10K images for testing, equally partitioned into 100 categories. With a relatively large number of categories, and only a few hundred training examples per class, CIFAR-100 represents a challenging classification task. CIFAR-10 contains 50K images for training and 10K images for testing, equally partitioned into 10 categories. It brings forth a balanced trade-off between the simplicity of SVHN and the complexity of CIFAR-100, and accordingly served as the central dataset throughout our experiments. Namely, all cross-validations were carried out on CIFAR-10 (with 10K training images held out for validation), with SVHN and CIFAR-100 used for final evaluation only. In terms of implementation, we have integrated SimNets into Caffe toolbox (), with the aim of making our code publicly available in the near future.

In all our experiments, we trained both SimNets and ConvNets by minimizing softmax loss using SGD with Nesterov acceleration (). Batch size, momentum, weight decay and learning rate were chosen through cross-validation, though we observed, at least for the case of SimNets, that the following choices consistently produced good results: batch size 128, momentum 0.9, weight decay 0.0001 and learning rate 0.01 decreasing by a factor of 10 after 200 and 250 epochs (out of 300 total). Unlike ConvNets which are mostly initialized randomly nowadays (), SimNets are naturally pre-trained using statistical estimation methods (sec. 5). For computational efficiency, we implemented stochastic versions of these algorithms. Unless otherwise stated, all reported SimNet results were obtained using its pre-training scheme.

2 Single layer SimNet

3 Two layer SimNet

The networks were initially evaluated on CIFAR-10. Training hyper-parameters for the SimNet were configured via cross-validation, whereas for Caffe ConvNet we used the values that come built-in to Caffe. After measuring CIFAR-10 test accuracies, the same settings (network architectures and training hyper-parameters) were used to evaluate test accuracies on SVHN. For evaluation of test accuracies on CIFAR-100, we again used the exact same settings as in CIFAR-10, but this time increased the number of output channels in both networks from 10 to 100. The results of this experiment are summarized in table 1. As can be seen, the SimNet is roughly twice as efficient as Caffe ConvNet, yet achieves significantly higher accuracies on the more challenging benchmarks (CIFAR-10 and CIFAR-100). On SVHN accuracies are comparable, the reason being that in this simple benchmark classification error is dominated by overfit, to which the enhanced expressiveness of SimNets does not contribute.

4 Three layer SimNet

In the previous experiments we have seen that SimNets are more accurate than ConvNets when networks are constrained to be compact, i.e. when classification run-time is limited. In such a setting, the lower approximation error of SimNets plays an important role. In contrast, when networks are over-specified (i.e. are much larger than necessary in order to model the problem at hand) – standard practice for achieving state of the art accuracy, the approximation error is virtually zero, and the advantage of the SimNet architecture fades. Moreover, the additional expressive power of SimNets could actually be a burden, as additional regularization for controlling overfit would be required. It is therefore of interest to explore the ability of SimNets to reach state of the art accuracy with over-specified networks. This is the aim of our third and final experiment, carried out on CIFAR-10.

As a final sanity check, we compared extremely compact versions of our three layer SimNet and Network in Network (NiN, ) We chose to work against NiN since it bears an architectural resemblance to our SimNet, thus it was clear how both networks can be made compact in an analogous way.. Specifically, we changed the number of channels in all layers of both networks to 10, and removed dropout (NiN) and multiplicative Gaussian noise (SimNet), leaving all other hyper-parameters intact. The resulting networks had only 5K parameters each, and required just 3.5M FLOPs to classify an image. With such limited resources we expect the SimNet to benefit from its inherent expressiveness, and indeed, it outperformed NiN significantly, providing 76.8% accuracy compared to 72.3% reached by NiN.

Conclusion

We presented a deep layered architecture called SimNets that generalizes convolutional neural networks. The architecture is driven by two operators: (i) the similarity operator, which is a generalization of the inner-product operator on which ConvNets are based, and (ii) the MEX operator, that can realize non-linear activation and pooling, but has additional capabilities that make SimNets a powerful generalization of ConvNets. An interesting property of the SimNet architecture is that applying its two operators in succession – similarity followed by MEX, results in what can be viewed as an artificial neuron in a high-dimensional feature space (sec. 3). This also holds for the more elaborate image processing SimNet incorporating locality, sharing and pooling (sec. 4.1).

We argue that a higher abstraction level for the basic network building blocks carries with it the advantage of obtaining higher accuracies with small networks, an important trait for mobile and real-time applications. Through a detailed set of experiments we validated the conjecture of higher accuracy for small networks, and we have also shown that SimNets can achieve state of the art accuracy in large-scale settings where computational efficiency is not a concern (and thus the higher abstraction per given network size is not an advantage).

Finally, the SimNet architecture is endowed with a natural pre-training scheme based on unlabeled data. Besides its aid in training, the scheme also has the potential of determining the number of channels in hidden layers based on statistical analysis of patterns generated in previous layers. This implies that the structure of SimNets can potentially be determined automatically based on (unlabeled) training data. Future work includes a study of this capability, and more generally, further analysis of probabilistic properties of SimNets and unsupervised/supervised algorithms derived thereof.

We thank Ronen Tamari for his dedicated contribution to the experiments. The work is partly funded by Intel grant ICRI-CI 9-2012-6133 and ISF grant 1790/12. Nadav Cohen is supported by a Google Fellowship in Machine Learning.

References

References