Greedy Layerwise Learning Can Scale to ImageNet
Eugene Belilovsky, Michael Eickenberg, Edouard Oyallon
Introduction
Deep Convolutional Neural Networks (CNNs) trained on large-scale supervised data via the back-propagation algorithm have become the dominant approach in most computer vision tasks (Krizhevsky et al., 2012). This has motivated successful applications of deep learning in other fields such as speech recognition (Chan et al., 2016), natural language processing (Vaswani et al., 2017), and reinforcement learning (Silver et al., 2017). Training procedures and architecture choices for deep CNNs have become more and more entrenched, but which of the standard components of modern pipelines are essential to the success of deep CNNs is not clear. Here we ask: do CNN layers need to be learned jointly to obtain high performance? We will show that even for the challenging ImageNet dataset the answer is no.
Supervised end-to-end learning is the standard approach to neural network optimization. However it has potential issues that can be valuable to consider. First, the use of a global objective means that the final functional behavior of individual intermediate layers of a deep network is only indirectly specified: it is unclear how the layers work together to achieve high-accuracy predictions. Several authors have suggested and shown empirically that CNNs learn to implement mechanisms that progressively induce invariance to complex, but irrelevant variability (Mallat, 2016; Yosinski et al., 2015) while increasing linear separability (Zeiler & Fergus, 2014; Oyallon, 2017; Jacobsen et al., 2018) of the data. Progressive linear separability has been shown empirically but it is unclear whether this is merely the consequence of other strategies implemented by CNNs, or if it is a sufficient condition for the high performance of these networks. Secondly, understanding the link between shallow Neural Networks (NNs) and deep NNs is difficult: while generalization, approximation, or optimization results (Barron, 1994; Bach, 2014; Venturi et al., 2018; Neyshabur et al., 2018; Pinkus, 1999) for 1-hidden layer NNs are available, the same studies conclude that multiple-hidden-layer NNs are much more difficult to tackle theoretically. Finally, end-to-end back-propagation can be inefficient (Jaderberg et al., 2016; Salimans et al., 2017) in terms of computation and memory resources and is considered not biologically plausible.
Sequential learning of CNN layers by solving shallow supervised learning problems is an alternative to end-to-end back-propagation. This classic (Ivakhnenko & Lapa, 1965) learning principle can directly specify the objective of every layer. It can encourage the refinement of specific properties of the representation (Greff et al., 2016), such as progressive linear separability. The development of theoretical tools for deep greedy methods can naturally draw from the theoretical understanding of shallow sub-problems. Indeed, (Arora et al., 2018; Bengio et al., 2006; Bach, 2014; Janzamin et al., 2015) show global optimal approximations, while other works have shown that networks based on sequential 1-hidden layer training can have a variety of guarantees under certain assumptions (Huang et al., 2017; Malach & Shalev-Shwartz, 2018; Arora et al., 2014): greedy layerwise methods could permit to cascade those results to bigger architectures. Finally, a greedy approach will rely much less on having access to a full gradient. This can have a number of benefits. From an algorithmic perspective, they do not require storing most of the intermediate activations nor to compute most intermediate gradients. This can be beneficial in memory-constrained settings. Unfortunately, prior work has not convincingly demonstrated that layer-wise training strategies can tackle the sort of large-scale problems that have brought deep learning into the spotlight.
Recently multiple works have demonstrated interest in determining whether alternative training methods (Xiao et al., 2019; Bartunov et al., 2018) can scale to large data-sets that have only been solved by deep learning. As is the case for many algorithms (not just training strategies) many of these alternative training strategies have been shown to work on smaller datasets (e.g. MNIST and CIFAR) but fail completely on large-scale datasets (e.g. ImageNet). Also, these works largely focus on avoiding the weight transport problem in backpropagation (Bartunov et al., 2018) while simple greedy layer-wise learning reduces the extent of this problem and should be considered as a potential baseline.
In this context, our contributions are as follows. (a) First, we design a simple and scalable supervised approach to learn layer-wise CNNs in Sec. 3. (b) Then, Sec. 4.1 demonstrates empirically that by sequentially solving 1-hidden layer problems, we can match the performance of the AlexNet on ImageNet. We motivate in Sec. 3.3 how this model can be connected to a body of theoretical work that tackles 1-hidden layer networks and their sequentially trained counterparts. (c) We show that layerwise trained layers exhibit a progressive linear separability property in Sec. 4.2. (d) In particular, we use this to help motivate learning layer-wise CNN layers via shallow -hidden layer auxiliary problems, with . Using this approach our sequentially trained -hidden layer models can reach the performance level of VGG models (Sec. 4.3) and end-to-end learning. (e) Finally, we suggest an approach to easily reduce the model size during training of these networks.
Related Work
Several authors have previously studied layerwise learning. In this section we review related works and re-emphasize the distinctions from our work.
Greedy unsupervised learning has been a popular topic of research in the past. Greedy unsupervised learning of deep generative models (Bengio et al., 2007; Hinton et al., 2006) was shown to be effective as an initialization for deep supervised architectures. Bengio et al. (2007) also considered supervised greedy layerwise learning as initialization of networks for subsequent end-to-end supervised learning, but this was not shown to be effective with the existing techniques at the time. Later work on large-scale supervised deep learning showed that modern training techniques permit avoiding layerwise initialization entirely (Krizhevsky et al., 2012). We emphasize that the supervised layerwise learning we consider is distinct from unsupervised layerwise learning. Moreover, here layerwise training is not studied as a pretraining strategy, but a training one.
Layerwise learning in the context of constructing supervised NNs has been attempted in several works. It was considered in multiple earlier works Ivakhnenko & Lapa (1965); Fahlman & Lebiere (1990b); Lengellé & Denoeux (1996) on very simple problems and in a climate where deep learning was not a dominant supervised learning approach. These works were aimed primarily at structure learning, building up architectures that allow the model to grow appropriately based on the data. Others works were motivated by the avoidance of difficulties with vanishing gradients. Similarly, Cortes et al. (2016) recently proposed a progressive learning method that builds a network such that the architecture can adapt to the problem, with theoretical contributions to structure learning, but not on problems where deep networks are unmatched in performance. Malach & Shalev-Shwartz (2018) also train a supervised network in a layerwise fashion, showing that their method provably generalizes for a restricted class of image models. However, the results of these model are not shown to be competitive with handcrafted approaches (Oyallon & Mallat, 2015). Similarly (Kulkarni & Karande, 2017; Marquez et al., 2018) revisit layerwise training, but in a limited experimental setting.
Huang et al. (2017) combined boosting theory with a residual architecture (He et al., 2016) to sequentially train layers. However, results are presented for limited datasets and indicate that the end-to-end approach is often needed ultimately to obtain competitive results. This proposed strategy does not clearly outperform simple non-deep-learning baselines. By contrast, we focus on settings where deep CNN based-approaches do not currently have competitors and rely on a simpler objective function, which is found to scale well and be competitive with end-to-end approaches.
Another related thread is methods which add layers to existing networks and then use end-to-end learning. These approaches usually have different goals from ours, such as stabilizing end-to-end learned models. Brock et al. (2017) builds a network in stages, where certain layers are progressively frozen, permitting faster training. Mosca & Magoulas (2017); Wang et al. (2017) propose methods that progressively stack layers, performing end-to-end learning on the resulting network at each step. A similar strategy was applied for training GANs in Karras et al. (2017). By the nature of our goals in this work, we never perform fine-tuning of the whole network. Several methods also consider auxiliary supervised objectives (Lee et al., 2015) to stabilize end-to-end learning, but this is different from the case where these objectives are not solved jointly.
Supervised Layerwise Training of CNNs
In this section we formalize the architecture, training algorithm, and the necessary notations and terminology. We focus on CNNs, with ReLU non-linearity denoted by . Sec. 3.1 describes a layerwise training scheme using a succession of auxiliary learning tasks. We add one layer at a time: the first layer of a -hidden layer CNN problem. Finally, we discuss the distinctions in varying .
Our architecture has blocks (see Fig. 1), which are trained in succession. From an input signal , an initial representation is propagated through convolutions, giving . Each feeds into an auxiliary classifier to obtain prediction , which computes an intermediate classification output. At depth , denote by a convolutional operator with parameters , an auxiliary classifier with all its parameters denoted , and a down-sampling operator. The parameters correspond to kernels with bias terms. Formally, from layer we iterate as follows:
where is the number of classes. For the pooling operator we choose the invertible downsampling operation described in Dinh et al. (2017), which consists in reorganizing the initial spatial channels into the 4 spatially decimated copies obtainable by spatial sub-sampling, reducing the resolution by a factor . We decided against strided pooling, average pooling, and the non-linear max-pooling, because these strongly encourage loss of information. As is standard practice, is applied at certain layers (), but not others (). The CNN classifier is given by:
2 Training by Auxiliary Problems
Our training procedure is layerwise: at depth , while keeping all other parameters fixed, is obtained via an auxiliary problem: optimizing to obtain the best training accuracy for auxiliary classifier . We formalize this idea for a training set : For a function parametrized by and a loss (e.g. cross entropy), we consider the classical minimization of the empirical risk: .
At depth , assume we have constructed the parameters . Our algorithm can produce samples . Taking , we employ an optimization procedure that aims to minimize the risk . This procedure (Alg. 1) consists in training (e.g. using SGD) the shallow CNN classifier on top of , to obtain the new parameter . Under mild conditions, it improves the training error at each layer as shown below:
A technical requirement for the actual optimization procedure is to not produce a worse objective than the initialization, which can be achieved by taking the best result along the optimization trajectory.
The cascade can inherit from the individual properties of each auxiliary problem. For instance, as is 1-Lipschitz, if each is 1-Lipschitz then so is w.r.t. . Another example is the nested objective defined by Alg. 1: the optimality of the solution will be largely governed by the optimality of the sub-problem solver. Specifically, if the auxiliary problem solution is close to optimal then the solution of Alg. 1 will be close to optimal.
with, and with , then, we prove by induction:
The proof can be found in the Appendix A. This demonstrates an example of how the training strategy can permit to extend results from shallow CNNs to deeper CNNs, in particular for . Applying an existing optimization strategy could give us a bound on the solution of the overall objective of Alg. 1, as we will discuss below.
3 Auxiliary Problems & The Properties They Induce
We now discuss the properties arising from the auxiliary problems. We start with , for which the auxiliary classifier consists of only the linear and operators. Thus, the optimization aims to obtain the weights of a 1-hidden layer NN. For this case, as discussed in Sec. 1, a variety of theoretical results exist (e.g. (Cybenko, 1989; Barron, 1994)). Moreover, (Arora et al., 2018; Ge et al., 2017; Du & Goel, 2018; Bach, 2014) proposed provable optimization strategies for this case. Thus the analysis and optimization of the 1-hidden layer problem is a case that is relatively well understood compared to deep counterparts. At the same time, as shown in Prop. 3.2, applying an existing optimization strategy could give us a bound on the solution of the overall objective of Alg. 1. To build intuition let us consider another example where the analysis can be simplified for training. Recall the classic least square estimator (Barron, 1994) of a 1-hidden layer network:
where is the function of interest. Following a suggestion from (Mallat, 2016) (detailed in Appendix A) we can state there exists a set and , where are the parameters of greedily trained layers (with width ) and is sigmoidal. For simplicity, let us assume that . It implies that at step of a greedy training procedure with , the corresponding sample loss is:
In particular, the right term is shaped as Eq. (4) and thus we can apply standard bounds available only for 1-hidden layer settings (Barron, 1994; Janzamin et al., 2015). In the case of jointly learning the layers we could not make this kind of formulation. For example if one now applied the algorithm of (Janzamin et al., 2015) and their Theorem 5 it would give the following risk bound:
where is obtained from an initial bounded distribution , is the Barron Constant (as described in (Lee et al., 2017)) of , a constant depending on and an estimation error. Furthermore, if we fit the layerwise using a method such as (Janzamin et al., 2015), we reach an approximate optimum for each layer given the state of the previous layer. Under Prop 3.2, we observe that small errors at each layer, even taken cumulatively, will not affect the final representation learned by the cascade. Observe that if decreases with , the approximation bound on will correspondingly improve.
Another view of the training using a standard classification objective: the optimization of the 1-hidden layer network will encourage the hidden layer outputs to make its classification output maximally linearly separable with respect to its inputs. Specializing Prop. 3.1 for this case shows that the layerwise procedure will try to progressively improve the linear separability. Progressive linear separation has been empirically studied in end-to-end CNNs (Zeiler & Fergus, 2014; Oyallon, 2017) as an indirect consequence, while the training permits us to study this basic principle more directly as the layer objective. Concurrent work (Elad et al., 2019) follows the same argument to use a layerwise training procedure to evaluate mutual information more directly.
Unique to our layerwise learning formulation, we consider the case where the auxiliary learning problem involves auxiliary hidden layers. We will interpret, and empirically verify, in Sec. 4.2 that this builds layers that are progressively better inputs to shallow CNNs. We will also show a link to building, in a more progressive manner, linearly separable layers. Considering only shallow (with respect to total depth) auxiliary problems (e.g. in our work) we can maintain several advantages. Indeed, optimization for shallower networks is generally easier, as we can for example diminish the vanishing gradient problem, reducing the need for identity loops or normalization techniques (He et al., 2016). Two and three hidden layer networks are also appealing for extending results from one hidden layer (Allen-Zhu et al., 2018) as they are the next natural member in the family of NNs.
Experiments and Discussion
We performed experiments on the large-scale ImageNet-1k (Russakovsky et al., 2015), a major catalyst for the recent popularity of deep learning, as well as the CIFAR-10 dataset. We study the classification performance of layerwise models with , comparing them to standard benchmarks and other sequential learning methods. Then we inspect the representations built through our auxiliary tasks and motivate the use of models learned with auxiliary hidden layers , which we subsequently evaluate at scale.
We highlight that many algorithms do not scale to large datasets (Bartunov et al., 2018) like ImageNet. For example a SOTA hand-crafted image descriptor combined with a linear model achieves on CIFAR-10 while only on ImageNet (Oyallon et al., 2018). 1-hidden layer CNNs from our experiments can obtain accuracy on CIFAR-10 while the same CNN results on ImageNet only gives accuracy. Alternative learning methods for deep networks can sometimes show similar behavior, highlighting the importance of assessing their scalability. For example Feedback Alignment (Lillicrap et al., 2016) which is able to achieve on CIFAR-10 compared to with backprop, on the 1000 class ImageNet obtains only accuracy compared to (Bartunov et al., 2018; Xiao et al., 2019). Thus, based on these observations and the sparse and small-scale experimental efforts of related works on greedy layerwise learning it is entirely unclear whether this family of approaches can work on large datasets like ImageNet and be used to construct useful models. Noting that ImageNet models not only represent benchmarks but have generic features (Yosinski et al., 2014) and thus AlexNet- and VGG-like accuracies for CNN’s on this dataset typically indicate representations that are generic enough to be useful for downstream applications.
We briefly introduce the datasets and preprocessing. CIFAR-10 consists of small RGB images with respectively and samples for training and testing. We use the standard data augmentation and optimize each layer with SGD using a momentum of 0.9 and a batch-size of 128. The initial learning rate is and we use the reduced schedule with decays of every epochs (Zagoruyko & Komodakis, 2016), for a total of 50 epochs in each layer. ImageNet consists of RGB images of varying size for training. Our data augmentation consists of random crops of size . At testing time, the image is rescaled to then cropped at size . We used SGD with momentum 0.9 for a batch size of 256. The initial learning rate is (He et al., 2016) and we use the reduced schedule with decays of every 20 epochs for 45 epochs. We use 4 GPUs to train our ImageNet models.
1 AlexNet Accuracy with 1-Hidden Layer Auxiliary Problems
We consider the atomic, layerwise CNN with which corresponds to solving a sequence of 1-hidden layer CNN problems. As discussed in Sec. 2, previous attempts at supervised layerwise training (Fahlman & Lebiere, 1990a; Arora et al., 2014; Huang et al., 2017; Malach & Shalev-Shwartz, 2018), which rely solely on sequential solving of shallow problems have yielded performance well below that of typical deep learning models even on the CIFAR-10 dataset. We show, surprisingly, that it is possible to go beyond the AlexNet performance barrier (Krizhevsky et al., 2012) without end-to-end backpropagation on ImageNet with elementary auxiliary problems. To emphasize the stability of the training process we do not apply any batch-normalization to this model.
We trained a model with layers, down-sampling at layers , and layer sizes starting at . We obtain and note that this accuracy is close to the AlexNet model performance (Krizhevsky et al., 2012) for CIFAR-10 (89.0%). Comparisons in Table 2 show that other sequentially trained 1-hidden layer networks have yielded performances that do not exceed those of the top hand-crafted methods or those using unsupervised learning. To the best of our knowledge they obtain accuracy (Huang et al., 2017). The end-to-end version of this model obtains , while using more GPU memory than training.
ImageNet.
Our model is trained with layers and downsampling operations at layers . Layer sizes start at . Our final trained model achieves 79.7% top-5 single crop accuracy on the validation set and 80.8% a weighted ensemble of the layer outputs. In addition to exceeding AlexNet, this model compares favorably to alternatives to end-to-end supervised CNNs including hand-crafted computer vision techniques (Sánchez et al., 2013). Full results are shown in Table 3 where we also highlight the substantial performance gaps to multiple alternative training methods. Figure 2 shows the per-layer classification accuracy. We observe a remarkably rapid progression (near linear in the first 5 layers) that takes the accuracy from to top-5. We leave a full theoretical explanation of this fast rate as an open question. Finally, in Appendix C.3 we demonstrate that this model maintains the transfer learning properties of deep networks trained on ImageNet.
We also note that our final training accuracy is relatively high for ImageNet (87%), which indicates that appropriate regularization may lead to a further improvement in test accuracy. We now look at empirical properties induced in the layers and subsequently evaluate the distinct .
2 Empirical Separability Properties
We study the intermediate representations generated by the layerwise learning procedure in terms of linear separability as well as separability by a more general set of classifiers. Our aims are (a) to determine empirically whether indeed progressively builds more and more linearly separable data representations and (b) to determine how linear separability of the representations evolves for networks constructed with auxiliary problems. Finally we ask whether the notion of building progressively better inputs to a linear model ( training) has an analogous counterpart for : building progressively better inputs for shallow CNNs (discussed in Sec 3.3).
We define linear separability of a representation as the maximum accuracy achievable by a linear classifier. Further we define the notion of CNN-p-separability as the accuracy achieved by a -layer CNN trained on top of the representation to be assessed.
As expected from Sec. 3.3, linear separability monotonically increases with depth for . Interestingly, linear separability also improves in the case of , even though it is not directly specified by the auxiliary problem objective. At earlier layers, linear separation capability of models trained with increases fastest as a function of layer depth compared to models trained with deeper auxiliary networks, but flattens out to a lower asymptotic linear separability at deeper layers. Thus, the simple principle of the objective that tries to produce the maximal linear separation at each layer might not be an optimal strategy for achieving progressive linear separation.
We also notice that the deeper the auxiliary classifier, the slower the increase in linear separability initially, but the higher the linear separability at deeper layers. From the two right diagrams we also find that the CNN--separability progressively improves - but much more so for trained networks. This shows that linear separability of a layer is not the sole criterion for rendering a representation a good "input" for a CNN. It further shows that our sequential training procedure for the case can build a representation that is progressively a better input to a shallow CNN.
3 Scaling up Layerwise CNNs with 2 and 3 Hidden Layer Auxiliary Problems
We now compare this training method to end-to-end learning directly on a common model, VGG-11. We use training for all but the final convolutional layer. We additionally reduce the spatial resolution by before applying the auxiliary convolutions. As in the last experiment we use a different final auxiliary network so that the model architecture matches VGG-11. The last convolutional layers’ auxiliary model simply becomes the final max-pooling and 2-hidden layer fully connected network of VGG. We train a baseline VGG-11 model using the same 45 epoch training schedule. The performance of the training strategy matches that of the end-to-end training, with the ensemble model being better. The main differences of this model to SimCNN, besides the final auxiliary, is the max-pooling and the input image starting from full size.
Reference models relying on residual connections and very deep networks have better performance than those considered here. We believe that one can extend layer-wise learning to these modern techniques. However, this is outside the scope of this work. Moreover, recent ImageNet models (after VGG) are developed in industry settings, with large-scale infrastructure available for architecture and hyper-parameter search. Better design of sub-problem optimization geared for this setting may further improve results.
We emphasize that this approach enables the training of larger layer-wise models than end-to-end ones on the same hardware. This suggests applications in fields with large models (e.g. 3-D vision and medical imaging). We also observed that using outputs of early layers that were not yet converged still permitted improvement in subsequent layers. This suggests that our work might allow an extension that solves the auxiliary problems in parallel to a certain degree.
Layerwise Model Compression
Conclusion
We have shown that an alternative to end-to-end learning of CNNs relying on simpler sub-problems and no feedback between layers can scale to large-scale benchmarks such as ImageNet and can be competitive with standard CNN baselines. We build these competitive models by training only shallow CNNs and using standard architectural elements (ReLU, convolution). Layer-wise training opens the door to applications such as larger models under memory constraints, model prototyping, joint model compression and training, parallelized training, and more stable training for challenging scenarios. Importantly, our results suggest a number of open questions regarding the mechanisms that underlie the success of CNNs and provide a major simplification for theoretical research aiming at analyzing high performance deep learning models. Future work can study whether the 1-hidden layer cascades objective can be better specified to more closely mimic the objective.
References
Appendix A Proof of Proposition
Assume that . Then there exists such that:
As , we simply have to chose such that . ∎
We will now show that given an optimization -optimal procedure for the sub-problem optimization, the optimization of Algo 1. can be used directly to obtain an error on the overall. solution. Denote the parameters the optimal solutions for Algo 1.
with, and with , then, we prove by induction:
First observe that by non expansivity. Thus, by induction, . Then, let us show that: by induction. Indeed, for :
As , the property is true for .
We first briefly discuss the result of Sec.2 of (Mallat, 2016). Let us introduce: . We introduce: . For , we define . Observe this defines indeed a function, and that:
Observe also that . The set is simply the set of samples which are well discriminated by the neural network .
Appendix B Additional Details on Imagenet Models and Performance
For ImageNet we report the improvement in accuracy obtained by adding layers in Figure 4 as seen by the auxiliary problem solutions. We observe that indeed the accuracy of the model on both the training and validation is able to improve from adding layers as discussed in depth in Section 4.2. We observe that also over-fits substantially, suggesting better regularization can help in this setting.
We provide a more explicit view of the network sizes in Tab. 4 and Tab. 5. We also show the number of parameters in the ImageNet networks in Tab. 7. Although some of the models are not as parameter efficient compared to the related ones in the literature, this was not a primary aim of the investigation in our experiments and thus we did not optimize the models for parameter efficiency (except explicitly at the end of Sec. 4.3), choosing our construction scheme for simplicity. We highlight that this is not a fundamental problem in two ways: (a) for the model we note that removing the last two layers reduces the size by , while the top 5 accuracy at the earlier J=6 layer is (versus ), see Figure 4 for detailed accuracies. (b) Our models for have most of their parameters in the final auxiliary network which is easy to correct for once care is applied to this specific point as at the end of Sec. 4.3. We note also that the model with , is actually more parameter efficient than those in the VGG family while having similar performance. We also point out that we use for simplicity the VGG style construction involving only convolutions and downsampling operations that only half the spatial resolution, which indeed has been shown to lead to relatively less parameter efficient architectures (He et al., 2016), using less uniform construction (larger filters and bigger pooling early on) can yield more parameter efficient models.
Appendix C Additional Studies
We report additional studies that elucidate the critical components of the system and demonstrate the transferability properties of the greedily learned features.
In our experiments we use primarily the invertible downsampling operator as the downsampling operation. This choice is to reduce architectural elements which may be inherently lossy such as average pooling. Compared to maxpooling operations it also helps to maintain the network in Sec. 4.1 as a pure ReLU network, which may aid in analysis as the maxpooling introduces an additional non-linearity. We show here the effects of using alternative downsampling approaches including: average pooling, maxpooling and strided convolution. On the CIFAR dataset in the setting of we find that they ultimately lead to very similar results with invertible downsampling being slightly better. This shows the method is rather general. In our experiments we follow the same setting described for CIFAR. The setting here uses and downsamplings at . The size is always halved in all cases and the downsampling operation and the output sizes of all networks are the same. Specifically the Average Pooling and Max Pooling use kernels and the strided convolution simply modifies the convolutions in use to have a stride of . Results are shown in Tab. 6.
C.2 Effect of Width
We report here an additional view of the aggregated results for linear separability discussed in Sec. 4.2. We observe that the trend of the aggregated diagram is similar when comparing only same sized models, with the primary differences in model sizes being increased accuracy.
C.3 Transfer Learning on Caltech-101
Deep CNNs such as AlexNet trained on Imagenet are well known to have generic properties for computer vision tasks, permitting transfer learning on many downstream applications. We briefly evaluate here if the imagenet model (Sec. 4.1) shares this generality on the Caltech-101 dataset. This dataset has 101 classes and we follow the same standard experimental protocol as (Zeiler & Fergus, 2014): 30 images per class are randomly selected, and the rest is used for testing. The average per class accuracy is reported using 10 random splits. As in (Zeiler & Fergus, 2014) we restrict ourselves to a linear model. We use a multinomial logistic regression applied on features from different layers including the final one. For the logistic regression we rely on the default hyperparameter settings for logistic regression of the sklearn package using the SAGA algorithm. We apply a linear averaging and PCA transform (for each fold) to reduce the dimensionality to in all cases. We find the results are similar to those reported in (Zeiler & Fergus, 2014) for their version of the AlexNet. This highlights the model has similar transfer properties and also shows similar progressive linear separability properties as end-to-end trained models.