Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
Thao Nguyen, Maithra Raghu, Simon Kornblith
Introduction
Deep neural network architectures are typically tailored to available computational resources by scaling their width and/or depth. Remarkably, this simple approach to model scaling can result in state-of-the-art networks for both high- and low-resource regimes Tan & Le (2019). However, despite the ubiquity of varying depth and width, there is limited understanding of how varying these properties affects the final model beyond its performance. Investigating this fundamental question is critical, especially with the continually increasing compute resources devoted to designing and training new network architectures.
More concretely, we can ask, how do depth and width affect the final learned representations? Do these different model architectures also learn different intermediate (hidden layer) features? Are there discernible differences in the outputs? In this paper, we study these core questions, through detailed analysis of a family of ResNet models with varying depths and widths trained on CIFAR-10 Krizhevsky et al. (2009), CIFAR-100 and ImageNet Deng et al. (2009).
We show that depth/width variations result in distinctive characteristics in the model internal representations, with resulting consequences for representations and outputs across different model initializations and architectures. Specifically, our contributions are as follows:
We develop a method based on centered kernel alignment (CKA) to efficiently measure the similarity of the hidden representations of wide and deep neural networks using minibatches.
We apply this method to different network architectures, finding that representations in wide or deep models exhibit a characteristic structure, which we term the block structure. We study how the block structure varies across different training runs, and uncover a connection between block structure and model overparametrization — block structure primarily appears in overparameterized models.
Through further analysis, we find that the block structure corresponds to hidden representations having a single principal component that explains the majority of the variance in the representation, which is preserved and propagated through the corresponding layers. We show that some hidden layers exhibiting the block structure can be pruned with minimal impact on performance.
With this insight on the representational structures within a single network, we turn to comparing representations across different architectures, finding that models without the block structure show reasonable representation similarity in corresponding layers, but block structure representations are unique to each model.
Finally, we look at how different depths and widths affect model outputs. We find that wide and deep models make systematically different mistakes at the level of individual examples. Specifically, on ImageNet, even when these networks achieve similar overall accuracy, wide networks perform slightly better on classes reflecting scenes, whereas deep networks are slightly more accurate on consumer goods.
Related Work
Neural network models of different depth and width have been studied through the lens of universal approximation theorems Cybenko (1989); Hornik (1991); Pinkus (1999); Lu et al. (2017); Hanin & Sellke (2017); Lin & Jegelka (2018) and functional expressivity Telgarsky (2015); Raghu et al. (2017b). However, this line of work only shows that such networks can be constructed, and provides neither a guarantee of learnability nor a characterization of their performance when trained on finite datasets. Other work has studied the behavior of neural networks in the infinite width limit by relating architectures to their corresponding kernels Matthews et al. (2018); Lee et al. (2018); Jacot et al. (2018), but substantial differences exist between behavior in this infinite width limit and the behavior of finite-width networks Novak et al. (2018); Wei et al. (2019); Chizat et al. (2019); Lewkowycz et al. (2020). In contrast to this theoretical work, we attempt to develop empirical understanding of the behavior of practical, finite-width neural network architectures after training on real-world data.
Previous empirical work has studied the effects of width and depth upon model accuracy in the context of convolutional neural network architecture design, finding that optimal accuracy is typically achieved by balancing width and depth Zagoruyko & Komodakis (2016); Tan & Le (2019). Further study of accuracy and error sets have been conducted in Hacohen & Weinshall (2020) (error sets over training), and Hooker et al. (2019) (error after pruning). Other work has demonstrated that it is often possible for narrower or shallower neural networks to attain similar accuracy to larger networks when the smaller networks are trained to mimic the larger networks’ predictions Ba & Caruana (2014); Romero et al. (2015). We instead seek to study the impact of width and depth on network internal representations and (per-example) outputs, by applying techniques for measuring similarity of neural network hidden representations Kornblith et al. (2019); Raghu et al. (2017a); Morcos et al. (2018). These techniques have been very successful in analyzing deep learning, from properties of neural network training Gotmare et al. (2018); Neyshabur et al. (2020), objectives Resnick et al. (2019); Thompson et al. (2019); Hermann & Lampinen (2020), and dynamics Maheswaranathan et al. (2019) to revealing hidden linguistic structure in large language models Bau et al. (2019); Kudugunta et al. (2019); Wu et al. (2019; 2020) and applications in neuroscience Shi et al. (2019); Li et al. (2019); Merel et al. (2019); Zhang & Bellec (2020) and medicine Raghu et al. (2019).
Experimental Setup and Background
Our goal is to understand the effects of depth and width on the function learned by the underlying neural network, in a setting representative of the high performance models used in practice. Reflecting this, our experimental setup consists of a family of ResNets He et al. (2016); Zagoruyko & Komodakis (2016) trained on standard image classification datasets CIFAR-10, CIFAR-100 and ImageNet.
For standard CIFAR ResNet architectures, the network’s layers are evenly divided between three stages (feature map sizes), with numbers of channels increasing by a factor of two from one stage to the next. We adjust the network’s width and depth by increasing the number of channels and layers respectively in each stage, following Zagoruyko & Komodakis (2016). For ImageNet ResNets, ResNet-50 and ResNet-101 architectures differ only by the number of layers in the third () stage. Thus, for experiments on ImageNet, we scale only the width or depth of layers in this stage. More details on training parameters, as well as the accuracies of all investigated models, can be found in Appendix B.
We observe that increasing depth and/or width indeed yields better-performing models. However, we will show in the following sections how they exhibit characteristic differences in internal representations and outputs, beyond their comparable accuracies.
Neural network hidden representations are challenging to analyze for several reasons including (i) their large size; (ii) their distributed nature, where important features in a layer may rely on multiple neurons; and (iii) lack of alignment between neurons in different layers. Centered kernel alignment (CKA) Kornblith et al. (2019); Cortes et al. (2012) addresses these challenges, providing a robust way to quantitatively study neural network representations by computing the similarity between pairs of activation matrices. Specifically, we use linear CKA, which Kornblith et al. (2019) have previously validated for this purpose, and adapt it so that it can be efficiently estimated using minibatches. We describe both the conventional and minibatch estimators of CKA below.
Kornblith et al. (2019) show that, when measured between layers of architecturally identical networks trained from different random initializations, linear CKA reliably identifies architecturally corresponding layers, whereas several other proposed representational similarity measures do not. However, naive computation of linear CKA requires maintaining the activations across the entire dataset in memory, which is challenging for wide and deep networks. To reduce memory consumption, we propose to compute linear CKA by averaging HSIC scores over minibatches:
This approach of estimating HSIC based on minibatches is equivalent to the bagging block HSIC approach of Yamada et al. (2018), and converges to the same value as if the entire dataset were considered as a single minibatch, as proven in Appendix A. We use minibatches of size obtained by iterating over the test dataset 10 times, sampling without replacement within each time.
Depth, Width and Model Internal Representations
We begin our study by investigating how the depth and width of a model architecture affects its internal representation structure. How do representations evolve through the hidden layers in different architectures? How similar are different hidden layer representations to each other? To answer these questions, we use the CKA representation similarity measure outlined in Section 3.1.
We find that as networks become wider and/or deeper, their representations show a characteristic block structure: many (almost) consecutive hidden layers that have highly similar representations. By training with reduced dataset size, we pinpoint a connection between block structure and model overparametrization — block structure emerges in models that have large capacity relative to the training dataset.
In Figure 1, we show the results of training ResNets of varying depths (top row) and widths (bottom row) on CIFAR-10. For each ResNet, we use CKA to compute the representation similarity of all pairs of layers within the same model. Note that the total number of layers is much greater than the stated depth of the ResNet, as the latter only accounts for the convolutional layers in the network but we include all intermediate representations. We can visualize the result as a heatmap, with the x and y axes representing the layers of the network, going from the input layer to the output layer.
The heatmaps start off as showing a checkerboard-like representation similarity structure, which arises because representations after residual connections are more similar to other post-residual representations than representations inside ResNet blocks. As the model gets wider or deeper, we see the emergence of a distinctive block structure — a considerable range of hidden layers that have very high representation similarity (seen as a yellow square on the heatmap). This block structure mostly appears in the later layers (the last two stages) of the network. We observe similar results in networks without residual connections (Appendix Figure C.1).
Block structure across random seeds: In Appendix Figure D.1, we plot CKA heatmaps across multiple random seeds of a deep network and a wide network. We observe that while the exact size and position of the block structure can vary, it is present across all training runs.
2 The Block Structure and Model Overparametrization
Having observed that the block structure emerges as models get deeper and/or wider (Figure 1), we next study whether block structure is a result of this increase in model capacity — namely, is block structure connected to the absolute model size, or to the size of the model relative to the size of the training data?
Commonly used neural networks have many more parameters than there are examples in their training sets. However, even within this overparameterized regime, larger networks frequently achieve higher performance on held out data Zagoruyko & Komodakis (2016); Tan & Le (2019). Thus, to explore the connection between relative model capacity and the block structure, we fix a model architecture, but decrease the training dataset size, which serves to inflate the relative model capacity.
The results of this experiment with varying network widths are shown in Figure 2, while the corresponding plot with varying network depths (which supports the same conclusions) can be found in Appendix Figure D.2. Each column of Figure 2 shows the internal representation structure of a fixed architecture as the amount of training data is reduced, and we can clearly see the emergence of the block structure in narrower (lower capacity) networks as less training data is used. Refer to Figures D.3 and D.4 in the Appendix for a similar set of experiments on CIFAR-100. Together, these observations indicate that block structure in the internal representations arises in models that are heavily overparameterized relative to the training dataset.
Probing the Block Structure
In the previous section, we show that wide and/or deep neural networks exhibit a block structure in the CKA heatmaps of their internal representations, and that this block structure arises from the large capacity of the models in relation to the learned task. While this latter result provides some insight into the block structure, there remains a key open question, which this section seeks to answer: what is happening to the neural network representations as they propagate through the block structure?
Through further analysis, we show that the block structure arises from the preservation and propagation of the first principal component of its constituent layer representations. Additional experiments with linear probes Alain & Bengio (2016) further support this conclusion and show that some layers that make up the block structure can be removed with minimal performance loss.
Figure 3 explores this relationship between the block structure and the first principal components of the corresponding layer representations, demonstrated on a deep network (left group) and a wide network (right group). By comparing the variance explained by the first principal component (bottom left) to the location of the block structure (top right) we observe that layers belonging to the block structure have a highly dominant first principal component. Cosine similarity of the first principal components across all pairs of layers (top left) also shows a similarity structure resembling the block structure (top right), further demonstrating that the principal component is preserved throughout the block structure. Finally, removing the first principal component from the representations nearly eliminates the block structure from the CKA heatmaps (bottom right). A full picture of how this process impacts models of increasing depth and width can be found in Appendix Figure D.6.
In contrast, for models that do not contain the block structure, we find that cosine similarity of the first principal components across all pairs of layers bears little resemblance to the representation similarity structure measured by CKA, and the fractions of variance explained by the first principal components across all layers are relatively small (see Appendix Figure D.7). Together these results demonstrate that the block structure arises from preserving and propagating the first principal component across its constituent layers.
Although layers inside the block structure have representations with high CKA and similar first principal components, each layer nonetheless computes a nonlinear transformation of its input. Appendix Figure D.8 shows that the sparsity of ReLU activations inside and outside of the block structure is similar. In particular, ReLU activations in the block structure are sometimes in the linear regime and sometimes in the saturating regime, just like activations elsewhere in the network.
2 Linear Probes and Collapsing the Block Structure
With the insight that the block structure is preserving key components of the representations, we next investigate how these preserved representations impact task performance throughout the network, and whether the block structure can be collapsed in a way that minimally affects performance.
In Figure 4, we train a linear probe Alain & Bengio (2016) for each layer of the network, which maps from the layer representation to the output classes. In models without the block structure (first 2 panes), we see a monotonic increase in accuracy throughout the network, but in models with the block structure (last 2 panes), linear probe accuracy shows little improvement inside the block structure. Comparing the accuracies of probes for layers pre- and post-residual connections, we find that these connections play an important role in preserving representations in the block structure.
Informed by these results, we proceed to pruning blocks one-by-one from the end of each residual stage, while keeping the residual connections intact, and find that there is little impact on test accuracy when blocks are dropped from the middle stage (Figure 5), unlike what happens in models without block structure. When compared across different seeds, the magnitude of the drop in accuracy appears to be connected to the size and the clarity of the block structure present. This result suggests that block structure could be an indication of redundant modules in model design, and that the similarity of its constituent layer representations could be leveraged for model compression.
Depth and Width Effects on Representations Across Models
The results of the previous sections help characterize effects of varying depth and width on a (single) model’s internal representations, specifically, the emergence of the block structure with increased capacity, and its impacts on how representations are propagated through the network. With these insights, we next look at how depth and width affect the hidden representations across models. Concretely, are learned representations similar across models of different architectures and different random initializations? How is this affected as model capacity is changed?
We begin by studying the variations in representations across different training runs of the same model architecture. Figure 6 illustrates CKA heatmaps for a smaller model (left), wide model (middle) and deep model (right), trained from random initializations. The smaller model does not have the block structure, and representations across seeds (off diagonal plots) exhibit the same grid-like similarity structure as within a single model. The wide and deep models show block structure in all their seeds (as seen in plots along the diagonal), and comparisons across seeds (off-diagonal plots) show that while layers not in the block structure exhibit some similarity, the block structure representations are highly dissimilar across models.
Appendix Figure E.1 shows results of comparing CKA across different architectures, controlled for accuracy. Wide and deep models without the block structure do exhibit representation similarity with each other, with corresponding layers broadly being of the same proportional depth in the model. However, similar to what we observe in Figure 6, the block structure representations remain unique to each model.
Depth, Width and Effects on Model Predictions
To conclude our investigation on the effects of depth and width, we turn to understanding how the characteristic properties of internal representations discussed in the previous sections influence the outputs of the model. How diverse are the predictions of different architectures? Are there examples that wide networks are more likely to do well on compared to deep networks, and vice versa?
By training populations of networks on CIFAR-10 and ImageNet, we find that there is considerable diversity in output predictions at the individual example level, and broadly, architectures that are more similar in structure have more similar output predictions. On ImageNet we also find that there are statistically significant differences in class-level error rates between wide and deep models, with the former exhibiting a small advantage in identifying classes corresponding to scenes over objects.
Figure 7a compares per-example accuracy for groups of 100 architecturally identical deep models (ResNet-62) and wide models (ResNet-14 (2)), all trained from different random initializations on CIFAR-10. Although the average accuracy of these groups is statistically indistinguishable, they tend to make different errors, and differences between groups are substantially larger than expected by chance (Figure 7b). Examples of images with large accuracy differences are shown in Appendix Figure F.1, while Appendix Figures F.2 and F.3 further explore patterns of example accuracy for networks of different depths and widths, respectively. As the architecture becomes wider or deeper, accuracy on many examples increases, and the effect is most pronounced for examples where smaller networks were often but not always correct. At the same time, there are examples that larger networks are less likely to get right than smaller networks. We show similar results for ImageNet networks in Appendix Figure F.4.
We next ask whether wide and deep ImageNet models have systematic differences in accuracy at the class level. As shown in Figure 7c, there are small but statistically significant differences in accuracy for classes (, Welch’s -test), accounting for 11% of the variance in the differences in example-level accuracy (see Appendix F.3). Three of the top 5 classes that are more likely to be correctly classified by wide models reflect scenes rather than objects (seashore, library, bookshop). Indeed, the wide architecture is significantly more accurate on the 68 ImageNet classes descending from “structure” or “geological formation” ( vs. , , Welch’s t-test). Looking at synsets containing ImageNet classes, the deep architecture is significantly more accurate on the 62 classes descending from “consumer goods” ( vs. , ; Table F.2). In other parts of the hierarchy, differences are smaller; for instance, both models achieve 81.6% accuracy on the 118 dog classes ().
Conclusion
In this work, we study the effects of width and depth on neural network representations. Through experiments on CIFAR-10, CIFAR-100 and ImageNet, we have demonstrated that as either width or depth increases relative to the size of the dataset, analysis of hidden representations reveals the emergence of a characteristic block structure that reflects the similarity of a dominant first principal component, propagated across many network hidden layers. Further analysis finds that while the block structure is unique to each model, other learned features are shared across different initializations and architectures, particularly across relative depths of the network. Despite these similarities in representational properties and performance of wide and deep networks, we nonetheless observe that width and depth have different effects on network predictions at the example and class levels. There remain interesting open questions on how the block structure arises through training, and using the insights on network depth and width to inform optimal task-specific model design.
Acknowledgements
We thank Gamaleldin Elsayed for helpful feedback on the manuscript.
References
Appendix
Let be the set of all 4-tuples of indices between 1 and m where each index occurs exactly once. As proven in Theorem 3 of Song et al. (2012), is a U-statistic:
where and the kernel of the U-statistic is defined in Song et al. (2012). Let be 1 if the 4-tuple of dataset indices is selected in minibatch and 0 otherwise. Then:
Taking the expectation with respect to , and noting that is independent of ,
Appendix B Training Details
Our CIFAR-10 and CIFAR-100 networks follow the same architecture as He et al. (2016); Zagoruyko & Komodakis (2016). We train a set of models where we fix the width multiplier of deep networks to 1 and experiment with models of depths 32, 44, 56, 110, 164. On CIFAR-100, the block structure only appears at a greater depth so we also include depths 218 and 224 in our investigation. For wide networks, we examine width multipliers of 1, 2, 4, 8 and 10 and depths of 14, 20, 26, and 38. We use SGD with momentum of 0.9, together with a cosine decay learning rate schedule and batch size of 128, to train each model for 300 epochs. Models are trained with standard CIFAR-10 data augmentation comprising random flips and translations of up to 4 pixels. Each depth and width configuration is trained with 10 different seeds for CKA analysis, and 200 seeds for model predictions comparison.
On ImageNet, we start with the ResNet-50 architecture and increase depth or width in the third stage only, following the scaling approach of He et al. (2016). We train for 120 epochs using SGD with momentum of 0.9 and a cosine decay learning rate schedule at a batch size of 256. We use 100 seeds for model prediction comparison.
For experiments with reduced dataset size, we subsample the training data from the original CIFAR training set by the corresponding proportion, keeping the number of samples for each class the same. All CKA results are then computed based on the full CIFAR test set.
Appendix C Block Structure in a Different Architecture
Appendix D Probing the Block Structure
Appendix E Representations Across Models
Appendix F Example- and Class-Level Accuracy Differences
F.2 Effect of Varying Width and Depth on ImageNet Predictions
F.3 Effect Sizes for Class-Level Effects
We measure how much of the difference between the example-level predictions of the wide and deep ImageNet ResNets in Section 7 could be explained by the classes to which they belonged by fitting a set of three models. Model A attempts to model whether each prediction was correct or incorrect as a linear combination factors corresponding to the example ID and whether the prediction came from a wide or deep model (statsmodels formula y ~ C(example_id) + C(wide_or_deep)). Model B includes a factor corresponding to the example ID as well as factors corresponding to the interaction between the class ID and the type of model the prediction came from (statsmodels formula y ~ C(example_id) + C(wide_or_deep) * C(class_id)). Model C includes the interaction between the example ID and the type of model, and corresponds to simply measuring the average accuracy separately for both types of models (statsmodels formula y ~ C(example_id) * C(wide_or_deep)). Note that model C is nested inside model B, which is nested inside model A.
This approach is analogous to the pseudo- of Efron (1978).
Finally, we can compare the AIC values of the logistic regression models, shown in the table below. Because the models are nested GLMs, we can also test for statistical significance using a test, which is highly significant for each pair of nested models. We do not report p-values because they are 0 to within machine precision.