Net2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networks
Ruth Fong, Andrea Vedaldi
Introduction
While deep neural networks keep setting new records in almost all problems in computer vision, our understanding of these black-box models remains very limited. Without developing such an understanding, it is difficult to characterize and work around the limitations of deep networks, and improvements may only come from intuition and trial-and-error.
For deep learning to mature, a much better theoretical and empirical understanding of deep networks is thus required. There are several questions that need answering, such as how a deep network is able to solve a problem such as classifying an image, or how it can generalize so well despite having access to limited training data in relation to its own capacity . In this paper, we ask in particular what a convolutional neural network has learned to do once training is complete. A neural network can be seen as a sequence of functions, each mapping an input image to some intermediate representation. While the final output of a network is usually easy to interpret (as it provides, hopefully, a solution to the task that the network was trained to solve), the meaning of the intermediate layers is far less clear. Understanding the information carried by these representations is a first step to understanding how these networks work.
Several authors have researched the possibility that individual filters in a deep network are responsible for capturing particular semantic concepts. The idea is that low-level primitives such as edges and textures are recognized by earlier layers, and more complex objects and scenes by deeper ones. An excellent representative of this line of research is the recent Network Dissection approach by . The authors of this paper introduce a new dataset, BRODEN, which contains pixel-level segmentation for hundreds of low- and high-level visual concepts, from textures to parts and objects. They then study the correlation between extremal filter responses and such concepts, seeking for filters that are strongly responsive for particular ones.
While this and similar studies did find clear correlations between feature responses and various concepts, such an interpretation has intrinsic limitations. This can be seen from a simple counting argument: the number of available feature channels is usually far smaller than the number of different concepts that a neural network may need to encode to interpret a complex visual scene. This suggests that, at the very least, the representation must use combinations of filter responses to represent concepts or, in other words, be at least in part distributed.
The goal of this paper is to go beyond looking at individual filters, and to study instead what information is captured by combinations of neural network filters. In this paper, we conduct a thorough analysis to investigate how semantic concepts, such as objects and their parts, are encoded by CNN filters. In order to make this analysis manageable, we introduce the Net2Vec framework (section 3), which aligns semantic concepts with filter activations. It does so via learned concept embeddings that are used to weight filter activations to perform semantic tasks like segmentation and classification. Our concept vectors can be used to investigate both quantitatively and qualitatively the “overlap” of filters and concepts. Our novelty lies in outlining methods that go beyond simply demonstrating that multiple filters better encode concepts that single ones to quantifying and describing how a concept is encoded. Principally, we gain unique, interpretive power by formulating concepts vectors as embeddings.
Using Net2Vec, we look first at two questions (section 4): (1) To what extent are individual filters sufficient to express a concept? Or, are multiple filters required to code for a single concept? (2) To what extent does a filter exclusively code for a single concept? Or, is a filter shared by many, diverse concepts? While answers to these questions depend on the specific filter or concept under consideration, we demonstrate how to quantify the “overlap” between filters and concepts and show that there are many cases in which both notions of exclusive overlap do not hold. That is, if we were to interpret semantic concepts and filter activations as corresponding set of images, in the resulting Venn’s diagram the sets would intersect partially but neither kind of set would contain or be contained by the other.
While quantifying the relationship between concepts and representation may seem an obvious aim, so far much of the research on explaining how concepts are encoded by deep networks roughly falls into two more qualitative categories: (1) Interpretable visualizations of how single filters encode semantic concepts; (2) Demonstrations of distributive encoding with limited explanatory power of how a concept is encoded. In this work, we present methods that seek to marry the interpretive benefits of single filter visualizations with quantitative demonstrations of how concepts are encoded across multiple filters (section 5).
As part of our analysis, we also highlight the problem with visualizing only the inputs that maximally activate a filter and propose evaluating the power of explanatory visualizations by how well they can explain the whole distribution of filter activations (section 5.1).
Related Work
Several methods have been proposed to explain what a single filter encodes by visualizing a real or generated input that most activates a filter; these techniques are often used to argue that single filters substantially encode a concept. In contrast, shows that visualizing the real image patches that most activate a layer’s filters after a random basis has been applied also yields semantically, coherent patches. visualize segmentation masks extracted from filter activations for the most confident or maximally activating images; they also evaluate their visualizations using human judgments.
demonstrates that most PASCAL classes require more than a few hidden units to perform classification well. Most similar to , concludes that only a few hidden units encode semantic concepts robustly by measuring the overlap between image patches that most activate a hidden unit with ground truth bounding boxes and collecting human judgments on whether such patches encode systematic concepts. compares using individual filter activations with using clusters of activations from all units in a layer and shows that their clusters yielded better parts detectors and qualitatively correlated well with semantic concepts. probes mid-layer filters by training linear classifiers on their activations and analyzing them at different layers and points of training.
Net2Vec
With our Net2Vec paradigm, we propose aligning concepts to filters in a CNN by (a) recording filter activations of a pre-trained network when probed by inputs from a reference, “probe” dataset and (b) learning how to weight the collected probe activations to perform various semantic tasks. In this way, for every concept in the probe dataset, a concept weight is learned for the task of recognizing that concept. The resulting weights can then be interpreted as concept embeddings and analyzed to understand how concepts are encoded. For example, the performance on semantic tasks when using learned concept weights that span all filters in a layer can be compared to when using only a single filter or subset of filters.
In the remainder of the section, we provide details for how we learn concept embeddings by learning to segment (3.1) and classify (3.2) concepts. We also outline how we compare embeddings arising from using only a restricted set of filters, including single filters. Before we do so, we briefly discuss the dataset used to learn concepts.
We build on the BRODEN dataset recently introduced by and use it to primarily probe AlexNet trained on the ImageNet dataset as a representative model for image classification. BRODEN contains over 60,000 images with pixel- and image-level annotations for 1197 concepts across 6 categories: scenes (468), objects (584), parts (234), materials (32), textures (47), and colors (11). We exclude 8 scene concepts for which there were no validation examples. Thus, of the 1189 concepts we consider, all had image-level annotations, but only 682 had segmentation annotations, as only image-level annotations are provided for scene and texture concepts. Note that our paradigm can be generalized to any probe dataset that contains pixel- or image-level annotations for concepts. To compare the effects of different architectures and supervision, we also probe VGG16 conv5_3 and GoogLeNet inception5b trained on ImageNet and Places365 as well as conv5 of the following self-supervised, AlexNet networks: tracking , audio , objectcentric , moving , and egomotion . Post-ReLU activations are used.
1 Concept Segmentation
In this section, we show how learning to segment concepts can be used to induce concept embeddings using either all the filters available in a CNN layer or just a single filter. We also show how embeddings can be used to quantify the degree of overlap between filter combinations and concepts. This task is performed on all Broden concepts with segmentation annotations, which excludes scene and texture concepts.
We start by considering single filter segmentation following ’s paradigm with three minor modifications, listed below. For every filter , let be its corresponding activation (at a given pixel location and for a given input image). The activation’s quantile is determined such that , and is computed with respect to the distribution of filter activations over all probe images and spatial locations; we use this cut-off point to match .
Images may contain any number of different concepts, indexed by . We use the symbol to denote the probe images that contain concept . To determine which filter best segments concept , we compute a set IoU score. This score is given by the formula
which computes the intersection over union (Jakkard index) difference between the binary segmentation masks produced by the filter and the ground-truth segmentation masks . Note that sets are merged for all images in the subset of the data, where . The best filter is then selected on the training set and the validation score IoU is reported.
We differ from in the following ways: (1) we threshold before upsampling, in order to more evenly compare to the method described below; (2) we bilinearly upsample without anchoring interpolants at the center of filter receptive fields to speed up the upsampling part of the experimental pipeline; and (3) we determine the best filter for a concept on the training split rather than whereas does not distinguish a training and validation set.
1.2 Concept Segmentation by Filter Combinations
Similar to the single filter case, for each concept the weights are learned on and the set IoU score computed on thresholded masks for is reported. In addition to evaluating on the set IoU score, per-image IoU scores are computed as well:
Note that choosing a single filter is analogous to setting to a one-hot vector, where for the selected filter and otherwise, recovering the single-filter segmenter of section 3.1.1, with the output rescaled by the sigmoid function (2).
For each concept , the segmentation concept weights are learned using SGD with momentum (lr , momentum , batch size , epochs) to minimize a per-pixel binary cross entropy loss weighted by the mean concept size, i.e. 1-:
where , , and , where is the number of foreground pixels for concept in the ground truth (g.t.) mask for and is the number of pixels in g.t. masks.
2 Concept Classification
As an alternate task to concept segmentation, the problem of classifying concept (i.e., to tell whether the concept occurs somewhere in the image) can be used to induce concept embeddings. In this case, we discuss first learning embeddings using generic filter combinations (3.2.1) and then reducing those to only use a small subset of filters (3.2.2).
where and denote the height and width respectively of layer ’s activation map .
For each concept , the training images are divided into the positive subset of images that contain concept and its complement of images that do not. While in general the positive and negative sets are unbalanced, during training, images from the two sets are sampled with equal probability in order to re-balance the data (supp. sec. 1.2). To evaluate performance, we calculate the classification accuracy over a balanced validation set.
2.2 Concept Classification by a Subset of Filters
Quantifying the Filter-Concept Overlap
We start by investigating a popular hypothesis: whether concepts are well represented by the activation of individual filters or not. In order to quantify this, we consider how our learned weights, which combine information from all filter activations in a layer, compare to a single filter when being used to perform segmentation and classification on BRODEN.
Figure 2 shows that, on average, using learned weights to combine filters outperforms using a single filter on both the segmentation and classification tasks (sections 3.1.1 and 3.2.2) when being evaluated on validation data. The improvements can be quite dramatic for some concepts and starts in conv1. For instance, even for simple concepts like colors, filter combinations outperform individual filters by up to (see supp. figs. 2-4 for graphs on the performance of individual concepts). This suggests that, even if filters specific to a concept can be found, these do not optimally encode or fully “overlap” with the concept. In line with the accepted notion that deep layers improve representational quality, task performance generally improves as the layer depth increases, with trends for the color concepts being the notable exception. Furthermore, the average performance varies significantly by concept category and consistently in both the single- and multi-filter classification plots (bottom). This suggests that certain concepts are less well-aligned via linear combination to the filter space.
To answer this question, we observe how varying the number of top conv5 filters, , from which we learn concept weights affects performance (section 3.2.2). Figure 3 shows that mean performance saturates at different for the various concept categories and tasks. For the classification task (right), most concept categories saturate by ; however, scenes reaches near optimal performance around , which is much more quickly than that of materials. For the segmentation task (left), performance peaks much earlier at for materials and parts, for objects, and for colors. We also observe performance drops after reaching optimal peaks for materials and parts in the segmentation class. This highlights that the segmentation task is challenging for those concept categories in particular (i.e., object parts are much smaller and harder to segment, materials are most different from network’s original ImageNet training examples of objects); with more filters to optimize for, learning is more unstable and more likely to reach a sub-optimal solution.
While on average our multi-filter approach significantly outperforms a single-filter approach on both segmentation and classification tasks (fig. 2), Table 1 shows that for around of concepts, this does not hold. For segmentation, this percentage increases with layer depth. Upon investigation, we discovered that the concepts for which our learned weights do not outperform the best filter either have very few examples for that concept, i.e. mostly which leads to overfitting; or are very small objects, of average size less than of an image, and thus training with the size weighted (4) loss is unstable and difficult, particularly at later layers where there is low spatial resolution. A similar analysis on the classification results shows that small concept dataset size is also causing overfitting in failure cases: Of the 133 conv5 failure cases, 103 had at most 20 positive training examples and all but one had less than 100 positive training examples (supplementary material figs. 7 and 8).
2 Are Filters Shared between Concepts?
Next, we investigate the extent to which a single filter is used to encode many concepts. Note that Figure 1 suggests that a single filter might be activated by different concepts; often, the different concepts a filter appears to be activated by are related by a latent concept that may or may not be human-interpretable, i.e., an ‘animal torso’ filter which also is involved in characterizing animals like ‘sheep’, ‘cow’, and ‘horse’ (fig. 4, supp. fig. 9).
Using the single best filters identified in both the segmentation and classification tasks, we explore how often a filter is selected as the best filter to encode a concept. Figure 5 shows the distribution of how many filters (y-axis) encode how many concepts (x-axis). Interestingly, around of conv1 filters (as well as several in all the other layers) were selected for encoding at least 20 and 30 concepts (# of concepts / # of conv1 filters = 10.7 and 18.6; supp. tbl. 1) for the segmentation and classification tasks respectively and a substantial portion of filters in each layer (except conv1 for the segmentation task) are never selected. The filters selected to encode numerous concepts are not exclusively “overlapped” by a single concept. The filters that were not selected to encode any concepts are likely not be involved in detecting highly discriminative features.
3 More Architectures, Datasets, and Tasks
Figure 6 shows segmentation (top) and classification (bottom) results when using AlexNet (AN) conv5, VGG16 (VGG) conv5_3, and GoogLeNet (GN) inception5b trained on both ImageNet (IN) and Places365 (P) as well as conv5 of these self-supervised (SS), AlexNet networks: tracking, audio, objectcentric, moving, and egomotion. GN performed worse than VGG because of its lower spatial resolution ( vs. ); GN-IN inception4e () outperforms VGG-IN conv5_3 (supp. fig. 11). In , GN detects scenes well, which we exclude due to lack of segmentation data. SS performance improves more than supervised networks (5-6x vs. 2-4x), suggesting that SS networks encode BRODEN concepts more distributedly.
Interpretability
In this section, we propose a new standard for visualizing non-extreme examples, show how the single- and multi-filter perspectives can be unified, and demonstrate how viewing concept weights as embeddings in filter space give us novel explanatory power.
Many visual explanation methods demonstrate their value by showing visualizations of inputs that maximally activate a filter, whether that be real, maximally-activating image patches ; learned, generated maximally-activated inputs ; or filter segmentation masks for maximally-activating images from a probe dataset .
While useful, these approaches fail to consider how visualizations differ across the distribution of examples. Figure 7 shows that using a single filter to segment concepts yields scores of for many examples; such examples are simply not considered by the set IoU metric. This often occurs because no activations survive the -thresholding step, which suggests that a single filter does not consistently fire strongly on a given concept.
We argue that a visualization technique should still work on and be informative for non-maximal examples. In Figure 8, we automatically select and visualize examples at each decile of the non-zero portion of the individual IoU distribution (fig. 7) using both learned concept weights and the best filters identified for each of the visualized categories. For ‘dog’ and ‘airplane’ visualizations using our weighted combination method, the predicted masks are informative and salient for most of the examples, even the lowest 10th percentile (leftmost column). Ideally, using this decile sampling method, the visualizations should appear salient even for examples from lower deciles. However, for examples using the best single filter (odd rows), the visualizations are not interpretable until higher deciles (rightmost columns). This is in contrast to the visually appealing, maximally activating examples shown in supp. fig. 13.
2 Unifying Single- & Multi-Filter Views
Figure 9 highlights that single filter performance is often strongly, linearly correlated with the learned weights , thereby showing that individual filter performance is indicative of how weighted it’d be in a linear filter combination. Visually, a filter’s set IoU score appears correlated with its associated weight value passed through a ReLU, i.e., . For each of the BRODEN segmentation concepts and each AlexNet layer, we computed the correlation between and . By conv3, around of segmentation concepts are significantly correlated (): conv1: 47.33%, conv2: 69.12%, conv3: 81.14%, conv4: 79.13%, conv5: 82.47%. Thus, we show how the single filter perspective can be unified with and utilized to explain the distributive perspective: we can quantify how much a single filter contributes to concept ’s encoding from either where is ’s learned weight vector or .
3 Explanatory Power via Concept Embeddings
Finally, the learned weights can be considered as embeddings, where each dimension corresponds to a filter. Then, we can leverage the rich literature on word embeddings derived from textual data to better understand which concepts are similar to each other in network space. To our knowledge, this is the first work that learns semantic embeddings aligned to the filter space of a network from visual data alone. (For this section, concept weights are normalized to be unit length, i.e., ).
Table 2 shows the five closest concepts in cosine distance, where denotes that is from and denotes that is from . These examples suggest that the embeddings from the segmentation and classification tasks capture slightly different relationships between concepts. Specifically, the nearby concepts in segmentation space appear to be similar-category objects (i.e., animals in the case of ‘cat’ and ‘horse’ being nearest to ‘dog’), whereas the nearby concepts in classification space appear to be concepts that are related compositionally (i.e., parts of an object in the case of ‘muzzle’ and ‘paw’ being nearest to ‘dog’). Note that ‘street’ and ‘bedroom’ are categorized as scenes and thus lack segmentation annotations.
Table 3 shows that we can also do vector arithmetic by adding and subtracting concept embeddings to get meaningful results. For instance, we observe an analogy relationship between ‘grass’‘green’ and ‘sky’‘blue’ and other coherent results, such as non-green, ‘ground’-like concepts for ‘grass’ minus ‘green’ and floral concepts for ‘tree’ minus ‘wood’. t-SNE visualizations and K-means clustering (see supp. table 2 and supp. figs. 16 and 17) also demonstrate that networks learn meaningful, semantic relationships between concepts.
Conclusion
We present a paradigm for learning concept embeddings that are aligned to a CNN layer’s filter space. Not only do we answer the binary questions, “does a single filter encode a concept fully and exclusively?,” we also introduce the idea of filter and concept “overlap” and outline methods for answering the scalar extension questions, “to what extent…?” We also propose a more fair standard for visualizing non-extreme examples and show how to explain distributed concept encodings via embeddings. While powerful and interpretable, our approach is limited by its linear nature; future work should explore non-linear ways concepts can be better aligned to the filter space. We gratefully acknowledge the support of the Rhodes Trust for Ruth Fong and ERC 677195-IDIU for Andrea Vedaldi.