Growing Interpretable Part Graphs on ConvNets via Multi-Shot Learning

Quanshi Zhang, Ruiming Cao, Ying Nian Wu, Song-Chun Zhu

Introduction

Convolutional neural networks (?; ?; ?) (CNNs) have achieved near human-level performance in object classification on some datasets. However, in real-world applications, we are still facing the following two important issues.

Firstly, given a CNN that is pre-trained for object classification, it is desirable to derive an interpretable graphical model to explain explicit semantics hidden inside the CNN. Based on the interpretable model, we can go beyond the detection of object bounding boxes, and discover an object’s latent structures with different part components from the pre-trained CNN representations.

Secondly, it is also desirable to learn from very few annotations. Unlike data-rich applications (e.g. pedestrian and vehicle detection), many visual tasks demand for modeling certain objects or certain object parts on the fly. For example, when people teach a robot to grasp the handle of a cup, they may not have enough time to annotate sufficient training samples of cup handles before the task. It is better to mine common knowledge of cup handles from a few examples on the fly.

Motivated by the above observations, in this paper, given a pre-trained CNN, we use very few (3–12) annotations to model a semantic part for the task of part localization. When a CNN is pre-trained using object-level annotations, we believe that its conv-layers have contained implicit representations of the objects. We call the implicit representations latent patterns, each corresponding to a component of the semantic part (namely a sub-part) or a contextual region w.r.t. the semantic part. For each semantic part, our goal is to mine latent patterns from the conv-layers related to this part. We use an And-Or graph (AOG) to organize the mined latent patterns to represent the semantic hierarchy of the part.

Input and output: Given a pre-trained CNN and a number of images for a certain category, we only annotate the semantic parts on a few images as input. We develop a method to grow a semantic And-or Graph (AOG) on the pre-trained CNN, which associates certain CNN units with the semantic part. Our method does not require massive annotations for learning, and can work with even a single part annotation. We can use the learned AOG to parse/localize object parts and their sub-parts for hierarchical object parsing.

Fig. 2 shows that the AOG has four layers. In the AOG, each OR node encodes its alternative representations as children, and each AND node is decomposed into its constituents.

Layer 1: the top OR node for semantic part describes the head of a sheep in Fig. 2. It lists a number of part templates as children.

Layer 2: AND nodes for part templates correspond to different poses or local appearances for the part, e.g. a black sheep head from a front view and a white sheep head from side view.

Layer 3: OR nodes for latent patterns describe sub-parts of the sheep head (e.g. a corner of the nose) or a contextual region (e.g. the neck region).

Layer 4: terminal nodes are CNN units. A latent pattern naturally corresponds to a certain range of units within a conv-slice. It selects a CNN unit within this range to account for local shape deformation of this pattern.

Learning method and key benefits: The basic idea for growing AOG is to define a metric to distinguish reliable latent patterns from noisy neural activations in the conv-layers. We expect latent patterns with high reliability to 1) consistently represent certain sub-parts on the annotated object samples, 2) frequently appear in unannotated objects, and 3) keep stable spatial relationship with other latent patterns. We mine reliable latent patterns to construct the AOG. This learning method is related to previous studies of pursuing AOGs, which mined hierarchical object structures from Gabor wavelets on edges (?) and HOG features (?). We extend such ideas to feature maps of neural networks.

Our method has the following three key benefits:

∙\bullet CNN semanticization: We semanticize the pre-trained CNN by connecting its units to an interpretable AOG. In recent years, people have shown a special interest in opening the black-box representation of the CNN. In this paper, we retrieve “implicit” patterns from the CNN, and use the AOG to associate each pattern with a certain “explicit” semantic part. We can regard the AOG as an interpretable representation of CNN patterns, which may contribute to the understanding of black-box knowledge organization in the CNN.

∙\bullet Multi-shot learning: The idea of pattern mining also enables multi-shot learning from small data. Conventional end-to-end learning usually requires a large number of annotations to learn/finetune networks. In contrast, in our learning scenario, all patterns in the CNN have been well pre-trained using object-level annotations. We only use very few (3–12) part annotations to retrieve certain latent patterns, instead of finetuning CNN parameters. For example, we use the annotation of a specific tiger head to mine latent patterns. The mined patterns are not over-fitted to the head annotation, but represent common head appearance among different tigers. Therefore, we can greatly reduce the number of part annotations for training.

∙\bullet Incremental learning: we can incrementally enrich the knowledge of semantic parts. Given a pre-trained CNN, we can incrementally grow new neural connections from CNN units to a new AOG, in order to represent a new semantic part. It is important to maintain the generality of the pre-trained CNN during the learning procedure. I.e. we do not change/fine-tune the original convolutional weights within the CNN, when we grow new AOGs. This allows us to continuously add new semantic parts to the same CNN, without worrying about the model drift problem.

Contributions of this paper are summarized as follows.

1) From the perspective of model learning, given a few part annotations, we propose a method to incrementally grow interpretable AOGs on a pre-trained CNN to gradually model semantic parts of the object.

2) From the perspective of knowledge transferring, our method semanticizes a CNN by mining reliable latent patterns from noisy neural responses of the CNN and associating the implicit patterns with explicit semantic parts.

3) To the best of our knowledge, we can regard our method as the first study to achieve weakly supervised (e.g. 3–12 annotations) learning for part localization. Our method exhibits superior localization performance in experiments (about 13%–107% improvement in part center prediction).

Related work

Long-term learning & short-term learning: As reported in (?), there are “two learning systems instantiated in mammalians:” 1) the neocortex gradually acquires sophisticated knowledge representation, and 2) the hippocampus quickly learns specifics of individual experiences. CNNs are typically trained using big data, and contain rich appearance patterns of objects. If one compares CNNs to the neocortex, then the fast retrieval of latent patterns related to a semantic part can be compared to the short-term learning in hippocampus.

Semantics in the CNN: In order to explore the hidden semantics in the CNN, many studies have focused on the visualization of CNN units (?; ?; ?; ?; ?) and analyzed their statistical features (?; ?; ?; ?). Liu et al. (?) extracted and visualized a subspace of CNN features.

Going beyond “passive” visualization, some studies “actively” extracted CNN units with certain semantics for different applications. Zhou et al. (?; ?) discovered latent “scene” semantics from CNN feature maps. Simon et al. discovered objects (?) in an unsupervised manner from CNN feature maps, and learned semantic parts in a supervised fashion (?). In our study, given very few part annotations, we mine CNN patterns that are related to the semantic part. Obtaining clear semantics makes it easier to transfer CNN patterns to other part-based tasks.

AOG for knowledge transfer: Transferring hidden patterns in the CNN to other tasks is important for neural networks. Typical research includes end-to-end fine-tuning and transferring CNN knowledge between different categories (?; ?) and/or datasets (?). In contrast, we believe that a good explanation and transparent representation of part knowledge will creates a new possibility of transferring part knowledge. As in (?; ?), the AOG is suitable to represent the semantic hierarchy, which enables semantic-level interactions between human and neural networks.

Modeling “objects” vs. modeling “parts” in un-/weakly-supervised learning: Generally speaking, in terms of un-/weakly-supervised learning, modeling parts is usually more challenging than modeling entire objects. Given image-level labels (without object bounding boxes), object discovery (?; ?; ?) can be achieved by identifying common foreground patterns from noisy background. Closed boundaries and common object structure are also strong prior knowledge for object discovery.

In contrast to objects, semantic parts are hardly distinguishable from other common foreground patterns in an unsupervised manner. Some parts (e.g. the abdomen) do not have shape boundaries to determine their shape extent. Inspired by graph mining (?; ?; ?), we mine common patterns from CNN activation maps in conv-layers to explain the part.

Part localization/detection vs. semanticizing CNN patterns: Part localization/detection is an important task in computer vision (?; ?; ?; ?). There are two key points to differentiate our study from conventional part-detection approaches. First, most methods for detection, such as the CNN and the DPM (?; ?; ?), limit their attention to the classification problem. In contrast, our effort is to clarify semantic meanings of implicit CNN patterns. Second, instead of summarizing knowledge from massive annotations, our method mines CNN semantics with very limited supervision.

And-Or graph for part parsing

In this section, we introduce the structure of the AOG and part parsing/localization based on the AOG. The AOG structure is suitable for clearly representing semantic hierarchy of a part. The method for mining latent patterns and building the AOG will be introduced in the next section. An AOG represents the semantic structure of a part at four layers.

Each OR node in the AOG represents a list of alternative appearance (or deformation) candidates. Each AND node is composed of a number of latent patterns to describe its sub-regions.

In Fig. 2, given CNN activation maps on an image IIBecause the CNN has demonstrated its superior performance in object detection, we assume that the target object can be well detected by the pre-trained CNN. Thus, to simplify the learning scenario, we crop II to only contain the object, resize it to the image size for CNN inputs, and only focus on the part localization task., we can use the AOG for part parsing. From a top-down perspective, the parsing procedure 1) identifies a part template for the semantic part; 2) parses an image region for the selected part template; 3) for each latent pattern under the part template, it selects a CNN unit within a certain deformation range to represent this pattern.

In this way, we select certain AOG nodes in a parse graph to explain sub-parts of the object (shown as red lines in Fig. 2). For each node VV in the parse graph, we parse an image region ΛV\Lambda_{V} within image IIImage regions of OR nodes are propagated from their children. Each terminal node has a fixed image region, and each part template (AND node) has a fixed region scale (will be introduced later). Thus, we only need infer the center position of each part template in (2) during part parsing.. We use SI(V)S_{I}(V) to denote an inference/parsing score, which measures the fitness between the parsed region ΛV\Lambda_{V} and VV (as well as the sub-AOG under VV).

Given an image II2 and an AOG, the actual parsing procedure is solved by dynamic programming in a bottom-up manner, as follows.

Terminal nodes (CNN units): We first focus on parsing configurations of terminal nodes. Terminal nodes under a latent pattern are displaced in location candidates of this latent pattern. Each terminal node VuntV^{\textrm{unt}} has a fixed image region ΛVunt\Lambda_{V^{\textrm{unt}}}: we propagate VuntV^{\textrm{unt}}’s receptive field back to the image plane as ΛVunt\Lambda_{V^{\textrm{unt}}}. We compute VuntV^{\textrm{unt}}’s inference score SI(Vunt)S_{I}(V^{\textrm{unt}}) based on both its neural response value and its displacement w.r.t. its parent (see appendix for detailsPlease see the section of appendix for details.).

Latent patterns: Then, we propagate parsing configurations from terminal nodes to latent patterns. Each latent pattern VlatV^{\textrm{lat}} is an OR node. VlatV^{\textrm{lat}} naturally corresponds to a square within a certain conv-slice in the output of a certain CNN conv-layer as its deformation rangeWe set a constant deformation range for each latent pattern, which potentially covers 75 ⁣× ⁣7575\!\times\!75 pxls on the image. Deformation ranges of different patterns in the same conv-slice may overlap.. VlatV^{\textrm{lat}} connects all the CNN units within the deformation range as children, which represent different deformation candidates. Given parsing configurations of its children CNN units as input, VlatV^{\textrm{lat}} selects the child V^unt\hat{V}^{\textrm{unt}} with the highest score as the true deformation configuration:

Part templates: Each part template VtmpV^{\textrm{tmp}} is an AND node, which uses its children (latent patterns) to represent its sub-part/contextual regions. Based on the relationship between VtmpV^{\textrm{tmp}} and its children, VtmpV^{\textrm{tmp}} uses its children’s parsing configurations to parse its own image region ΛVtmp\Lambda_{V^{\textrm{tmp}}}. Given parsing scores of children, VtmpV^{\textrm{tmp}} computes the image region Λ^Vtmp\hat{\Lambda}_{V^{\textrm{tmp}}} that maximizes its inference score.

Just like typical part models (e.g. DPMs), the AND node uses each child’s region VlatV^{\textrm{lat}} to infer its own region. SI(Vlat)S_{I}(V^{\textrm{lat}}) measures the score of each child, and Sinf(ΛVtmp∣Λ^Vlat)S^{\textrm{inf}}(\Lambda_{V^{\textrm{tmp}}}|\hat{\Lambda}_{V^{\textrm{lat}}}) measures spatial compatibility between VtmpV^{\textrm{tmp}} and each child VlatV^{\textrm{lat}} in region parsing (see the appendix for formulations).

Semantic part: Finally, we propagate parsing configurations to the top node VsemV^{\textrm{sem}}. VsemV^{\textrm{sem}} is an OR node. It contains a list of alternative templates for the part. Just like OR nodes of latent patterns, VsemV^{\textrm{sem}} selects the child V^tmp\hat{V}^{\textrm{tmp}} with the highest score as the true parsing configuration:

Learning: growing an And-Or graph

The basic idea of AOG growing is to distinguish reliable latent patterns from noisy neural responses in conv-layers and use reliable latent patterns to construct the AOG.

Training data: Let I{\bf I} denote an image set for a target category. Among all objects in I{\bf I}, we label bounding boxes of the semantic part in a small number of images, Iant ⁣= ⁣{I1,I2,…,IM}⊂I{\bf I}^{\textrm{ant}}\!=\!\{I_{1},I_{2},\ldots,I_{M}\}\subset{\bf I}. In addition, we manually define a number of templates for the part. Thus, for each I∈IantI\in{\bf I}^{\textrm{ant}}, we annotate (ΛVsem∗,Vtmp∗)(\Lambda_{V^{\textrm{sem}}}^{*},V^{\textrm{tmp}*}), where ΛVsem∗\Lambda_{V^{\textrm{sem}}}^{*} denotes the ground-truth bounding box of the part in II, and Vtmp∗V^{\textrm{tmp}*} specifies the ground-truth template ID for the part.

Which AOG parameters to learn: We can use human annotations to define the first two layers of the AOG. If people specify a total of mm different part templates during the annotation process, correspondingly, we can directly connect the top node with mm part templates {Vtmp∗}\{V^{\textrm{tmp}*}\} as children. For each part template VtmpV^{\textrm{tmp}}, we fix a constant scale for its region ΛVtmp\Lambda_{V^{\textrm{tmp}}}. I.e. if there are nn ground-truth part boxes that are labeled for VtmpV^{\textrm{tmp}}, we compute the average scale among the nn part boxes as the constant scale for ΛVtmp\Lambda_{V^{\textrm{tmp}}}.

Thus, the key to AOG construction is to mine children latent patterns for each part template. We need to mine latent patterns from a total of KK conv-layers. We select nkn_{k} latent patterns from the kk-th (k=1,2,…,Kk=1,2,\ldots,K) conv-layer, where KK and {nk}\{n_{k}\} are hyper-parameters. Let each latent pattern VlatV^{\textrm{lat}} in the kk-th conv-layer correspond to a square deformation range5, which is located in the DVlatD_{V^{\textrm{lat}}}-th conv-slice of the conv-layer. P‾Vlat\overline{\bf P}_{V^{\textrm{lat}}} denotes the center of the range. As analyzed in the appendix, we only need to estimate the parameters of DVlat,P‾VlatD_{V^{\textrm{lat}}},\overline{\bf P}_{V^{\textrm{lat}}} for VlatV^{\textrm{lat}}.

How to learn: We mine the latent patterns by estimating their best locations DVlat,P‾Vlat∈θD_{V^{\textrm{lat}}},\overline{\bf P}_{V^{\textrm{lat}}}\in{\boldsymbol{\theta}} that maximize the following objective function.

where θ{\boldsymbol{\theta}} is the set of AOG parameters. First, let us focus on the first half of the equation, which learns from part annotations. Given annotations (ΛVsem∗,Vtmp∗)(\Lambda_{V^{\textrm{sem}}}^{*},V^{\textrm{tmp}*}) on II, SI(Vsem)S_{I}(V^{\textrm{sem}}) denotes the parsing score of the part. ∥P^Vsem−PVsem∗∥\|\hat{\bf P}_{V^{\textrm{sem}}}-{\bf P}^{*}_{V^{\textrm{sem}}}\| measures localization error between the parsed part region P^Vsem\hat{{\bf P}}_{V^{\textrm{sem}}} and the ground truth PVsem∗{\bf P}^{*}_{V^{\textrm{sem}}}. We ignore the small probability of the AOG assigning an annotated image with an incorrect part template to simplify the computation of parsing scores, i.e. SI(Vsem)≈SI(Vtmp∗)S_{I}(V^{\textrm{sem}})\approx S_{I}(V^{\textrm{tmp}*}).

The second half of (4) learns from objects without part annotations. We formulate S^{\textrm{unsup}}_{I^{\prime}}\!(V^{\textrm{lat}})\!=\!\lambda^{\textrm{unsup}}\!\big{[}\!S_{I^{\prime}}^{\textrm{rsp}}(\hat{V}^{\textrm{unt}})\!+\!S_{I^{\prime}}^{\textrm{loc}}\!(V^{\textrm{unt}})\!-\!\lambda^{\textrm{close}}\!\|\Delta{\bf P}_{V^{\textrm{lat}}}\|^{2}\!\big{]}, where latent pattern VlatV^{\textrm{lat}} selects CNN unit V^unt\hat{V}^{\textrm{unt}} as its deformation configuration on I′I^{\prime}. The first term SI′rsp(V^unt)S_{I^{\prime}}^{\textrm{rsp}}(\hat{V}^{\textrm{unt}}) denotes the neural response of the CNN unit V^unt\hat{V}^{\textrm{unt}}. The second term SI′loc(Vunt)=−λloc∥P^Vunt−P‾Vlat∥2S_{I^{\prime}}^{\textrm{loc}}(V^{\textrm{unt}})=-\lambda^{\textrm{loc}}\|\hat{\bf P}_{V^{\textrm{unt}}}-\overline{\bf P}_{V^{\textrm{lat}}}\|^{2} measures the deformation level of the latent pattern. The third term measures the spatial closeness between the latent pattern and its parent VtmpV^{\textrm{tmp}}. We assume that 1) latent patterns that frequently appear among unannotated objects may potentially represent stable sub-parts and should have higher priorities; and that 2) latent patterns spatially closer to VtmpV^{\textrm{tmp}} are usually more reliable. Please see the appendix for details of SI′rsp(V^unt)S_{I^{\prime}}^{\textrm{rsp}}(\hat{V}^{\textrm{unt}}) and scalar weights of λunsup\lambda^{\textrm{unsup}}, λclose\lambda^{\textrm{close}}, and λloc\lambda^{\textrm{loc}}.

When we set λVtmp∗\lambda_{V^{\textrm{tmp}*}} to a constant λinf∑k=1Knk\lambda^{\textrm{inf}}\sum_{k=1}^{K}n_{k}, we can transform the learning objective in (4) as follows.

where Score(Vlat) ⁣= ⁣meanI∈IVtmp[SI(Vlat)+Sinf(ΛVsem∗∣Λ^Vlat)]Score(V^{\textrm{lat}})\!=\!{\textrm{mean}}_{I\in{\bf I}_{V^{\textrm{tmp}}}}[S_{I}(V^{\textrm{lat}})+S^{\textrm{inf}}(\Lambda^{*}_{V^{\textrm{sem}}}|\hat{\Lambda}_{V^{\textrm{lat}}})] +meanI′∈ISI′unsup(Vlat)+{\textrm{mean}}_{I^{\prime}\in{\bf I}}S^{\textrm{unsup}}_{I^{\prime}}(V^{\textrm{lat}}). θVtmp⊂θ{\boldsymbol{\theta}}_{V^{\textrm{tmp}}}\subset{\boldsymbol{\theta}} denotes the parameters for the sub-AOG of VtmpV^{\textrm{tmp}}. We use IVtmp⊂Iant{\bf I}_{V^{\textrm{tmp}}}\subset{\bf I}^{\textrm{ant}} to denote the subset of images that are annotated with VtmpV^{\textrm{tmp}} as the ground-truth part template.

Learning the sub-AOG for each part template: Based on (5), we can mine the sub-AOG for each part template VtmpV^{\textrm{tmp}}, which uses this template’s own annotations on images I∈IVtmp⊂IantI\in{\bf I}_{V^{\textrm{tmp}}}\subset{\bf I}^{\textrm{ant}}, as follows.

1) We first enumerate all possible latent patterns corresponding to the kk-th CNN conv-layer (k=1,…,Kk=1,\ldots,K), by sampling all pattern locations w.r.t. DVlatD_{V^{\textrm{lat}}} and P‾Vlat\overline{\bf P}_{V^{\textrm{lat}}}.

2) Then, we sequentially compute Λ^Vlat\hat{\Lambda}_{V^{\textrm{lat}}} and Score(Vlat)Score(V^{\textrm{lat}}) for each latent pattern.

3) Finally, we sequentially select a total of nkn_{k} latent patterns. In each step, we select V^lat ⁣= ⁣arg⁡ ⁣max⁡VlatΔL\hat{V}^{\textrm{lat}}\!=\!{\arg\!\max}_{V^{\textrm{lat}}}\Delta{\bf L}. I.e. we select latent patterns with top-ranked values of Score(Vlat)Score(V^{\textrm{lat}}) as VtmpV^{\textrm{tmp}}’s children.

Experiments

We chose the 16-layer VGG network (VGG-16) (?) that was pre-trained using the 1.3M images in the ImageNet ILSVRC 2012 dataset (?) for object classification. Then, given a target category, we used images in this category to fine-tune the original VGG-16 (based on the loss for classifying target objects and background). VGG-16 has 13 conv-layers and 3 fully connected layers. We chose the last 9 (from the 5-th to the 13-th) conv-layers as valid conv-layers, from which we selected units to build the AOG.

Note that during the learning process, we applied the following two techniques to further refine the AOG model. First, multiple latent patterns in the same conv-slice may have similar positions P‾Vlat\overline{\bf P}_{V^{\textrm{lat}}}, and their deformation ranges may highly overlap with each other. Thus, we selected the latent pattern with the highest Score(Vlat)Score(V^{\textrm{lat}}) within each small range of ϵ×ϵ\epsilon\times\epsilon in this conv-slice, and removed other nearby patterns to obtain a spare AOG structure. Second, for each VtmpV^{\textrm{tmp}}, we estimated nkn_{k}, i.e. the best number of latent patterns in conv-layer kk. We assumed that scores of all the latent patterns in the kk-th conv-layer follow the distribution of Score(Vlat)∼αexp⁡[−(βrank)0.5]+γScore(V^{\textrm{lat}})\sim\alpha\exp[-(\beta{rank})^{0.5}]+\gamma, where rankrank denotes the score rank of VlatV^{\textrm{lat}}. We found that when we set nk=⌈0.5/β⌉n_{k}=\lceil 0.5/\beta\rceil, the AOG usually had reliable performance.

Datasets

We tested our method on three benchmark datasets: the PASCAL VOC Part Dataset (?), the CUB200-2011 dataset (?), and the ILSVRC 2013 DET dataset (?). Just like in most part-localization studies (?), we also selected six animal categories—bird, cat, cow, dog, horse, and sheep—from the PASCAL Part Dataset for evaluation, which prevalently contain non-rigid shape deformation. The CUB200-2011 dataset contains 11.8K images of 200 bird species. As in (?; ?), we regarded these images as a single bird category by ignoring the species labels. All the above seven categories have ground-truth annotations of the head (it is the forehead part in the CUB200-2011 dataset) and torso/back. Thus, for each category, we learned two AOGs to model its head and torso/back, respectively.

In order to provide a more comprehensive evaluation of part localization, we built a larger object-part dataset based on the off-the-shelf ILSVRC 2013 DET dataset. We used 30 animal categories among all the 200 categories in the ILSVRC 2013 DET dataset. We annotated bounding boxes for the heads and front legs/feet in these animals as two common semantic parts for evaluation. In Experiments, we annotated 3–12 boxes for each part to build the AOG, and we used the rest images in the dataset as testing images.

Two experiments on multi-shot learning

We applied our method to all animal categories in the above three benchmark datasets. We designed two experiments to test our method in the scenarios of (1 ⁣× ⁣3)(1\!\times\!3)-shot learning and (4 ⁣× ⁣3)(4\!\times\!3)-shot learning, respectively. We applied the learned AOGs to part localization for evaluation.

Exp. 1, three-shot AOG construction: For each semantic part of an object category, we learn three different part templates. We annotated a single bounding box for each part template. Thus, we used a total of three annotations to build the AOG for this part.

Exp. 2, AOG construction with more annotations: We continuously added more part annotations to check the performance changes. Just as in Experiment 1, each part contains the same three part templates. For each part template, we annotated four parts in four different object images to build the corresponding AOG.

Baselines

We compared our method with the following nine baselines. The first baseline was the fast-RCNN (?). We directly used the fast-RCNN to detect the target parts on objects. To enable a fair comparison, we learned the fast-RCNN by first fine-tuning the VGG-16 network of the fast-RCNN using all object images in the target category and then training the fast-RCNN using the part annotations. The second baseline was the strongly supervised DPM (SS-DPM) (?), which was trained with part annotations for part localization. The third baseline was proposed in (?), which trained a DPM component for each object pose to localize object parts (namely, PL-DPM). We used the graphical model proposed in (?) as the fourth baseline for part localization (PL-Graph). The fifth baseline, namely CNN-PDD, was proposed by (?), which selected certain conv-slices (channels) of the CNN to represent the target object part. The sixth baseline (VGG-PDD-finetuned) was an extension of CNN-PDD, which was conducted based the VGG-16 network that was pre-fine-tuned using object images in the target category. Because in the scope of weakly supervised learning, “simple” methods are usually insensitive to the over-fitting problem, we designed the last three baselines as follows. Given the pre-trained VGG-16 network that was used in our method, we directly used this network to extract fc7 features from image patches of the annotated parts, and learned a linear SVM and a RBF SVM to classify target parts and background. Then, given a testing image, the three baselines brutely searched part candidates from the image, and used the linear SVM, the RBF SVM, and the nearest-neighbor strategy, respectively, to detect the best part. All the baselines were conducted using the same set of annotations for a fair comparison.

Evaluation metric

As mentioned in (?), a fair evaluation of part localization requires to remove the factors of object detection. Therefore, we used object bounding boxes to crop objects from the original images as the testing samples. Note that detection-based baselines (e.g. fast-RCNN, PL-Graph) may produce several bounding boxes for the part. Just as in (?; ?), we took the most confident bounding box per image as the localization result. Given localization results of a part in a certain category, we used three evaluation metrics. 1) Part detection: a true part detection was identified based on the widely used “IOU ≥0.5\geq 0.5” criterion (?); the part detection rate of this category was computed. 2) Center prediction: as in (?), if the predicted part center was localized inside the true part bounding box, we considered it a correct center prediction; otherwise not. The average center prediction rate was computed among all objects in the category for evaluation. 3) The normalized distance in (?) is a standard metric to evaluate localization accuracy on the CUB200-2011 dataset. Because object parts may not have clear boundaries (e.g. the forehead of the bird), center prediction and normalized distance are more often used for evaluation of part localization.

Results and quantitative analysis

Table 1 lists the average children number of an AOG node at different layers. Fig. 4 shows the positions of the extracted latent pattern nodes, and part-localization results based on the AOGs. Given an image, we also used latent patterns in the AOG to reconstruct the corresponding semantic part based on the technique of (?), in order to show the interpretability of the AOG.

In Tables 2, 3, 4, 5, and 6, we compared the performance of different baselines. Our method exhibited much better performance than other baselines that suffered from over-fitting problems. In Fig. 3, we showed the performance curve when we increased the annotation number from 3 to 12. Note that the 12-shot learning only improved about 0.9%–2.9% of center prediction over the 3-shot learning. This demonstrated that our method was efficient in mining CNN semantics, and the CNN units related to each part template had been roughly mined using just three annotations. In fact, we can further improve the performance by defining more part templates, rather than by annotating more part boxes for existing part templates.

Conclusions and discussion

In this paper, we have presented a method for incrementally growing new neural connections on a pre-trained CNN to encode new semantic parts in a four-layer AOG. Given an demand for modeling a semantic part on the fly, our method can be conducted with a small number of part annotations (even a single box annotation for each part template). In addition, our method semanticizes CNN units by associating them with certain semantic parts, and builds an AOG as a interpretable model to explain the semantic hierarchy of CNN units.

Because we reduce high-dimensional CNN activations to low-dimensional representation of parts/sub-parts, our method has high robustness and efficiency in multi-shot learning, and has exhibited superior performance to other baselines.

Acknowledgement

This study is supported by MURI project N00014-16-1-2007 and DARPA SIMPLEX project N66001-15-C-4035.

Appendix

In general, we use the notation of PV{\bf P}_{V} to denote the central position of an image region ΛV\Lambda_{V} as follows.

Each latent pattern VlatV^{\textrm{lat}} is defined by its location parameters {LVlat,DVlat,P‾Vlat,ΔPVlat}⊂θ\{L_{V^{\textrm{lat}}},D_{V^{\textrm{lat}}},\overline{\bf P}_{V^{\textrm{lat}}},\Delta{\bf P}_{V^{\textrm{lat}}}\}\subset{\boldsymbol{\theta}}, where θ{\boldsymbol{\theta}} is the set of AOG parameters. It means that a latent pattern VlatV^{\textrm{lat}} uses a square5 within the DVlatD_{V^{\textrm{lat}}}-th conv-slice/channel in the output of the LVlatL_{V^{\textrm{lat}}}-th CNN conv-layer as its deformation range. Each VlatV^{\textrm{lat}} in the kk-th conv-layer has a fixed value of LVlat=kL_{V^{\textrm{lat}}}=k. ΔPVlat\Delta{\bf P}_{V^{\textrm{lat}}} is used to compute Sinf(ΛVtmp∣Λ^Vlat)S^{\textrm{inf}}(\Lambda_{V^{\textrm{tmp}}}|\hat{\Lambda}_{V^{\textrm{lat}}}). Given parameter P‾Vlat\overline{\bf P}_{V^{\textrm{lat}}}, the displacement ΔPVlat\Delta{\bf P}_{V^{\textrm{lat}}} can be estimated as ΔPVlat ⁣= ⁣PVtmp∗−P‾Vlat\Delta{\bf P}_{V^{\textrm{lat}}}\!=\!{\bf P}^{*}_{V^{\textrm{tmp}}}-\overline{\bf P}_{V^{\textrm{lat}}}, where P‾Vtmp∗\overline{\bf P}^{*}_{V^{\textrm{tmp}}} denotes the average position among all ground-truth parts that are annotated for VtmpV^{\textrm{tmp}}. As a result, for each latent pattern VlatV^{\textrm{lat}}, we only need to learn its conv-slice DVlat∈θD_{V^{\textrm{lat}}}\in{\boldsymbol{\theta}} and central position P‾Vlat∈θ\overline{\bf P}_{V^{\textrm{lat}}}\in{\boldsymbol{\theta}}.

Scores of terminal nodes

The inference score for each terminal node VuntV^{\textrm{unt}} under a latent pattern VlatV^{\textrm{lat}} is formulated as

The score of SI(Vunt)S_{I}(V^{\textrm{unt}}) consists of the following three terms: 1) SIrsp(Vunt)S_{I}^{\textrm{rsp}}(V^{\textrm{unt}}) denotes the response value of the unit VuntV^{\textrm{unt}}, when we input image II into the CNN. X(Vunt)X(V^{\textrm{unt}}) denotes the normalized response value of VuntV^{\textrm{unt}}; Snone=−3S_{none}=-3 is set for non-activated units. 2) When the parent VlatV^{\textrm{lat}} selects VuntV^{\textrm{unt}} as its location inference (i.e. Λ^Vlat←ΛVunt\hat{\Lambda}_{V^{\textrm{lat}}}\leftarrow\Lambda_{V^{\textrm{unt}}}), SIloc(Vunt)S_{I}^{\textrm{loc}}(V^{\textrm{unt}}) measures the deformation level between VuntV^{\textrm{unt}}’s location PVunt{\bf P}_{V^{\textrm{unt}}} and VlatV^{\textrm{lat}}’s ideal location P‾Vlat\overline{\bf P}_{V^{\textrm{lat}}}. 3) SIpair(Vunt)S_{I}^{\textrm{pair}}(V^{\textrm{unt}}) indicates the spatial compatibility between neighboring latent patterns: we model the pairwise spatial relationship between latent patterns in the upper conv-layer and those in the current conv-layer. For each VuntV^{\textrm{unt}} (with its parent VlatV^{\textrm{lat}}) in conv-layer LVlatL_{V^{\textrm{lat}}}, we select 15 nearest latent patterns in conv-layer LVlat+1L_{V^{\textrm{lat}}}+1, w.r.t. ∥P‾Vlat−P‾Vupperlat∥\|\overline{\bf P}_{V^{\textrm{lat}}}-\overline{\bf P}_{V^{\textrm{lat}}_{\textrm{upper}}}\|, as the neighboring latent patterns. We set constant weights λrsp=1.5,λloc=1/3,λpair=10.0\lambda^{\textrm{rsp}}=1.5,\lambda^{\textrm{loc}}=1/3,\lambda^{\textrm{pair}}=10.0, λunsup=5.0\lambda^{\textrm{unsup}}=5.0, and λclose=0.4\lambda^{\textrm{close}}=0.4 for all categories. Based on the above design, we first infer latent patterns corresponding to high conv-layers, and use the inference results to select units in low conv-layers.

Scores of AND nodes

where we set d ⁣= ⁣37d\!=\!37 pxls and λinf=5.0\lambda^{\textrm{inf}}=5.0.

References