Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

Si Liu, Xiaodan Liang, Luoqi Liu, Xiaohui Shen, Jianchao Yang, Changsheng Xu, Liang Lin, Xiaochun Cao, Shuicheng Yan

Introduction

Human parsing, namely partitioning the human body into several semantic regions (e.g., hat, left/right leg, glasses and upper-body clothes), has drawn much attention in recent years and serves as the basis for many high-level applications, such as clothing classification and retrieval .

Several parametric and non-parametric human parsing methods are proposed and show very promising performances on the human parsing task. On one hand, parametric methods learn knowledge, such as the appearance of regions of different semantic labels and the structural relationship among different labels, from annotated images. Those methods usually rely on the manually designed structural models , which may not fit specific data well and thus achieve only suboptimal performance. Moreover, for another set of new training data and semantic labels, new models have to be designed/retrained, which makes those parametric models impractical because new clothing styles may come out quite often. On the other hand, the non-parametric methods can flexibly use the newly annotated images on the fly and address the issues of parametric model, which are more appealing for practical applications . This kind of methods usually firstly build the pixel-level , superpixel-level or hypothesis-level matching between a testing image and the annotated images in a corpus, then transfer the labels from the manually annotated images to the testing image based on the matching outputs, and finally fuse the transferred labels by heuristic aggregation schemes (typically majority voting). However, the quality of matching is usually limited by the lack of explicit semantic meaning of the bottom-up superpixels or hypotheses.

The above-mentioned parametric and non-parametric human parsing methods rely on the hand designed pipelines composed of multiple sequential components, e.g., hand-crafted feature extraction, bottom-up over-segmentation, human pose estimation, manually designed complex model structure. Therefore, the possibly bad performance of each component may become the bottleneck of the performance of the whole pipeline. For example, the human pose estimation, an important component in the above pipeline, itself is quite a challenging task. Such a sequential processing strategy usually makes the whole pipeline mostly suboptimal. Instead of a combination of multiple sequential steps, several Convolutional Neural Network (CNN) based methods are proposed for an end-to-end image parsing . However, these deep models cannot be easily updated to incorporate new semantic labels.

To address these issues, we propose a quasi-parametric human parsing framework, which inherits the merits of both the parametric and non-parametric models. The proposed end-to-end framework is able to take full advantage of the supervision information from annotation training data, and meanwhile is easy to extend for new added labels. The core part of the proposed framework is a specially designed Matching Convolutional Neural Network (M-CNN) to match any semantic region of a KNN image (also denoted as KNN region in this paper) to the testing image.

As shown in Fig. 1, we first apply the human detection to a testing image and obtain the human centric image. Then the K Nearest Neighbours (KNN) images of the test image is retrieved from the annotated/manually-parsed image corpus. Each KNN image derives several semantic regions, which are generated by masking out the background region with the mean image of the image corpus. Three KNN regions for hat, skirt and pants are shown in Fig. 1. Then, the pair of the test image and each of the KNN regions is fed into the proposed M-CNN to estimate their matching confidence and displacements. The matching confidence measures how the KNN region matches the input image while the displacements describe the coordinate translations between the KNN region and the matched region in the testing image. The matching confidences are then averaged over all KNN regions and thresholded to predict whether a specific label is present. For the labels predicted to be present, such as hat and skirt in Fig. 1, the corresponding label maps can be transferred from the KNN regions to the testing image based on the estimated displacements. For the labels predicted as invisible, such as pants in Fig. 1, no transferred label map is generated. Then, all matched regions of a specific label are combined to produce a probability map. Finally, the probability maps for all labels are refined by a superpixel smoothing step to get the final parsing result.

Reliable matching between an input image and a KNN region is challenging, because the matching needs to handle the large spatial variance of semantic regions. For example, the bags can be placed on the left, right or in front of the human body. The proposed M-CNN is able achieve accurate multi-ranged matching. As shown in Fig. 2, M-CNN contains three paths, i.e., two single image convolutional paths and a cross image convolutional path. The single image convolutional path receives the input image or a particular KNN region, and produces its discriminative hierarchical feature representations layer by layer. The cross image convolutional path embeds cross image filters into every convolutional layer to characterize the multi-ranged matching. The cross image filters are applied to all feature maps in previous convolutional layers, including the single image feature maps and cross image feature maps. Because the scale of receptive fields of the feature maps increase when tracing up the M-CNN, the cross image matching filters capture the displacements from the near-range to the far-range. Therefore the feature maps from the cross image convolutional path can well represent the displacements. Because the feature maps generated by the two single image convolutional paths are excellent feature representations, their absolute difference maps are calculated as another measurement of the displacements. The difference maps are combined with the cross image feature maps and then link to the subsequent fully connected layers. Finally, the matching confidence and displacements are regressed. Since the M-CNN targets at matching an input image and any KNN region of any semantic label, it can work even if new semantic labels are included. Instead of training a M-CNN for each label, we train a unified M-CNN for all KNN regions of all labels.

Comprehensive evaluations over a large dataset with 7,7007,700 annotated human images well demonstrate the effectiveness of our quasi-parametric framework. The major contributions are summarized as follows:

We build a novel deep quasi-parametric human parsing framework. It can learn from annotated data and also flexibly use newly annotated (possibly uncommon) images.

We propose a Matching Convolutional Neural Network (M-CNN) to match a semantic region of a KNN image to a testing image. The novel cross image filters are embedded into different convolutional layers, each aiming to capture a particular range of displacements.

We integrate all the step-by-step components (over-segmentation, pose estimation, feature extraction, label modeling, etc.) in traditional pipelines into one unified end-to-end deep CNN framework.

Related Work

In this section, we sequentially review the parametric human parsing methods, non-parametric methods and deep learning based methods.

For parametric human parsing, Yamaguchi et al. proposed to boost the human parsing with pixel-level classification by relying on human pose estimation. To capture more complex contextual information, Dong et al. designed an And-or Graph structure to model the correlations of a group of parselets, and their extension work unified the human parsing and pose estimation in one framework. The image co-segmentation and region co-labeling for human parsing were also used to capture the correlations between different human images . In addition, Liu et al. utilized user-generated category tags to build a human parser. For general image parsing, Tighe et al. proposed a segmentation by detection approach. Firstly, the bounding boxes of the objects are estimated by exemplar SVM , based on which the segmentation masks are transferred from the image corpus to the input image. In general, the power of existing parametric methods is largely limited by the suboptimal performance of many hand designed intermediate components, such as pose estimation, and also cannot be easily extended to parse new labels.

In non-parametric human parsing, pixels , superpixels and object proposals were used to facilitate non-parametric image parsing. Specifically, the model of Yamaguchi et al. transferred parsing masks from retrieved examples to the query image. Their label transferring is based on superpixels, which are generated by over-segmentation based on the low level appearance cues and therefore lack semantic meaning. Liu et al. used SIFT Flows to build the pixel-pixel correspondence and the dense deformation field between images. However, the optimization problem for finding the SIFT flow is rather complex and expensive to solve. Recently, Long et al. proved the better performance of convolutional activation features over traditional features, such as SIFT, for tasks requiring correspondence. Overall, the non-parametric methods are limited by the inaccurate matching, which results in the noises/outliers during the label transferring.

Our quasi-parametric model integrates the advantages of both parametric models and non-parametric models by the proposed M-CNN. There exist some works on semantic segmentation with CNN architectures. Girshick et al. and its extension work proposed to classify the candidate regions by CNN for semantic segmentation. Wang et al. presented a joint task learning framework, in which the object localization task and the object segmentation task are tackled collaboratively via CNN. Farabet et al. trained a multi-scale CNN from raw pixels to extract deep features for assigning the label to each pixel. The recurrent CNN was proposed to speed up scene parsing and achieved the state-of-the-art performance. Our M-CNN inherits the merit of existing CNN parsing models in our single image convolutional path. It differs from all existing CNN based parsing models in that we handle a pair of images instead of a single image and we incorporate cross image filters to specifically characterize multi-ranged matching. Last but not least, M-CNN can effortlessly handle new semantic labels.

Quasi-parametric Human Parsing

For each human image, we first retrieve its KNN images from the annotated image corpus (Sec. 3.1). Then, M-CNN predicts the matching confidence and displacements between the input image and a semantic region from one KNN image, based on which a label map is generated (Sec. 3.2). Finally, all label maps are fed into a post processing procedure to produce the parsing result (Sec. 3.3).

For each of the input images, we use the human detection algorithm to detect the human body. The resulting human centric image II is then rescaled to 227×227×3227\times 227\times 3. We then extract a global 4,0964,096-dimensional feature from the penultimate fully-connected layer in the pre-trained CNN model trained on ILSVRC 2012 classification dataset based on the Krizhevsky architecture . Its KNN images G={g1,g2,...,gK}G=\left\{{{g_{1}},{g_{2}},...,{g_{K}}}\right\} are retrieved from the image corpus based on the deep features.

2 Matching Convolutional Neural Network

Input, Output and Loss Function: Given a label l∈{1,...,L}l\in\{1,...,L\}, where LL is the total number of labels, the input image II and each KNN regions gklg_{kl} from KNN image gkg_{k} form a pair and are fed into the M-CNN. Note that gkg_{k} is from the image corpus and thus its label map is known. gklg_{kl} is generated by keeping the regions of the label ll in gkg_{k} and masking out other regions by the mean image calculated from the image corpus. If gkg_{k} does not contain label ll, gklg_{kl} is exactly the mean image.

where φ(⋅)\varphi\left(\cdot\right) is the label set contained in the specific image. The first term in Eq. (1) is the loss for matching confidence while the second term corresponds to the loss for displacements. We penalize the displacements loss when both the KNN image gkg_{k} and the input image II contain the label ll. Then the losses of all training pairs are summed and the parameters are learned by back propagation.

Architecture: Since the KNN images are retrieved based on the global appearance similarity, the KNN region of each label may locate quite differently in images. For example, in Fig. 3, the bags can be placed on the left side or right side or in front of the human body. M-CNN is designed to estimate the multi-ranged matching by embedding the cross image matching filters in different convolutional layers. As shown in Fig. 2, M-CNN contains two kinds of paths, i.e., two single image convolutional paths and one cross image convolutional path the outputs of these three paths are further fused to estimate the matching confidence and displacements. The single image path aims for hierarchical feature representation while the cross image convolutional path estimates the displacements between the input pair.

Single Image Convolutional Path: We have two instantiations of the single image path in the top and bottom row of Fig. 2, each of which separately processes II or gklg_{kl}. They share the same architecture and extract the hierchical feature representations of II or gklg_{kl}. The outputs are their respective feature maps in “conv5”. In this path, the single image filters of the next convolutional layers are connected to those feature maps in the previous layer, shown as the green dashed line in Fig. 2. The ReLU non-linearity is applied to the output of every convolutional layer. The sizes of feature maps are gradually reduced by using the stride of 22 for all the convolutional layers. The most important difference between M-CNN and the infrastructure in is that M-CNN removes the pooling layer. Although pooling is useful for enhancing translation invariance for object recognition, it loses precise spatial information that is necessary for accurately predicting the locations of the labels . The details about the network parameters, such as image/feature map sizes, kernel size/numbers are shown in Fig. 2. The powerful representation capability of the single image convolutional path lays the foundation for the accurate estimation of the matching confidence and displacements.

Cross Image Convolutional Path: The cross image convolutional path lies in the middle row of Fig. 2. It outputs the cross image feature maps in “conv5” layer. The mm-th cross image feature map in the jj-th layer xj,mC{x_{j,m}^{C}} is generated by convolving the corresponding matching filter (including three components fj,p,mI{f_{j,p,m}^{I}}, fj,q,mR{f_{j,q,m}^{R}} and fj,t,mC{f_{j,t,m}^{C}}) with both singe image (xj−1,pI{x_{j-1,p}^{{I}}}, xj−1,qR{x_{j-1,q}^{{R}}}) and cross image feature maps (xj−1,tC{x_{j-1,t}^{{C}}}) in the j−1j-1-th layer. The fj,p,mI{f_{j,p,m}^{I}} component links the pp-th (out of all PP) input image feature map xj−1,pI{x_{j-1,p}^{{I}}} in the j−1j-1-th layer to xj,mC{x_{j,m}^{C}}. Analogously, the fj,q,mR{f_{j,q,m}^{R}} component links the qq-th (out of all QQ) KNN region feature map xj−1,qR{x_{j-1,q}^{{R}}} to xj,mC{x_{j,m}^{C}}. Moreover, the fj,t,mC{f_{j,t,m}^{C}} component links the tt-th (out of all TT) cross image feature map xj−1,tC{x_{j-1,t}^{{C}}} to xj,mC{x_{j,m}^{C}}. The operation of the matching filters from one layer to the next is shown as the purple dashed line in Fig. 2. Mathematically, the cross feature map xj,mC{x_{j,m}^{C}} is calculated by:

where ∗* denotes convolution and bj,m{b_{j,m}} is the bias for the mm-th output map. max⁡(0,⋅)\max(0,\cdot) is the non-linear activation function, and is operated element-wisely. From Eq. (2), we can see that the cross image feature map in the next layer is calculated by considering both single and cross image feature maps in the previous layer, and thus the displacements between the input image II and KNN region gklg_{kl} can be effectively estimated. Note that along with the M-CNN, the receptive fields of different layers of the single image and cross image convolutional path gradually increase . In this way, multi-ranged matchings can be achieved.

Finally, we fuse the feature maps from two single image paths and one cross image path. More specifically, the absolute differences of the feature maps of the input image and the KNN region (from the single image convolutional path) are first calculated and then are stacked with the output of the cross image convolutional path. Our fusion is applied on the feature maps. It is different from “Siamese” architecture which calculates the absolute differences of the fully-connected representations. Experiments show the our fusion strategy outperforms “Siamese” by keeping more spatial information.

3 Post Processing

Given the matching confidences and displacements estimated by M-CNN, the parsing result of the input image can be calculated as follows. Firstly, the confidence of II containing the ll-th label is calculated by averaging matching confidence ckl{c_{kl}} for all KNN regions satisfying l∈φ(gk)l\in\varphi\left({{g_{k}}}\right). If the confidence is greater than a threshold ξ1{\xi_{1}}, the label ll is predicted as visible in the input image, otherwise predicted as invisible. Secondly, we estimate the locations of the visible labels. More specifically, the coordinates of the matching region in the input image II is calculated based on the matching displacements tkl{t_{kl}} and the ground-truth coordinates of gklg_{kl}. Then, we morph the associated ground-truth label mask of gklg_{kl} into matched region in II. In this way, we get a probability map MlM_{l} of II for each label l∈[1,L]l\in[1,L]. We pixel-wisely max all MlM_{l} for all the labels and get the foreground probability. The pixels with the probability larger than a threshold ξ2{\xi_{2}} are regarded as the rough foreground, while the remaining are the rough background. The rough foreground and background are further eroded by a filter size 1010 to produce the final foreground and background seeds. Based on the seeds, we can obtain the background probability by the algorithm . The obtained background probability is combined with the foreground probability map MlM_{l}, l∈[1,L]l\in[1,L], based on which, we can get an initial human parsing results with the pixel-wise Maximum a Posterior Probability (MAP) assignment. Finally, to respect boundaries of actual semantic labels, we further over-segment II using the entropy rate based segmentation algorithm and assign the label of the superpixel by the majority of its covered pixels’ initial parsing results.

Experiments

Datasets: We use the dataset in pixel-wisely labeled by the 1818 categories defined by Daily Photos dataset . The dataset contains 7,7007,700 images (6,0006,000 for training, 700700 for validation and 1,0001,000 for testing). We adopt four evaluation metrics, i.e., accuracy, average precision, average recall, and average F-1 scores over pixels .

Training Image Pairs Generation: To reduce over-fitting in the model training and partially address the detection error, we enlarge the cropped human centric images region by 11 and 1.21.2 times. We also horizontally mirror the images. In short, each image has 44 variations and training data can be greatly augmented.

For each of 6,0006,000 training images, 5050 KNN images are retrieved from the image corpus. Each training image and each KNN region form a training pair. After unevenly sampling of the training pairs to balance different labels, we finally have 55 million pairs, which even outnumbers the that of ILSVRC2012 . We shuffle the training pair in order to increase the diversity of each epoch.

Implementation Details: We implement the M-CNN under the Caffe framework and train it using stochastic gradient descent with a batch size of 128128 examples, momentum of 0.90.9, and weight decay of 0.00050.0005. We use an equal learning rate for all layers. The learning rate is adjusted manually by dividing 1010 when the validation error rate stops decreasing with the current learning rate. The learning rate is initialized at 0.00050.0005. We train M-CNN for roughly 5050 epochs, which takes 1111 to 1212 days on one NVIDIA GTX TITAN 6GB GPU. In the training phase, we first calculate the element-wise mean and variance of the matching confidence and displacements of the whole image corpus, and element-wisely normalize the training output by the mean and variance. In the testing phase, we project the matching confidence and displacements estimated by M-CNN to their absolute values by the mean and variance. In the post processing step, the thresholds ξ1{\xi_{1}} and ξ2{\xi_{2}} are set to be 0.80.8 and 0.50.5. The number of KNN regions for each input image is set as 99.

2 Results and Analysis

Comparison with The State-of-the-arts: We compare our M-CNN based quasi-parametric human parsing framework with two state-of-the-arts: Yamaguchi et al. and PaperDoll . We use their publicly available codes and train their models with the same 6,0006,000 training images as our method for fair comparison. We do not compare with Dong et al. because their code is not publicly available and their method is reported to be slower than ours.

The average results for all labels are in Table 1. The methods of Yamaguchi et al. and the PaperDoll trained on the same 6,0006,000 training images and tested on the same 1,0001,000 images as M-CNN, and their average F1F_{1} scores achieve 41.80%41.80\% and 44.76%44.76\%. Our “M-CNN” significantly outperforms these two baselines by over 21.01%21.01\% for Yamaguchi et al. and 18.05%18.05\% for PaperDoll . “M-CNN” also gives a huge boost in foreground accuracy: the two baselines achieve 55.59%55.59\% for Yamaguchi et al. and 62.18%62.18\% for PaperDoll while “M-CNN” obtains 73.98%73.98\%. “M-CNN” also obtains much higher precision (64.56% vs 37.54% for and 52.75% for ) as well as higher recall (65.17% vs 51.05% for and 49.43% for ). This verifies the effectiveness of our end-to-end M-CNN based quasi-parametric framework.

We also present the F1-scores for each label in Table 2. Generally, “M-CNN” shows much higher performance than the baselines. In terms of predicting labels for small semantic regions such as hat, belt, bag and scarf, our method achieves a large gain, e.g. 43.38% vs 11.43% and 2.95% for scarf, 57.87% vs 24.53% , 30.52% for bag and 38.45% vs 14.68% and 16.94% for belt. It demonstrates that our quasi-parametric network can effectively capture the internal relations between the labels and robustly predict the label masks with various clothing styles and poses.

Ablation of Our Networks: We also extensively explore different CNN architectures to demonstrate the effectiveness of each component in M-CNN more transparently. The architecture of M-CNN is shown in Fig. 2 and other variants are constructed by gradually adding/eliminating the cross image filters in different layers. M-CNN contains 44 cross image matching filters from layers conv2conv2 to conv5conv5. “M-CNN (cross 5,4,3,2,1)” is obtained by adding 11×11×611\times 11\times 6 cross image filters in the “conv1” layer to the “M-CNN”. The added matching filters are applied on the stacked image composed of the RGB channels of input image and KNN region. In addition, we continue to remove the cross image matching filters layer by layer, producing “M-CNN (cross 5,4,3)”, “M-CNN (cross 5,4)” , “M-CNN (cross 5)” and “M-CNN(w/o cross)”. Note that no matching filters are used in the “M-CNN(w/o cross)” architecture, where “w/o” stands for without. For fair comparison, we keep the number of feature maps of each convolutinal layer unchanged for different M-CNN variations. Therefore, the number of removed cross image filters is evenly added to the corresponding two single image layers. For example, the number of feature maps for both single and cross image convolutional paths are all 3030 in “conv2”. After we remove the cross matching filters to derive “M-CNN (cross 5,4,3)”, “M-CNN (cross 5,4)” and “M-CNN (cross 5)”, the number of feature maps in the single image convolutional path is set as 4545. In addition, we compare with a classic CNN based verification architecture named “Siamese” which is a composite structure of two identical sub-networks. The outputs of the two sub-networks are fully-connected linear layer responses, whose absolute differences are calculated as features to fit the final matching confidence and displacements. Finally, we compare with“M-CNN w/o ss ” which is the same with M-CNN except that the superpixel smoothing processing step is skipped.

The average scores over all labels in Table 1 offer following observations. Incrementally adding cross image matching filters into more convolutional layers produces 4 variations of “M-CNN”, including “M-CNN (w/o cross)”, “M-CNN (cross 5)”, “M-CNN (cross 5,4)”, “M-CNN (cross 5,4,3)” and “M-CNN”. Their F1F_{1} scores increase from 56.99%56.99\%, 58.07%58.07\%, 60.36%60.36\% to 62.81%62.81\%. The highest F1F_{1} score is 62.81%62.81\% which is reached by “M-CNN”. The gradually improving performance validates that inserting more cross image matching into multiple convolutional layers can help achieve better matching. Adding a cross image matching kernel in the “conv1” layer drops the F1F_{1} score from 62.81%62.81\% to 61.53%61.53\%. The reason for the relatively lower result is that the receptive field corresponding to the first cross image matching kernel is small and only involves certain part of a semantic label. However the targets of this work, i.e., the matching confidence and displacements, are defined on the semantic label level and thus are beyond the receptive field of the cross image matching kernel inserted in the “conv1” layer. When the M-CNN grows deeper, the receptive fields become much larger, with a higher probability to cover the semantic label, which can faciliate estimating the semantic label-level displacements. In addition, “M-CNN (w/o cross)” performs better than “Siamese” because our ultimate task is to estimate displacements. “M-CNN (w/o cross)” calculates the difference between two “conv5” layer feature maps while “Siamese” calculates the differences of two fully connected features, where spatial structures in the 22-dim images are partially lost. The lower F1F_{1} score of “M-CNN w/o ss” compared with “M-CNN” proves that the adopted superpixel smoothing technique can better preserve the boundary information although it is a simple and fast voting of pixels’ labels. The superior performance of “M-CNN w/o ss” than the state-of-the-arts shows that our M-CNN has the capability of directly predicting more reliable label masks even without the post processing step.

Sensitivity to the Number of KK: In Fig. 5, we report the performance of our human parsing method with respect to different numbers of KNN images. We find that M-CNN reaches the highest F1F_{1} score 63.58%63.58\% when 99 KNN regions are considered. When only 11 KNN region is considered, the performance is still quite competitive (56.92%56.92\%).

Qualitative Parsing Results Comparison: Fig. 4 shows the comparison between M-CNN and PaperDoll . The results demonstrate that our method can successfully predict the label maps with small regions, which can be attributed to the reliable label transferring from the KNN regions. For example, the bag, scarf and hat in three images in the first column are successfully located by M-CNN but totally missed by PaperDoll. Another example is that M-CNN successfully finds the small sunglasses in the first row, which are missed by PaperDoll. In addition, it can be observed that the results from M-CNN are robust to pose variations. As shown in the bottom row, M-CNN can accurately estimate the locations of left and right arms, while PaperDoll cannot. The superior performance is because our method is an end-to-end model while PaperDoll relies on a separate pose estimation preprocessing step. Another observation of Fig. 4 is that our segmented regions are more complete while the PaperDoll regions are fragmented, such as the lower right result. That is because PaperDoll transfers labels based on oversegments which lack explicit semantic meaning.

Conclusion and Future Work

In this work, we tackle the human parsing problem by proposing a new quasi-parametric model. Our unified end-to-end quasi-parametric framework inherits the merits of both parametric and non-parametric parsing methodologies. It takes advantage of the supervision from annotated data and can be easily extended to newly annotated images and labels. To characterize the multi-ranged matching, we propose a Matching Convolutional Neural Network, which contains two single image convolutional paths for better feature representation and a cross image convolutional path where cross image matching filters are embedded into the convolutional layers. Extensive experimental results clearly demonstrate significant performance gain from the quasi-parametric model over the state-of-the-arts. In the future, we will extend the framework to other exemplar-based tasks, such as face parsing. Moreover, we plan to use other more power network structure, e.g, GoogLeNet .

Acknowledgement

This work is supported by National Natural Science Foundation of China (No.61422213, 61332012, 61328205), and 100 Talents Programme of The Chinese Academy of Sciences. This work is partly supported by gift funds from Adobe, National High-tech R&\&D Program of China (Y2W0012102Y2W0012102), the Hi-Tech Research and Development Program of China (no.2013AA013801), Guangdong Natural Science Foundation (no.S2013050014548), and Program of Guangzhou Zhujiang Star of Science and Technology (no.2013J2200067).

References