Deep Semantic Matching with Foreground Detection and Cycle-Consistency

Yun-Chun Chen, Po-Hsiang Huang, Li-Yu Yu, Jia-Bin Huang, Ming-Hsuan Yang, Yen-Yu Lin

Introduction

Dense correspondence matching is an important and active research topic in computer vision. Optical flow estimation and stereo matching aim to estimate per-pixel correspondences to match across images depicting the same scene or object instance. While correspondence estimation has been extensively studied, there has been a growing trend to extend the idea of matching the same objects across images to matching images covering different instances of an object category. This progress not only attracts substantial attention but also facilitates many real-world applications ranging from object recognition , object co-segmentation , to 3D reconstruction . However, due to the presence of background clutter, ambiguity induced by large intra-class variations, and the limited scalability of obtaining large-scale datasets with manually annotated correspondences, semantic matching remains challenging.

Conventional methods for semantic matching rely on hand-crafted descriptors such as SIFT or HOG as well as an effective geometric regularizer. However, these hand-crafted descriptors cannot be adapted to the given visual domains, leading to a sub-optimal performance of semantic matching. Driven by the recent success of convolutional neural networks (CNNs), several learning-based approaches have been proposed for tackling the problem of semantic matching . While promising results have been shown, these approaches still suffer from the following limitations. The methods in require a vast amount of supervised data for training the network. Collecting large-scale and diverse data, however, is expensive and labor-intensive. While weakly supervised methods such as have been recently proposed to relax the issue, these approaches implicitly match the background features from both images to be similar. Thus, they often suffer from the unfavorable effect of background clutter.

In this paper, we address these challenges by performing foreground detection and enforcing cycle consistency constraints in semantic matching. To suppress the negative impacts caused by background clutter, we develop a foreground detection module that allows the model to exclude background regions and focus on matching the detected foreground regions. Therefore, the effect of background clutter can be alleviated. To address the matching difficulties caused by complex appearance and large intra-class variations, we focus on filtering out geometrically inconsistent correspondences. Our key insight is that correct correspondence should be cycle-consistent meaning that when matching a particular point from one image to the other and then performing reverse matching, we should arrive at the same spot. To exploit this property, we introduce a cycle-consistency loss that provides additional supervisory signals for network training. We further extend this idea to explore transitivity consistency across multiple images. We build upon the model by Rocco et al. for a weakly-supervised and end-to-end trainable network and evaluate the effectiveness of the proposed approach on three benchmarks. Experimental results demonstrate that our approach improves the baseline model , as shown in Fig. 1, and performs favorably against the state-of-the-art methods.

Our contributions are summarized as follows. First, we present a weakly-supervised learning framework that integrates foreground detection into semantic matching. With a module for explicit foreground detection, the proposed network suppresses the unfavorable effect of background clutter. Second, our model implicitly tackles the ambiguity induced by a vast matching space via inferring bi-directional geometric transformations during matching. With these transformations, we explicitly enforce the inferred geometric transformations to be cycle-consistent by introducing the forward-backward consistency loss. In addition, we explore the property of transitivity consistency and introduce the transitivity consistency loss to further enhance the matching performance. We train our network with the image pairs of the PF-PASCAL dataset . We then evaluate the proposed model on several benchmark datasets for semantic matching, including the PF-PASCAL , PF-WILLOW , and TSS datasets. Extensive comparisons with existing semantic matching algorithms demonstrate that the proposed approach achieves the state-of-the-art performance.

Related Work

Semantic matching has been extensively studied in the literature. Here, we review several related topics.

Conventional methods for semantic matching leverage hand-crafted descriptors such as SIFT or HOG along with geometric matching models. These methods find keypoint correspondences across images through energy minimization. The SIFT Flow method aligns two images with SIFT features using a similar formulation as an optical flow algorithm. Kim et al. compute dense correspondence efficiently using the deformable spatial pyramid. Ham et al. use the object proposals as the matching primitives and leverage the HOG descriptor to establish semantic correspondence. With the use of object proposals, the Proposal Flow method is robust to scaling and background clutter. Taniai et al. propose a hierarchical Markov random field model to jointly perform object co-segmentation and dense correspondence. However, the aforementioned methods adopt hand-crafted descriptors, which are pre-defined and not optimized for the given images.

Convolutional neural networks have been successfully applied to semantic matching. Choy et al. propose the universal correspondence network (UCN) and a correspondence contrastive loss for network training. The UCN method adopts a convolutional spatial transformer for feature transformations, making their method robust to scaling and rotations. Kim et al. propose the fully convolutional self-similarity (FCSS) descriptor and integrate the descriptor into the Proposal Flow framework for image matching. The SCNet method learns a geometrically plausible model for semantic correspondence by incorporating geometric consistency constraints into its loss function. While the methods in employ trainable descriptors for semantic correspondence, the feature matching is learned at the object-proposal level. Consequently, these methods are not end-to-end trainable since a fusion step is required to produce the final results. Rocco et al. present an end-to-end trainable CNN architecture for estimating parametric geometric transformations. While these methods perform better than those based on hand-crafted features, the dependence on supervised training data with manually labeled keypoint correspondences limits the applicability.

Several recent CNN-based methods have carried out weakly supervised semantic correspondence. The AnchorNet learns a set of filters whose response is geometrically consistent across different object instances. The AnchorNet model, however, is not end-to-end trainable due to the use of the hand-engineered alignment model. The WarpNet learns fine-grained image matching with small-scale and pose variations via aligning objects across images through known deformation. Inspired by the inlier scoring procedure of RANSAC, Rocco et al. propose an end-to-end trainable alignment network which computes dense semantic correspondence while aligning two images.

Our proposed method differs from these methods in two aspects. First, our approach further takes into account foreground detection. Our network learns feature embedding to enhance inter-image foreground similarity while alleviating the unfavorable effects caused by complex background. Second, our model simultaneously infers bi-directional transformations. We explicitly enforce cycle-consistent constraints on the predicted transformations, resulting in more accurate and consistent matching results.

Exploiting cycle consistency property to regularize learning has been extensively studied. In the context of motion analysis, computing bi-directional optical flow has been shown to be useful to reason about occlusion for learning optical flow and enforcing temporal consistency . In the context of image-to-image translation, enforcing cycle consistency enables learning mapping between domains with unpaired data . In the context of unsupervised domain adaptation, exploiting cycle consistency constraints allows the model to produce consistent task predictions across domains . In the context of visual recognition , enforcing consistency constraints allows the model to be more robust to resolution variations . Several methods exploit the idea of cycle consistency for semantic matching. Zhou et al. tackle the problem of matching multiple images by jointly optimizing feature matching and enforcing cycle consistency. The FlowWeb method learns image alignment by establishing globally-consistent dense correspondences with cycle consistency constraints. However, these methods employ hand-engineered descriptors which cannot adapt to an arbitrary object category given for matching. Zhou et al. establish dense correspondences by using an additional 3D CAD model to form a cross-instance loop between synthetic data and real images. However, the cycle consistency loss in requires four images at a time. In contrast, we develop two loss functions to enforce cycle consistency and do not need additional data to guide the training. Experimental results demonstrate that by exploiting cycle consistency constraints, the proposed method produces consistent matching results and improves the performance.

Proposed Algorithm

In this section, we first provide an overview of our approach. We then describe each loss in our objective function in detail and the implementation details.

Let D={Ii}i=1N\mathcal{D}=\{I_{i}\}_{i=1}^{N} denote a set of images consists of instances of the same object category, where IiI_{i} is the ithi^{th} image and NN is the number of images. Our goal is to learn a CNN-based model that can estimate the keypoint correspondences between each image pair (IA,IB)(I_{A},I_{B}) in D\mathcal{D} without knowing the object class in advance. Our formulation for semantic matching is weakly-supervised since training our model requires only weak image-level supervision in the form of training image pairs containing common objects. No ground truth keypoint correspondences are used.

To accomplish this task, we present an end-to-end trainable network which is composed of two modules: 1) the feature extractor F\mathcal{F} and 2) the transformation predictor G\mathcal{G}. The feature extractor F\mathcal{F} extracts features for each image in a given image pair. The transformation predictor G\mathcal{G} predicts the transformation that warps an image so that the warped image can better align the other image.

2 Objective function

where λC\lambda_{C} and λT\lambda_{T} are hyper-parameters used to control the relative importance of the respective loss functions. Below we outline the details of each loss function.

3 Foreground-guided matching loss

Note that both the correlation maps SABS_{AB} and SBAS_{BA} are compiled through a rectified linear unit (ReLU) to eliminate negative matching values in advance. Therefore, the value of the estimated foreground masks at each pixel will be bounded between and 11. Intuitively, the mask MA(p)M_{A}(\mathbf{p}) has a low value (i.e., location p\mathbf{p} is likely to belong to background) if none of the feature vectors in fBf_{B} matches well with fA(p)f_{A}(\mathbf{p}). The mask MBM_{B} can be obtained following a similar procedure.

With the estimated geometric transformation TABT_{AB}, we can identify and filter out geometrically inconsistent correspondences. Consider a correspondence with endpoints (p∈PA,q∈PB)(\mathbf{p}\in\mathcal{P}_{A},\mathbf{q}\in\mathcal{P}_{B}), where PA\mathcal{P}_{A} and PB\mathcal{P}_{B} are the sets of all spatial coordinates of fAf_{A} and fBf_{B}, respectively. The distance ∥TAB(p)−q∥\|T_{AB}(\mathbf{p})-\mathbf{q}\| represents the projection error of this correspondence with respect to transformation TABT_{AB}. Following Rocco et al. , we introduce a correspondence mask mAm_{A} to determine if the correspondences are geometrically consistent with transformation TABT_{AB}. Specifically, mAm_{A} is of the form

where φ=1\varphi=1 is the number of pixels.

Given the geometric transformation TABT_{AB} and the correspondence mask mAm_{A}, we compute matching score of each spatial location p∈PA\mathbf{p}\in\mathcal{P}_{A} as

4 Cycle consistency

where ∥TBA(TAB(p))−p∥\|T_{BA}(T_{AB}(\mathbf{p}))-\mathbf{p}\| is the reprojection error between coordinate p\mathbf{p} and the reprojected coordinate TBA(TAB(p))T_{BA}(T_{AB}(\mathbf{\mathbf{p}})).

4.2 Transitivity consistency loss.

5 Network selection and initialization

Experimental Results

Experiments are conducted in this section. Here, we first describe the implementation details and the experimental setting. We evaluate and compare the proposed approach with the state-of-the-art, following analyzing the relative contributions of individual components in the proposed model.

We implement our model using PyTorch. We use the training and validation image pairs from the PF-PASCAL dataset . All images are resized to the resolution of 240×240240\times 240. We perform data augmentation by horizontal flipping, random cropping the input images, and swapping the order of images in the image pair. We train our model using the ADAM optimizer with an initial learning rate of 5×10−85\times 10^{-8}. For transitivity consistency loss, the input triplets are randomly selected within a mini-batch. We sample 10×10=10010\times 10=100 spatial coordinates for computing the forward-backward consistency loss and the transitivity consistency loss. The training process takes about 2 hours on a single NVIDIA GeForce GTX 1080 GPU.

2 Evaluation metric and datasets

We conduct the evaluation on the PF-PASCAL , PF-WILLOW , and TSS benchmark datasets.

We evaluate the performance of the proposed method on a semantic correspondence task. To assess the performance, we adopt the percentage of correct keypoints (PCK) metric which measures the percentage of keypoints whose reprojection errors are below the given threshold. The reprojection error is the Euclidean distance d(ϕ(p),p∗)d(\phi(\mathbf{p}),\mathbf{p}^{*}) between the locations of the warped keypoint ϕ(p)\phi(\mathbf{p}) and the ground truth keypoint p∗\mathbf{p}^{*}. The threshold is defined as τ⋅max⁡(h,w)\tau\cdot\max(h,w) where hh and ww are the height and width of the annotated object bounding box on the image, respectively.

2.2 PF-PASCAL [8].

The PF-PASCAL dataset is selected from the PASCAL 2011 keypoint annotations containing 1,351 semantically related image pairs from 20 object categories. For images of a category, they contain different object instances of that category with similar poses but different appearances. In addition, the presence of background clutter makes it a challenging dataset on semantic matching. We divide the dataset into 735 pairs for training, 308 pairs for validation, and 308 pairs for testing. Manually annotated correspondences are provided for each image pairs. However, under the weakly supervised setting, we do not use the keypoint annotations for training. The annotations are used only for evaluation. We compute the PCK for each object category with τ\tau equals to 0.1.

2.3 PF-WILLOW [8].

The PF-WILLOW dataset is composed of 100 images with 900 image pairs divided into four semantically related subsets: car, duck, motorbike, and wine bottle. Each subset contains images with large intra-class variations and background clutters. For each image, there are 10 keypoint annotations. We follow Han et al. and compute the PCK at three different thresholds with τ\tau equals to 0.05, 0.1, and 0.15, respectively.

2.4 TSS [35].

The TSS dataset comprises 400 semantically related image pairs divided into three groups, including FG3DCar, JODS, and PASCAL. FG3DCar contains 195 image pairs of automobiles. JODS is composed of 81 image pairs of airplanes, cars, and horses. There are 124 image pairs of trains, cars, buses, bikes, and motorbikes form the group of PASCAL. Ground truth flows for each image pair are provided. Following Taniai et al., we compute the PCK over foreground object by setting τ\tau to 0.05.

3 Experimental results on the PF-PASCAL dataset

In the following, we compare the performance of the proposed method with the state-of-the-art approaches. Note that many of the existing methods require manually annotated correspondences while our model can be trained using only image-level supervision.

We compare our method with the Proposal Flow , the UCN , different versions of the SCNet , the CNNGeo with different feature extractors, and a weakly supervised approach proposed by Rocco et al. . Table 1 presents the experimental results for the PF-PASCAL dataset. Our results show that the proposed approach compares favorably against state-of-the-art methods, achieving an overall PCK of 78.0% (outperforming the previous best method by 3.2%). The advantage of incorporating foreground detection and enforcing cycle consistency constraints can be observed by comparing our method with ResNet-101+CNNGeo(W) since both methods utilize the same feature descriptor and are trained with image-level supervision only.

Fig. 3 presents the qualitative results of semantic correspondence on the PF-PASCAL dataset. To further highlight the importance of each component of the proposed method, we present an ablation study of our method.

3.2 Ablation study.

The ablation study shows that all of the proposed components play crucial roles in producing accurate matching results. From Fig. 2, we observe that the proposed method outperforms the best competitor with a significant margin at multiple thresholds.

4 Experimental results on the PF-WILLOW and TSS datasets

To evaluate the generalization capability, we apply the learned model trained on the PF-PASCAL dataset to test directly on the PF-WILLOW and TSS datasets without finetuning on these two datasets.

Table 4 presents the quantitative results for the PF-WILLOW dataset. We compare the performance with several recent methods as well as conventional approaches using hand-crafted features. The results are directly taken from except . For and , we run the code provided by the authors to obtain the results. From Table 4, we observe that our model achieves the state-of-the-art performance with all three thresholds.

4.2 Results on the TSS dataset.

We also evaluate the performance on the TSS dataset. Table 4 presents the quantitative results. We observe that our method achieves the state-of-the-art performance on two of the three groups of the TSS dataset: FG3DCar and JODS. Our results are slightly worse than that in in the PASCAL group. However, the method in uses additional images from the PASCAL VOC 2007 dataset. We report their results for completeness. Under the same experimental settings, the proposed method performs favorably against existing approaches.

Conclusions

In this work, we present an effective approach to improve semantic matching. The core technical novelty of our approach lies in the explicit modeling of a foreground detection module to suppress the effect of background clutter and exploiting the cycle consistency constraints so that the predicted geometric transformations are geometrically plausible and consistent across multiple images. The network training requires only training image pairs with image-level supervision and thus significantly alleviates the cost of constructing and labeling large-scale training datasets. Experimental results demonstrate that our approach performs favorably against existing semantic matching algorithms on several standard benchmarks. Moving forward, we believe that the semantic matching network can be further integrated to other computer vision tasks, e.g., supporting 3D semantic object reconstruction and fine-grained visual recognition.

This work is supported in part by Ministry of Science and Technology under grants MOST 105-2221-E-001-030-MY2 and MOST 107-2628-E-001-005-MY3.

References