SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Juhong Min, Jongmin Lee, Jean Ponce, Minsu Cho

Motivation

The problem of semantic correspondence aims at establishing visual correspondences between images depicting different instances of the same object or scene category . Unlike other conventional problems of visual correspondence such as stereo matching, optical flow, and wide-baseline matching, it inherently involves a variety of intra-class variations, which makes the problem notoriously challenging. With growing interest in semantic correspondence, several annotated benchmarks are now available. Due to the high expense of ground-truth annotations for semantic correspondence, early benchmarks only support indirect evaluation using a surrogate evaluation metric rather than direct matching accuracy. For example, the Caltech-101 dataset in provides binary mask annotations of objects of interest for 1,515 pairs of images and the accuracy of mask transfer is evaluated as a rough approximation to that of matching. Recently, Ham et al. and Taniai et al. have introduced datasets with ground-truth correspondences. Since then, PF-WILLOW and PF-PASCAL have been used for evaluation in many papers. They contain 900 and 1,300 image pairs, respectively, with keypoint annotations for semantic parts.

All previous datasets, however, have several drawbacks: First, the amount of data is not sufficient to train and test a large model. Second, image pairs do not display much variability in viewpoint, scale, occlusion, and truncation. Third, the annotations are often limited to either keypoints or object segmentation masks, which hinders in-depth analysis. Fourth, the datasets have no clear splits for training, validation, and testing. Due to this, recent evaluations in have been done with different dataset splits of PF-PASCAL. Furthermore, the splits are disjoint in terms of image pairs, but not images: some images are shared between training and testing data.

To resolve these issues, we introduce a new dataset, SPair-71k, consisting of total 70,958 pairs of images from PASCAL 3D+ and PASCAL VOC 2012 . The dataset is significantly larger with rich annotations and clearly organized for learning. In particular, several types of useful annotations are available: keypoints of semantic parts, object segmentation masks, bounding boxes, view-point, scale, truncation, and occlusion differences for image pairs, etc. Figure 1 shows the dataset statistics in pie chart forms and sample image pairs with their annotations. Our benchmark is available online at http://cvlab.postech.ac.kr/research/SPair-71k/.

∗This article is extended from section 4 of our recent paper to provide the details of the dataset and more results.

Dataset generation and annotation

We have created the SPair-71k dataset using 1,800 images from 18 categories of PASCAL VOC . Specifically, we extract 1,000 images from 10 rigid categories of PASCAL 3D+ (aeroplane, bike, boat, bottle, bus, car, chair, motorbike, train, tv/monitor) and 800 images of 8 non-rigid categories of PASCAL VOC 2012 (bird, cat, cow, dog, horse, person, potted plant, sheep). Note that we do not use ‘dining table’ and ‘sofa’ categories present in PASCAL VOC as they usually appear as background and their semantic keypoints are too ambiguous to localize properly. The images are selected to cover diverse viewpoints of each category as much as possible. For the selected 1,800 images, we manually annotated keypoints and generate 70,958 pairs of images with pair-level annotations as followsWhen annotating keypoints, we treat ‘bottle’, ‘potted plant’, ‘train’ and ‘tv/monitor’ as flat instances as it is hard to discriminate between front/back and left/right of the instances due to their cylindrical shapes..

Image-level annotations. Keypoints for each object category are carefully selected and annotated according to three keypoint selection criteria: (1) each keypoint should describe an object’s part shared across instances of the same object category, (2) keypoints of an object category should be distinct from each other, and (3) keypoints of an object category should spread over the whole object. The number of selected keypoints varies from 9 (potted plant) to 30 (car) across categories; occluded or truncated keypoints are not annotated, thus varying from 3 to 30 across instances in practice. Azimuths for 10 rigid categories are directly obtained from PASCAL 3D+ and quantized to one of the eight angular bins while azimuth bins for the other 8 non-rigid categories are manually annotated. Bounding box, segmentation mask, truncation and occlusion labels are retrieved from PASCAL VOC 2012 where both truncation and occlusion labels are binary indicators, i.e., whether an instance in the image is truncated (occluded) or not.

Image-level splits. In order to build disjoint splits of image pairs for training, validation, and testing, we first create corresponding splits of images before generating pairs. 100 images of each object category are divided into three splits with an approximate ratio of 5:2:3 so that the images of each split spreads over quantized azimuth values, thus obtaining 997, 322, and 481 images for training, validation, and testing splits, respectively. Figure 2 shows category-wise circular histograms of azimuth values for the splits. Images of each split are then used to generate image pairs with pair-level annotations as follows.

Pair-level annotations. We build pair-level annotations using keypoints, azimuth bins, bounding boxes, truncations, and occlusions annotated in images. Common keypoints in two images are used as keypoints of pair-level annotations, i.e., keypoint correspondences. If there are no common keypoints between the two, the pair is excluded. View-point differences are divided into three levels of ‘easy’, ‘medium’, and ‘hard’; a pair is marked as ‘easy’, ‘medium’, and ‘hard’ if the difference between azimuth bin indexes of the two instances falls in the range of {0, 1}, {2, 3}, and {4}, respectively. Scale differences are labeled levels of ‘easy’, ‘medium’, and ‘hard’; a pair is labeled ‘easy’, ‘medium’, and ‘hard’ if the area ratio of corresponding object bounding boxes falls in the range of [1,2), [2, 4), [4, ∞\infty]. Each pair is also annotated for both truncation and occlusion levels with ‘none’, ‘source only’, ‘target only’, and ‘both’. See Table 1 and 2 for details.

Finally, we obtain the SPair-71k dataset of 70,958 image pairs in total, which consists of 53,340 for training, 5,384 for validation, and 12,234 for testing, respectively.

Baseline results on SPair-71k

We evaluate recent state-of-the-art methods on SPair-71k to provide baseline results for further research. For each method in comparison, we run two versions of each model: an original trained model provided by the authors and a model further finetuned by ourselves using SPair-71k train/val set. The results are shown in Table 3. We fail to successfully train the method of on SPair-71k so that their performances drop when trained. We guess that their original learning objectives for weakly-supervised learning is fragile in presence of large view-point differences as in SPair-71k. We leave this issue for further investigation and will update the results at our benchmark page.

Analysis by variation factors. Each image pair in SPair-71k has annotations of difficulty levels for four variation factors (i.e., view-point, scale, truncation, and occlusion) between corresponding instances of the same category; ‘easy’, ‘medium’, or ‘hard’ is annotated for view-point and scale changes, while ‘none’, ‘source’, ‘target’ or ‘both’ is annotated for truncation and occlusion. PCK analysis of the models using these annotations are summarized in Table 4. The results show that all the models perform better given pairs with less variation, and that view-point and scale changes significantly affect the performances.

Impact of individual variations. The results in Table 4 does not clearly demonstrate an impact of each individual variation because the four types of variations co-exist in a pair and interfere with each other when evaluated. To measure an impact of each variation individually, we need to control the other variations to remain fixed. To this end, we evaluate the performances of different levels of a specific variation factor while fixing the levels of the other variations as easy (view-point and scale) and none (truncation and occlusion). For example, when evaluating the performances varying view-point levels, we only use pairs that are labeled ‘easy’ scale, ‘none’ truncation, and ‘none’ occlusion. The results are summarized in Table 5. It shows that the performances of local-region-matching methods is more robust to view-point variation compared to global-image-alignment methods ; the performance of the image alignment models drops more quickly than those of the region matching ones. In terms of scale changes, truncation, and occlusion, however, we find no significant difference in performance drop between the methods. While both truncation and occlusion clearly degrade the performances, the impacts are less than view-point and scale variations.

Conclusion

In this paper, we have presented a large-scale benchmark dataset, SPair-71k, which consists of 71k image pairs for semantic correspondence. Compared to previous datasets, it contains a significantly large number of image pairs with diverse variations in view-point, scale, truncation and occlusion, thus generalizing the problem of visual correspondence by reflecting real-world scenarios. Moreover, its rich annotations including object bounding boxes, keypoint correspondences, variation factors, azimuths, and object segmentation masks will be useful for future research on semantic correspondence and its joint problems.

Acknowledgements. This work is supported by Samsung Advanced Institute of Technology (SAIT) and Basic Science Research Program (NRF-2017R1E1A1A01077999), and also in part by the Inria/NYU collaboration and the Louis Vuitton/ENS chair on artificial intelligence.

References