FreeSOLO: Learning to Segment Objects without Annotations
Xinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz, Anima Anandkumar, Chunhua Shen, Jose M. Alvarez
Introduction
Instance segmentation is a fundamental computer vision task that requires recognizing the objects in an image and segmenting each of them at the pixel level. Instance segmentation subsumes object detection, as bounding box can be thought of as a coarse parametric representation of a segmentation mask. Therefore, it is a more demanding and challenging task than object detection by requiring both instance-level and pixel-level predictions. Recently, significant progress fcis; he2017mask; de2017semantic; chen2019hybrid; yolact; CondInst; wang2021SOLO has been made to address the instance segmentation task. However, the dense prediction nature of the task requires rich and expensive annotations during training. Weakly-supervised instance segmentation methods are thus proposed to relax the annotation requirements KhorevaBH0S17; hsu2019weakly; liu2020leveraging; tian2020boxinst; cheng2021pointly; lan2021discobox. Latest methods such as BoxInst tian2020boxinst and DiscoBox lan2021discobox have significantly closed the gap to fully supervised methods. However, their competitive result still relies on box or point annotations that contain strong localization information.
In this work, we explore learning class-agnostic instance segmentation without any annotations. The work here is built upon our recent work of SOLO wang2021SOLO, a simple yet strong instance segmentation framework, and the self-supervised dense feature learning method of DenseCL wang2020DenseCL. SOLO adopts a one-stage design, which contains a category branch and a mask branch to encode the object category information and segmentation proposals, respectively. Our main intuition is that this “top-down meets bottom-up” design allows us to unify pixel grouping, object localization and feature pre-training in a fully self-supervised manner.
Our proposed framework, FreeSOLO, contains two major pillars: Free Mask and Self-supervised SOLO, as shown in Figure 1. Specifically, Free Mask contains self-supervised design elements that promote objectness in network attention. It contains a “query-key” attention design, where the queries and keys are constructed from self-supervised features. The method takes the cosine similarity between each query with all the keys, thus obtaining a set of query-conditioned (seeded) attention maps as coarse masks. The coarse masks are ranked and filtered by their maskness scores, followed by non-maximum suppression (NMS) to further remove the redundant masks. Self-Supervised SOLO then takes the coarse masks as pseudo-labels to train a SOLO model. Since the coarse masks can be inaccurate, Self-Supervised SOLO contains a weakly-supervised design to better accommodate the label noise. This is followed by a self-training strategy to further refine mask quality and to improve accuracy. Our network design is almost the same as SOLO with minimal modifications, thus leading to simple and fast inference process.
FreeSOLO provides an effective solution to the challenging problem of self-supervised instance segmentation. With the bounding boxes obtained from the predicted masks, FreeSOLO also shows significant advantage as an unsupervised object discovery method. In addition to the above roles, we further consider FreeSOLO as a strong self-supervised pretext task for instance segmentation by jointly learning object-level and pixel-level representations. Compared to pre-training for image classification moco; simclr; byol, object detection updetr; detcon2021 and semantic segmentation PinheiroABGC20; ChaitanyaEKK20, pre-training for instance segmentation is still under-studied. General instance segmentation requires not only localizing objects at the pixel level, but also recognizing their semantic categories. Interestingly, the design of FreeSOLO allows us to directly learn object-level semantic representations in an unsupervised manner. Upon completing the pre-training, all the learned parameters except for the last classification layer can be used to initialize the supervised instance segmentation models to improve accuracy.
Our contributions can be summarized as follows.
We propose the Free Mask approach, which leverages the specific design of SOLO to effectively extract coarse object masks and semantic embeddings in an unsupervised manner.
We further propose Self-Supervised SOLO, which takes the coarse masks and semantic embeddings from Free Mask and trains the SOLO instance segmentation model, with several novel design elements to overcome label noise in the coarse masks.
With the above methods, FreeSOLO presents a simple and effective framework that demonstrates unsupervised instance segmentation successfully for the first time. Notably, it outperforms some proposal generation methods that use manual annotations. FreeSOLO also outperforms state-of-the-art methods for unsupervised object detection/discovery by a significant margin (relative +100% in COCO AP).
In addition, FreeSOLO serves as a strong self-supervised pretext task for representation learning for instance segmentation. For example, when fine-tuning on COCO dataset with 5% labeled masks, FreeSOLO outperforms DenseCL wang2020DenseCL by +9.8% AP.
Related Work
Instance segmentation. Instance segmentation has attracted much attention in recent years. Most existing works focus on learning instance segmentation with full annotations. Top-down methods fcis; he2017mask; panet; chen2019hybrid solve the problem from the perspective of object detection, i.e., detecting the bounding box of objects first and then segmenting the object in the box. Bottom-up methods associativeembedding; de2017semantic; SGN17; Gao_2019_ICCV view the task as a label-then-cluster problem, e.g., by learning per-pixel embeddings first and then clustering them into groups. Some recent methods yolact; chen2020blendmask; CondInst; wang2020solo; wang2020solov2 seek a combination of top-down and bottom-up approaches to perform faster inference and better segmentation. Among these methods, SOLO has shown a promising speed/accuracy trade-off with a very simple architecture. A few works explore learning instance segmentation with weak annotations, e.g., image-level and box-level labels KhorevaBH0S17; Zhou2018PRM; hsu2019weakly; tian2020boxinst. To the best of our knowledge, none have additionally explored learning instance segmentation without any labels at all.
In particular, BoxInst tian2020boxinst attains strong instance segmentation results using box annotations only, demonstrating that instance segmentation may not necessarily be more difficult to solve than box-level object detection. We move one-step forward by reporting strong instance segmentation results in an unsupervised setting, without any annotations.
Self-supervised learning. To learn a good visual representation from unlabeled data, a wide range of pretext tasks have been explored, e.g., colorization zhang2016colorful, inpainting inpainting16, jigsaw puzzles jigsaw and orientation discrimination gidaris2018rotations. The breakthroughs came from the contrastive learning methods, e.g., SimCLR simclr and MoCo moco that perform an instance discrimination pretext task wu2018unsupervised. Besides pre-training for image classification byol; swav20; simsiam21, some recent works wang2020DenseCL; xie2020propagate; detcon2021; xie2021detco; xiao2021region design self-supervised pre-training methods for dense prediction tasks, e.g., object detection and semantic segmentation. Different from them, our method can not only learn intermediate representations, but also train instance segmenters, which can segment objects in the wild. Our FreeSOLO naturally serves as a strong pretext task for learning representations for instance segmentation. The pre-trained model can be seamlessly transferred to supervised fine-tuning and can achieve significant gains compared to existing pre-training methods.
Unsupervised object discovery. A wide range of approaches have been proposed for unsupervised object discovery, including statistical topic discovery models Sivic2005DiscoveringOC; RussellFESZ06, link analysis technique KimT09, clustering by composition FaktorI14, and part-based matching ChoKSP15. Some recent works VoBCHLPP19; vo2020toward formulate object discovery as an optimization problem. LOD lod further proposes to formulate unsupervised object discovery as a ranking problem. Yet, the existing methods have achieved limited success in challenging and complicated scenes. Furthermore, most of these methods can only find coarse bounding boxes of objects. By contrast, our method discovers and localizes objects in the wild with pixel-wise segmentation masks. With bounding boxes obtained from predicted masks, FreeSOLO outperforms the state-of-the-art unsupervised object discovery methods by a large margin.
Unsupervised segmentation. To remove the dependency on manual supervision, some object co-segmentation works JoulinBP10; HsuLC18; ChenL0H21 make a strong assumption about the image collection, i.e., to segment common repeated objects in a collection of images. Besides, there are a few works JiVH19; HwangYSCYZC19; maskcontrast that explore unsupervised semantic segmentation. Some JiVH19 only deal with simple scenarios, and some HwangYSCYZC19; maskcontrast still require a salient object estimator or boundary annotations. In addition, the key difference lies in the task. Instead of semantic segmentation, our method solves the harder problem of instance segmentation, i.e., to segment each object individually.
Method
Background. We briefly introduce the supervised instance segmentation method SOLO wang2021SOLO. SOLO shows that instance segmentation can be solved by directly mapping an input image to the desired object categories and instance masks using fully convolutional networks (FCNs), eliminating the need for bounding box detection or grouping via post-processing. Its main idea is to formulate instance segmentation into two simultaneous category-aware pixel-level prediction problems. It conceptually divides the input image into grids. A grid cell is responsible for predicting the semantic category as well as the segmentation mask for an object whose center falls into that grid cell. The model consists of two branches, i.e., a category branch and a mask branch. The category branch predicts the semantic categories. The mask branch generates sized masks, one corresponding to each grid cell. Specifically, the dynamic SOLO variant employs dynamic convolutions to separately predict the mask kernels and mask features. The mask features are then convolved with the predicted mask kernels to generate the masks. This operation can be written as:
where is the convolution kernel, and denotes the score maps for all the masks. is then normalized via a operation, and input to mask NMS to form the final object masks.
We propose a novel framework for self-supervised instance segmentation, termed FreeSOLO. FreeSOLO does not require any type of annotations, neither pixel-level nor image-level labels, and simply uses a collection of unlabeled images for training. Its overall pipeline is illustrated in Figure 1. We first propose the Free Mask approach to generate segmentation masks from a self-supervised pre-trained model. For each unlabeled image, the coarse object masks can be generated fast with simple operations, e.g., at 21 FPS on a V100 GPU with a ResNet-50-based backbone. We further propose Self-Supervised SOLO, which trains the SOLO-based instance segmenter using the coarse masks and semantic embeddings from Free Mask, with several novel design elements including weaky-supervised design, self-training, and semantic embedding learning.
With FreeSOLO, we obtain an instance segmentation model given only unlabeled images. In addition to unsupervised instance segmentation itself, the well-trained model serves as a strong pre-trained model for downstream fine-tuning. All its parameters except the last classification layer can be transferred to supervised instance segmentation as a strong initialization.
2 Free Mask
The score maps are then normalized as soft masks by shifting the scores to the range $\tt masknessN\tau$. We then sort the binary masks by their maskness scores and remove the redundant masks via mask non-maximum-suppression (NMS). The overall process can be formulated as:
where denotes the object masks that Free Mask outputs.
Self-supervised pre-training. Free Mask uses a pre-trained backbone via self-supervision as the starting point. We propose to leverage the self-supervised model pre-trained with dense correspondence. Specifically, we find that dense contrastive learning wang2020DenseCL achieves considerably better results with our Free Mask approach, compared to the conventional self-supervised learning by global image-level contrasting. This can be attributed to the similar objective of Free Mask and dense contrastive learning. Here we briefly introduce how the dense contrastive learning is performed. It optimizes a pairwise (dis)similarity loss at the level of local features between two views of the input image. A local feature vector, i.e., a query vector, should be similar to the corresponding positive key in the other view while being dissimilar to other negative keys. Observe that this is also aligned with Equation (2) where the cosine similarity between a query and the keys is evaluated. This also explains why Free Mask extracts reasonable masks. We believe that there could be even better pre-training methods for Free Mask, e.g., those which tackle how to learn fine-grained representations at higher resolutions to generate better masks. We leave this for future research.
Pyramid queries. When constructing the queries from , we design a pyramid queries method to generate masks for instances at different scales. Specifically, we set a list of scale factors, e.g., , when downsampling , thus leading to a list of at different scales from large to small. All pyramid queries are flattened and concatenated together as the final .
Maskness score. A scoring function is required for evaluating the quality of each generated coarse mask, which cannot be learned from annotations. We use the non-parametric maskness method wang2020solo, i.e., , to obtain the confidence score of an extracted mask. Here denotes the number of foreground pixels of the soft mask , i.e., the pixels that have values greater than threshold . Intuitively, this score weighs more heavily on masks that have high confidence on foreground pixels and down weights masks with uncertain foreground pixels.
Unified with SOLO. We can see that the pipeline in Equation (4) is unified with that of SOLO, as introduced in the above background section. They both go through FCN, dynamic convolution, normalization and NMS operations to generate object masks. However, the two are proposed to solve different problems. The latter aims to learn instance segmentation with rich annotated data, while the former is for segmenting objects in unlabeled images. This provides a unifying perspective on segmenting objects in images.
3 Self-Supervised SOLO
We aim to train the SOLO-based instance segmenter using the segmentation masks and semantic embeddings, i.e., feature embeddings with high-level semantics, from Free Mask. We separately introduce the methods for learning with coarse masks, self-training, and the semantic representation learning.
Learning with coarse masks. In SOLO, the Dice loss vnet is used to supervise the predicted masks with their ground truth labels. However, this is not ideally suited for our case of learning with noisy masks. As the masks are coarse, directly using them as ground-truth masks can lead to unsatisfactory results. We propose to use the coarse masks as a type of weak annotation and perform weakly supervised instance segmentation with them.
Inspired by the latest weakly-supervised method of BoxInst tian2020boxinst, we project the predicted masks and the coarse masks on to the -axis and the -axis via a operation along each axis. The model is supervised to minimize the discrepancy between the projections of predicted masks and the coarse masks. The loss term can be defined as:
where is the Dice loss, and are the predicted mask and the coarse mask. and denote the operations along each axis.
We further propose to project the predicted and coarse masks onto the and axes via an operation along each axis. The motivation is that the operation may emphasize outlier segmentations in coarse masks, while the operation de-emphasize the outliers. In addition, operation preserves solid shape of the object mask, which can benefit the training. The loss term can be written as:
where and denote the operation along each axis. We also employ a pairwise affinity loss tian2020boxinst to leverage the prior that the proximal pixels are likely to be in the same class, i.e., foreground or background, if they have similar colors in the raw image.
Overall, the total loss for mask prediction can be formulated as:
where acts as the weight to balance the various loss terms.
Self-training. With our carefully-designed loss function, we are able to train a SOLO-based instance segmenter with the free and noisy coarse masks. As shown in Figure 1, the object masks predicted by the instance segmenter are considerably better than the original coarse masks from Free Mask, which is also validated by the boosted accuracy in Table . As such, we propose to perform self-training with the initially trained instance segmenter to further improve accuracy. We input unlabeled images into the instance segmenter and collect their predicted object masks. The low-confidence predictions are removed and the remaining ones are treated as a new set of coarse masks. We again train an instance segmenter with the unlabeled images and the new masks, using the loss function in Equation (7). Performing self-training once already brings clear improvements and more iterations do not provide additional gains.
The total loss for the category branch can be written as:
where acts as the weight to balance the two terms. Overall, we train the instance segmenter with a combination of and , corresponding to the losses for the mask branch and category branch, respectively.
Experiments
Technical details. For Free Mask, the shorter side of the input image is set to 800 pixels. Threshold is set to . DenseCL wang2020DenseCL with a pre-trained ResNet-50 resnet architecture is adopted as the backbone unless specified. Matrix NMS wang2020solo is used for mask NMS. After NMS, we filter out the low-quality masks with a maskness threshold of 0.7. When training the SOLO model, we initialize the backbone with the pre-trained model used in Free Mask. We set the and parameters to and , respectively. We employ the simple copy-paste strategy copypaste for data augmentation. During self-training, we set the confidence threshold for removing the low-confidence predictions to .
Datasets. For FreeSOLO, we use the images in COCO and COCO coco as the set of unlabeled images, containing a total of 241k images. These unlabeled images are input to Free Mask and are used to train the instance segmenter. The self-supervised backbone in Free Mask is pre-trained on ImageNet with 1.28 million unlabeled images. We further employ COCO , UVO uvo, and PASCAL VOC voc datasets for evaluation.
Evaluation protocol. We evaluate self-supervised instance segmentation with the standard COCO protocol. We report class-agnostic COCO mask average precision (AP) and average recall (AR) on k split, which is averaged over 10 intersection-over-union (IoU) thresholds evenly-spaced between and . AP considers recall and precision simultaneously, which computes the average precision value for recall values over 0 to 1. AR allows redundant or random detection results, as it computes the maximum recall given a fixed number of detections per image.
To compare with unsupervised object detection methods, we convert the masks to boxes and report the box AP on both the COCO , COCO , and VOC . We further evaluate the pre-trained model by fine-tuning with annotations. Specifically, we fine-tune the instance segmenter on COCO and evaluate on COCO . We provide two settings, i.e., limited fully annotated images, and limited segmentation masks (see Appendix A.2). Mask AP averaged across all 10 IoU thresholds and all 80 categories is reported.
2 Main Results
Self-supervised instance segmentation. For evaluating the self-supervised instance segmenter, we first provide qualitative results to show how FreeSOLO performs at the task of class-agnostic instance segmentation. As shown in Figure , without any annotations, FreeSOLO is able to segment object instances of many different categories. To provide a quantitative comparison with previous methods, we report the results of unsupervised class-agnostic instance segmentation in Table 1 and Table 2. As there is no reported result for this new problem, we evaluate a few popular segmentation proposal methods on this benchmark. Among the compared methods, MCG mcg uses the annotated BSDS500 dataset MartinFTM01 for training a boundary detector, and COB cob trains its hierarchies and combinatorial grouping on PASCAL Context dataset MottaghiCLCLFUY14. By contrast, our FreeSOLO method achieves better results without any annotations. We further compare against the supervised methods trained with full annotations. It is worth noting that FreeSOLO even performs closely to the fully supervised Mask R-CNN he2017mask trained on the LVIS dataset lvis2019, e.g., 4.8% vs 6.8% AP on the UVO dataset.
Self-supervised object detection. By converting the masks into boxes, our self-supervised instance segmenter naturally serves as a self-supervised object detector as well. We report the results of class-agnostic object detection on COCO benchmark in Table 3. Our method shows significantly superior performance. To compare with existing object discovery methods, we also evaluate FreeSOLO on VOC and COCO for multi-object discovery. As shown in Table 4, our method largely outperforms the state-of-the-art object discovery methods, including a concurrent work lost. Its relative improvements are up to 100% on the COCO dataset.
Supervised fine-tuning. In addition to evaluating the self-supervised instance segmenter directly, we also evaluate the performance of our approach in a supervised setting by fine-tuning the self-supervised instance segmenter with annotations. As shown in Table 5, FreeSOLO pre-training outperforms ImageNet supervised pre-training by 4.0% AP when using 5% COCO training images. The gains over the state-of-the-art self-supervised pre-training methods are also clear, e.g., 2.0% AP better than DenseCL wang2020DenseCL.
To further compare the pre-training methods with different amount of mask annotations, in Table 6, we conduct fine-tuning experiments with only limited masks available. When fine-tuning with 5% masks, FreeSOLO achieves significant gains of 9.8% AP over supervised pre-training. These fine-tuning experiments demonstrate that FreeSOLO serves as a strong instance segmentation pre-training method, outperforming both the supervised and state-of-the-art self-supervised pre-training methods.
3 Ablation Study
We conduct ablation experiments to show how each component contributes to FreeSOLO. The ablation studies are performed on the COCO split.
Free Mask with different pre-trained backbones. In Table , we show how Free Mask performs with different pre-trained backbones. The conventional self-supervised learning methods that contrast the global representations of image pairs, e.g., SimCLR and MoCo-v2, show worse results compared to supervised ImageNet pre-training. The self-supervised learning methods that consider dense correspondence, e.g., EsViT and DenesCL, yield better results than those that do not. DenseCL shows the best results compared to both supervised and other self-supervised methods. This aligns with our hypothesis in Section 3.2 that DenseCL’s objective is consistent with Free Mask’s. We provide some visualizations of Free Mask in Figure 3.
Pyramid queries. We compare different scales of the queries used in Free Mask in Table . A smaller scale is better for large objects but worse for medium and small objects. A large scale is just the opposite. Pyramid queries with scales yield the best results.
Loss functions. In Table , we compare our weakly-supervised design against the full mask supervision, i.e., the original Dice loss used in SOLO computed with the full masks. Directly using the coarse masks to provide full supervision to the instance segmenter leads to unsatisfactory results. Our weakly-supervised loss outperforms the original full mask loss by a large margin. In Table , we study the mask loss terms in Equation (7). The performance drops sharply when learning without , i.e., with only the projection loss from operation and pairwise loss as in tian2020boxinst. The model even collapses to only segmenting the contours when trained longer (Figure 4). Our method tackles this problem by leveraging the projection from operation, which not only preserves the shape but is also less sensitive to outlier pixels.
Self-training. Our method performs self-training by selecting high-confidence predictions of the self-supervised instance segmenter and training the instance segmenter again with them. We compare the results of performing different iterations of self-training in Table . ‘’ refers to the initial coarse masks. Zero iteration refers to learning from the coarse masks without self-training. We show that performing self-training once already brings clear improvements, but additional iterations do not provide additional gains.
Semantic embedding. To validate the effectiveness of the semantic embedding learning, in Table we compare the models trained with or without the semantic embedding loss defined in Equation (8). The models are fine-tuned with 10% of fully annotated COCO images. It shows that the semantic embedding loss yields clear improvements when fine-tuning instance segmentation with annotations.
Discussion and Conclusion
In this work, we have developed a simple and effective self-supervised instance segmentation framework FreeSOLO. FreeSOLO enables learning to segment objects without any annotations, neither pixel-level nor image-level labels. We hope that its novel design elements provide insights for future works on unsupervised visual learning, e.g., unsupervised panoptic segmentation, and beyond.
Limitations. Without category labels, our self-supervised instance segmenter cannot predict the categories of the detected objects, but generate class-agnostic object masks. There is still a large gap between our self-supervised model and the supervised one trained with rich annotations. Our method could fail in some scenarios (Figure 5). We believe there is plenty of room to improve based on our method.
Broader impacts. This work shows that one can learn a class-agnostic instance segmenter without any annotations. In the future, there is a chance for self-supervised segmenter to reach or even outperform the supervised model trained with manual annotations, which may eliminate the need for annotating masks or boxes for common objects. We expect that the proposed technique can be used to largely reduce data annotation effort for a few instance-level recognition tasks in computer vision.
References
Appendix
Appendix A Additional implementation details
While evaluating the performance of class-agnostic instance segmentation, we also report the results of an easier protocol AP∗ which evaluates medium and large objects. Here, only the objects with area greater than are considered and their mask AP∗ with an IoU threshold of is computed. AP and AP are also reported for medium and large objects, i.e., objects with area in the range of and those with area greater than , respectively. The results of MCG and COB are computed using the official segmentation masks.
A.2 Supervised fine-tuning
We evaluate the pre-trained instance segmentation model by fine-tuning it with manual annotations. Specifically, we fine-tune a dynamic SOLO model (aka SOLOv2) on COCO and evaluate on COCO . Synchronized batch normalization is used in the backbone along with FPN fpn during training. We provide two training settings, i.e., limited fully annotated images, and limited segmentation masks.
Limited images. For the experiments with limited images, we use 5% and 10% images from COCO , which corresponds to 6k and 12k fully annotated images, respectively. We fine-tune the instance segmenter initialized with the pre-trained model for 20k iterations with an initial learning rate of 0.01, which is then divided by 10 at 12k and 18k iterations.
Limited masks. For the experiments with limited masks, we use 5% and 10% segmentation masks from COCO . In this setting, only 5% and 10% of the images have mask annotations, i.e., 6k and 12k images, respectively. Specifically, we use all the class labels to supervise the category branch, but only use a part of the annotated masks to supervise the mask branch. The model is trained for 90k iterations with the standard schedule.
A.3 Training details
For the self-supervised pre-trained backbones, we use the official models trained on ImageNet without labels for 200 epochs. For FreeSOLO, we use the images in COCO and COCO as the set of unlabeled images, containing a total of 241k images. We use ResNet-50 as the backbone for all the fine-tuning experiments and ablation study and use ResNet-101 for other results and visualizations. We train for 30k iterations on 8 GPUs with a total of 32 images per mini-batch. The learning rate is set to 0.0025. In the self-training, we repeat the schedule once and train for another 30k iterations.
Copy-paste augmentation. For a pair of images in a batch, we randomly select objects from one image and paste them at random locations on the other image. These objects are not pasted if they have a high overlap (IoU ) with existing objects.
Appendix B Additional results
We report the results of an easier protocol AP∗ which evaluates medium and large objects in Table S1. As shown, the gains over MCG and COB are larger, especially for the large objects.
Appendix C Additional visualizations
In this section, we provide additional visualizations of FreeSOLO. We show qualitative results of our method for the task of class-agnostic instance segmentation in Figure S1. In Figure S2, we provide more qualitative comparison of FreeSOLO with and without the . As shown in Figure S3, we further show that FreeSOLO can even produce more precise segmentation results than manual annotations at some object boundaries, which indicates FreeSOLO’s great potential for tasks such as auto-labeling.