Cut and Learn for Unsupervised Object Detection and Instance Segmentation

Xudong Wang, Rohit Girdhar, Stella X. Yu, Ishan Misra

Introduction

Object localization is a critical task in computer vision that enables AI systems to perceive, reason, plan and act in an object-centric manner. Training models for localization require special annotations like object boxes, masks, localized points, etc. which are both difficult and resource intensive to collect. Without accounting for overhead, annotating ∼\sim164K images in the COCO dataset with masks for just 80 classes took more than 28K human hours of annotation time. In this work, we study unsupervised object detection and instance segmentation models that can be trained without any human labels. Our key insight is that simple probing and training mechanisms can amplify the innate localization ability of self-supervised models , leading to state-of-the-art unsupervised zero-shot detectors.

Our method Cut-and-LEaRn (CutLER) consists of three simple, architecture- and data-agnostic mechanisms. Consistent with prior self-supervised learning methods , CutLER is trained exclusively on unlabeled ImageNet data without needing additional training data, but contrary to these methods, CutLER can be directly employed to perform complex segmentation and detection tasks over a wide range of domains. First, we propose MaskCut that can automatically produce multiple initial coarse masks for each image, using the pretrained self-supervised features. Second, we propose a simple loss dropping strategy to train detectors using the coarse masks while being robust to objects missed by MaskCut. Finally, we observe that despite learning from these coarse masks, the detectors ‘clean’ the ground truth and produce masks (and boxes) that are better than the coarse masks used to train them. Therefore, we further show that multiple rounds of self-training on the models’ own predictions allow it to evolve from capturing the similarity of local pixels to capturing the global geometry of the object, thus producing finer segmentation masks.

Prior work shows that a self-supervised vision transformer (ViT) can automatically learn patch-wise features that detect a single salient object in an image . However, unlike CutLER, such salient object detection methods only locate a single, usually the most prominent, object and cannot be used for real world images containing multiple objects. While some recent methods, e.g., FreeSOLO and DETReg , also aim at unsupervised multi-object detection (or multi-object discovery), they rely on a particular detection architecture, e.g., SOLO-v2 or DDETR . Additionally, apart from self-supervised features trained on ImageNet , the current state-of-the-art methods FreeSOLO and MaskDistill also require ‘in-domain’ unlabeled data for model training.

In contrast, CutLER works with various detection architectures and can be trained solely on ImageNet, without requiring in-domain unlabeled data. Thus, during model training, CutLER does not see any images from any target dataset and yields a zero-shot model capable of detecting and segmenting multiple objects in diverse domains.

Features of CutLER. 1) Simplicity: CutLER is simple to train and agnostic to the choice of detection and backbone architectures. Thus, it can be integrated effortlessly into existing object detection and instance segmentation works. 2) Zero-shot detector: CutLER trained solely on ImageNet shows strong zero-shot performance on 11 different benchmarks where it outperforms prior work trained with additional in-domain data. We double the AP50box{}^{\text{box}}_{50} performance on 10 of these benchmarks, as shown in Fig. 1, and even outperform supervised detectors on the UVO video instance segmentation benchmark. 3) Robustness: CutLER exhibits strong robustness against domain shifts when tested on images from different domains such as video frames, sketches, paintings, clip arts, etc. 4) Pretraining for supervised detection: CutLER can also serve as a pretrained model for training fully supervised object detection and instance segmentation models and improves performance on COCO, including on few-shot object detection benchmarks.

Related Work

Self-supervised feature learning involves inferring the patterns within the large-scale unlabeled data without using human-annotated labels. Contrastive learning based methods learn such representations that similar samples or various augmentations of the same instance are close to each other, while dissimilar instances are far apart. Similarity-based self-supervised learning methods learn representations via minimizing the distance between different augmentations of the same instance and use only positive sample pairs. Clustering-based feature learning automatically discovers the natural grouping of data in the latent representation space. Recently, have shown that masked autoencoders, which learn representations via masking out a large random subset of image patches and reconstructing the missing pixels or patches , are scalable self-supervised learners for computer vision .

In contrast to these unsupervised representation learning efforts, our work aims to automatically discover natural pixel groupings and locate instances within each image.

Unsupervised object detection and instance segmentation. The main comparisons to previous works are listed in Table 1 and are elaborated as follows:

DINO observes that the underlying semantic segmentation of images can emerge from the self-supervised Vision Transformer (ViT) , which does not appear explicitly in either supervised ViT or ConvNets . Based on this observation, LOST and TokenCut leverage self-supervised ViT features and propose to segment one single salient object from each image based on a graph that is constructed with DINO’s patch features.

These previous works either can not detect more than one object from each image, e.g., DINO and TokenCut, or can not improve the quality of features for better transfer to downstream detection and segmentation tasks, e.g., TokenCut and LOST. Unlike these works, CutLER can locate multiple objects and serve as a pretrained model for label-efficient and fully-supervised learning.

FreeSOLO achieves unsupervised instance segmentation by extracting coarse object masks in an unsupervised manner, followed by mask refinement through a self-training procedure. While FreeSOLO’s FreeMask stage can generate multiple coarse masks per image, the quality of these masks is often rather low . MaskDistill distills class-agnostic initial masks from the affinity graph produced by a self-supervised DINO . However, it utilizes one single mask per image in the distillation stage, which greatly limits the model’s ability to detect multiple objects.

By contrast, the initial masks generated by our MaskCut are usually better in quality and quantity than the initial masks used by . Therefore, CutLER achieves 2\times\raise 0.73193pt\hbox{\scriptstyle\sim}4×\times higher APbox{}^{\text{box}} and APmask{}^{\text{mask}} than FreeSOLO and MaskDistill on almost all experimented detection and segmentation benchmarks, even when FreeSOLO and MaskDistill are trained and tested on the same domain.

Method

We tackle the problem of unsupervised object detection and segmentation with a simple cut-and-learn pipeline. Our method builds upon insights from recent work , showing that self-supervised representations can discover objects. While these methods often find a single object per image, we propose a simple approach that can discover multiple objects and significantly improves segmentation and detection performance. The overview of our cut-and-learn pipeline is illustrated in Fig. 2. First, we propose MaskCut that generates multiple binary masks per image using self-supervised features from DINO (Sec. 3.2). Second, we show a dynamic loss dropping strategy, called DropLoss, that can learn a detector from MaskCut’s initial masks while encouraging the model to explore objects missed by MaskCut (Sec. 3.3); Third, we further improve the performance of our method through multiple rounds of self-training (Sec. 3.4).

Normalized Cuts (NCut) treats the image segmentation problem as a graph partitioning task . We construct a fully connected undirected graph via representing each image as a node. Each pair of nodes is connected by edges with weights WijW_{ij} that measure the similarity of the connected nodes. NCut minimizes the cost of partitioning the graph into two sub-graphs, i.e., a bipartition, by solving a generalized eigenvalue system

for finding the eigenvector xx that corresponds to the second smallest eigenvalue λ\lambda, where DD is a N ⁣× ⁣NN\!\times\!N diagonal matrix with d(i)=∑jWijd(i)=\sum_{j}W_{ij} and WW is a N ⁣× ⁣NN\!\times\!N symmetrical matrix.

DINO and TokenCut. DINO finds that the self-supervised ViT can automatically learn a certain degree of perceptual grouping of image patches. TokenCut leverages the DINO features for NCut and obtaining foreground/background segments in an image. The authors use the similarity of the patches in the DINO feature space as the similarity weight WijW_{ij} in NCut. Specifically, following multiple recent methods , we use the cosine similarity of ‘key’ features from the last attention layer of DINO-pretrained model, i.e., Wij ⁣= ⁣KiKj∥Ki∥2∥Kj∥2W_{ij}\!=\!\frac{K_{i}K_{j}}{\|K_{i}\|_{2}\|K_{j}\|_{2}} where KiK_{i} is the ‘key’ feature of patch ii, and solve Eq. 1 for finding the second smallest eigenvector xx.

A limitation of TokenCut is that it only computes a single binary mask for an image and thus only finds one object per image. Although we can use the other N ⁣− ⁣2N\!-\!2 smallest eigenvectors to locate more than one instance, this significantly degrades the performance for multi-object discovery, as demonstrated in Sec. 5.

2 MaskCut for Discovering Multiple Objects

As we discussed in Sec. 3.1, vanilla NCut is limited to discovering a single object in an image. We propose MaskCut that extends NCut to discover multiple objects per image by iteratively applying NCut to a masked similarity matrix (illustrated in Fig. 3). After getting the bipartition xtx^{t} from NCut at stage tt, we get two disjoint groups of patches and construct a binary mask MtM^{t}, where

To determine which group corresponds to the foreground, we make use of two criteria: 1) intuitively, the foreground patches should be more prominent than background patches . Therefore, the foreground mask should contain the patch corresponding to the maximum absolute value in the second smallest eigenvector MtM^{t}; 2) we incorporate a simple but empirically effective object-centric prior : the foreground set should contain less than two of the four corners. We reverse the partitioning of the foreground and background, i.e., Mijt ⁣= ⁣1 ⁣− ⁣MijtM^{t}_{ij}\!=\!1\!-\!M^{t}_{ij}, if the criteria 1 is not satisfied while the current foreground set contains two corners or the criteria 2 is not satisfied. In practice, we also set all Wij ⁣< ⁣τncutW_{ij}\!<\!\tau^{\text{ncut}} to 1e−51e^{-5} and Wij ⁣≥ ⁣τncutW_{ij}\!\geq\!\tau^{\text{ncut}} to 1.

where M^ijs ⁣= ⁣1 ⁣− ⁣Mijs\hat{M}^{s}_{ij}\!=\!1\!-\!{M}^{s}_{ij}. Using the updated Wijt+1W^{t+1}_{ij}, we repeat Eqs. 1 and 2 to get a mask Mt+1M^{t+1}. We repeat this process tt times and set t ⁣= ⁣3t\!=\!3 by default.

3 DropLoss for Exploring Image Regions

A standard detection loss penalizes predicted regions rir_{i} that do not overlap with the ‘ground-truth’. Since the ‘ground-truth’ masks given by MaskCut may miss instances, the standard loss does not enable the detector to discover new instances not labeled in the ‘ground-truth’. Therefore, we propose to ignore the loss of predicted regions rir_{i} that have a small overlap with the ‘ground-truth’. More specifically, during training, we drop the loss for each predicted region rir_{i} that has a maximum overlap of τIoU\tau^{\text{IoU}} with any of the ‘ground-truth’ instances:

where IoUimax\text{IoU}_{i}^{\text{max}} denotes the maximum IoU with all ‘ground-truth’ for rir_{i} and Lvanilla\mathcal{L}_{\text{vanilla}} refers to the vanilla loss function of detectors. Ldrop\mathcal{L}_{\text{drop}} does not penalize the model for detecting objects missed in the ‘ground-truth’ and thus encourages the exploration of different image regions. In practice, we use a low threshold τIoU=0.01\tau^{\text{IoU}}=0.01.

4 Multi-Round Self-Training

Empirically, we find that despite learning from the coarse masks obtained by MaskCut, detection models ‘clean’ the ground truth and produce masks (and boxes) that are better than the initial coarse masks used for training. The detectors refine mask quality, and our DropLoss strategy encourages them to discover new object masks. Thus, we leverage this property and use multiple rounds of self-training to improve the detector’s performance.

5 Implementation Details

Training data. We only use the images from the ImageNet dataset (1.3 million images) for all parts of the CutLER model and do not use any type of annotations either for training or any supervised pretrained models.

MaskCut. We use MaskCut with three stages on images resized to 480×\times480 pixels and compute a patch-wise affinity matrix using the ViT-B/8 DINO model. We use Conditional Random Field (CRF) to post-process the masks and compute their bounding boxes.

Detector. While CutLER is agnostic to the underlying detector, we use popular Mask R-CNN and Cascade Mask R-CNN for all experiments, and use Cascade Mask R-CNN by default, unless otherwise noted. We train the detector on ImageNet with initial masks and bounding boxes for 160160K iterations with a batch size of 16. When training the detectors with a ResNet-50 backbone , we initialize the model with the weights of a self-supervised pretrained DINO model. We explored other pre-trained models, including MoCo-v2 , SwAV , and CLD , and found that they gave similar detection performance. We also leverage the copy-paste augmentation during the model training process. Rather than using the vanilla copy-paste augmentation, to improve the model’s ability to segment small objects, we randomly downsample the mask with a scalar uniformly sampled between 0.3 and 1.0. We then optimize the detector for 160160K iterations using SGD with a learning rate of 0.005, which is decreased by 5 after 8080K iterations, and a batch size of 16. We apply a weight decay of 5 ⁣× ⁣10−55\!\times\!10^{-5} and a momentum of 0.9.

Self-training. We initialize the detection model in each stage using the weights from the previous stage. We optimize the detector using SGD with a learning rate of 0.01 for 80KK iterations. Since the self-training stage can provide a sufficient number of pseudo-masks for model training, we don’t use the DropLoss during the self-training stages.

We provide more details on model implementation and training in Sec. A.1.

Experiments

We evaluate CutLER on various detection and segmentation benchmarks. In Sec. 4.1, we show that CutLER can discover objects without any supervision on completely unseen images. Despite being evaluated in a zero-shot manner on eleven benchmarks, CutLER outperforms prior methods that use in-domain training data. Sec. 4.2 shows that finetuning CutLER further improves detection performance, outperforming prior work like MoCo-V2 and FreeSOLO.

We conduct extensive experiments on eleven different datasets, covering various object categories, image styles, video frames, resolutions, camera angles, etc. to verify the effectiveness of CutLER as a universal unsupervised object detection and segmentation method. We describe the different datasets used for zero-shot evaluation in detail in Sec. A.2. CutLER is trained solely using images from ImageNet and evaluated in a zero-shot manner on all downstream datasets without finetuning on any labels or data.

Evaluating unsupervised object detectors poses two unique challenges. First, since the model is trained without any notion of semantic classes, it cannot be evaluated using the class-aware detection setup. Thus, like prior work we use the class-agnostic detection evaluation. Second, object detection datasets often only annotate a subset of the objects in the images. For example, while COCO and LVIS use the same images, COCO only labels 80 object classes, and LVIS labels 1203 object classes. In this partially labeled setup, Average Recall (AR) is a valuable metric for unsupervised detection as it does not penalize the models for detecting novel objects unlabeled in the dataset. Thus, we additionally report AR for all datasets.

Zero-shot detection on 11 benchmarks. We evaluate CutLER on a variety of datasets and report the detection performance using AP50box{}^{\text{box}}_{50} and AR100box{}^{\text{box}}_{100} metrics in Fig. 1 and Table 2. CutLER uses a smaller model size and less training data than prior work. Compared to the previous SOTA approach, FreeSOLO with a backbone of ResNet101, CutLER, with the smaller ResNet50 backbone, significantly outperforms it in each of these benchmarks spanning various image distributions, more than doubling performance on 10 of them. Also note that, FreeSOLO requires FreeMask pre-training using approximately 1.3M ImageNet images and model fine-tuning using additional data in test benchmarks.

We observe that on different domains, e.g. watercolor or frames from videos (UVO dataset), CutLER improves performance by over 4×4\times and 2×2\times, respectively. Fig. 1 shows some qualitative examples of CutLER’s predictions.

Detailed comparisons on COCO20K and COCO. Table 3 presents detailed detection and segmentation evaluations (also referred to as ‘multi-object’ discovery) on two popular benchmarks: COCO val2017 and COCO 20K, which contains a subset of 20K images of COCO . CutLER consistently surpasses prior works by a large margin (often gets 2∼\scriptstyle\sim3×\times higher AP) on both the segmentation and detection tasks. Although CutLER is not trained on any images from COCO, it surpasses existing methods trained on COCO by more than 10% in terms of AP50mask{}^{\text{mask}}_{50} and AP50box{}^{\text{box}}_{50}.

Fig. 4 shows the qualitative comparisons between and our CutLER on COCO val2017, along with human annotations. Surprisingly, CutLER can often detect novel instances that human annotators miss.

We present detailed comparisons on COCO 20K, COCO val2017 and LVIS benchmarks in Sec. A.3.

Detailed comparisons on UVO and VOC. For a comprehensive comparison with existing unsupervised multi-object detection methods, we report the results for UVO val and VOC trainval07 . Table 4 shows that CutLER yields significant performance gains over previous SOTA, obtaining over 3×\times higher AP, with the most considerable improvement coming from APL. On UVO, Table 5 shows that CutLER more than quadruples the AP of previous SOTA and almost triples the AP50box{}^{\text{box}}_{50}. Our AP50mask{}^{\text{mask}}_{50} is even 4.8% higher than the fully-supervised SOLOv2 trained on LVIS with 100% annotations, significantly narrowing the gap between supervised and unsupervised learning.

2 Label-Efficient and Fully-Supervised Learning

We now evaluate CutLER as a pretraining method for training object detection and instance segmentation models. While CutLER can discover objects without any supervision, finetuning it on a target dataset aligns the model output to the same set of objects labeled in the dataset.

Setup. We use CutLER to initialize a standard Cascade Mask R-CNN detector with a ResNet50 . Prior work uses more advanced detectors, SOLOv2 used in and DDETR used in , that perform better. However, we choose Cascade Mask R-CNN for its simplicity and show in Sec. 5 that CutLER’s performance improves with stronger detectors. We train the detector on the COCO dataset using the bounding box and instance mask labels. To evaluate label efficiency, we subsample the training set to create subsets with varying proportions of labeled images. We train the detector, initialized with CutLER, on each of these subsets. As a baseline, we follow the settings from MoCo-v2 and train the same detection architecture initialized with a MoCo-v2 ResNet50 model, given its strong performance on object detection tasks. Both MoCo-v2 and our models are trained for the 1×1\times schedule using Detectron2 , except for extremely low-shot settings with 1% or 2% labels. Following previous works , when training with 1% or 2% labels, we train both MoCo-v2 and our model for 3,600 iterations with a batch size of 16.

Results. Fig. 5 shows the results of fine-tuning the detector on different subsets of COCO. When tested with low-shot settings, e.g., 2% and 5% labeled data, our approach achieves 5.4% and 7.3% higher APbox{}^{\text{box}} than the MoCo-v2 baseline, respectively. Even when training with full annotations, CutLER still consistently gives more than 2% improvements, outperforming MoCo-v2 for both object detection and segmentation. More impressively, CutLER outperforms prior SOTA methods - FreeSOLO and DETReg despite using an older detection architecture.

Ablations

We analyze the design decisions in CutLER. We use similar settings to Sec. 4 and train CutLER only on ImageNet. We use the Cascade Mask R-CNN detection architecture and evaluate our model primarily on the COCO and UVO unsupervised detection benchmarks. All ablation studies are conducted without self-training unless otherwise noted.

We analyze the main components of CutLER and report their relative contribution in Table 6. We report results on the popular COCO dataset and a densely annotated video instance segmentation dataset UVO . We also report the performance of running TokenCut on the COCO dataset. Next, we use TokenCut’s official codes to generate masks on ImageNet and use them for training a Cascade Mask R-CNN . This base model provides substantial gains over just using TokenCut on COCO. We add each of our proposed components to this strong base model. Using MaskCut increases AP50mask{}_{50}^{\text{mask}} and APmask{}^{\text{mask}} by 4.7% and 2.7%, respectively. Also, the improvements to AP50mask{}_{50}^{\text{mask}} is larger for densely annotated dataset UVO, i.e. 4.7% vs. 2.7%. These results prove that MaskCut’s ability to segment multiple instances per image is vital for densely annotated datasets. Adding DropLoss brings another 1.6% and 0.9% improvements to AP50mask{}_{50}^{\text{mask}} for UVO and COCO, respectively. Multi-round of self-training increases the quantity and quality of pseudo-masks, leading to 1.3% improvements. These results show that each simple proposed component is critical for strong performance.

Comparison with TokenCut.

TokenCut is also a zero-shot segmentation method. However, it only segments a single instance per image, as discussed in Sec. 3.1. In order to generate more than one segmentation mask per image, we use a modified TokenCut by using more of the smaller eigenvectors and combining all produced masks. Table 7 shows the object detection performance on COCO’s validation set for vanilla TokenCut, our modified TokenCut and CutLER. Although using more eigenvectors increases the recall AR100box{}_{100}^{\text{box}}, it significantly reduces the precision APbox{}^{\text{box}}. CutLER not only improves the average recall AR100box{}_{100}^{\text{box}} by 4×\times but also surpasses TokenCut’s average precision APbox{}^{\text{box}} by 4.8×\times, i.e. 480% relative improvements.

Design choices in MaskCut and DropLoss

and their impact on the final localization performance is presented in Table 8. We first study the effect of the image size used by MaskCut for generating the initial masks. As expected, LABEL:tab:ablate_image_size shows that MaskCut benefits from using higher resolution images presumably as it provides a higher resolution similarity between pixels. We pick a resolution of 480px for a better trade-off between the speed of MaskCut and its performance. In LABEL:tab:ablate_tau_ncut, we study the effect of the threshold used in MaskCut for producing a binary WW matrix (Sec. 3.2). Overall, CutLER seems to be robust to the threshold values. We understand the impact of the number of masks per image generated by MaskCut in LABEL:tab:ablate_num_masks. Increasing the number improves the performance of the resulting CutLER models. This shows that MaskCut generates high-quality masks that directly impact the overall performance. Finally, in LABEL:tab:ablate_tau_droploss, we vary the IOU threshold used for DropLoss. With a high threshold, we ignore the loss for a higher number of predicted regions while encouraging the model to explore. 0.010.01 works best for the trade-off between exploration and detection performance.

Self-training

and its impact on the final performance is analyzed in Table 9. Self-training consistently improves performance across the UVO and COCO benchmarks and all metrics. UVO, which has dense object annotations, benefits more from the multi-round of self-training. By default, CutLER uses 3 rounds of self-training. Fig. 6 shows qualitative examples of how self-training improves both the quality of predictions and the number of objects predicted.

Generalization to different detection architectures.

We use different detector architectures for training CutLER and measure their performance in Table 10. We observe that CutLER works with various architectures, and its performance is improved with stronger architectures.

Impact of the pretraining dataset.

We now study the impact of the dataset used for 1) pretraining the self-supervised DINO model and 2) training the CutLER model. The commonly used ImageNet dataset has a well-known object-centric bias which may affect the unsupervised detection performance. Thus, we also use YFCC , a non-object-centric dataset. We control for the number of images in both ImageNet and YFCC for a fair comparison and use them for training DINO and CutLER. As Table 11 shows, CutLER’s performance on COCO is robust to the choice of object-centric or non-object-centric datasets as long as the same dataset is used to train DINO and CutLER. This shows the generalization of CutLER to different data distributions. However, training DINO and CutLER with different data leads to worse performance, suggesting the importance of using the same image distribution for learning both DINO and CutLER models.

Summary

Object localization is a fundamental task in computer vision. In this paper, we have shown that a simple yet effective cut-and-learn approach can achieve extraordinary performance on challenging object detection and instance segmentation tasks without needing to train with human annotations. As a zero-shot unsupervised detector, CutLER, trained solely on ImageNet, outperforms the detection performance of previous works by over 2.7×\times on 11 benchmarks across various domains.

References

Appendix A Appendix

While CutLER is agnostic to the underlying detector, we use popular Mask R-CNN and Cascade Mask R-CNN for all experiments, and use Cascade Mask R-CNN by default, unless otherwise noted. We train the detector on ImageNet with initial masks and bounding boxes for 160160K iterations with a batch size of 16. When training the detectors with a ResNet-50 backbone , we initialize the model with the weights of a self-supervised pretrained DINO model. We explored other pre-trained models, including MoCo-v2 , SwAV , and CLD , and found that they give similar detection performance. Therefore, we initialize model weights with DINO by default.

We also leverage the copy-paste augmentation during the model training process. Rather than using the vanilla copy-paste augmentation to improve the model’s ability to segment small objects, we randomly downsample the mask with a scalar uniformly sampled between 0.3 and 1.0. We then optimize the detector for 160K160K iterations using SGD with a learning rate of 0.005, which is decreased by 5 after 80K80K iterations and a batch size of 16. We apply a weight decay of 5 ⁣× ⁣10−55\!\times\!10^{-5} and a momentum of 0.9.

For the multi-round of self-training, in each stage, we initialize the detection model using the weights from the previous stage. We optimize the detector using SGD with a learning rate of 0.01 for 80KK iterations. Since the self-training stage can provide a sufficient number of pseudo-masks for model training, we don’t use the exploration loss during the self-training stage.

A.2 Datasets used for zero-shot evaluation

COCO and COCO20K is a large-scale object detection and instance segmentation dataset, containing about 115KK and 5KK images in the training and validation split, respectively. Additionally, COCO has an unannotated split of 123KK images. We test our model in a class-agnostic manner on COCO val2017 and COCO 20K, without fine-tuning on any images in COCO. COCO 20KK is a subset of the COCO trainval2014 , containing 19817 randomly sampled images, used as a benchmark in . We report class-agnostic COCO style averaged precision and averaged recall for object detection and segmentation tasks.

Pascal VOC is another popular benchmark for object dtetection. We evaluate our model on its trainval07 split in COCO style evaluation matrics.

UVO . Unidentified Video Objects (UVO) is an exhaustively annotated dataset for video object detection and instance segmentation. We evaluate our model on UVO val by frame-by-frame inference and report results in COCO style evaluation matrics.

LVIS collected 2.2 million high-quality instance segmentation masks for over 1000 entry-level object categories, which naturally constitutes the long-tailed data distribution. We report class-agnostic object detection and instance segmentation results on LVIS val split, containing about 5KK images.

CrossDomain contains three subsets of watercolor, clipart, and comics, in which objects are depicted in watercolor, sketch and painting styles, respectively. We evaluate our model on all annotated images from these three datasets, i.e., traintest.

Objects365 V2 presents a supervised object detection benchmark with a focus on diverse objects in the wild. We evaluate CutLER on the 80K images from its val split.

OpenImages V6 unifies image classification, object detection, and instance segmentation, visual relationship detection, etc. in one dataset. We evaluate CutLER on its 42K images from the val split.

KITTI presents a dataset captured from cameras mounted on mobile vehicles used for autonomous driving research. We evaluate CutLER on 7521 images from KITTI’s trainval split.

We provide the summary of these datasets used for zero-shot evaluation in Table 12.

A.3 Additional results for zero-shot detection & segmentation

In this section, we use official COCO API and provide more results with standard COCO metrics, including AP across various IoU thresholds - AP (averaged over IoU thresholds from 0.5 to 0.95 with a step size of 0.05), AP50 (IoU@0.50.5) and AP75 (IoU@0.750.75), and AP across scales - APS{}_{\text{S}} (small objects), APM{}_{\text{M}} (medium objects) and APL{}_{\text{L}} (large objects). We provide detailed results on all these benchmarks listed in Table 12 and report these results in Table 13. We report the performance of object detection for all datasets. In addition, for those datasets that provide annotations for instance segmentation, we also present the performance of the instance segmentation task. It is worth noting that on these datasets without segmentation labels, CutLER can still predict instance segmentation masks, but since we do not have ground truth masks to be compared, we cannot evaluate the results.

A.4 CutLER vs. Selective Search

Selective Search is a popular unsupervised object discovery method, used in many early state-of-the-art detectors such as R-CNN and Fast R-CNN . However, generating possible object locations with sliding windows greatly reduces inference speed (please refer to for more details on selective search). We compare CutLER’s performance to selective search in Fig. 7 and observe that CutLER provides a significant improvement in both precision and recall, which indicates that CutLER is a better performing unsupervised method for region proposal generation with real-time inference speed.

A.5 Training details for label-efficient and fully-supervised learning

We train the detector on the COCO dataset using the bounding box, and instance mask labels. To evaluate label efficiency, we subsample the training set to create subsets with varying proportions of labeled image We train the detector, initialized with CutLER, on each of these subsets.

As a baseline, we follow the settings from MoCo-v2 and train the same detection architecture initialized with a MoCo-v2 ResNet50 model, given its strong performance on object detection tasks. MoCo-v2 and our models use the same training pipeline and hyper-parameters and are trained for the 1×1\times schedule using Detectron2 , except for extremely low-shot settings with 1% or 2% labels. Following previous works , when training with 1% or 2% labels, we train both MoCo-v2 and our model for 3,600 iterations with a batch size of 16.

Our detector weights are initialized with ImageNet-1K pre-trained CutLER, except for the weights of the final bounding box prediction layer and the last layer of the mask prediction head, which are randomly initialized with values taken from a normal distribution. For experiments on COCO with labeling ratios below 50%, during model training, we use a batch size of 16, and learning rates of 0.04 and 0.08 for model weights loaded from the pre-trained CutLER and randomly initialized, respectively. For experiments on COCO with labeling ratios between 50% and 100%, the learning rates of all layers decay by a factor of 2.

For a fair comparison, baselines and CutLER use the same hyper-parameters and settings.

A.6 More visualizations

We provide more qualitative visualizations of CutLER’s zero-shot predictions in Fig. 9.