Cut and Learn for Unsupervised Object Detection and Instance Segmentation
Xudong Wang, Rohit Girdhar, Stella X. Yu, Ishan Misra
Introduction
Object localization is a critical task in computer vision that enables AI systems to perceive, reason, plan and act in an object-centric manner. Training models for localization require special annotations like object boxes, masks, localized points, etc. which are both difficult and resource intensive to collect. Without accounting for overhead, annotating 164K images in the COCO dataset with masks for just 80 classes took more than 28K human hours of annotation time. In this work, we study unsupervised object detection and instance segmentation models that can be trained without any human labels. Our key insight is that simple probing and training mechanisms can amplify the innate localization ability of self-supervised models , leading to state-of-the-art unsupervised zero-shot detectors.
Our method Cut-and-LEaRn (CutLER) consists of three simple, architecture- and data-agnostic mechanisms. Consistent with prior self-supervised learning methods , CutLER is trained exclusively on unlabeled ImageNet data without needing additional training data, but contrary to these methods, CutLER can be directly employed to perform complex segmentation and detection tasks over a wide range of domains. First, we propose MaskCut that can automatically produce multiple initial coarse masks for each image, using the pretrained self-supervised features. Second, we propose a simple loss dropping strategy to train detectors using the coarse masks while being robust to objects missed by MaskCut. Finally, we observe that despite learning from these coarse masks, the detectors ‘clean’ the ground truth and produce masks (and boxes) that are better than the coarse masks used to train them. Therefore, we further show that multiple rounds of self-training on the models’ own predictions allow it to evolve from capturing the similarity of local pixels to capturing the global geometry of the object, thus producing finer segmentation masks.
Prior work shows that a self-supervised vision transformer (ViT) can automatically learn patch-wise features that detect a single salient object in an image . However, unlike CutLER, such salient object detection methods only locate a single, usually the most prominent, object and cannot be used for real world images containing multiple objects. While some recent methods, e.g., FreeSOLO and DETReg , also aim at unsupervised multi-object detection (or multi-object discovery), they rely on a particular detection architecture, e.g., SOLO-v2 or DDETR . Additionally, apart from self-supervised features trained on ImageNet , the current state-of-the-art methods FreeSOLO and MaskDistill also require ‘in-domain’ unlabeled data for model training.
In contrast, CutLER works with various detection architectures and can be trained solely on ImageNet, without requiring in-domain unlabeled data. Thus, during model training, CutLER does not see any images from any target dataset and yields a zero-shot model capable of detecting and segmenting multiple objects in diverse domains.
Features of CutLER. 1) Simplicity: CutLER is simple to train and agnostic to the choice of detection and backbone architectures. Thus, it can be integrated effortlessly into existing object detection and instance segmentation works. 2) Zero-shot detector: CutLER trained solely on ImageNet shows strong zero-shot performance on 11 different benchmarks where it outperforms prior work trained with additional in-domain data. We double the AP performance on 10 of these benchmarks, as shown in Fig. 1, and even outperform supervised detectors on the UVO video instance segmentation benchmark. 3) Robustness: CutLER exhibits strong robustness against domain shifts when tested on images from different domains such as video frames, sketches, paintings, clip arts, etc. 4) Pretraining for supervised detection: CutLER can also serve as a pretrained model for training fully supervised object detection and instance segmentation models and improves performance on COCO, including on few-shot object detection benchmarks.
Related Work
Self-supervised feature learning involves inferring the patterns within the large-scale unlabeled data without using human-annotated labels. Contrastive learning based methods learn such representations that similar samples or various augmentations of the same instance are close to each other, while dissimilar instances are far apart. Similarity-based self-supervised learning methods learn representations via minimizing the distance between different augmentations of the same instance and use only positive sample pairs. Clustering-based feature learning automatically discovers the natural grouping of data in the latent representation space. Recently, have shown that masked autoencoders, which learn representations via masking out a large random subset of image patches and reconstructing the missing pixels or patches , are scalable self-supervised learners for computer vision .
In contrast to these unsupervised representation learning efforts, our work aims to automatically discover natural pixel groupings and locate instances within each image.
Unsupervised object detection and instance segmentation. The main comparisons to previous works are listed in Table 1 and are elaborated as follows:
DINO observes that the underlying semantic segmentation of images can emerge from the self-supervised Vision Transformer (ViT) , which does not appear explicitly in either supervised ViT or ConvNets . Based on this observation, LOST and TokenCut leverage self-supervised ViT features and propose to segment one single salient object from each image based on a graph that is constructed with DINO’s patch features.
These previous works either can not detect more than one object from each image, e.g., DINO and TokenCut, or can not improve the quality of features for better transfer to downstream detection and segmentation tasks, e.g., TokenCut and LOST. Unlike these works, CutLER can locate multiple objects and serve as a pretrained model for label-efficient and fully-supervised learning.
FreeSOLO achieves unsupervised instance segmentation by extracting coarse object masks in an unsupervised manner, followed by mask refinement through a self-training procedure. While FreeSOLO’s FreeMask stage can generate multiple coarse masks per image, the quality of these masks is often rather low . MaskDistill distills class-agnostic initial masks from the affinity graph produced by a self-supervised DINO . However, it utilizes one single mask per image in the distillation stage, which greatly limits the model’s ability to detect multiple objects.
By contrast, the initial masks generated by our MaskCut are usually better in quality and quantity than the initial masks used by . Therefore, CutLER achieves 2\times\raise 0.73193pt\hbox{\scriptstyle\sim}4 higher AP and AP than FreeSOLO and MaskDistill on almost all experimented detection and segmentation benchmarks, even when FreeSOLO and MaskDistill are trained and tested on the same domain.
Method
We tackle the problem of unsupervised object detection and segmentation with a simple cut-and-learn pipeline. Our method builds upon insights from recent work , showing that self-supervised representations can discover objects. While these methods often find a single object per image, we propose a simple approach that can discover multiple objects and significantly improves segmentation and detection performance. The overview of our cut-and-learn pipeline is illustrated in Fig. 2. First, we propose MaskCut that generates multiple binary masks per image using self-supervised features from DINO (Sec. 3.2). Second, we show a dynamic loss dropping strategy, called DropLoss, that can learn a detector from MaskCut’s initial masks while encouraging the model to explore objects missed by MaskCut (Sec. 3.3); Third, we further improve the performance of our method through multiple rounds of self-training (Sec. 3.4).
Normalized Cuts (NCut) treats the image segmentation problem as a graph partitioning task . We construct a fully connected undirected graph via representing each image as a node. Each pair of nodes is connected by edges with weights that measure the similarity of the connected nodes. NCut minimizes the cost of partitioning the graph into two sub-graphs, i.e., a bipartition, by solving a generalized eigenvalue system
for finding the eigenvector that corresponds to the second smallest eigenvalue , where is a diagonal matrix with and is a symmetrical matrix.
DINO and TokenCut. DINO finds that the self-supervised ViT can automatically learn a certain degree of perceptual grouping of image patches. TokenCut leverages the DINO features for NCut and obtaining foreground/background segments in an image. The authors use the similarity of the patches in the DINO feature space as the similarity weight in NCut. Specifically, following multiple recent methods , we use the cosine similarity of ‘key’ features from the last attention layer of DINO-pretrained model, i.e., where is the ‘key’ feature of patch , and solve Eq. 1 for finding the second smallest eigenvector .
A limitation of TokenCut is that it only computes a single binary mask for an image and thus only finds one object per image. Although we can use the other smallest eigenvectors to locate more than one instance, this significantly degrades the performance for multi-object discovery, as demonstrated in Sec. 5.
2 MaskCut for Discovering Multiple Objects
As we discussed in Sec. 3.1, vanilla NCut is limited to discovering a single object in an image. We propose MaskCut that extends NCut to discover multiple objects per image by iteratively applying NCut to a masked similarity matrix (illustrated in Fig. 3). After getting the bipartition from NCut at stage , we get two disjoint groups of patches and construct a binary mask , where
To determine which group corresponds to the foreground, we make use of two criteria: 1) intuitively, the foreground patches should be more prominent than background patches . Therefore, the foreground mask should contain the patch corresponding to the maximum absolute value in the second smallest eigenvector ; 2) we incorporate a simple but empirically effective object-centric prior : the foreground set should contain less than two of the four corners. We reverse the partitioning of the foreground and background, i.e., , if the criteria 1 is not satisfied while the current foreground set contains two corners or the criteria 2 is not satisfied. In practice, we also set all to and to 1.
where . Using the updated , we repeat Eqs. 1 and 2 to get a mask . We repeat this process times and set by default.
3 DropLoss for Exploring Image Regions
A standard detection loss penalizes predicted regions that do not overlap with the ‘ground-truth’. Since the ‘ground-truth’ masks given by MaskCut may miss instances, the standard loss does not enable the detector to discover new instances not labeled in the ‘ground-truth’. Therefore, we propose to ignore the loss of predicted regions that have a small overlap with the ‘ground-truth’. More specifically, during training, we drop the loss for each predicted region that has a maximum overlap of with any of the ‘ground-truth’ instances:
where denotes the maximum IoU with all ‘ground-truth’ for and refers to the vanilla loss function of detectors. does not penalize the model for detecting objects missed in the ‘ground-truth’ and thus encourages the exploration of different image regions. In practice, we use a low threshold .
4 Multi-Round Self-Training
Empirically, we find that despite learning from the coarse masks obtained by MaskCut, detection models ‘clean’ the ground truth and produce masks (and boxes) that are better than the initial coarse masks used for training. The detectors refine mask quality, and our DropLoss strategy encourages them to discover new object masks. Thus, we leverage this property and use multiple rounds of self-training to improve the detector’s performance.
5 Implementation Details
Training data. We only use the images from the ImageNet dataset (1.3 million images) for all parts of the CutLER model and do not use any type of annotations either for training or any supervised pretrained models.
MaskCut. We use MaskCut with three stages on images resized to 480480 pixels and compute a patch-wise affinity matrix using the ViT-B/8 DINO model. We use Conditional Random Field (CRF) to post-process the masks and compute their bounding boxes.
Detector. While CutLER is agnostic to the underlying detector, we use popular Mask R-CNN and Cascade Mask R-CNN for all experiments, and use Cascade Mask R-CNN by default, unless otherwise noted. We train the detector on ImageNet with initial masks and bounding boxes for K iterations with a batch size of 16. When training the detectors with a ResNet-50 backbone , we initialize the model with the weights of a self-supervised pretrained DINO model. We explored other pre-trained models, including MoCo-v2 , SwAV , and CLD , and found that they gave similar detection performance. We also leverage the copy-paste augmentation during the model training process. Rather than using the vanilla copy-paste augmentation, to improve the model’s ability to segment small objects, we randomly downsample the mask with a scalar uniformly sampled between 0.3 and 1.0. We then optimize the detector for K iterations using SGD with a learning rate of 0.005, which is decreased by 5 after K iterations, and a batch size of 16. We apply a weight decay of and a momentum of 0.9.
Self-training. We initialize the detection model in each stage using the weights from the previous stage. We optimize the detector using SGD with a learning rate of 0.01 for 80 iterations. Since the self-training stage can provide a sufficient number of pseudo-masks for model training, we don’t use the DropLoss during the self-training stages.
We provide more details on model implementation and training in Sec. A.1.
Experiments
We evaluate CutLER on various detection and segmentation benchmarks. In Sec. 4.1, we show that CutLER can discover objects without any supervision on completely unseen images. Despite being evaluated in a zero-shot manner on eleven benchmarks, CutLER outperforms prior methods that use in-domain training data. Sec. 4.2 shows that finetuning CutLER further improves detection performance, outperforming prior work like MoCo-V2 and FreeSOLO.
We conduct extensive experiments on eleven different datasets, covering various object categories, image styles, video frames, resolutions, camera angles, etc. to verify the effectiveness of CutLER as a universal unsupervised object detection and segmentation method. We describe the different datasets used for zero-shot evaluation in detail in Sec. A.2. CutLER is trained solely using images from ImageNet and evaluated in a zero-shot manner on all downstream datasets without finetuning on any labels or data.
Evaluating unsupervised object detectors poses two unique challenges. First, since the model is trained without any notion of semantic classes, it cannot be evaluated using the class-aware detection setup. Thus, like prior work we use the class-agnostic detection evaluation. Second, object detection datasets often only annotate a subset of the objects in the images. For example, while COCO and LVIS use the same images, COCO only labels 80 object classes, and LVIS labels 1203 object classes. In this partially labeled setup, Average Recall (AR) is a valuable metric for unsupervised detection as it does not penalize the models for detecting novel objects unlabeled in the dataset. Thus, we additionally report AR for all datasets.
Zero-shot detection on 11 benchmarks. We evaluate CutLER on a variety of datasets and report the detection performance using AP and AR metrics in Fig. 1 and Table 2. CutLER uses a smaller model size and less training data than prior work. Compared to the previous SOTA approach, FreeSOLO with a backbone of ResNet101, CutLER, with the smaller ResNet50 backbone, significantly outperforms it in each of these benchmarks spanning various image distributions, more than doubling performance on 10 of them. Also note that, FreeSOLO requires FreeMask pre-training using approximately 1.3M ImageNet images and model fine-tuning using additional data in test benchmarks.
We observe that on different domains, e.g. watercolor or frames from videos (UVO dataset), CutLER improves performance by over and , respectively. Fig. 1 shows some qualitative examples of CutLER’s predictions.
Detailed comparisons on COCO20K and COCO. Table 3 presents detailed detection and segmentation evaluations (also referred to as ‘multi-object’ discovery) on two popular benchmarks: COCO val2017 and COCO 20K, which contains a subset of 20K images of COCO . CutLER consistently surpasses prior works by a large margin (often gets 23 higher AP) on both the segmentation and detection tasks. Although CutLER is not trained on any images from COCO, it surpasses existing methods trained on COCO by more than 10% in terms of AP and AP.
Fig. 4 shows the qualitative comparisons between and our CutLER on COCO val2017, along with human annotations. Surprisingly, CutLER can often detect novel instances that human annotators miss.
We present detailed comparisons on COCO 20K, COCO val2017 and LVIS benchmarks in Sec. A.3.
Detailed comparisons on UVO and VOC. For a comprehensive comparison with existing unsupervised multi-object detection methods, we report the results for UVO val and VOC trainval07 . Table 4 shows that CutLER yields significant performance gains over previous SOTA, obtaining over 3 higher AP, with the most considerable improvement coming from APL. On UVO, Table 5 shows that CutLER more than quadruples the AP of previous SOTA and almost triples the AP. Our AP is even 4.8% higher than the fully-supervised SOLOv2 trained on LVIS with 100% annotations, significantly narrowing the gap between supervised and unsupervised learning.
2 Label-Efficient and Fully-Supervised Learning
We now evaluate CutLER as a pretraining method for training object detection and instance segmentation models. While CutLER can discover objects without any supervision, finetuning it on a target dataset aligns the model output to the same set of objects labeled in the dataset.
Setup. We use CutLER to initialize a standard Cascade Mask R-CNN detector with a ResNet50 . Prior work uses more advanced detectors, SOLOv2 used in and DDETR used in , that perform better. However, we choose Cascade Mask R-CNN for its simplicity and show in Sec. 5 that CutLER’s performance improves with stronger detectors. We train the detector on the COCO dataset using the bounding box and instance mask labels. To evaluate label efficiency, we subsample the training set to create subsets with varying proportions of labeled images. We train the detector, initialized with CutLER, on each of these subsets. As a baseline, we follow the settings from MoCo-v2 and train the same detection architecture initialized with a MoCo-v2 ResNet50 model, given its strong performance on object detection tasks. Both MoCo-v2 and our models are trained for the schedule using Detectron2 , except for extremely low-shot settings with 1% or 2% labels. Following previous works , when training with 1% or 2% labels, we train both MoCo-v2 and our model for 3,600 iterations with a batch size of 16.
Results. Fig. 5 shows the results of fine-tuning the detector on different subsets of COCO. When tested with low-shot settings, e.g., 2% and 5% labeled data, our approach achieves 5.4% and 7.3% higher AP than the MoCo-v2 baseline, respectively. Even when training with full annotations, CutLER still consistently gives more than 2% improvements, outperforming MoCo-v2 for both object detection and segmentation. More impressively, CutLER outperforms prior SOTA methods - FreeSOLO and DETReg despite using an older detection architecture.
Ablations
We analyze the design decisions in CutLER. We use similar settings to Sec. 4 and train CutLER only on ImageNet. We use the Cascade Mask R-CNN detection architecture and evaluate our model primarily on the COCO and UVO unsupervised detection benchmarks. All ablation studies are conducted without self-training unless otherwise noted.
We analyze the main components of CutLER and report their relative contribution in Table 6. We report results on the popular COCO dataset and a densely annotated video instance segmentation dataset UVO . We also report the performance of running TokenCut on the COCO dataset. Next, we use TokenCut’s official codes to generate masks on ImageNet and use them for training a Cascade Mask R-CNN . This base model provides substantial gains over just using TokenCut on COCO. We add each of our proposed components to this strong base model. Using MaskCut increases AP and AP by 4.7% and 2.7%, respectively. Also, the improvements to AP is larger for densely annotated dataset UVO, i.e. 4.7% vs. 2.7%. These results prove that MaskCut’s ability to segment multiple instances per image is vital for densely annotated datasets. Adding DropLoss brings another 1.6% and 0.9% improvements to AP for UVO and COCO, respectively. Multi-round of self-training increases the quantity and quality of pseudo-masks, leading to 1.3% improvements. These results show that each simple proposed component is critical for strong performance.
Comparison with TokenCut.
TokenCut is also a zero-shot segmentation method. However, it only segments a single instance per image, as discussed in Sec. 3.1. In order to generate more than one segmentation mask per image, we use a modified TokenCut by using more of the smaller eigenvectors and combining all produced masks. Table 7 shows the object detection performance on COCO’s validation set for vanilla TokenCut, our modified TokenCut and CutLER. Although using more eigenvectors increases the recall AR, it significantly reduces the precision AP. CutLER not only improves the average recall AR by 4 but also surpasses TokenCut’s average precision AP by 4.8, i.e. 480% relative improvements.
Design choices in MaskCut and DropLoss
and their impact on the final localization performance is presented in Table 8. We first study the effect of the image size used by MaskCut for generating the initial masks. As expected, LABEL:tab:ablate_image_size shows that MaskCut benefits from using higher resolution images presumably as it provides a higher resolution similarity between pixels. We pick a resolution of 480px for a better trade-off between the speed of MaskCut and its performance. In LABEL:tab:ablate_tau_ncut, we study the effect of the threshold used in MaskCut for producing a binary matrix (Sec. 3.2). Overall, CutLER seems to be robust to the threshold values. We understand the impact of the number of masks per image generated by MaskCut in LABEL:tab:ablate_num_masks. Increasing the number improves the performance of the resulting CutLER models. This shows that MaskCut generates high-quality masks that directly impact the overall performance. Finally, in LABEL:tab:ablate_tau_droploss, we vary the IOU threshold used for DropLoss. With a high threshold, we ignore the loss for a higher number of predicted regions while encouraging the model to explore. works best for the trade-off between exploration and detection performance.
Self-training
and its impact on the final performance is analyzed in Table 9. Self-training consistently improves performance across the UVO and COCO benchmarks and all metrics. UVO, which has dense object annotations, benefits more from the multi-round of self-training. By default, CutLER uses 3 rounds of self-training. Fig. 6 shows qualitative examples of how self-training improves both the quality of predictions and the number of objects predicted.
Generalization to different detection architectures.
We use different detector architectures for training CutLER and measure their performance in Table 10. We observe that CutLER works with various architectures, and its performance is improved with stronger architectures.
Impact of the pretraining dataset.
We now study the impact of the dataset used for 1) pretraining the self-supervised DINO model and 2) training the CutLER model. The commonly used ImageNet dataset has a well-known object-centric bias which may affect the unsupervised detection performance. Thus, we also use YFCC , a non-object-centric dataset. We control for the number of images in both ImageNet and YFCC for a fair comparison and use them for training DINO and CutLER. As Table 11 shows, CutLER’s performance on COCO is robust to the choice of object-centric or non-object-centric datasets as long as the same dataset is used to train DINO and CutLER. This shows the generalization of CutLER to different data distributions. However, training DINO and CutLER with different data leads to worse performance, suggesting the importance of using the same image distribution for learning both DINO and CutLER models.
Summary
Object localization is a fundamental task in computer vision. In this paper, we have shown that a simple yet effective cut-and-learn approach can achieve extraordinary performance on challenging object detection and instance segmentation tasks without needing to train with human annotations. As a zero-shot unsupervised detector, CutLER, trained solely on ImageNet, outperforms the detection performance of previous works by over 2.7 on 11 benchmarks across various domains.
References
Appendix A Appendix
While CutLER is agnostic to the underlying detector, we use popular Mask R-CNN and Cascade Mask R-CNN for all experiments, and use Cascade Mask R-CNN by default, unless otherwise noted. We train the detector on ImageNet with initial masks and bounding boxes for K iterations with a batch size of 16. When training the detectors with a ResNet-50 backbone , we initialize the model with the weights of a self-supervised pretrained DINO model. We explored other pre-trained models, including MoCo-v2 , SwAV , and CLD , and found that they give similar detection performance. Therefore, we initialize model weights with DINO by default.
We also leverage the copy-paste augmentation during the model training process. Rather than using the vanilla copy-paste augmentation to improve the model’s ability to segment small objects, we randomly downsample the mask with a scalar uniformly sampled between 0.3 and 1.0. We then optimize the detector for iterations using SGD with a learning rate of 0.005, which is decreased by 5 after iterations and a batch size of 16. We apply a weight decay of and a momentum of 0.9.
For the multi-round of self-training, in each stage, we initialize the detection model using the weights from the previous stage. We optimize the detector using SGD with a learning rate of 0.01 for 80 iterations. Since the self-training stage can provide a sufficient number of pseudo-masks for model training, we don’t use the exploration loss during the self-training stage.
A.2 Datasets used for zero-shot evaluation
COCO and COCO20K is a large-scale object detection and instance segmentation dataset, containing about 115 and 5 images in the training and validation split, respectively. Additionally, COCO has an unannotated split of 123 images. We test our model in a class-agnostic manner on COCO val2017 and COCO 20K, without fine-tuning on any images in COCO. COCO 20 is a subset of the COCO trainval2014 , containing 19817 randomly sampled images, used as a benchmark in . We report class-agnostic COCO style averaged precision and averaged recall for object detection and segmentation tasks.
Pascal VOC is another popular benchmark for object dtetection. We evaluate our model on its trainval07 split in COCO style evaluation matrics.
UVO . Unidentified Video Objects (UVO) is an exhaustively annotated dataset for video object detection and instance segmentation. We evaluate our model on UVO val by frame-by-frame inference and report results in COCO style evaluation matrics.
LVIS collected 2.2 million high-quality instance segmentation masks for over 1000 entry-level object categories, which naturally constitutes the long-tailed data distribution. We report class-agnostic object detection and instance segmentation results on LVIS val split, containing about 5 images.
CrossDomain contains three subsets of watercolor, clipart, and comics, in which objects are depicted in watercolor, sketch and painting styles, respectively. We evaluate our model on all annotated images from these three datasets, i.e., traintest.
Objects365 V2 presents a supervised object detection benchmark with a focus on diverse objects in the wild. We evaluate CutLER on the 80K images from its val split.
OpenImages V6 unifies image classification, object detection, and instance segmentation, visual relationship detection, etc. in one dataset. We evaluate CutLER on its 42K images from the val split.
KITTI presents a dataset captured from cameras mounted on mobile vehicles used for autonomous driving research. We evaluate CutLER on 7521 images from KITTI’s trainval split.
We provide the summary of these datasets used for zero-shot evaluation in Table 12.
A.3 Additional results for zero-shot detection & segmentation
In this section, we use official COCO API and provide more results with standard COCO metrics, including AP across various IoU thresholds - AP (averaged over IoU thresholds from 0.5 to 0.95 with a step size of 0.05), AP50 (IoU@) and AP75 (IoU@), and AP across scales - AP (small objects), AP (medium objects) and AP (large objects). We provide detailed results on all these benchmarks listed in Table 12 and report these results in Table 13. We report the performance of object detection for all datasets. In addition, for those datasets that provide annotations for instance segmentation, we also present the performance of the instance segmentation task. It is worth noting that on these datasets without segmentation labels, CutLER can still predict instance segmentation masks, but since we do not have ground truth masks to be compared, we cannot evaluate the results.
A.4 CutLER vs. Selective Search
Selective Search is a popular unsupervised object discovery method, used in many early state-of-the-art detectors such as R-CNN and Fast R-CNN . However, generating possible object locations with sliding windows greatly reduces inference speed (please refer to for more details on selective search). We compare CutLER’s performance to selective search in Fig. 7 and observe that CutLER provides a significant improvement in both precision and recall, which indicates that CutLER is a better performing unsupervised method for region proposal generation with real-time inference speed.
A.5 Training details for label-efficient and fully-supervised learning
We train the detector on the COCO dataset using the bounding box, and instance mask labels. To evaluate label efficiency, we subsample the training set to create subsets with varying proportions of labeled image We train the detector, initialized with CutLER, on each of these subsets.
As a baseline, we follow the settings from MoCo-v2 and train the same detection architecture initialized with a MoCo-v2 ResNet50 model, given its strong performance on object detection tasks. MoCo-v2 and our models use the same training pipeline and hyper-parameters and are trained for the schedule using Detectron2 , except for extremely low-shot settings with 1% or 2% labels. Following previous works , when training with 1% or 2% labels, we train both MoCo-v2 and our model for 3,600 iterations with a batch size of 16.
Our detector weights are initialized with ImageNet-1K pre-trained CutLER, except for the weights of the final bounding box prediction layer and the last layer of the mask prediction head, which are randomly initialized with values taken from a normal distribution. For experiments on COCO with labeling ratios below 50%, during model training, we use a batch size of 16, and learning rates of 0.04 and 0.08 for model weights loaded from the pre-trained CutLER and randomly initialized, respectively. For experiments on COCO with labeling ratios between 50% and 100%, the learning rates of all layers decay by a factor of 2.
For a fair comparison, baselines and CutLER use the same hyper-parameters and settings.
A.6 More visualizations
We provide more qualitative visualizations of CutLER’s zero-shot predictions in Fig. 9.