Localizing Objects with Self-Supervised Transformers and no Labels
Oriane Siméoni, Gilles Puy, Huy V. Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, Jean Ponce
Introduction
Object detectors are now part of critical systems, such as autonomous vehicles. However, to reach a high level of performance, they are trained on a vast amount of costly annotated data. Various approaches have been proposed to reduce these costs, such as semi-supervision , weak supervision , active-learning and self-supervision with task fine-tuning.
We consider here the extreme case of localizing objects in images without any annotation. Early works investigate regions proposals based on saliency or intra-image similarity , i.e., only between patches within the considered image (not across the image collection). However, these proposals have low precision and are produced in large quantities only to reduce the search space in other tasks, such as supervised or weakly-supervised object detection. Often using region proposals as input, unsupervised object discovery leverages information from the entire image collection and explores inter-image similarities to localize objects in an unsupervised fashion, e.g., with probabilistic matching , principal component analysis , optimization and ranking . However, because of the quadratic complexity of region comparison among images, together with the high number of region proposals for a single image, these methods hardly scale to large datasets. Other approaches do not require annotations but exploit extra modalities, e.g., audio or LiDAR .
We propose here a simple approach to localize objects in an image, that we then apply to unsupervised object discovery. Our localization method stays at the level of a single image, rather than exploring inter-image similarity, which makes it linear w.r.t. the number of images and thus highly scalable. For this, we leverage high-quality features obtained from a visual transformer pre-trained with DINO self-supervision . Concretely, we divide the image of interest into equal-sized patches and feed it to the DINO model. Instead of focusing on the CLS token, we propose to use the key component of the last attention layer for computing the similarities between the different patches. In doing so, we are able to localize a part of an object by selecting the patch with the least number of similar patches, here called the seed. The justification for this seed selection criterion is based on the empirical observation that patches of foreground objects are less correlated than patches corresponding to background. We add to this initial seed other patches that are highly correlated to it and thus likely to be part of the same object, a process which we call seed expansion. Finally, we construct a binary object segmentation mask by computing the similarities of each image patch to the selected seed patches and infer the bounding box of an object as the box that tightly encloses the largest connected component in this mask that contains the initial seed. In following this simple method, we not only outperform methods for region proposals but also those for single-object discovery. Even more, by training an off-the-self class-agnostic object detector using our localized boxes as ground-truth boxes, we are able to derive a much more accurate object localization model that is actually able to detect multiple objects in an image. We call this task unsupervised class-agnostic object detection (which may resort to self-supervision despite being called unsupervised). Finally, by using clustering techniques to group the localized objects into visual consistent classes, we are able to train class-aware object detectors without any human supervision, but using instead the predicted object locations and their cluster ids as ground-truth annotations. We call this task unsupervised (class-aware) object detection. We show that the predictions of our unsupervised detection model for certain clusters correlate very well with labelled semantic classes in the dataset and reach for them detection results competitive to object detectors trained with weak supervision .
Our main contributions are as follows: (1) we show how to extract relevant features from a self-supervised pre-trained vision transformer and use the patch correlations within an image to propose a simple single-object localization method with linear complexity w.r.t. to dataset size; (2) we leverage it to train both class-agnostic and class-aware unsupervised object detectors able to accurately localize multiple object per image and, in the class-aware case, group them to semantically-coherent classes; (3) we outperform the state of the art in unsupervised object discovery with a significant margin.
Related work
Region proposal methods generate in an unsupervised way numerous class-agnostic bounding boxes with high recall but low precision, to speed-up sliding window search. From supervised pre-trained networks, objects can emerge by masking the input , interpreting neurons or from saliency maps . Weakly-supervised object detection (WSOD) uses image-level labels without bounding boxes to learn to detect objects. The different instances of WSOD (each with specific assumptions on the availability and amount of image-level and box-level annotations) are often addressed as semi-supervised learning and leverage self-training . Recent work replaces manual annotations with automatic supervision from a different modality, e.g., LiDAR or audio . In contrast, we do not use any annotations or other modalities at any stage: we extract object candidates from the activations of a self-supervised pre-trained network, compute pseudo-labels and then train an object detector.
Object discovery.
Given a collection of images, object discovery groups images depicting similar objects, and then localizes objects within these images. Early works focus mostly on the first task and to, a lesser extent, on localization . On the contrary, shift focus on the second task and achieve good object localization on image collections in the wild. However, casting object discovery as the selection of recurring visual patterns across an image collection involves expensive computation and only is able to scale to large datasets. Our work also discovers object locations but does not consider inter-image similarity. Instead, we rely on the power of self-supervised transformer features and only consider intra-image similarity. Consequently, our method can localize objects in a single image with little computation. Close to ours, is also able to localize objects from a single image by exploiting scale-invariant features. Finally, some works on object discovery attempt to simultaneously learn an image representation and to decompose images into object masks. These works, however, are only evaluated on image collections of very simple geometric objects.
Transformers.
In this work, we leverage transformer representations to address object discovery. Self-attention layers have been previously integrated into CNNs , yet transformers for vision are very recent and still in an incipient stage. Findings on training heuristics and architecture design are released at high pace. Early adaptations of transformers to different tasks (e.g., image classification , retrieval , object detection and semantic segmentation ) have demonstrated their utility and potential for vision. Meanwhile, several works attempt to better understand this new family of models from various perspectives . Interestingly, transformers have been shown to be less biased towards textures than CNNs , hinting that their features encapsulate more object-aware representations. These findings motivate us to study manners of localizing objects from transformer features.
Self-supervised learning (SSL)
is a powerful training scheme to learn useful representations without human annotations. It does so via a pretext learning task for which the supervision signal comes from the data itself . SSL pre-trained networks have been shown to outperform ImageNet pre-trained networks on several computer vision tasks, in particular object detection . For transformers, SSL methods also work well , bringing a few interesting side-effects. In particular, DINO feature activations appear to contain explicit information about the semantic segmentation of objects in an image. In the same spirit, we extract another kind of transformer features to build our object localization.
Proposed approach
Our method exploits image representations extracted by a vision transformer. In this section, we first recall how such representations are obtained, then present our method.
Self-attention.
Features for object localization.
2 Finding objects with LOST
Our seed selection strategy is based on the assumptions that (a) regions/patches within objects correlate more with each other than with background patches and vice versa, and (b) an individual object covers less area than the background. Consequently, a patch with little correlation in the image has higher chances to belong to an object.
To compute the patch correlations, we rely on the distinctiveness of self-supervised transformer features, which is particularly noticeable when using transformer’s keys. We empirically observe that using these tranformer features as patch representation meets assumption (a) in practice: patches in an object correlate positively with each other but negatively with patches in the background. Therefore, based on assumption (b), we select the first seed by picking the patch with the smallest number of positive correlations with other patches.
Concretely, we build a patch similarity graph per image, represented by the binary symmetric adjacency matrix such that
In other words, two nodes are connected by an undirected edge if their features are positively correlated. Then, we select the initial seed as a patch with the lowest degree :
We show in Figure 2 examples of seeds selected in four different images. A representation of the degree map for each of these images is also presented. We remark that the patches with lowest degrees are the most likely to fall in an object. Finally, we also observe in this figure that the few patches that correlate positively with are also likely to belong to an object.
Seed expansion.
Once the initial seed is selected, the second step consists in selecting patches correlated with the seed that are also likely to fall in the object. Again, we achieve this step relying on the empirical observations that pixels within an object tend to be positively correlated and to have a small degree in . We select the next best seeds after as the pixels that are positively correlated with : within , the patches with the lowest degree. (In case of patches with equal degrees, we break ties arbitrarily to ensure that .) Note that and a typical value for is .
Box extraction.
The last step consists in computing a mask by comparing the seed features in with all the image features. The entry of the mask satisfies
In other words, a patch is considered as part of an object if, on average, its feature positively correlates with the features of the patches in . To remove the last spurious correlated patches, we finally select the connected component in that contains the initial seed and use the bounding box of this component as the detected object. An illustration of the detected boxes before and after seed expansion is provided in Figure 3.
3 Towards unsupervised object detection
We exploit the accurate single-object localization of LOST for training object detection models without any human supervision. Starting from a set of unlabeled images, each one assumed to contain at least one prominent object, we extract one bounding box per image using LOST. Then, we train off-the-shelf object detectors using these pseudo-annotated boxes. We explore two scenarios: class-agnostic and (pseudo) class-aware training of object detectors.
A class-agnostic detection model localizes salient objects in an image without predicting nor caring about their semantic category. We train such a detector by assigning the same “foreground” category to all the boxes produced by LOST, which we call “pseudo-boxes” afterwards, as they are obtained with no supervision. Unlike LOST, the trained detector can localize multiple objects per image, even if it was trained on a dataset containing only one pseudo-box annotation per image. The experiments confirm that the trained detector can output multiple detections and the quantitative results (Table 1) show that this trained detector is in fact even better than LOST in terms of localization accuracy.
Class-aware detection (OD).
We now consider a typical detector that both localizes objects and recognizes their semantic category. To train such a detector, apart from LOST’s pseudo-boxes, we also need a class label for each of these boxes. In order to remain fully-unsupervised, we discover visually-consistent object categories using K-means clustering. For each image, we crop the object detected by LOST, resize the cropped image to , feed this image in the DINO pre-trained transformer, and extract the CLS token at the last layer. The set of CLS tokens are clustered using K-means and the cluster index is used as a pseudo-label for training the detector. At evaluation time, we match these pseudo-labels to the ground-truth class labels using the Hungarian algorithm , which give names to pseudo-labels.
Experiments
We explore in this section three variants of the object localization problem, in order of increasing complexity: (1) localizing one salient object in each image (single-object discovery) in §4.2, (2) using the corresponding bounding boxes as ground-truth to train a binary classifier for foreground object detection (unsupervised class-agnostic object detection), and (3) using clustering to capture an unsupervised notion of object categories, and detect the corresponding instances (unsupervised object detection). Both are discussed in §4.3. None of the building blocks of this pipeline uses any annotation, just a large number of unlabelled images to sequentially train, in a self-supervised way, the DINO transformer, the class agnostic foreground/background classifier, and finally the classifier using the cluster identifier as labels. Also, we provide more qualitative results in supplementary.
Unless otherwise specified, we use the ViT-S model introduced in , which follows the architecture of DEiT-S . It is trained using DINO , with a patch size of and the keys (without the entry corresponding to the CLS token) of the last layer as input features , with which we achieve the best results. Results obtained alternatively with the attention, the queries and values are presented and discussed in the supplementary material. For comparison, we also present results using the base version of ViT (ViT-B), ViT-S with a patch size of , as well as with features of the last convolutional layer of a dilated ResNet-50 and of a VGG16 pre-trained either following DINO, or in a supervised fashion on Imagenet .
Datasets.
We evaluate the performance of our approach on the three variants of object localization on VOC07 trainval+test, VOC12 trainval and COCO_20K . VOC07 and VOC12 are commonly used benchmarks for object detection . COCO_20k is a subset of the COCO2014 trainval dataset , consisting of 19817 randomly chosen images, used as a benchmark in . When evaluating results on the unsupervised object discovery task, we follow a common practice and evaluate scores on the trainval set of the different datasets. Such an evaluation is possible as the task is fully unsupervised. We follow the same principle for the unsupervised class-agnostic task: we generate boxes on VOC07 trainval, VOC12 trainval and COCO_20k, use them to train a class-agnostic detector, and then evaluate again on these datasets (against ground-truth boxes this time). For unsupervised class-aware object detection, we generate boxes and train the detector on VOC07 trainval and/or VOC12 trainval, but evaluate the detector on the VOC07 test set to facilitate comparisons to weakly-supervised object detection methods. Note that for unsupervised object discovery, some previous works evaluate on subsets of VOC07 trainval and VOC12 trainval. For completeness, we present the object discovery performance of our method on these reduced datasets in the supplemental material.
2 Application to unsupervised object discovery
Similar to methods for unsupervised single-object discovery, LOST produces one box for each image. It therefore can be directly evaluated for this task. Following , we use the Correct Localization (CorLoc) metric, i.e., the percentage of correct boxes, where a predicted box is considered correct if it has an intersection over union (IoU) score superior to with one of the labeled object bounding boxes.
In Table 1, we present the CorLoc of our method, in comparison to state-of-the-art object discovery methods and region proposals .
Despite its simplicity, we see that LOST outperforms the other methods by large margins. We also compare against an adapted version of the segmentation method proposed in . Concretely, we extract the self-attention of the CLS query at the last layer of the transformer, create a binary mask where the largest entries of this self-attention are set to , retrieve the largest spatially-connected component from this binary mask, and use the bounding box of this component as the detected object. This method returns one box per self-attention head and we report results obtained with the best performing head over the entire dataset, noted as DINO-seg. LOST improves over DINO-seg by 8 to 17 of CorLoc points, demonstrating the efficacy of our approach for object localization based on self-supervised pre-trained transformer features.
Finally, we also evaluate our unsupervised class-agnostic detector (denoted by ‘+ CAD’) for single-object discovery. To this end, we return for each image the box that the detector assigns the highest score. It can be seen that training a class-agnostic detector on LOST’s outputs further improves the performance by 4 to 7 CorLoc points. In total, our method surpasses the prior state of the art by at least CorLoc points on each evaluated dataset.
Impact of the backbone architecture.
Table 2 studies the effect of the backbone on LOST. We see that transformer representations are better suited for our method (best results with ViT-S/16). In contrast, our performance using the DINO-pre-trained ResNet-50 is significantly lower. It indicates that the performance of our method is not only due to the contributions of self-supervision but also to the property and quality of the specific features we extract.
3 Unsupervised object detection
Here we explore the application of LOST in unsupervised object detection. To that end, we use LOST’s pseudo-boxes to train a Faster R-CNN model on the datasets. We measure detection performance using the Average Precision at IoU 0.5 metric (AP@0.5), which is commonly used in the PASCAL detection benchmark. As Faster R-CNN backbone, we use a ResNet50 pre-trained with DINO self-supervision, thus making our training pipeline fully-unsupervised. We trained the Faster R-CNN models using the detectron2 implementation (more details in the supplementary material).
To generate pseudo-labels for the class-aware detectors, we apply K-means clustering on DINO-ViT-S tokens using as many clusters as the number of different classes in the dataset. Since the cluster-based pseudo-labels are “anonymous”, to evaluate the detection results we must map the clusters to the ground-truth classes. Following prior work in image clustering , we use Hungarian matching for that. We stress that this matching is only for reporting evaluation results; we do not use any human labels during training.
Unsupervised class-aware detection.
Table 3 provides results of unsupervised class-aware object detectors trained with LOST (entry ‘LOST + OD’). We are not aware of any prior work that addresses unsupervised object detection on real-world images of complex scenes, as those in PASCAL, that does not use extra modalities. We could not compare to as we focus on image-only benchmarks.
We see that, although fully-unsupervised, our method learns to accurately detect several object classes. For example, detection performance for classes “aeroplane”, “bus”, “dog”, “horse” and “train” is more than , and for “cat” it reaches . Even more so, for some classes our method achieves better AP than the weakly-supervised methods WSDDN and PCL , which require image-wise human labels. Although the results are not entirely comparable due to backbone differences between our method and the weakly-supervised ones (self-supervised ResNet50 vs. supervised VGG16), they still demonstrate the efficacy of our method in unsupervised object detection, which is an extremely hard and ill-posed task.
We also evaluate the AP of our pseudo-boxes (with their assigned cluster id as pseudo-labels) when generated for VOC07 test (entry ‘LOST pseudo-boxes’). Evidently, training the detector on pseudo-boxes leads to a significantly higher AP than the initial pseudo-boxes.
Finally, switching our pseudo-boxes with those of rOSD for the detector training (adding pseudo-labels to rOSD pseudo-boxes by clustering DINO features in exactly the same way as in our method) leads to performance degradation (entry ‘rOSD + OD’).
Unsupervised class-agnostic detection.
In Table 4, we report class-agnostic detection results obtained using pseudo-boxes from our method (‘LOST + CAD’) as well as from rOSD (‘rOSD + CAD’) and LOD (‘LOD + CAD’). As we see, our method leads to a significantly better detection performance. We also report detection results using the Selective Search and EdgeBox proposal algorithms, which perform worse than our method.
4 Limitations and future work
Despite the good performance of LOST, it exhibits some limitations.
LOST, as it stands, can separate same-class instances that do not overlap (as it only keeps the connected component of the initial seed to create a box), but it is not designed to separate instances when overlapping. This is actually a challenging problem, related to the difference between supervised semantic and instance segmentation methods, which, as far as we know, is an open problem in the absence of any supervision. A potential lead could be to use a matching algorithm such as Probabilistic Hough Matching to separate instances within image regions found in multiple images.
Another issue is when an object covers most of the image. It violates our second assumption for the initial seed selection (expressed in subsection 3.2) that an individual object covers less area than the background, thus possibly causing the seed to fall in the background instead of a foreground object. Ideally, we would like to filter out such failure cases, e.g., by using the attention maps of the CLS token. We leave this as future work.
Conclusion
We have presented LOST, a simple, yet effective method for localizing objects in images without any labels, by leveraging self-supervised pre-trained transformer features . Despite its simplicity, LOST outperforms state-of-the-art methods in object discovery by large margins. Having high precision, the boxes found by LOST can be used as pseudo ground truth for training a class-agnostic detector which further improves the object discovery performance. LOST boxes can also be used to train an unsupervised object detector that yields competitive results compared to weakly-supervised counterparts for several classes.
Future work will be dedicated to investigate other applications of LOST boxes, e.g., high-quality region proposals for object detection tasks, and the power of self-supervised transformer features for unsupervised object segmentation.
Acknowledgments and Disclosure of Funding
This work was supported in part by the Inria/NYU collaboration, the Louis Vuitton/ENS chair on artificial intelligence and the French government under management of Agence Nationale de la Recherche as part of the “Investissements d’avenir” program, reference ANR19-P3IA-0001 (PRAIRIE 3IA Institute). Huy V. Vo was supported in part by a Valeo/Prairie CIFRE PhD Fellowship.
References
Appendix A Ablation Study
As explained in subsection 3.2 of the main paper, we chose to use the keys of the last attention layer as patch features in LOST. As we will see here, this choice provides the best localization performance among other alternatives. Specifically, in the first section of Table 5, we report the performance of LOST when using as patch features either the keys , the queries , or the values of the attention layer. We see that, when using the queries or the values , LOST’s performance deteriorates by at least 11 CorLoc points compared to using the keys .
Another way to measure the similarity between two patches in a transformer architecture is to use the scalar product between the queries and the keys. We thus test substituting
Results in Table 5 show that all these alternatives using queries and keys yield results that are not as good as when using the keys as patch features.
A.2 Importance of the seed expansion step
We analyse here the importance of the seed expansion step that is controlled by . The seed expansion step allows us to enlarge the region of interest so as to include all the parts of an object and not only the part localized from the first, initial seed.
Table 5 presents the impact of the parameter , which corresponds to the maximum number of patches that can be used to construct the mask . We notice that, without seed expansion (i.e., ), there is a drastic drop in the localization performance. The performance then improves when increasing to - with a slight decrease at .
Visualizations of results with and are presented in the Figure 3 of the main paper and Figure 4 here. We see that the boxes in yellow obtained with are small and localized on probably what is the most discriminative part of the objects. Increasing permits us to increase the size of the box and localize the object better. We also present in Figure 5 cases of failures where the seed expansion step is either insufficient to localize the whole object or yields a box containing multiple objects.
A.3 Analysis of DINO-seg
In this section, we investigate alternative setups of the baseline DINO-seg which is based on the work of Caron et al. . They are presented in Table 6.
First, instead of using the best attention head over the entire dataset (as we did in the main paper), we evaluate the localization accuracy of DINO-seg for each one of the 6 available heads. We find out that one head in particular, namely head 4, captures objects well, whilst results with other heads are much lower. Due to its superior performance, in the main paper we report DINO-seg using head 4.
We also explore dynamically selecting one box per image among boxes corresponding to the different heads using some heuristics. We report the two variants that gave the best results. In the first variant, we consider selecting the box corresponding to the head with the biggest connected component (‘DINO-seg BCC’). However, it yields worse results than with head 4. We also try selecting, over the 6 boxes of the different heads, the box that has the highest average IoU overlap with the remaining 5 boxes (‘DINO-seg HAIoU’). It improves over DINO-seg [head 4] by 1 point on both VOC07 and VOC12. However, as shown in Table 6, it still performs significantly worse than LOST in this single-object discovery task.
A.4 Impact of the number of clusters on class-aware detection training
For the unsupervised class-aware detection experiments of the main paper, we assume that we know the exact number of object classes present in the used dataset, i.e., in the VOC dataset, and use the same number of K-means clusters. Here we only assume that we have a rough estimate of the number of classes and study the impact of the requested number of clusters on the performances of the unsupervised detector.
To that end, in Table 7, we provide the mean AP across all the VOC classes when using , , and clusters. For the case when we use more clusters than the classes of the VOC dataset, Hungarian matching, which is used for reporting the AP results, will map to the VOC classes only the most fitted available clusters. Thus, when reporting the per-class AP results, we ignore the detections in these unmatched clusters (since they have not been mapped to any ground-truth class).
In Table 7, we observe that our unsupervised detector achieves good results for all the numbers of clusters. Interestingly, for and clusters there is a noticeable performance improvement. Similar findings have been observed on prior clustering work .
A.5 Impact of the non-determinism of the K-means clustering
We investigate the impact of the randomness in the K-means clustering on the results of the object detector. To that end, we repeat 4 times, using different random seeds, the unsupervised class-aware object detection experiment with LOST + OD† (using the model trained on the union of VOC07 and VOC12 trainval sets, cf. Table 3 in subsection 4.3 of the main paper). We obtain a standard deviation of 0.8 for the AP@0.5 %, which shows that the method is fairly insensitive to the randomness of the clustering method.
Appendix B More quantitative results and comparisons
For completeness, we present in Table 8 results on the datasets used in previous object discovery works . In particular, we evaluate our method on the datasets VOC07noh and VOC12noh datasets (also named VOC_all in literature). They are subsets of the trainval set of the well-known PASCAL VOC 2007 and PASCAL VOC 2012 datasets containing 3550 and 7838 images respectively. These subsets exclude all images containing only objects annotated as “hard” or “truncated” and all boxes annotated as “hard” or “truncated”.
B.2 Multi-object discovery results
We compare in Table 9 the object discovery performance of different methods in the setting where multiple regions are returned per image. This setting has been explored in and .
Following , instead of considering the object recall (detection rate) for a given number of predicted regions per image, as in , we consider as a metric a form of Average Precision adapted to the task, that we name here “odAP”. It is the average of the AP of predicted objects for each number of predicted regions, from one to the maximum number of ground-truth objects in an image in the dataset. This odAP metric thus does not depend on the number of detections per image and remains related to AP, which is a standard metric for object detection. actually uses two variants of this metric: odAP50, where a prediction is correct if its intersection-over-union (IoU) with one of the ground-truth boxes is at least , and odAP@, the average odAP value at 10 equally-spaced values of the IoU threshold between and .
As LOST only returns one region per image, we only consider here LOST + CAD, which is the output of a class-agnostic detector (CAD) trained with LOST boxes, and we compare it to other existing approaches. It can be seen that LOST + CAD outperforms significantly all the previous methods, including the class-agnostic detector trained with LOD boxes (LOD + CAD).
B.3 Image nearest neighbor retrieval
Following LOD , we use LOST box descriptors to find images that are similar to each other (image neighbors) in the image collection.
To this end, each image is represented by the CLS descriptors of its LOST box and the cosine similarity between these descriptors is used to define a similarity between the images. Then, for each image, the top images with the highest similarity are chosen as its neighbors. Similar to LOD , we choose and use CorRet as the evaluation metric, defined as the average percentage of the retrieved image neighbors that are actual neighbors (i.e., that contain objects of the same category) in the ground-truth image graph over all images.
We compare the performance of our method in this task with rOSD and LOD in Table 11. We see that LOST boxes, when represented by DINO features, yield the better CorRet score compared to . When VGG16 features are used, LOST is behind LOD but better than rOSD .
B.4 Using DINO features
We are aware that, in Table 1 of the main paper, we compare our method using a transformer backbone to methods based on a VGG16 pre-trained on ImageNet models. For a fair comparison, we investigate here the state-of-the-art LOD method when adapted to use the transformers features.
LOD uses the algorithm from rOSD to generate region proposals from CNN features for their pipeline, but we observe that this algorithm does not yield good proposals with transformer features. We therefore run LOD with edgeboxes and use DINO features, extracted with ROIPool , to represent these proposals. We present the results on VOC07_trainval, VOC12_trainval and COCO20k_trainval dataset in Table 11.
Our results in Table 2 of the main paper show that a direct adaption of LOST, designed by analysing the properties of transformers features, to CNN features yields worse performance. Conversely, as we see in Table 11 here, adapting algorithms developed using properties of CNN features to transformer features is also not direct. Nevertheless, the number of design choices to adapt these algorithms to new types of features is vast and we do not exclude that some design choices might improve the results even further, e.g., by exploiting together CNN and transformer features.
B.5 Using supervised pre-training.
We test LOST but this time using a transformer pre-trained under full supervision on ImageNet. We use the model provided by DeiT .
With this model, LOST achieves a CorLoc of % which is significantly worse than the results obtained with the DINO self-supervised pre-trained model. We remark that a similar observation was made for DINO , where the segmentation performance obtained with the model trained under full supervision yields significantly worse results than when using DINO’s model. It is unclear, however, if this difference of performance can be attributed to the properties of the self-supervision loss or to the more aggressive data augmentation used during DINO pre-training.
Appendix C More visualizations (single- and multi-object discovery)
We present in Figures 4-9 additional qualitative results of our method.
Figure 4 and Figure 5 are discussed in the subsection A.2.
Figure 6 and Figure 7 show successful examples of LOST + CAD in VOC07_trainval and COCO20k_trainval datasets. It can be seen that it is able to localize multiple objects in the same image.
Figure 8 and Figure 9 present results obtained with LOST + OD on the VOC07 and COCO datasets respectively. They show the localization predictions with their predicted pseudo-classes. Each pseudo-class is assigned a different color. In Figure 9, the “person” objects are assigned three different pseudo-classes; those failures show the difficulty to assign the same class to “person” in very different positions.
Appendix D Training details of the Faster R-CNN detection models
In the main paper, we explore the application of LOST in unsupervised object detection by using its pseudo-boxes as ground truth for training Faster R-CNN detection models.
For the implementation of the Faster R-CNN detector, we use the R50-C4 model of Detectron2 that relies on a ResNet-50 backbone. In our experiments, this ResNet-50 backbone is pre-trained with DINO self-supervision. Then, to train the Faster R-CNN model on the considered dataset, we use the protocol and most hyper-parameters from He et al. .
In details, we train with mini-batches of size across GPUs using SyncBatchNorm to finetune BatchNorm parameters, as well as adding an extra BatchNorm layer for the RoI head after conv5, i.e., Res5ROIHeadsExtraNorm layer in Detectron2. During training, the learning rate is first warmed-up for steps to and then reduced by a factor of after K and K training steps. We use in total K training steps for all the experiments, except when training class-agnostic detectors on the pseudo-boxes of the VOC07 trainval set, in which case we use K steps. For all experiments, during training, we freeze the first two convolutional blocks of ResNet-50, i.e., conv1 and conv2 in Detectron2.