Unsupervised Object Localization: Observing the Background to Discover Objects

Oriane Siméoni, Chloé Sekkat, Gilles Puy, Antonin Vobecky, Éloi Zablocki, Patrick Pérez

Introduction

The task of object localization — either performed by detecting or segmenting objects — is required in many safety-critical systems such as self-driving cars. Today’s best methods train large deep models on large sets of labeled data . To mitigate such needs in annotation, it is possible to use strategies such as semi-supervised , weakly-supervised and active learning .

In this work, we consider the unsupervised object localization task, which consists in discovering objects in an image with no human-made annotation. This task has recently received a lot of attention as it is a solution to detect objects in a scene with no prior about what they should look like or which category they should belong to. Early works exploit hand-crafted features and inter-image information but hardly scale to large datasets. Recent works leverage strong self-supervised features learned using pretext tasks: localize a single object per image just by exploiting a similarity graph at the level of an image; proposes to combine different self-supervised representations, in an ensemble-fashion, and trains a model to learn the concept of object, the same as what does. However, most of these methods make assumptions about what an object is. For example, assume that an image contains more background pixels than object pixels, while discards masks that fill in the width of an image. Such hypotheses restrict objects one can find.

In this work, we propose to tackle the problem the other way around: we make no assumptions about objects but focus instead on the concept of background. Then, we use the idea that a pixel not belonging to the background is likely to belong to an object. Doing so, we do not need to make hypotheses about the number or the size of objects in order to find them. Our method, named FOUND, is cheap both at training and inference time.

We start by computing a rough estimate of the background mask; this step works by mining a first patch that likely belongs to the background. To do this, we leverage attention maps in a self-supervised transformer and select one of the patches that received the least attention. Then the background mask incorporates patches similar to this mined one. One of our contributions is a reweighting scheme to reduce the effect of noisy attention maps based on the sparsity concept. In the second step, we use the fact that the complement of this background mask provides an approximate estimation of the localization of the objects. This estimate is refined by training a single conv1×1conv1\times 1 layer on top of the frozen self-supervised transformer, using only the masks computed in the first step, an edge-preserving filter, and a self-labelling procedure. We show that this cheap method allows us to reach state-of-the-art results in the tasks of saliency detection, unsupervised object discovery and semantic segmentation retrieval.

We propose to think about the object discovery problem upside-down, and to look for what is not background instead of directly looking for objects.

We propose a new way to exploit already self-trained features and show that they allow us to discover the concept of background.

We show that the use of attention heads can be improved by integrating a weighting scheme based on attention sparsity.

We propose a lightweight model composed only of a single conv1×1conv1\times 1 layer and show that there is no need to train a large segmenter for the task.

We demonstrate that our model performs well on unsupervised saliency detection, unsupervised object discovery and unsupervised semantic segmentation retrieval tasks. We reach state-of-the-art results in all tasks with a method much faster and lighter than competing ones.

Related work

In self-supervised learning, a model is trained to solve a pretext task (e.g., jigsaw solving, colorization, or rotation prediction) on unlabeled data . Recently, with the surge of Vision Transformers (ViT) that stand out compared to convolutional networks, one can obtain rich, and dense descriptors of image patches with models trained in a self-supervised fashion on massive amounts of data . For example, DINO employs a teacher-student framework where the two networks see different and randomly transformed input parts and the student network learns to predict the mean-centered output of the teacher network. In MAE , patches of the input image are randomly masked and the pretext task aims at learning to reconstruct the missing pixels by auto-encoding. In these works, it has been shown that the representations of the self-attention maps of the ViTs contain interesting localization information , which have led recent methods to exploit these properties in several downstream tasks as unsupervised object discovery or semantic segmentation . In this paper, we build upon such self-supervised features to partition background and foreground patches. Arguably, learning self-supervised representation on unlabaled Imagenet — a curated dataset — induces a certain supervision. We leave for future work using models trained on less curated and more heterogenous datasets.

Unsupervised object localization.

Localizing objects within images without any supervision is in the literature traditionally addressed by two distinct branches: 1) unsupervised saliency detection methods find binary masks of objects while 2) unsupervised object detection seeks for bounding boxes around objects . Unsupervised saliency detection has been approached with hand-crafted methods , generative adversarial models , or, closer to us, by refining noisy labels . The first attempts in an unsupervised object discovery have often used region proposals as input. These works explored a collection of images and inter-image information using methods such as principal component analysis , optimization or ranking .

Recently, these historically distinct tasks have been tackled jointly in unified frameworks building on the advent of aforementioned self-trained dense visual features . Given an image, these methods create a weighted graph where each node is a patch, and edges represent the similarity between the patches. Foreground objects are segmented by leveraging this similarity. In particular, LOST uses this graph to mine an object seed as the patch with the least connection to other patches and expands the zone of interest to all connected similar patches afterwards. Building on LOST, TokenCut and Deep Spectral Methods refine this result by using a normalized graph-cut to separate an object from the highly connected patches, which most likely depict the background.

Another line of methods proposes to compute mask proposals that are later refined. SelfMask explores the use of multiple self-supervised features as the input of a spectral clustering algorithm. FreeSOLO proposes FreeMask that generates correlation maps which are then ranked and filtered by a maskness score. DINOSAUR performs representation learning by separating the features of an image and reconstructing them into individual objects or parts.

It should be noted that these prior works make strong underlying assumptions about what an object is. This includes priors about the contrast , the size , the centerness , the shape or boundary of the sought object. Instead, in our work, by looking for the background, we do not need to make any assumptions about the presence or number of objects.

Learning to generalize through training.

While we build our seed masks from single-image information, we refine these masks in a self-training step that leverages information shared across the whole image collection. This self-training step aims at improving the quality of predictions by propagating and refining the initial seed of pseudo-annotations to a large set of unlabeled instances. Early works in unsupervised saliency detection learn a deep unsupervised saliency network from noisy predictions obtained from handcrafted methods . After clustering self-supervised features, train a Class-Agnostic Detection (CAD) network over predicted pseudo-boxes and show that this trained detector can smooth out poor discoveries, therefore boosting results. Similarly, in semantic segmentation, FreeSOLO and COMUS feed coarse masks to train a segmentation model on these pseudo masks .

When propagating and refining pseudo-labels through the dataset with training, previous methods generally employ heavy training procedures involving learning several millions of parameters. Instead, our self-training step is extremely lightweight and fast as it is only composed of one layer of 1×\times1 convolutions and a two-epoch training scheme.

Our method FOUND

In this work, we tackle the unsupervised object localization task by considering the problem upside-down. Our approach consists of two stages. First, we propose to look for patches corresponding to the background in order to highlight patches that are likely objects (Sec. 3.1). Then, starting from these coarse masks, we design a fast and lightweight self-supervised learning scheme to refine them (Sec. 3.2). An overview of FOUND is shown in Fig. 2.

To identify the background, we start by identifying one patch which likely belongs to the background. This patch, called the background seed, is defined as the patch with the least attention in A\mathbf{A} — a patch which the model has learned to not give too much attention to. This seed is the sths^{\rm th} patch, where

In the equation above, Api\mathbf{A}_{pi} is the attention score between the CLS token and the pthp^{\rm th} patch in the ithi^{\rm th} attention head.

Reweighting the attention heads.

When observing the hh different attention maps in A\mathbf{A}, we notice that the background appears more or less clearly in the different heads. Therefore, we propose to weight each head differently in Eq. 1. We exploit the sparsity of the attention map to compute these weights since the background appears better in a sparse attention map (as illustrated in the supplementary materials). Inspired by , we compute the sparsity SiS_{i} of each map by counting the number of attention values above a certain threshold μ>0\mu>0:

and reweight each attention map in (1) by

Notice that wiw_{i} increases when the sparsity SiS_{i} decreases, i.e., when we visually observe a clearer separation of the background from the foreground thanks to sparser attention maps. Finally, Eq. 1 becomes

Discovery of the background.

2 Refining masks with self-training

The proposed background discovery method described above is able to segment a good portion of the background but the corresponding masks are still far from perfect, as observed in Fig. 3. To improve them, and therefore to better segment the foreground objects, we propose a very simple refinement step of learning a lightweight segmentation head in a self-supervised fashion.

The segmentation head consists of a single 1×11\times 1 convolution. For each patch, it compresses DINO frozen patch features into a scalar, which is passed into a sigmoid function to encode the probability of the patch belonging to the foreground. We stress that, unlike most recent works, we do not train a heavy segmentation backbone or a detection model . This aspect brings considerable training and inference efficiency both in terms of time and memory, as studied in Sec. 4.4.

The segmentation head is trained in a self-supervised fashion. The general idea is that the model learns to predict a smoothed version of the complement of the coarse background masks and its own prediction, such that it quickly converges to refined masks. We describe it formally below.

Self-training is done thanks to two losses with distinct roles. The first objective consists of initializing and guiding the predictions toward the coarse background masks. The second objective aims at smoothing and refining predictions.

Experiments

In this section, we make several experiments to assess the quality of FOUND. We first evaluate it on the tasks of unsupervised object discovery (Sec. 4.1), unsupervised saliency detection (Sec. 4.2), and unsupervised semantic segmentation retrieval (Sec. 4.3). Besides, we compare training/inference costs of the different methods in Sec. 4.4, discuss qualitative results in Sec. 4.5, and measure the impact of different components of our method in Sec. 4.6.

In all experiments, we use a ViT-S/8 architecture pre-trained with . Following , we use the key features of the last attention layer as F\mathbf{F} and we use τ=0.3\tau=0.3 in the background discovery step. The parameter μ\mu in Eq. 2 is computed per image as the overall mean attention over all heads. We use the coarse masks as pseudo ground-truth for m=100m=100 iterations before refining the predictions directly. We balance the losses by setting λ=1.5\lambda=1.5. Similar to , we train FOUND on DUTS-TR (10,553 images) for 500500 iterations with a batch of 5050 images — corresponding to a bit more than 22 epochs. We follow a similar training protocol as SelfMask : we use random scaling with a range of [0.1,3.0][0.1,3.0] followed by image resizing to (224,224)(224,224) and Gaussian blurring applied with probability 0.50.5. We use the parameters of the bilateral solver as provided by .

In our evaluation, we consider two protocols: ‘FOUND – single’ and ‘FOUND – multi’. In the ‘single’ mode, we select the biggest connected component in M\mathbf{M}. In the ‘multi’ mode, we consider the mask as is — with all detected objects. Additionally, when applying the bilateral solver ζ()\zeta(), we extract, similarly, either the biggest connected component (single), or all connected components (multi). When not specified, we are using the ‘multi’ setup.

1 Unsupervised object discovery

We first evaluate our method on the task of unsupervised object discovery. We follow the common practice and use the trainval sets of PASCAL VOC07 & VOC12 datasets and COCO20k (a subset of 19,81719,817 randomly chosen images from the COCO2014 trainval dataset following ). As in , we report results with the Correct Localization (CorLoc) metric. It measures the percentage of correct boxes, i.e., predicted boxes having an intersection-over-union greater than 0.50.5 with one of the ground-truth boxes.

In Tab. 1, we compare FOUND – single (no bilateral solver) to methods with no learning phase (LOST , TokenCut , DSS ), and to methods with a learning phase (SelfMask , FreeSOLO , and DINOSAUR ). FreeSOLO predicts multiple instance masks per image and, as such, we propose to merge all instances into a single mask, this gave us the best results. Other choices are discussed in the supplementary materials. For SelfMask , if the mask contains multiple connected components, only the largest one is considered.

We show that FOUND achieves state-of-the-art results on 22 out of the 33 datasets while being much cheaper to train. Indeed, the best method, DINOSAUR, achieves results significantly better than all others on COCO20k, but performs representation learning at a much higher training cost (as discussed in Sec. 4.4). We note that it also achieves worse results than our method on VOC12 (-5.75.7pt). We discuss qualitative results in Sec. 4.5 and in the Supplemental.

2 Unsupervised saliency detection

We then consider the unsupervised saliency detection task, which is typically evaluated on a collection of datasets depicting a large variety of objects in different backgrounds. To compare to previous works, we evaluate on three popular saliency datasets: DUT-OMRON (5,168 images), DUTS-TE (5,019 images), ECSSD (1,000 images). We report results in terms of intersection-over-union (IoU), pixel accuracy (Acc) and maximal FβF_{\beta} score (max FβF_{\beta}) with β2=0.3\beta^{2}=0.3 following (additional details are given in the supplementary materials).

Tab. 2 presents our results compared to state-of-the-art methods, including LOST, DeepSpectralMethods (denoted DSS in the table), TokenCut and the trained SelfMask . When no bilateral solver is used, we observe that our method outperforms all methods, showing that we trained a good saliency estimator which produces high quality object masks. With the application of the bilateral solver, we reach the same or better scores than the other methods, except for the IoU on DUT-OMRON. We observed that the bilateral solver sometimes amplifies the under-segmentation observed in the input mask (visual examples can be found in the supplementary materials). Correcting this behaviour is left for future work.

3 Semantic Segmentation Retrieval

In this section, we test our method on the task of unsupervised semantic segmentation retrieval on the PASCAL VOC12 dataset in order to evaluate the quality of the predicted saliency masks. We follow a protocol proposed by and compare to related methods whose code is available online, namely to TokenCut , SelfMask and FreeSOLO . We also include a comparison to MaskContrast , which takes the opposite approach to ours as it trains the feature representations while having a frozen pre-trained saliency predictor. We consider two different evaluation setups. First, (a) we assume that the predicted mask depicts a single object. For FreeSOLO , which generates several instances per image, we tried several combinations and merged all instances into a single one or consider only the largest instance (noted “largest inst.”). (b) We test the multiple-instances setting, which is more fair to FreeSOLO, and allows us to evaluate the ability of FOUND to separate objects. In this setup, we consider each instance of FreeSOLO as an object. For all other methods, we compute the connected components in the mask outputs, and each component is then treated as an object (we discard those smaller than 1%1\% of an input image size).

Given an object mask, we compute a per-object feature vector averaged over the corresponding pixels. We apply this procedure both in the train and val splits. We use a ViT-S/8 trained using DINO as a feature extractor for FOUND, TokenCut, SelfMask, and FreeSOLO. MaskContrast uses its own optimized feature extractor. Finally, we find the nearest neighbors of each object of the val set to objects in the train set and assign them the corresponding ground-truth label. We measure the mean Intersection-over-Union (mIoU) between the predictions and ground truths.

Results in Tab. 3 are given for both setups and are computed either over 77 (bus, airplane, car, person, cat, cow and bottle) or all 2121 classes of the VOC dataset, following . We can observe that FOUND outperforms all methods in both cases by a consistent margin. Results also confirm SelfMask as a strong competitor that is however outperformed by FOUND across all considered setups with gaps between 1.3 and 2.2 mIoU points, excepting the single saliency with 7 classes evaluation where SelfMask surpasses FOUND by 0.5 point. Improvements of FOUND over TokenCut and FreeSOLO can be explained because TokenCut localizes only a single object per image and FreeSOLO finds objects that are often not considered as so in the dataset. We continue the discussion in Sec. 4.5.

4 Comparison of method costs

We compare FOUND to methods that either do or do not include training, and that have very different costs at inference time. In this section, we highlight the advantage of our method in terms of complexity and speed. FOUND is a segmenter head composed of just 770770 parameters, trained over 22 epochs on DUTS-TR on a single GPU, and which can infer at 8080 FPS, including the forward pass through DINO, on a V100 GPU. We summarize key numbers in Tab. 4.

First, regarding methods with no training, requires the costly computation of an eigenvector on the Laplacian matrix of the affinity graph, therefore making the method rather slow (0.4 FPS). For the same reasons, runs at equivalent speed to . LOST is almost as fast as us but achieves much lower performance, as seen before.

Second, regarding methods that include training, SelfMask trains a model of ≈36\approx 36M parameters over 1212 epochs on DUTS-TR , by exploiting 2727 mask proposals generated using three different backbones, thus making the training considerably more expensive than ours. FreeSOLO proposes a faster mask proposal extraction step using a DenseCL model based on a ResNet backbone. It then trains a SOLO model (≈65\approx 65M learnable parameters) for in total 6060k iterations on 8 GPUs, making it much more expensive to train compared to us.

5 Qualitative results

We show visualizations of saliency masks predicted by FOUND and related methods in Fig. 4. We notice that FreeSolo and SelfMask tend to oversegment the objects in all examples, while FOUND yields masks much more accurate with respect to the ground truth. Regarding TokenCut , we observe, in the last row of the figure, that it segments just a part of one chair, while FOUND segments all the chair rather accurately. These examples illustrate the efficiency of our method in dealing with multiple objects.

6 Ablation study

We present in Tab. 5 an ablation study of our method on the saliency dataset ECSSD — more can be found in Appendix. We measure scores on the unsupervised saliency detection task following the protocol detailed in Sec. 4.2.

We evaluate our background discovery method (Sec. 3.1) with and without the attention head reweighting scheme (column R in Tab. 5). We can observe that the reweighting boosts results up to 11pt when evaluated in a multi-setup mode. We also compare results with and without the application of the post-processing bilateral solver, noted ζ()p\zeta()_{p}, and observe that the refined masks yield better results by 33pts of IoU in the “single” setting. Such improvements (visualized in Fig. 3 and the supplementary materials) are significant. Overall, our background discovery method (Sec. 3.1) already achieves decent results, particularly when considering the single setup. As discussed before and observed in Fig. 3, our coarse maps cover several objects and do not focus only on the most salient one.

The impact of learning

In the same table, we present results obtained after the training of the single conv1×1conv1\times 1 layer. Training over coarse masks provides a significant boost of more than 1515 IoU pts in the multi setup. This shows that the model learns the concept of foreground objects and smooth results over the dataset. Using the bilateral solver in Eq. 6-7, noted ζ()t\zeta()_{t}, further improves results by 1.7 IoU pts and by an additional .6 pts when also applied as post-processing.

Discussion

In this work, we address the problem of unsupervised object localization, that we propose to attack sideways: we look first for the scene background — using self-supervised features — instead of looking for the objects directly. Putting this simple idea at work, we extract coarse masks that encompass most of the background, their complements thus highlighting objects. Using the inverse of the background masks, we train a lightweight segmenter head made of only 770770 learned parameters, which runs at 8080 FPS at inference time — including the forward pass through the backbone — and reaches state-of-the-art results in unsupervised object discovery, unsupervised saliency detection, and unsupervised instance segmentation retrieval.

This work was supported by the HPC resources of GENCI-IDRIS in France under the 2021 grant AD011013413, and by the ANR grant MultiTrans (ANR-21-CE23-0032), It was also supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90140) and by CTU Student Grant SGS21184OHK33T37.

References

Appendix A Extra details

During training, ζ()\zeta() is applied at the image resolution. To do so, masks are upsampled to the original image size and the output refined masks are downsampled to the feature map size. The model is trained with the AdamW optimizer provided by PyTorch, with an initial learning rate of 5e ⁣− ⁣25\text{e}\!-\!2. We use a simple step scheduler which applies a decay of 0.950.95 every 5050 iterations.

A.2 Unsupervised saliency detection

We detail here the different metrics used in the task of unsupervised saliency detection.

is the maximum FβF_{\beta} over various masks which have been binarized using different thresholds. Formally, FβF_{\beta} is the harmonic mean of precision (P) and recall (R) between a binary mask MM and the ground-truth mask GG, i.e.,

where β2\beta^{2} is the precision weight, set at 0.30.3 following . The max FβF_{\beta} is computed by taking a soft predicted mask Mp∈M_{p}\in and binarizing it using 255255 different thresholds between and 254254; max FβF_{\beta} is then the maximum value of FβF_{\beta} among all the generated binary masks, taken over the whole dataset (single optimal threshold). We noticed in SelfMask’s code that the maximal FβF_{\beta} is computed with an optimal threshold found for each image rather than over the whole dataset. For this reason, and for a fair comparison, we do not report this original max FβF_{\beta} in our unsupervised saliency detection table.

The Intersection-over-Union​​​

measures the overlap between foreground regions of a predicted binary mask and the ground-truth mask, averaged over the entire dataset.

The pixel accuracy metric​​​

measures the pixel-wise accuracy between a predicted binary mask M∈{0,1}H×WM\in\{0,1\}^{H\times W} and the corresponding ground-truth mask G∈{0,1}H×WG\in\{0,1\}^{H\times W}. Formally, it can be defined as:

with δ\delta being the Kroneker-delta function and GijG_{ij}, MijM_{ij} being the value of the ground-truth and predicted masks at position (i,j)∈{1⋯H}×{1…W}(i,j)\in\{1\cdots H\}\times\{1\ldots W\}.

A.3 Different setups for FreeSOLO

FreeSOLO is a class-agnostic instance segmentation method and outputs several instance masks per image, making it different to other baselines. In order to compare it to our method, we use the code provided online. We follow the original paper to get the prediction masks, i.e., we apply matrix non-maximum suppresion (NMS) and keep masks with a maskness score above 0.70.7.

We present in Sec. 4.1 of the main paper our unsupervised object discovery protocol. The extraction of the single object box is straightforward for all methods but FreeSOLO . For this method we have considered three setups: (a) merging all instance masks into a single one; (b) keeping only the mask with the highest maskness score; (c) keeping only the mask containing the largest connected component. Best results were achieved with (a) and are reported in the main paper.

Semantic segmentation retrieval

We have performed similar tests with FreeSOLO in the semantic segmentation retrieval task. Additionally to the evaluation setups described in the main paper, we have experimented using two or more instances but without improvements of the results.

A.4 Semantic segmentation retrieval

These prototypes are first extracted for all train samples and serve as an index for retrieval. Then, to get a label for each val sample, we compute the sample prototype, find nearest neighbors in the train prototypes, and assign it the corresponding label.

Appendix B Sensitivity to masking method

We investigate here the impact of the background parameter τ\tau on final results. We report in Fig. 5 saliency detection results. We observe that FOUND is stable to changes of τ ⁣∈ ⁣[0.1,0.5]\tau\!\in\![0.1,0.5], with saliency scores varying by at most 0.2 percentage pts on DUT-OMRON and not at all on ECSSD.

B.2 Using masks from other methods

We investigate here the performance of our method when considering different mask generators. In particular, we consider the well-known object discovery methods TokenCut and LOST with which we extract the masks Mf\mathbf{M}^{f} that are then refined in our training process (following Sec. 3.2 of the main paper). We present the corresponding unsupervised object discovery results in Tab. 6. They show that our method is agnostic to the mask generator but still performs slightly better with our foreground masks — the complement of the background masks described in Sec. 3.1. It is also to be noted that our method is much faster than TokenCut because we do not need the computation of eigenvectors.

Appendix C Additional qualitative results

We present in this section more visualizations of FOUND results, first on more challenging images (Sec. C.1) and at the different step of our process (Sec. C.2). We then motivate the interest of reweighting the transformer heads (Sec. C.3) via visual illustration. Following we show examples where the application of the bilateral solver impacts negatively the results (Sec. C.4) and some more general failure cases of FOUND (Sec. C.5). We finally provide example of discovered objects as performed in the task of unsupervised object discovery (Sec. C.6).

We present in Fig. 6 some results of FOUND random images taken from the Internet. These results show the ability of FOUND to discover multiple and diverse objects, both in terms of classes and scales. In particular, dinosaurs and spaceships are not depicted in ImageNet nor DUT-TR and yet FOUND can detect them, showing the ability to discover objects which “are not background.” Moreover, the selected images here are non-object centric and out-of-domain showing the capacity of FOUND to go behond ImageNet-like images.

C.2 Visualization of masks at different steps

We provide in Fig. 7 additional visualizations of the masks generated at different steps of our method. We can observe that each step brings an improvement over the previous one. The right-most column presents the final output of FOUND without any refinement.

C.3 Reweighting the attention heads

We provide in Fig. 8 a visualization of the self-attention maps extracted from the last layer of our model. We show the self-attention obtained over the six heads; we can observe that the 4th4^{\rm th} head is noisy. When looking for the background seed, we are looking for the pixel with least attention. Our reweighting scheme helps in reducing the weight given to such noisy heads automatically and improves results, as shown in Tab. 5 of the main paper.

C.4 Potential negative effect of the bilateral solver

While the application of ζ()\zeta(), the bilateral solver , improves results in general (see Fig. 7), there are cases where ζ()\zeta() actually degrades the mask quality. We show examples of such cases in Fig. 9 both on coarse masks (rows 1 and 2) and on the final outputs (rows 3 and 4). We can observe that the function amplifies the under-segmentation, e.g., on the hat and the leopard head and legs (row 1 and 2). Moreover, long and thin segments can disappear, e.g., human and animal legs or arms (row 3). Correcting this behaviour would help improving our training and is left for future work.

C.5 Examples of failures cases

We show some failure cases of FOUND in Fig. 10. For these cases, we also present the results obtained with one of the best competitor: SelfMask . We observe that night or dark scenes are challenging (first two rows). Our method tends to under-segment objects but SelfMask has also difficulties in segmenting correctly the main objects in these situation. FOUND, just like SelfMask, is also not robust to reflection on water (third row). Finally, we observe that both methods fail to segment the hair in the fourth columns.

C.6 Unsupervised object discovery results

We present in Fig. 11, qualitative results for the unsupervised single object discovery task (no refinement is applied to the masks). We draw the extracted bounding box on top of the corresponding predicted mask. The conclusions here are similar to those discussed in the main paper. Overall our method segments the objects of interest better and provides cleaner boundaries.