Unsupervised Salient Object Detection with Spectral Cluster Voting

Gyungin Shin, Samuel Albanie, Weidi Xie

Introduction

Salient object detection (SOD) In contrast to object detection (which aims to localise and recognise objects with bounding boxes), salient object detection aims to segment foreground objects by predicting pixel-wise masks for them., which aims to group pixels that attract human visual attention, has been extensively studied in the field of computer vision due to its wide range of applications such as photo cropping , re-targeting in images and video .

In the literature, early work tackled this problem by utilising low-level features (e.g., colour ) together with priors on salient regions in an object such as contrast priors , boundary priors and centre priors . Recent SOD models have approached this task from the perspective of representation learning, typically training deep neural networks (DNNs) on a large-scale dataset with manual annotations. However, the scalability of such supervised learning approach is limited because it is a costly process to collect ground-truth mask annotations.

To overcome the necessity of large-scale human annotation, many unsupervised methods for saliency detection/object segmentation have recently been proposed . Despite these efforts, the gap between unsupervised and fully supervised SOD methods remains significant.

Interestingly, however, it has been noted that recent self-supervised models such as DINO exhibit significant potential for object segmentation despite the fact that their training objective does not explicitly encourage pixel grouping. The focus of this work is to leverage this observation to propose a simple yet effective mechanism for extracting object regions from self-supervised features that can be employed for the task of unsupervised salient object detection.

To this end, we explore the use of spectral clustering , a classical graph-theoretic clustering algorithm, and find that it can generate useful segmentation candidates across a range of self-supervised features (i.e., DINO , MoCov2 , and SwAV ). Motivated by this finding, we propose a simple winner-takes-all voting strategy to select the salient object masks among a collection of clusters produced by repeated applications of spectral clustering to self-supervised features. In this work, we use the terms cluster and mask interchangeably. Specifically, by mask, we mean a (one-hot) mask which encodes the spatial extent of a cluster. In particular, we base our voting strategy on two priors: The first is a framing prior that a salient object should not occupy the full image height or width; The second is a distinctiveness prior that assumes that salient regions are sufficiently distinctive that they will appear as clusters among an appropriately constructed collection of redundant re-clusterings of the data. We then show that the selected salient masks can be employed as pseudo-labels to train a saliency estimation network that achieves state-of-the-art results on a variety of benchmarks.

In summary, we make the following contributions: (i) We revisit spectral clustering and highlight its benefits over kk-means clustering as a proposal mechanism to identify object regions from self-supervised features on three popular salient object detection (SOD) datasets; (ii) We propose an effective voting strategy to select the most salient object mask in an image among multiple segmentations generated from different self-supervised features by leveraging saliency priors; (iii) Using the salient masks as pseudo ground truth masks (pseudo-masks), we train an object segmentation model, SelfMask, that outperforms the previous unsupervised saliency detection approaches on three SOD benchmarks.

Related work

Our work relates to two themes in the literature: self-supervised representation learning and unsupervised saliency detection. Each of these is discussed next.

There has been a great deal of interest in self-supervised approaches to learning visual representations that obviate the requirement for labels. These include techniques for solving proxy tasks, such as predicting patch locations , patch discrimination , grouping through co-occurrence , colourisation , jigsaw puzzles , common fate principles , clustering , and instance discrimination . Our work is inspired by recent work exploring the use of self-supervised transformers , and by the analysis provided by who noted that the self-attention of Vision Transformers (ViT) are capable of highlighting spatially coherent object regions in their input. While prior work has demonstrated that self-supervised pretraining can be effective for semantic segmentation when coupled with end-to-end supervised metric learning on the target dataset, we instead seek a simple way to exploit self-supervised features for object segmentation without annotation.

2 Unsupervised saliency detection

Supervised object segmentation requires pixel-wise annotations which are time-consuming to acquire. Seeking to avoid this cost, many attempts have been made to solve the task in an unsupervised fashion. Prior to the dominance of deep neural networks, a broad range of handcrafted methods were proposed based on one or more priors relating to foreground regions within an image such as the contrast prior , centre prior , and boundary prior . However, these handcrafted approaches suffer from poor performance relative to recent DNN-based models, described next.

Generative models. A common approach for DNN-based unsupervised object segmentation is to utilise generative adversarial networks (GANs) . Specifically, given an image, a generator is adversarially trained to produce an object mask which will be used to composite a realistic image by copying the corresponding object region in the image into a background which is either synthesised or taken from a different image . An alternative family of approaches aims to discover a direction in the latent space of a pre-trained GAN that can be used to segment foreground and background regions . Then, a saliency detector is trained on a synthetic data set composed of pairs of images together with their foreground masks generated via the discovered latent space structure. In contrast, we seek to exploit representations learned via self-supervision by discriminative, rather than generative, models.

Noisy supervision with pseudo-labels. More closely related to our work, the use of weak pseudo saliency masks for training a DNN has been proposed. SBF , the first attempt to train saliency detector without human annotation, proposed to train a model with superpixel-level pseudo-masks generated by fusing weak saliency maps from multiple unsupervised methods (\ie, ). Similarly, USD aims to learn from diverse noisy pseudo-labels obtained via distinct unsupervised handcrafted methods in such a way that a saliency detector trained with the pseudo-masks can predict a saliency map free from label noise. DeepUSPS proposed to refine the pseudo-masks for images produced by handcrafted saliency methods, by training segmentation networks via a self-supervised iterative refinement process. The refined pseudo-masks are then combined from different handcrafted methods to train a final segmentation network. In contrast to the methods above, we use neither handcrafted saliency methods nor an iterative refinement strategy, resulting in a simpler learning framework.

Object segmentation properties of self-supervised vision transformers. Another line of related work has sought to investigate the observation that self-supervised ViTs exhibit object segmentation potential. LOST propose to pick a seed patch from such a ViT that is likely to contain part of a foreground object, and then expand the seed patch to different patches sharing a high similarity with the seed patch. Concurrent work, TokenCut , proposes to use Normalised Cuts to segment the salient object among the final layer self-attention key features of ViT. SelfMask differs from the TokenCut approach to salient object detection in two key ways: (1) While we similarly employ spectral clustering as part of our pipeline, we demonstrate the significant additional value of integrating cues from diverse re-clusterings via voting to bootstrap a pseudo-labelling process; (2) Thanks to the flexibility of our clustering approach, we are able to leverage saliency cues from self-supervised convolutional neural networks (CNNs) as well as ViT architectures, and show the benefits of doing so. We compare our approach with theirs in section 4.

Method

In this section, we begin by formalising the problem scenario (Sec. 3.1) and briefly summarise spectral clustering (Sec. 3.2). Then, we introduce our approach to address unsupervised salient object detection by selecting pseudo ground truth masks via spectral clustering, and train a saliency prediction network, called SelfMask (Sec. 3.3).

Traditionally this has been treated as a clustering problem where the key challenge lies in designing effective features for accurately describing salient regions. In this work, we instead look for a simple yet effective solution by leveraging self-supervised visual representations.

2 Segmentation with spectral clustering

Conceptually, segmentation is obtained via spectral clustering with pixel-wise image features by projecting the features onto a representation space such that image partitions can be decided by directly comparing similarities between their corresponding features.

where the degree of a vertex fi∈V\mathbf{f}_{i}\in\mathbf{V} is defined as di=∑j=1Nwijd_{i}=\sum_{j=1}^{N}w_{ij}, and the degree matrix D\mathbf{D} is defined with the degrees d1,…,dnd_{1},\dots,d_{n} on the diagonal. Given the adjacency matrix W\mathbf{W} and the degree matrix D\mathbf{D}, the (un-normalised) graph Laplacian L\mathbf{L} is defined as:

Given L\mathbf{L}, we can solve the generalised eigenproblem:

Finally, a set of clusters C\mathcal{C} is obtained by running kk-means algorithm on the row vectors of a matrix U\mathbf{U} (see supplementary for more detail), producing regions with all pixels from the corresponding cluster. Note that, at this stage, the resulting clusters are composed of both object and background masks—the object mask itself will be selected by our selection strategy described next.

3 Supervision with pseudo-mask from spectral clusters

Here, we first introduce our voting strategy for selecting a salient mask from a set of spectral clusters from different features and multiple kk values (\ie, cluster numbers), which utilises a framing prior and distinctiveness prior. Then, we describe our model, SelfMask, which is trained by using the selected salient masks as pseudo-masks for supervision.

To choose the salient object among the mixture of foreground and background masks generated by spectral clustering, we propose a voting strategy based on two observations: (1) The spatial extent of an object rarely occupies the entire height and width of an image. (2) Salient object regions are likely to appear in multiple clusters across different self-supervised features as well as with different cluster numbers kk. In other words, among kk clusters from an application of clustering, we assume that at least one cluster encodes an object region within the image and that this holds for different features (\eg, DINO, MoCov2, or SwAV). We call these priors the framing prior and distinctiveness prior, respectively. Note that the framing prior bears a resemblance to the centre prior , which states that salient objects are likely to be located near the center of an image, and the boundary prior , which presumes that the foreground object rarely touches the boundary of an image. However, the framing prior differs from these priors in that it is not related to a location of an object but rather the scale of an object within an image.

To use the framing prior and distinctiveness prior in practice, we first form a candidate set of masks by repeatedly applying spectral clustering to different features with multiple kk values. Then, we treat masks whose spatial extent is as long as the width or height of the image as background masks, and eliminate them from the candidate set. Finally, we employ winner-takes-all voting: we pick the mask with the highest average pair-wise similarity w.r.t. IoU among all remaining masks as the final mask for salient objects (bottom of Fig. 2). There are two edge cases in the background elimination process we handle explicitly: (i) when no masks are left in the candidate set and (ii) when only two masks are left, sharing the same IoU. For the former case, we simply keep all the masks in the candidate set as every mask highlights regions spanning the spatial extent of the image, breaking the assumption of the framing prior. For the latter, we break ties randomly to pick one of the two masks.

3.2 SelfMask

Here, we describe our model architecture, training, and inference procedure.

Architecture. We base our salient object detection network, called SelfMask, on a variant of the MaskFormer architecture which was originally proposed for the semantic/instance segmentation task.

Objective. For training, we employ two objective functions: a mask loss and a ranking loss, denoted by Lmask\mathcal{L}_{\text{mask}} and Lrank\mathcal{L}_{\text{rank}}, respectively. Given nqn_{q} mask predictions for an image from the model, we encourage all predictions to be similar to the pseudo-mask. Specifically, following , we use the Dice coefficient , which considers the class-imbalance between foreground and background regions within an image, as the mask loss. It is worth noting that, unlike , we do not include the focal loss in the mask loss since we find that it hinders convergence.

To decide which prediction best highlights the salient region in the image among the proposed candidates when nq>1n_{q}>1, we rank the predicted masks based on their objectness score. Specifically, we first re-order the indices of the predicted masks by their mask loss in ascending order such that

for any i<ji<j, where Mpseudo\mathbf{M}_{\text{pseudo}} denotes the target pseudo-mask for the image. Then, we enforce the objectness score oio_{i} of the mask Mi\mathbf{M}_{i} to be higher than the scores ojo_{j} for any j>ij>i. As a consequence the model is encouraged to produce a higher score for a predicted mask that more closely resembles the pseudo-mask than other predictions. We instantiate this ranking loss as a hinge loss :

Overall, our final objective function is as follows:

where λ\lambda is a weighting factor, which is set to 1.0 across our experiments. Following , we compute the loss for outputs from each layer of the transformer decoder.

Inference. During inference, given nqn_{q} predicted masks for an image and their objectness score, we pick the mask with the highest score as the salient object detection and binarise it with a fixed threshold of 0.5.

Experiments

In this section, we first describe the datasets used in our experiments (Sec. 4.1) and provide implementation details (Sec. 4.2). We then conduct ablation study (Sec. 4.3) and report our results for salient object detection (Sec. 4.4).

We use DUTS-TR , which contains 10,553 images, to train our model with the pseudo-masks generated by following Sec. 3.3. We emphasize that only images are used for generating pseudo-masks and training, without the corresponding labels. For our ablation study and comparison to previous work, we consider five popular saliency datasets including DUT-OMRON , which comprises 5,168 images of varied content with ground-truth pixel level masks; DUTS-TE , containing 5,019 images selected from the SUN dataset and ImageNet test set ; ECSSD which contains 1,000 images that were selected to represent complex scenes; HKU-IS which consists of 4,447 scene images with foreground/background sharing the similar appearances; SOD which contains 300 images with many images having multiple salient objects.

2 Implementation details

Networks. We use the ViT-S/8 architecture for the encoder, a bilinear upsampler with a scale factor of 2 for the pixel decoder, and 6 transformer layers for the transformer decoder. For the MLP applied to per-mask queries (with a dimensionality of 384) that outputs a scalar value for the objectness score, we use three fully-connected layers with a ReLU activation between them. We set the same number of units for the hidden nodes as the input (\ie, 384) and output a single value followed by a sigmoid.

Training details. We train our models for 12 epochs and optimise all parameters including the backbone encoder using AdamW with a learning rate of 6e-6 and the Poly learning rate policy . For data augmentation, we use random scaling with a scale range of [0.1, 1.0], random cropping to a size of 224×\times224 and random horizontal flipping with a probability of 0.5. In addition, photometric transformations include random color jittering, color dropping, and Gaussian blurring are applied. We run each model with three different seeds and report the average.

Metrics. In our experiments, we report intersection-over-union (IoU), pixel accuracy (Acc) and maximal FβF_{\beta} score (max FβF_{\beta}) with β2\beta^{2} set to 0.3 following . Please refer to the supplementary for more details on these metrics.

3 Ablation study

In this section, we first conduct experiments to compare the effectiveness of spectral clustering and kk-means when applied to self-supervised image encoders. Next, we quantitatively verify the performance of our winner-takes-all voting strategy for foreground mask selection and compare it to different saliency selection methods. Lastly, we investigate the effect of different number of queries on SelfMask.

We compare spectral clustering against a kk-means clustering baseline on three salient object detection benchmarks. As the resulting segmentations from each algorithm are agnostic to foreground/background regions, we consider a best-case evaluation for both algorithms. In detail, given a groundtruth mask, we pick the cluster with the highest IoU w.r.t. the groundtruth. Such IoUs act as an upper-bound score among the clusters (\ie, average best overlap in ).

To account for the effect of the cluster number kk on performance of the clustering algorithms, we consider different kk values from 2 to 4 and average the results. For the full results with each kk value, please refer to the supplementary.

As shown in Tab. 1, we observe that object masks from spectral clustering consistently outperform kk-means masks by a large margin. Interestingly, however, when using fully-supervised image encoders, the performance gain of spectral clustering diminishes (Tab. 2). These findings boil down to a simple summary: while using self-supervised visual representations for grouping, spectral clustering is considerably superior to kk-means, regardless of the choice of encoder architecture and self-supervised learning algorithm.

3.2 Voting for salient object masks

Here, we conduct experiments to assess our voting method for selecting a foreground mask among the mask candidates. Since these experiments include ablations across hyperparameter choices, we conduct them on the HKU-IS and SOD , rather than the benchmarks used to compare to prior work.

In detail, we construct an initial corpus of mask candidates by clustering different self-supervised features with different number of clusters, as described in Sec. 3.2. For this, we build the candidate set with spectral clusters from different combinations of self-supervised features, \eg, MoCov2/DINO or SwAV/MoCov2. It is important to do so to allow our voting-based method to leverage the distinctiveness prior across different features. In addition, for each combination, we experiment with 3 different kk value settings: kk=22, {2,3}\{2,3\} or {2,3,4}\{2,3,4\} to account for various object scales, \eg, a lower kk tends to cover large regions, while a higher kk segments smaller objects. We then evaluate the selected masks on the HKU-IS set. For reference we also compute an upper bound IoU, which is computed in a similar way as done in the previous section.

Effectiveness of clustering various self-supervised models with different number of clusters. As shown in Tab. 3, we make two observations: (i) IoU of both selected masks and upper bound masks, denoted by pseudo-mask and UB improves by increasing kk across all feature combinations; (ii) using all three features (\ie, DINO, MoCov2 and SwAV) results in better pseudo-masks than using two of the three features (\eg, MoCov2 and SwAV). These support the distinctiveness prior, which assumes that at least one cluster represents a foreground region, and its application to our voting-based saliency selection method.

Effectiveness of the proposed voting approaches. We further validate the effectiveness of our voting scheme by comparing to different selection methods, \ie, random selection and a centre prior -based strategy. Specifically, we form a candidate set using DINO/MoCov2/SwAV features with k={2,3,4}k=\{2,3,4\}, which amounts to 27 masks in total. For the random strategy, we simply pick one of the masks uniformly from the candidates. For the centre prior selection strategy, we choose a mask whose average Euclidean distance to the image centre from its constituent pixel locations is lowest. In addition, we consider each method with or without utilising the framing prior to assess the influence of filtering out background mask. We evaluate the IoU of each case on the HKU-IS and SOD benchmarks.

As can be seen in Tab. 4, deploying the framing prior boosts IoU in all considered selection methods, with the proposed voting selection method performing best. The framing prior plays a crucial role in the voting process: voting without this prior performs much worse than its counterparts in both random and center-based selections on HKU-IS, and performs similarly to the random strategy on SOD. This is caused in large part by mistakes when selecting background masks as salient objects.

3.3 The influence of the number of queries

As described in Sec. 3.3, we train SelfMask using the selected salient masks as pseudo-masks. Here, we investigate the effect of the number of queries nqn_{q} in the Transformer decoder. For this, we train our model with nqn_{q}={5, 10, 20, 50, 100} on DUTS-TR and evaluate performance on the HKU-IS benchmark in terms of maxFβF_{\beta} for two settings, (i) using ground-truth masks to pick the best mask out of nqn_{q} mask predictions, denoted as the SelfMask upper bound (UB); (ii) taking the mask with the highest object score, denoted SelfMask.

As shown in Figure 4, both SelfMask and SelfMask UB are fairly robust to the number of queries, initially increasing slightly with this hyperparameter (\ie, predictions) before degrading after 20 queries. We conjecture that this is because a handful of queries are enough to localise the salient objects, while further predictions may make it challenging to appropriately rank the objectness of each prediction. For this reason, in the section that follows, we consider SelfMask with 20 queries, and pick the query with the highest objectness as our prediction during inference.

4 Comparison to state-of-the-art unsupervised saliency detection methods

To compare with existing works on unsupervised SOD, we evaluate on three popular SOD benchmarks in terms of Acc., IoU, and maxFβF_{\beta}. Following , we also report results after post-processing predictions with the bilateral solver . As shown in Tab. 5, while the pseudo-masks from spectral cluster voting already perform reasonably well compared to previous models, our self-trained model outperforms all existing approaches on all benchmarks. This suggests both that the model can learn to generalise effectively from noisy masks, and that the objectness score trained with the ranking loss is effective for picking the best salient mask.

5 Broader Impact

This work contributes a new framework to deliver performant unsupervised salient object detection. As such, it offers the potential to underpin a range of societally beneficial applications that are bottlenecked by annotation costs. These include improved low-cost medical image segmentation, crop measurement from aerial imagery, and wildlife monitoring. However, low-cost segmentation is a powerful dual-use technology, and we caution against its deployment as a tool for unlawful surveillance and oppression.

Conclusion

In this work, we address the challenging problem of unsupervised salient object detection (SOD). For this, we first observe that self-supervised features exhibit significantly greater object segmentation potential with spectral clustering than with kk-means. Inspired by this observation, we extract foreground regions among multiple masks generated with multiple types of features, and varying cluster numbers based on winner-takes-all voting. By using the selected masks as pseudo-masks, we train a saliency detection network and show promising results compared to previous unsupervised methods on various SOD benchmarks.

Acknowledgements. GS is supported by AI Factory, Inc. in Korea. WX is supported by Visual AI (EP/T028572/1). SA would like to thank Z. Novak and N. Novak for enabling his contribution. GS would like to thank Jaesung Huh for proof-reading.

References

Appendix A Normalised spectral clustering algorithm

Here, we describe the normalised spectral clustering algorithm used to generate pseudo-masks for our model in Alg. 1.

Note that the adjacency matrix W\mathbf{W} is computed using Eqn. 2 of the main paper, given the dense features from a visual encoder described next.

Appendix B Visual encoder

Our approach utilises image representations learned by either convolution-based or transformer-based architectures to which spectral clustering will be applied. Here, we first briefly review how these feature representations are computed with each model.

where the parameters for the CNNs are omitted for simplicity.

B.2 Transformer-based visual encoder

where FTransformer\mathbf{F}_{\text{Transformer}} denotes the dense features from a transformer-based encoder.

where xi\mathbf{x}_{i} denotes the iith patch.

Appendix C Descriptions of evaluation metrics

In the following, we describe the metrics used for evaluation:

FβF_{\beta} is the harmonic mean of precision and recall between a ground-truth G∈{0,1}H×WG\in\{0,1\}^{H\times W} and a binarised mask M∈{0,1}H×WM\in\{0,1\}^{H\times W}:

where β2\beta^{2} denotes a weight of precision.Precision=tptp+fp=\frac{tp}{tp+fp} and Recall =tptp+fn=\frac{tp}{tp+fn} where tptp, fpfp, and fnfn represent true-positive, false-positive, and false-negative, respectively. Following previous work , we set β2\beta^{2} to 0.3, putting more weight on precision. We use FβF_{\beta} to compute the maximal-FβF_{\beta}, described next.

maximal-FβF_{\beta} (maxFβF_{\beta}) is a maximum score of FβF_{\beta} among multiple masks binarised with different thresholds. Specifically, given a non-binarised mask prediction with its value between , it computes FβF_{\beta} from 255 binarised masks, each of which is thresholded by an integer among {0,...,254}\{0,...,254\} and takes the maximum FβF_{\beta} value for the result.

Intersection-over-union (IoU) is the size of overlapped foreground regions between a ground-truth GG and a binarised mask prediction MM divided by the total size of foreground regions from GG and MM.

Accuracy (Acc) is a metric that measures pixel-wise accuracy based on a ground-truth mask GG and a binarised mask prediction MM:

where δ\delta denotes the Kroneker-delta.

Appendix D Comparison between k𝑘k-means and spectral clustering

In Sec. 4.3 of the main paper, we show the performance of kk-means and spectral clustering applied to different architectures (i.e., ResNet50 and ViT-S/{8, 16}) and features (i.e., fully- and self-superivsed features) averaged over kk={2,3,4}\{2,3,4\} on the three saliency datasets. Here, we show the full results for each kk in Tab. 6. For the description, please refer to Sec. 4.3 of the main paper.

Appendix E Visualisation of failure cases

In Fig. 5, we visualise some failure predictions from our model on the DUT-OMRON and DUTS-TE datasets.

We notice there are two typical failure cases. First, when a salient object is of small scale, the model tends to undersegment it and prefers the large salient object. For instance, as shown by the top left example in Fig 5, the whole bed is segmented, rather than the pillow; Second, when there are more than one salient region in the image, our model may only segment one of them. For example, as shown by the middle right example in Fig 5, both screen and seats can be thought of as a salient region while the model only highlights only the latter. We conjecture that these cases are caused by a bias of the dataset (\ie, DUTS-TR ) on which the model is trained. That is, the training images likely to contain large salient regions composed of either an object or objects sharing a semantic meaning, thus discouraging the model from predicting a small salient region or more than one object with different semantics even if all the objects can be regarded salient.