Group Collaborative Learning for Co-Salient Object Detection
Qi Fan, Deng-Ping Fan, Huazhu Fu, Chi Keung Tang, Ling Shao, Yu-Wing Tai
Introduction
Co-salient object detection (CoSOD) targets at detecting common salient objects sharing the same attributes given a group of relevant images. CoSOD is more challenging than the standard salient object detection (SOD) task, because CoSOD needs to distinguish co-occurring objects across multiple images in presence of other objects. That is, both intra-class compactness and inter-class separability should be simultaneously maximized. With this favorable feature CoSOD is thus often employed as a pre-processing step for various computer vision tasks, such as image retrieval , image quality assessment , collection-based crops , co-segmentation , semantic segmentation , image surveillance , video analysis , video co-localization , \etc.
Previous works attempt to leverage the consistency among relevant images to facilitate CoSOD within an image group by exploring different shared cues or semantic connections . Some of them use predicted saliency maps by computing various inter-image cues to discover co-salient objects. Other works exploit a unified network to jointly optimize co-saliency information and saliency maps.
Despite their promising results, most current models only extract their CoSOD representations in an individual group, which introduces a number of limitations. First, images from the same group contain similar foregrounds (\ie, co-salient objects) only provide positive relations while lacking the negative relations between different objects. Training the model only using positive pairs may lead to overfitting and result in ambiguous results for outlier images. Moreover, the number of images in a group is typically limited (20 to 40 images for most CoSOD datasets), so using a single group cannot provide enough information for learning a discriminative representation. Finally, individual groups also fall short in offering high-level semantic information, which is necessary for distinguishing noisy objects during inference in complex real-world scenarios.
To address the above issues, we propose a novel group collaborative learning framework (GCoNet) to mine the semantic correlation between different image groups. The proposed GCoNet consists of three important components: group affinity module (GAM), group collaborating module (GCM) and auxiliary classification module (ACM), which simultaneously learn the intra-group compactness and inter-group separability. The GAM makes the network learn the consensus feature within the same image group, while the GCM discriminates target attributes between different groups, thus enabling the model to be trained on the existing large-scale SOD datasets.Note that the existing CoSOD datasets altogether contain about 6k images, while there are more than 12 SOD datasets, containing about 60k images. It may partially alleviate the insufficient training data issue in co-salient object detection. We further improve the feature representation at a global semantic level through our ACM on each image to learn a better embedding space. In summary, our contributions are:
We introduce a novel group collaborative learning strategy to address the CoSOD problem, and validate its effectiveness with extensive ablation studies.
We design a novel unified Group Collaborative Learning Network (GCoNet) for CoSOD by simultaneously considering intra-group compactness and inter-group separability to mine the consensus representation.
Our group affinity module (GAM) and group collaborating module (GCM) collaborate with each other to achieve better intra- and inter-group collaborative learning. The auxiliary classification module (ACM) further promotes learning at a global semantic level.
Extensive experiments on three challenging CoSOD benchmarks, \ie, CoCA, CoSOD3k, and Cosal2015, show that our GCoNet achieves the new state-of-the-art. Furthermore, we present two downstream applications based on our technical contributions, \ie, co-segmentation and co-localization.
Related Work
The traditional salient object detection task targets at directly segmenting salient object in each image separately, while CoSOD aims to segment the common salient objects across several relevant images. Previous works mainly exploit inter-image cues to detect co-salient objects. Early CoSOD methods explore the inter-image correspondence between image-pairs or a group of relevant images based on shallow handcrafted descriptors . They employ different approaches to mine the inter-image relationships using constraints or heuristic characteristics. Several studies attempt to capture the inter-image constraints by employing an efficient manifold ranking scheme to obtain guided saliency maps, or using a global association constraint with clustering , or translational alignment . Other works attempt to formulate the semantic attributes shared among images in a group from the high-level features in the heuristic characteristics, using multiple saliency cues and self-adaptive weights , regional histograms and constrasts , metric learning by optimizing a new objective function , or pairwise similarity ranking and linear programming .
Recently deep-based models simultaneously explore the intra- and inter-image consistency in a supervised manner with different approaches, such as graph convolution networks (GCN) , self-learning methods , inter-image co-attention with PCA projection or recurrent units , correlation techniques , quality measurement , or co-clustering . Some methods exploit multi-task learning to simultaneously optimize the co-saliency detection and co-segmentation or co-peak search . Other works explore hierachical features from multi-scale , multi-stage , or multi-layer features. Another notable research line is to explore group-wise semantic representation (consensus) which is used to detect co-salient regions for each image. There are different methods to capture the discriminative semantic representation, such as group attentional semantic aggregation , gradient feedback , co-category association , united fully convolutional network , or integrated multilayer graph . Methods are proposed to solve the CoSOD problem in a semi-supervised or unsupervised manner , and studies are availalbe on co-saliency detection from a single image.
Previous works have focused on intra-group (intra- and inter-image) cues for capturing common attributes of co-salient objects. The inter-group information has received less attention, although CODW focuses on visually similar neighbor. Recently Zhang et al. utilized a jigsaw training to implicitly exploit other images to facilitate group training. But their model still targets intra-group learning. Our method differs from existing models in the exploration of inter-group relations for discriminating feature learning at a group level explicitly and semantically.
Group Collaborative Learning Network
Given a group of relevant images containing common salient objects of a certain class, CoSOD aims to detect them simultaneously and output the co-saliency maps. Unlike existing CoSOD methods which only depend on the information within the image group, we propose a novel group collaborative learning network (GCoNet) to mine the consensus representations at both intra- and inter-group level.
2 Group Affinity Module
Intuitively, common objects from the same class always share some similarity in appearance and have high similarity in features, which have been widely employed in many tasks. Inspired by self-supervised video tracking methods , which propagate the segmentation masks of target objects based on the pixel-wise correspondences between two adjacent frames, we extend this idea to the CoSOD task by computing the global affinity among all images in a group.
The global affinity module focuses on capturing the commonality among co-salient objects within the same group and therefore improves the intra-group compactness of the consensus representation. Such intra-group compactness alleviates the disturbance of co-occurring noise and enables the model to concentrate on the co-salient regions. This allows the shared attributes of co-salient objects to be better captured and therefore results in better consensus representation. The obtained attention consensus is combined with the original feature maps through depth-wise correlation to achieve efficient information association. The resulting feature maps are fed to the decoder to predict co-saliency maps for each image. The loss function is:
where is the soft IoU loss and denotes the ground-truth label for each image in the group.
3 Group collaborating module (GCM)
Most CoSOD methods tend to focus on the intra-group compactness of the consensus, but the inter-group separability is equally crucial for distinguishing distracting objects, especially when processing complex images with more than one salient objects. To enhance the discriminative representations between different groups, we propose a simple but effective module, \ie, the GCM, by learning to encode the inter-group separability.
Given two image groups with the corresponding features and attention consensus obtained from the GAM, we apply an intra- and inter-group cross-multiplication. Specifically, the intra-group multiplication deals with the features and their consensus: and for the intra-group collaboration, while the inter-group multiplication acts on the features and consensus of different groups, \ie, and , to express the inter-group interaction. The intra-group representation is exploited to predict the co-saliency maps, and the inter-group representation is employed to provide a consensus with group separability. Specifically, we feed to a small convolutional network with an upsampling layer and produce the saliency map and . with different supervision signals: we use ground-truth labels to supervise , while all-zero maps are used for . The loss function is:
where is the focal loss , is the ground-truth, is the all-zero map and denotes the concatenation operation.
Our GCM thus encourages the consensus to distinguish different groups with high inter-group separability to identify distractors in complex environment. Another advantage is that this module enables the model to be trained on the existing SOD datasets, whose images typically contain only one dominating object. We can discard this module during inference without introducing additional computational overhead.
4 Auxiliary Classification Module (ACM)
To obtain more discriminative features for consensus, we also introduce an ACM to facilitate high-level semantic representation learning. Specifically, we add a classification predictor with a global average pooling layer and one fully connected layer to the backbone to classify to the corresponding class . In the Euclidean feature space, the classification supervision can separate classes by introducing a large margin, and cluster samples belonging to the same class. Therefore, it enables the model to generate more representative features and benefits the consensus learning for intra-group compactness and inter-group separability. The loss function is:
where is the cross-entropy loss and is the ground-truth class label.
5 End-to-end Training
During training, the GAM, GCM, and ACM are jointly trained with the backbone in an end-to-end manner. The whole framework is optimized by integrating all the aforementioned loss functions:
where , , and are hyperparameter weights to balance the loss functions.
Experiments
We use VGG-16 with Feature Pyramid Network (FPN) as our backbone. For fair comparison, we follow GICD and use the DUTS dataset as our training set. The group labels derived from GICD are used to group the images during training. In each training episode, we randomly pick two different groups with 16 samplesDue to limited computing resource. The larger the better. in each group to train the network. The images are all resized to 224x224 for training and testing, and the output saliency maps are resized to the original size for evaluation. The network is trained over epochs in total with the Adam optimizer. The initial learning rate is set to , and . The whole training takes around four hours and the inference speed on the image pair groupsCoSOD task works for image groups. Therefore we use the basic image pair group to evaluate the speed rather than the single image. is . The platform for training and inference is equipped with Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz and a Nvidia GeForce GTX 1080Ti.
2 Evaluation Datasets and Metrics
We employ three challenging datasets for evaluation: CoCA , CoSOD3k , and Cosal2015 . The last is a large dataset widely used in the evaluation of CoSOD methods. The first two were recently proposed for challenging real-world co-saliency evaluation, with the images usually containing multiple common and non-common objects against a complex background. Following the advice of recent large-scale benchmark work , we do not use iCoseg and MSRC for evaluation, because tbey usually provide only one salient object in an image and are not very suitable for evaluating CoSOD models. We use maximum E-measure , S-measure , maximum F-measure , and mean absolute error (MAE) to evaluate methods in our experiments. Evaluation toolbox: https://github.com/DengPingFan/CoSODToolbox.
3 Ablation Studies
In this section, we study the effectiveness of each component in our approach (Table 1) and investigate how they contribute to a good consensus feature.
Effectiveness of GAM. The global co-attention module is a fundamental component of our model, which is designed to capture the common attributes of co-salient objects in an image group for better intra-group compactness. Compared to the baseline model with only the vanilla consensus extracted by an average pooling operation, GAM improves the performance on all metrics and datasets. To get a deeper understanding of our GAM module, we visualize the learned attention masks in Figure 5. We find that our global co-attention effectively alleviates the influence of co-occurring noise and focuses on co-salient regions in the image groups, \eg, in both the monkey and bicycle groups, there are some co-occurring persons in some images, but our GAM is not adversely influenced. The global view of GAM enables the most common objects to be detected, while the local pair-wise co-attention cannot distinguish them in the local view.
Effectiveness of GCM. The group collaborating module is designed to enable the consensus inter-group separability to distinguish distracting objects from non-common objects. After equip the model with GCM, significant performance improvement (ID-1 versus ID-3) is obtained in Table 1 especially on the challenging CoCA dataset whose images usually contain multiple uncommon and common objects. To investigate the consensus characteristics when the model is trained with the GCM, we visualize the consensus using t-SNE on the CoCA dataset, and compare with the vanilla consensus without the GCM. As shown in Figure 1, the vanilla consensuses (top: other method) tend to cluster together, even if they belong to different groups, resulting in ambiguous co-saliency detection, especially for objects belonging to similar but different groups. In contrast, the consensuses trained with the GCM (bottom: our method) is more diverse with a higher group variance (), for more effective inter-group separability.
Effectiveness of ACM. As shown in Table 1, the classification module introduces better backbone features for the consensus with the auxiliary classification supervision. The ACM improves the baseline performance on all metrics and datasets. This cost-free improvement does not change the network architecture and does not introduce extra computational overhead, thus has substantial potential to other models and tasks to take advantage of the multi-task learning and more representative features.
4 Competing Methods
Since not all CoSOD models have publicly released codes, we only compare our GCoNet with one representative traditional algorithm (CBCD) and five deep-based CoSOD models, including GWD , RCAN , CSMG , GICD , and CoEGNet . Following the current state-of-the-art model , we also compare with four cutting-edge deep salient object detection (SOD)SOD methods can also be directly applied to the CoSOD task. models: BASNet , PoolNet , EGNet and SCRN . More complete leaderboard can be found in recent standard benchmark works .
Table 2 tabulates the quantitative results of our model and state-of-the-art methods. Our model outperforms all of them in all metrics, especially on the challenging CoCA and CoSOD3k datasets. Among these three datasets, CoCA is the most challenging, since the images typically contain other multiple objects in addition to the co-salient objects which are even smaller in size. Our model capitalizes on our better consensus and significantly outperforms other methods especially the SOD methods which are trapped in distinguishing many distracting objects instead. CoSOD3k has similar attributes, and our model still performs much better than other models on this dataset. Cosal2015 is the easiest dataset because its images typically only contain one co-salient object, and therefore the SOD algorithms can easily handle this dataset. Our model cannot take full advantage of the better consensus on this dataset and the improvement is not as significant as on other datasets.
Qualitative Results.
Figure 6 shows the saliency maps generated by different methods for qualitative comparison. In these difficult examples, each image contains other multiple objects in addition to the co-salient objects. As aforementioned, the SOD methods can only detect salient objects and fail to distinguish co-salient objects due to their intrinsic limitation. The CoSOD methods perform better than the SOD methods owing to their consensus for distinguishing co-salient regions. However, limited by the their weak consensus, they are still unable to handle the challenging cases. Our model introduces an effective consensus through optimizing intra-group compactness and inter-group separability, and therefore performs much better on detecting co-salient objects.
Discussion of Module Cooperation
Our three modules are closely interdependent and mutually reinforced for improving co-saliency detection performance. Combining the GAM and GCM can significantly improve the performance compared to the individual modules. Without the GAM the vanilla consensus is not robust against noise caused by uncommon objects and background, and the low-quality consensus cannot take full advantage of the GCM which heavily relies on the consensus for distinguishing different objects. On the other hand, although the consensus can capture common attributes with the help of the GAM, it is difficult to distinguish different groups without the GCM especially for similar groups. Overall, the GAM produces better consensus with high intra-group compactness to detect co-saliency objects, while the GCM further endows the consensus with inter-group separability for better discriminative ability. Adding ACM, the consensus can benefit from more representative features leveraged by the multi-task learning.
Figure 7 qualitatively analyse their cooperation. The baseline model detects uncommon objects, while the GAM and GCM can slightly ameliorate their adverse influence. When combining the GAM and GCM, the model can effectively capture co-salient objects with the ACM further boosting the co-salient object detection result.
Downstream Applications
Here, we show how the extracted co-saliency map can be utilized to generate high-quality segmentation masks for selected closely related downstream image processing tasks.
Application #1: Content-Aware Co-Segmentation. Co-saliency maps have been previously used in pre-processing for unsupervised object segmentation. In our implementation, we first manually select a group of images from the internet by keyword search . Then, co-saliency maps are generated by our GCoNet to automatically mine the salient content of the specific group. Similar to Cheng et al. , we also utilize GrabCut to obtain the final segmentation results. To initialize GrabCut, we simply choose adaptive threshold to binarize the saliency maps. Figure 8 shows the results of the content-aware object co-segmentation which should benefit existing e-commerce applications requiring background replacement.
Application #2: Automatic Thumbnails. The idea of paired-image thumbnails is derived from the seminal work . With the same goalNote that Jacobs et al.’s work is limited to the case of image pairs. , we present a CNN-based photographic triage application which is valuable for sharing images with friends on the website. As shown in Figure 9, we first generate the yellow box based on the co-saliency map obtained by our GCoNet. Then, we simply enlarge the yellow box to get a larger red box. Finally, we adopt the collection-aware crops technique to produce the results (2nd row).
Conclusion
In this paper, we investigate a novel group collaborative learning framework (GCoNet) for CoSOD. We find that group-level consensus can introduce effective semantic information to benefit the representation of both the intra-group compactness and inter-group separability for CoSOD. Our experiments quantitatively and qualitatively demonstrate the advantage of our GCoNet which outperforms existing state-of-the-art models. In addition, our GCoNet achieves real-time speed (16ms) which can greatly benefit many applications such as co-segmentation, co-localization, and among others.