Uncertainty-aware Joint Salient Object and Camouflaged Object Detection
Aixuan Li, Jing Zhang, Yunqiu Lv, Bowen Liu, Tong Zhang, Yuchao Dai
Introduction
Visual salient object detection (SOD) aims to localize the most salient region(s) of the images that attract human attention. To be qualified as a “salient” object, one should have high contrast compared with its global and local context. The camouflaged objects oppositely usually share similar structure or texture information with the environment, which try hard to fade themselves into the local context. In this way, the SOD models are designed based on both global contrast and local contrast, while the COD models usually avoid searching the camouflaged objects in those salient regions. We notice that a higher level of saliency indicates a lower level of camouflage and vice versa as shown in Fig. 1. This observation shows that the salient objects and camouflaged objects are two contradicting categories of objects. However, there still exist objects that are both salient and camouflaged, \egthe polar bear in the middle of Fig. 1, which indicates that these two tasks are partially positively related at the dataset level.
Existing SOD models mainly focus on two directions: 1) building effective saliency network for accurate saliency detection with pixel-wise accuracy constraint; and 2) designing appropriate loss functions to achieve structure-preserving saliency detection. The former digs into network structure, while the latter cares more about network loss function. We argue that an effective training dataset can lead to more performance gain in addition to network structure design or loss function, it’s the training data that is regressed.
One typical solution to explore the training dataset is data augmentation, which usually involves linear or non-linear transformation of the dataset. We find that, although performance improvement can be obtained with some basic data augmentation, \egimage flipping, rotation, cropping, and \etc, none of these methods are specially designed for saliency detection. As a context-based task, a more effective data augmentation technique should be context-aware. For SOD, the salient objects are those that can be easily detected, or the high-contrast objects as shown in Fig. 1. We intend to augment the dataset to include lower contrast samples. Considering the partial positively related attribute of SOD and COD at dataset level, we intend to design a joint learning framework to learn both tasks and select easy samples from COD (\egthe polar bear) as hard samples for SOD, achieving contrast-level data augmentation.
Joint training is mostly designed for positively related tasks . In contrary, we propose to integrate two contradicting tasks (SOD and COD) into one network with a “Similarity measure” module as shown in Fig. 2. The basic assumption of our similarity measure module is that the activated regions of the same image for the two tasks should be different, leading to latent features apart from each other. To this end, we introduce the third dataset, \egPASCAL VOC 2007 images in particular, to our framework serving as the connection modeling dataset. The goal of these extra images is to achieve similarity measures and force the two tasks to focus on different regions of the image.
Moreover, as shown in Fig. 1, the salient object is salient in both local and global contexts, while the camouflaged object is hiding in its local context. Due to the high contrast, the local context of the salient objects is easier to model than that of the camouflaged objects, as there exists no clear boundary between camouflaged objects and their surrounding. By jointly training a salient object detection network and camouflaged object detection network, the salient object branch can learn precise local context information for accurate camouflaged object detection.
Lastly, we observe two types of uncertainty while labeling the dataset for each task. For salient object detection, the subjective nature of saliency leads to ambiguity of prediction, as shown in Fig. 5(a). For camouflaged object detection, the uncertainty comes from the difficulty in fully annotating the camouflaged objects as they usually share similar color or texture with the environment, as shown in Fig. 5(c). We then introduce adversarial training to explicitly model the confidence of network predictions, and estimate model uncertainty for both tasks.
We summarize our main contributions as: 1) We introduce the first joint salient object detection and camouflaged object detection network within an adversarial learning framework to explicitly model prediction uncertainty of each task. 2) We design the similarity measure module to explicitly model the “contradicting” attributes of the two tasks. 3) We present a data interaction strategy and treat easy samples from camouflage dataset as hard samples for saliency detection, achieving a robust saliency model.
Related Work
Salient Object Detection Models Existing deep saliency detection models are mainly designed to achieve structure-preserving saliency predictions. introduced auxiliary edge detection branch to produce a saliency map with precise structure information. Wei et al. presented structure-aware loss function to penalize prediction along object edges. Wu et al. designed a cascade partial decoder to achieve accurate saliency detection with finer detailed information. Feng et al. proposed a boundary-aware mechanism to improve the accuracy of network prediction on the boundary. There also exist salient object detection models that benefit from data of other sources. integrated fixation prediction and salient object detection in a unified framework to explore the connections of the two related tasks. Zeng et al. presented to jointly learn a weakly supervised semantic segmentation and fully supervised salient object detection model to benefit from both tasks.
Camouflaged Object Detection Models Camouflage models are designed to discover the camouflaged object(s) hidden in the surrounding. Different from salient objects, which are those attracting human attention, camouflaged objects are those trying to decrease the conspicuousness. The concept of camouflage is usually associated with context . Cuthill et al. presented that an effective camouflage includes two mechanisms: the background pattern matching, where the color is similar to the environment, and the disruptive coloration, which usually involves bright colors along edge, and makes the boundary between camouflaged objects and the background unnoticeable. Bhajantri et al. utilized co-occurrence matrix to detect defective. Pike et al. combined several salient visual features to quantify camouflage, which could simulate the visual mechanism of a predator. In the field of deep learning, Fan et al. proposed the first publicly available camouflage deep network with the largest camouflaged object training set.
Multi-task Learning The basic assumption behind multi-task learning is that there exists shared information among different tasks. In this way, multi-task learning is widely used to extract complementary information about positively related tasks. Kalogeiton et al. jointly detected objects and actions in a video scene. Zhen et al. designed a joint semantic segmentation and boundary detection framework by iteratively fusing feature maps generated for each task with pyramid context module. In order to solve the problem of insufficient supervision in semantic alignment and object landmark detection, Jeon et al. designed a joint loss function to impose constraints between tasks, and only reliable matched pairs were used to improve the model robustness with weak supervision. Joung et al. solved the problem of object viewpoint changes in 3D object detection and viewpoint estimation with a cylindrical convolutional network, which obtains view-specific features with structural information at each viewpoint for both two tasks. Luo et al. presented a multitask framework for referring expression comprehension and segmentation.
Adversarial Learning Adversarial learning is an effective solution to improve the robustness of the deep neural network. In fully supervised models, adversarial learning can measure higher-order inconsistencies between labels and predictions . Specifically, Li et al. introduced generic noise to destroy adversarial perturbation of the input image for salient object detection. built the Generative Adversarial Network (GAN) based saliency detection network, where the fully connected discriminator was used to distinguish the real saliency map (ground truth) from the fake saliency map (prediction). Similarly, Jiang et al. designed a GAN based framework for RGB-D saliency detection to solve the cross-modality detection problem. In the case of insufficient annotations, adversarial learning can serve as guidance to select confident samples or generate new samples. Souly et al. used adversarial learning to generate fake images with image-level labels and noise, which in turn can make the real samples gathering in feature space and improve accuracy for weakly supervised semantic segmentation. treated adversarial learning as a confidence measure to obtain the trustworthy regions of semantic segmentation predictions of unlabeled data for semi-supervised learning.
Different from existing multi-task learning frameworks that mainly benefit from positively related tasks, we instead build the connections of two contradicting tasks within a joint learning framework by explicitly modeling the contradicting attributes with a similarity measure module.
Our Method
We design an uncertainty-aware joint learning framework as shown in Fig. 2 to learn SOD and COD in a unified framework. Firstly, as a data augmentation technique, we select a group of easy samples from the COD training dataset to achieve robust SOD. Then, we present the “Similarity measure” module to explicitly model the “contradicting” attributes of the two tasks. Lastly, we introduce our uncertainty-aware adversarial training network to produce interpretable predictions during testing, and achieve higher-order similarity measure during training.
As shown in Fig. 1, there exist samples in the COD dataset that are both salient and camouflaged. We argue that those samples can be treated as hard samples for salient object detection to achieve robust learning. To select those samples from the COD dataset, we resort to Mean Absolute Error (MAE), and select samples in COD dataset which achieve the smallest MAE by testing it using a trained SOD model . Specifically, for camouflaged object detection training dataset , where indexes the images, and is the size of camouflaged object detection training set. We defined the trained SOD model as . Then we obtain saliency prediction of the images in as , where is the saliency prediction in COD training dataset. We assume that easy samples for COD can be treated as hard samples for SOD as shown in Fig. 1. Then we select samples of the smallest MAE in , and replace it with randomly selected samples in our SOD training dataset as a data augmentation technique. We show the selected samples in Fig. 3, which clearly illustrates the partially positive connection of the two tasks at the dataset level.
2 Contradicting modeling
Similar as above, let’s define our camouflaged object detection training dataset as and augmented salient object detection training dataset as , where indexes images, and are the image and ground truth pair, and are size of the camouflage training set and saliency training set respectively. Based on the “contradicting” attribute of SOD and COD, we design a “Similarity measure” module in Fig. 2 to explicitly model the connection of the two tasks.
Specifically, we introduce another set of images from PASCAL VOC 2007 dataset as “connection modeling” dataset , from which we extract the camouflaged feature and the salient feature. With the three datasets (COD dataset , augmented SOD dataset and connection modeling dataset ), our contradicting modeling framework uses the “Feature encoder” module to extract both camouflage feature and saliency feature, and then use the “Similarity measure” module to model connection of the two tasks with the connection modeling dataset.
Feature Encoder We design both the saliency encoder and camouflage encoder network with the same backbone network, \egthe ResNet50 , where and are network parameter sets of each of them respectively. Initially, the ResNet50 backbone network has four groupsWe define feature maps of the same spatial size as same group. of convolutional layers of channel size 256, 512, 1024 and 2048 respectively. We then define the output features of both encoders as and , where is feature map of the -th group.
Similarity Measure Different from the feature encoder module, which takes images from and as input to produce task-specific feature maps, the similarity measure module takes the connection modeling data as input to model the connection of SOD and COD, where is parameter set of the similarity measure module. Given the trained saliency encoder and camouflage encoder , the saliency feature and camouflage feature of images are and respectively. We then concatenate each of above two features channel-wise and feed them to the same fully connected layer to obtain the latent saliency feature and latent camouflage feature of as and respectively. Empirically, we set the dimension of the latent space as . For the same image in , we assume that the SOD network and COD network should focus on different regions, leading to different feature representation. Then, we choose the cosine similarity to measure the difference between the saliency feature and the camouflage feature in latent space, and define the latent space loss as:
In Fig. 4, we show the activation region (the processed predictions) of the same image from both the saliency encoder (first row) and camouflage encoder (second row). Specifically, given same image , we compute it’s camouflage map and saliency map, and highlight the detected foreground region in red. Fig. 4 shows that the two encoders focus on different regions of the image, where the saliency encoder pays more attention to the region that stand out from the context, and the camouflage encoder focuses more on the hidden object with similar color or structure as the background, which is consistent with our assumption that these two tasks are contradicting with each other in general.
3 Uncertainty-aware adversarial learning
As shown in Fig. 5, for the SOD dataset, the uncertainty comes from the ambiguity of saliency ((a) (b)), and for the COD dataset, the uncertainty results from the difficulty of annotation ((c) (d)). \egthe ball in the orange rectangle (a) can be defined as salient, but it’s background in (b). The orange region in (c) belongs to the camouflaged object, while it’s too similar to the background, making it very difficult to create the accurate annotation. We then introduce an uncertainty-aware adversarial training strategy to model the task-specific uncertainty in our joint learning framework, which includes a “Prediction decoder” module to produce task-related predictions, a “Confidence estimation” module to estimate uncertainty of each prediction, and an adversarial learning strategy for robust model training.
Prediction decoder As shown in Fig. 2, we design a shared decoder structure for joint SOD and COD learning. We argue that the different “Feature encoder” modules can generate task-specific features for COD images and SOD images. Then the “Prediction decoder” module aims to integrate the task-specific feature with their corresponding lower level feature to produce predictions. Specifically, given task-specific feature and from the saliency encoder and camouflage encoder respectively, the prediction decoder produces saliency map and camouflage map , where is parameter set of the prediction decoder module. Specifically, we design a top-down connection network with the residual channel attention module to extract finer features. Furthermore, the dual attention module is adopted to effectively fuse higher level semantic information with the lower level structure information to obtain the initial predictions:
where , is the convolutional layer of output channel size , is the channel-wise concatenation operation, and we upsample the features to the same spatial size before concatenation. is the classification layer of kernel size , which maps the feature map to one channel prediction for each task. Then we add a refined structure to the decoder network in order to obtain a detailed prediction and , we use the holistic attention module to integrate features:
where and is the ResNet50 backbone convolutional layers of channel size 1024 and 2048 respectively. Then, we obtain the task-specific predictions and :
where the feature , and the feature .
Confidence Estimation As discussed above, uncertainty exists in both SOD and COD dataset. We introduce the “Confidence estimation” module to explicitly model the confidence of network predictions. Specifically, we design a fully convolutional discriminator network to evaluate confidence of the predictions from the “Prediction decoder” module. The fully convolutional discriminator network consists of five convolution layers as shown in Table 1, and produce a one-channel confidence map, where is the network parameter set. Note that, we have the batch normalization and leaky relu layers after the first four convolutional layers. aims to produce all-zero output with prediction or as input, and all-one matrix with ground truth as input.
Adversarial Learning We introduce adversarial learning to learn both the “Prediction decoder” and the “Confidence estimation” modules.
For the prediction decoder module, we first have the task-specific loss function to learn each task. Specifically, we adopt the structure-aware loss function for both SOD and COD, and define the loss function as:
where is the edge-aware weight, which is defined as , is the cross-entropy loss, is the boundary-IOU loss , which is defined as:
where , and .
And we have the structure-aware loss function for SOD and COD as:
To achieve adversarial learning, we have adversarial loss for both SOD and COD, which is defined as cross-entropy loss between network prediction and the pre-defined “real” indicator as:
respectvely for each task, where is an all-one matrix. In this way, the discriminator takes model prediction as input, and tries to recognize it as real ground truth.
For the confidence estimation module, similar to the typical definition of discriminator in GAN , we want it to clearly distinguish model prediction and ground truth map. Then, the adversarial loss for the confidence estimation module of the SOD task is defined as:
We then have the adversarial loss for the COD task as:
As shown in Fig. 2, the two tasks have separate encoder, shared decoder as well as the shared confidence estimation module. Given camouflaged object detection training images in , and salient object detection training images in , as well as the connection modeling images in , we first compute the similarity measure of as Eq. (1) and update the feature encoder ( and ) and similarity measure module (). Then, we train the adversarial learning for the saliency generator branch (saliency encoder and prediction decoder) with loss function:
where is a trade-off parameter, and empirically we set as . Similarly, we train the adversarial learning for the camouflage generator branch (camouflage encoder and prediction decoder) with loss function:
where we set in our experiments.
Then, we train the confidence estimation module of parameter set with the loss function:
Algorithm 1 is our complete training algorithm.
Experimental Results
Dataset: For salient object detection, we train our model using the augmented DUTS training dataset , and testing on six other testing dataset, including the DUTS testing dataset, ECSSD , DUT , HKU-IS , THUR and SOC testing dataset . For camouflaged object detection, we train our model using COD10K training set , and test on three camouflaged object detection testing sets, including CAMO , CHAMELEON , and COD10K.
Evaluation Metrics: We use four evaluation metrics to evaluate the performance of both SOD and COD models, including Mean Absolute Error, Mean F-measure, Mean E-measure and S-measure .
Training details: We train our model in Pytorch with ResNet50 as backbone as shown in Fig. 2. Both the encoders for saliency and camouflage branches are initialized with ResNet50 trained on ImageNet, and other newly added layers are randomly initialized. We resize all the images and ground truth to . The maximum iteration is 36000, and we iteratively update 3 times the saliency branch and then one time the camouflage branch. The initial learning rate is 2.5e-5. We adopt the “step” learning rate decay policy, and set the decay step as 24000 iteration, and decay rate as 0.1. The whole training takes 8 hours with batch size 15 on an NVIDIA GeForce RTX 2080 GPU.
2 SOD performance comparison
We compare performance of our SOD branch with eleven SOTA SOD models as shown in Table 2. One observation from Table 2 is that the structure-preserving strategy is widely used in the state-of-the-art saliency detection models, \egSCRN , F3Net , ITSD , and it can indeed improve model performance. Table 2 shows that we achieve 5/6 best performance, except on SOC testing dataset . The main reason is that there exists texture images in SOC, which may be treated as camouflaged object, thus influence our performance. We will investigate this issue further. Further, we show predictions of ours and SOTA models in Fig. 6, where the “Uncertainty” is obtained based on the prediction from the discriminator. Specifically, we define magnitude of gradient of the discriminator output as uncertainty following . Fig. 6 shows that we produce both accurate prediction and reasonable uncertainty estimation, where the brighter area of the uncertainty map indicates the less confident region.
3 COD performance comparison
As there exists only one open source deep camouflaged object detection network (SINet in particular), we re-train existing saliency detection models with the camouflaged object detection training dataset , and test on the existing camouflaged object detection testing set. The performance of these models is shown in Table 3 and Fig. 7. The consistent best performance of our camouflage model further illustrates effectiveness of the joint learning framework. Moreover, the produced uncertainty map clearly represents model confidence towards the current prediction, leading to interpretable prediction for the downstream tasks.
4 Ablation study
We present extra experiments to fully explore our model, and show performance in Table 4 and Table 5.
Train each task separately: We use the same “Feature encoder” and “Prediction decoder” in Fig. 2 to train the SOD network and the COD network separately, and show their performance as “ASOD” and “SCOD” respectively. We also train the saliency model using the original DUTS training dataset , and show its performance as “SSOD”. We notice a consistent performance improvement of “ASOD” compared with “SSOD”, which indicates the effectiveness of our contrast-level data augmentation technique. The comparable performance of “ASOD” and “SCOD” with their corresponding SOTA models further prove superior performance of our new network structure.
Joint training of SOD and COD: We train the “Feature encoder” and “Prediction decoder” within a joint learning pipeline. The performance is shown as “JSOD1” and “JCOD1” for the saliency detection task and camouflaged object detection task respectively. The improved performance of “JSOD1” and “JCOD1” compared with “ASOD” and “SCOD” indicates that the joint learning framework can further boost performance of each task.
Joint training of SOD and COD with similarity measure: We add the task connection constraint to the joint learning framework, \egthe similarity measure module in particular, and show performance as “JSOD2” and “JCOD2” respectively. In general, we can observe improved performance, especially for the COD10K dataset , which verifies effectiveness of our similarity measure module.
Uncertianty-aware joint training of SOD and COD: Based on the joint learning framework, we include the adversarial learning pipeline to our network, and show performance as “JSOD3” and “JCOD3”. We observe relative comparable performance with the adversarial learning framework. This mainly lies in the difficulty in training the adversarial learning branch. We provide the fully convolutional discriminator in Table 1, and set loss for the adversarial learning empirically. A better solution could be searching for a more effective discriminator and weight, which will be our future direction.
5 Hyper-parameters analysis
In our joint learning framework, we have several hyper-parameters that influence our performance, including the maximum iteration, the interval to iteratively train the saliency and camouflage branch, the base learning rate, the dimension of the latent space, the weight of both adversarial loss and latent loss. We found that due to the different dataset sizes and convergence rates, the maximum iteration has a great impact on the COD task. The SOD training dataset contains 10,553 images, which is 2.5 times of the COD dataset (the training dataset size of COD is 4,040). To avoid overfitting on COD, we iteratively update three times the saliency branch and then one time the camouflage branch. For the similarity measure module, we find that the PASCAL VOC 2007 dataset contains some samples that are both salient and camouflaged, which is contradicting with the goal of similarity measure. Thus, we train the similarity measure models every 400 iterations, and set the weight of latent loss as 0.1. For the “Confidence estimation” module, we observe that a large weight of the adversarial loss in Eq. (15) and Eq. (16), \eg, may destroy the prediction, especially for the SOD branch. Our main goal of using the adversarial learning is to provide reasonable uncertainty estimation. In this case, we set the adversarial loss weight as a relative small number, \eg0.01 in this paper, to achieve trade-off between model performance and effective uncertainty estimation.
Conclusion
In this paper, we have proposed the first joint salient object detection and camouflaged object detection network within an uncertainty-aware framework. First, we showed that the easy samples in COD dataset could be used as hard samples for SOD to learn robust SOD model. Second, by considering the contradicting attributes of these two tasks, we presented a similarity measure module to explicitly build the task connection with the extra connection modeling dataset. Lastly, we presented an adversarial learning network to explicitly model the confidence of network predictions to address the uncertainty in SOD and COD annotations. Experimental results on six benchmark SOD datasets and three benchmark COD datasets demonstrate the effectiveness of our joint learning solution.
Acknowledgements
This research was supported in part by National Natural Science Foundation of China (61871325), National Key Research and Development Program of China (2018AAA0102803), CSIRO’s Machine Learning and Artificial Intelligence Future Science Platform (MLAI FSP), and the Swiss National Science Foundation via the Sinergia grant CRSII5-180359. We would like to thank the anonymous reviewers for their useful feedbacks.