Synthetic Data Supervised Salient Object Detection

Zhenyu Wu, Lin Wang, Wei Wang, Tengfei Shi, Chenglizhao Chen, Aimin Hao, Shuo Li

Introduction

Salient object detection (SOD) aims to segment interesting objects that attract human attention in an image. As a fundamental tool, it can be leveraged to various applications including scene understanding (Zhou et al., 2021), semantic segmentation (Zhou et al., 2020) and image editing (Jiang et al., 2021; Cheng et al., 2010). Recently, SOD has achieved significant progress (Fang et al., 2021; Liu et al., 2021b; Wang et al., 2021; Tang et al., 2021; Zhao et al., 2021; Zhang et al., 2021b; Wu et al., 2022) due to the development of deep model. However, deep networks are extremely data-hungry, typically requiring pixel-level humanannotated datasets to achieve high performance (see Fig. 1.a). Labeling large-scale datasets with pixel-level annotations for SOD is very time-consuming, e.g., generally more than five people were asked to annotate the same image to guarantee the label consistency and another ten viewers were asked to cross-check the quality of annotations in the SOC dataset (Fan et al., 2018).

To alleviate the dependency on pixel-wise annotation, many weakly-supervised SOD methods (Zeng et al., 2019; Li et al., 2018; Wang et al., 2017) have been devised. Typically, image-level labels (see Fig. 1.c) are utilized in (Li et al., 2018; Wang et al., 2017) for saliency localization, and then iteratively finetune their models with predicted saliency maps. Additionally, scribble annotations (see Fig. 1.b) has been proposed recently in (Zhang et al., 2020c) to reduce the uncertainty of image-level labels. Although these methods are free of pixel-level annotations, they suffer from various disadvantages, including low prediction accuracy, complex training strategy, dedicated network architecture, and extra data information (e.g., edge) to obtain high-quality saliency maps.

In this paper, we propose a new paradigm SODGAN (see Fig. 1.d) for SOD, which can generate infinite high-quality image-mask pairs with a few labeled data to replace the human-labeled DUTS-TR (Wang et al., 2017) dataset. Concretely, our SODGAN has three stages: Stage 1. Learning a few-shot saliency mask generator to synthesize image-synchronous mask, while utilizing the existing generative adversarial networks (BigGAN (Brock et al., 2018)) to generate realistic images. Stage 2. Selecting high-quality image-mask pairs from the synthetic data pool. Stage 3. Training a saliency network on these filtered image-mask pairs. However, there are three main challenges with this approach: 1) Lacking pixel-wise labeled data as the training dataset to learn a segmentor because BigGAN was trained on the ImageNet that was designed to classification tasks without the pixel-level label. 2) Discovering a meaningful direction in GAN latent space to disentangle foreground saliency objects from backgrounds is nontrivial, which often requires domain knowledge and laborious engineering. 3) Low-quality image-mask pairs exist in the synthesized datasets.

To tackle these three challenges, first, we present a diffusion embedding network (DEN) (see Sec. 3.2) to utilize the existing well-annotated dataset (i.e., DUTS-TR), which can infer the image’s latent code that match with the ImageNet latent code space; thus, the existing labeled DUTS-TR dataset can provide the pixel-wise label for ImageNet. Second, in contrast to the existing works (Shen et al., 2020; Goetschalckx et al., 2019; Plumerault et al., 2019) focusing on latent space, we propose a few-shot saliency mask generator to automatically discover meaningful directions in the GANs feature space (see Sec. 3.3), which can synthesize infinite high-quality image synchronized saliency masks with a few labeled data. Third, we propose a quality-aware discriminator (see Sec. 3.4) to select high-quality synthesized image-mask pairs from the noisy synthetic data pool, improving the quality of synthetic data.

Our SODGAN has several desirable properties. a) Fewer labels. Our approach eliminates large-scale pixel-level supervision requiring only a few labeled data, which reduces the annotation costs. b) High performance. We demonstrate that the saliency model trained on synthetic data directly generated from GANs achieves an average 98.4%98.4\% F-measure of the saliency model trained on the DUTS-TR dataset. Moreover, our SODGAN achieves new SOTA performance in semi/weakly-supervised methods, and even outperforms some fully supervised methods. c) Generality. The synthetic data can be used to train any off-the-shelf SOD model without the need of special architectures, showing strong generalization capabilities on the real test datasets. We summarize the key contributions as follows:

For the first time, our SODGAN tackles SOD with synthetic data directly generated from the generative model, which opens up a new research paradigm for semi-supervised SOD and significantly reduces the annotation costs.

Our proposed the DEN can address manifold mismatch and is tractable for the latent code generation, better matching with the ImageNet latent space.

Our lightweight few-shot saliency mask generator can synthesize infinite accurate image-synchronous saliency masks with a few labeled data.

Our proposed quality-aware discriminator can select highquality synthesized image-mask pairs from the noisy synthetic data pool, improving the quality of synthetic data.

Related Work

Semi/Weakly-supervised SOD Approaches. With recent advances in semi/weakly-supervised learning, a few existing works exploit the potential of training saliency detectors on image-level (Zeng et al., 2019; Li et al., 2018; Wang et al., 2017), region-level (Yu et al., 2021; Zhang et al., 2020c, b), and limited pixel-level (Zhang et al., 2020a; Wu et al., 2020; Yan et al., 2019; Zhou et al., 2018) labeled data to relax the dependency of manually annotated pixel-level saliency masks. For image-level supervision, these approaches (Zeng et al., 2019; Li et al., 2018; Wang et al., 2017) follow the same technical route, i.e., producing initial saliency maps with image-level labels and then further refining it via iterative training. Recently, scribble annotation was proposed in (Zhang et al., 2020c; Yu et al., 2021), but it requires large-scale scribble annotations (10553 images) and extra data information (e.g., edge) to recover integral object structure. Differences. Distinct from all these works, our approach provides a new paradigm for semi-supervised SOD. In particular, we introduce SODGAN, a generative model, which can generate infinite high-quality image-mask pairs requiring minimal manual intervention. These generated pairs can then be used for training any existing SOD approaches.

Latent Interpretability of GANs. The previous works have shown that the GANs latent spaces are endowed with human-interpretable semantic arithmetic. A line of recent works (Shen et al., 2020; Goetschalckx et al., 2019; Plumerault et al., 2019; Shen and Zhou, 2021; Cherepkov et al., 2021; Yang et al., 2021) employ explicit human-provided supervision to identify interpretable directions in the latent space. For instance, (Goetschalckx et al., 2019; Shen et al., 2020) use the classifiers pretrained on the CelebA (Liu et al., 2015) dataset to produce pseudo labels for the generated images and their latent codes. Another active line of study on GANs (Abdal et al., 2021; Chen et al., 2019; Bielski and Favaro, 2019; Melas-Kyriazi et al., 2021; Voynov et al., 2021; Zhang et al., 2021a; Tritrong et al., 2021) targets the object segmentation task. (Abdal et al., 2021) and (Chen et al., 2019) are based on the idea of decomposing the generative process in a layer-wise fashion. Other works (Bielski and Favaro, 2019; Melas-Kyriazi et al., 2021; Voynov et al., 2021) exploit the idea that the object’s location or appearance can be perturbed without affecting image realism. Differences. In contrast to existing works manipulating the latent space, our approach is able to discover interpretable directions in the GANs features space, which allows complete control over the diversity of object categories and can automatically find the expected directions.

Method

As shown in Fig. 2, our SODGAN is composed of the diffusion embedding network DEN(⋅)DEN(\cdot), the mask synthesis network Gmask(⋅)G_{mask}(\cdot), the quality-aware discriminator Dq(⋅)D_{q}(\cdot), the image synthesis network Gimage(⋅)G_{image}(\cdot), and the image reconstruction discriminator Dr(⋅)D_{r}(\cdot). (1) The proposed DEN(⋅)DEN(\cdot) aims to address the lacking of pixel-wise labels in ImageNet, which is designed for recognition tasks without segmentation groundtruth. Our DEN(⋅)DEN(\cdot) can utilize existing labeled datasets DUTS-TR (Wang et al., 2017), and gradually turn it into a unique latent space Z+Z+ that matches with the ImageNet latent code space. (2) The proposed Gmask(⋅)G_{mask}(\cdot) is to discover meaningful directions in the GANs feature space, synthesizing image synchronized saliency mask. Our Gmask(⋅)G_{mask}(\cdot) is build on top of the Gimage(⋅)G_{image}(\cdot) architecture augmented with a few-shot saliency mask generation branch, (3) Our Dq(⋅)D_{q}(\cdot) is designed to select high-quality synthesized image-mask pairs from noisy synthetic data pool. The Gimage(⋅)G_{image}(\cdot) can be any off-the-shelf GANs models, and the Dr(⋅)D_{r}(\cdot) is the corresponding real/fake discriminator. Here, we demonstrate our approach using BigGAN (Brock et al., 2018), a class-conditional GANs trained on ImageNet (Deng et al., 2009). In our SODGAN, the proposed DEN(⋅)DEN(\cdot), Gmask(⋅)G_{mask}(\cdot) and Dq(⋅)D_{q}(\cdot) are trainable while the other components remain fixed.

2. Diffusion Embedding Network

Our DEN(⋅)DEN(\cdot) is to address the lacking of pixel-wise label in ImageNet, which is designed for recognition tasks without segmentation groundtruth, better matching with ImageNet latent code space. Previous work (Zhang et al., 2021a) addresses this issue by manually labeling a handful of sampled images, which is labor-consuming. An alternative idea is to utilize the existing labeled datasets (e.g., DUTS-TR) by using variational autoencoder (VAEs). However, the standard VAEs, with a Euclidean latent space, is structurally incapable of capturing topological properties of certain datasets, which is called manifold mismatch (Falorsi et al., 2018).

To address these challenges, we developed the diffusion embedding network DEN(⋅)DEN(\cdot) to utilize the existing labeled datasets with pixel-wise annotation (e.g., DUTS-TR), which allows for an arbitrarily closed manifold as a latent space and captures the underlying geometrical structure. The proposed DEN(⋅)DEN(\cdot) can gradually turn an image into a unique latent code z+z^{+} that better matches with ImageNet latent code space. Concretely, our DEN(⋅)DEN(\cdot) are latent variable models of the forms pθ(x0)=∫pθ(x0:T)dx1:Tp_{\theta}(x_{0})=\int p_{\theta}(x_{0:T})dx_{1:T}, where x1,...,xTx_{1},...,x_{T} are intermediate latent codes and x0∼q(x0)x_{0}\sim q(x_{0}) is the initial image. The joint distribution pθ(x0:T)p_{\theta}(x_{0:T}) is the embedding process, and it is defined as the Markov chain with learned Gaussian transitions p(xT)=N(xT;0,1)p(x_{T})=\mathcal{N}(x_{T};0,1):

The difference between our DEN(⋅)DEN(\cdot) and VAEs is that the approximate posterior q(x1:T∣x0)q(x_{1:T}|x_{0}), which is called the diffusion process, is fixed to a Markov chain that progressively adds Gaussian noise to the image in line with variance schedule β1,...,βT\beta_{1},...,\beta_{T}:

A desirable property of the diffusion process is that it admits sampling xtx_{t} at a arbitrary timestep tt in closed form:

where α=1−βt\alpha=1-\beta_{t} and α^=∏s=1tαs\hat{\alpha}=\prod_{s=1}^{t}\alpha_{s}. The reconstruction loss is to optimize the variational bound on negative log likelihood:

where DKL(⋅)D_{KL}(\cdot) is the KL divergence. The adversarial loss can be defined as:

where Ddut\mathcal{D}_{dut} is the DUTS-TR dataset. Note that our DEN(⋅)DEN(\cdot) doesn’t need special architectures. Here we adopt the MobileNetV3 (Howard et al., 2019) architecture for the diffusion model. In this way, given an image, our DEN(⋅)DEN(\cdot) can infer its latent code z+z^{+} that matches with the ImageNet latent code space, and find its groundtruth yy in the DUTS-TR.

Summarized Advantages: 1) Gaussian noise has the effect of filling low density regions in the original data distribution; thus, our DEN(⋅)DEN(\cdot) can obtain more training signal to improve latent distributions that faster converge to the true data distribution. 2) Our DEN(⋅)DEN(\cdot) is capable of capturing topological properties of certain datasets that better match with ImageNet latent code space.

3. Few-shot Saliency Mask Generator

Omni-Attentive Feature Fusion. Previous work (Yang et al., 2021) has demonstrated that in GANs feature space, low-level features contain local information like texture and color while high-level features capture global information, such as the style and layout of objects. To fully take advantage of the multi-level features, we proposed a novel omni-attentive feature fusion module, as depicted in Fig. 2 (bottomleft). To ensure the spatial alignment, we first upsample all feature maps {f0,f1,...,fl}\{f_{0},f_{1},...,f_{l}\} to the highest output resolution 256×256256\times 256, and then concatenate them along the channel dimension to obtain an aggregated feature f∗f^{*}:

where U(⋅)U(\cdot) denotes upsample operation, Conv⁡1×1(⋅)\operatorname{Conv}_{1\times 1}(\cdot) is 1×11\times 1 convolutional operation for reducing channel dimension, and Cat⁡(⋅)\operatorname{Cat}(\cdot) stands for concatenation. To better fusion the global and local contexts, we introduce the omni-attention OA(⋅){OA}(\cdot) module, including local attention LA(⋅)LA(\cdot) and global attention GA(⋅)GA(\cdot):

where the GAP(⋅)GAP(\cdot) is global average pooling, and PWC(⋅)PWC(\cdot) is the 1×11\times 1 point-wise convolution for reducing the parameters. Fig. 3 shows the visualized omni-attention maps. The aggregated feature f′f^{\prime} can be obtained by multipling with the OA(f∗)OA(f^{*}):

where the ⊗\otimes is the element-wise multiplication operator.

To train Gmask(⋅)G_{mask}(\cdot), we need to collect a small training set Dm={(z1+,y1),...,((zk+,yk))}\mathcal{D}_{m}=\{(z^{+}_{1},y_{1}),...,((z^{+}_{k},y_{k}))\}, where yiy_{i} is selected from the DUTS-TR. Specifically, we use the state-of-the-art (SOTA) image classification model CoAtNet (Dai et al., 2021) to classify the DUTS-TR, which can be divided into 522 categories. We then randomly select a pair of (z+,y)(z^{+},y) for each class, forming a small training set with 522 images. We then train the proposed Gmask(⋅)G_{mask}(\cdot) by using activation features z+z^{+} and the corresponding pixel-wise annotations. The training objective is

LS\mathcal{L}_{S} is the supervised loss on labeled images with a combination of cross entropy and dice loss, defined as:

where the HH and WW are the height and width of the image respectively, and y^ij\hat{y}_{ij} is the prediction probability at position (i,j)(i,j). The quality-aware discriminator loss LDq\mathcal{L}_{D_{q}} is:

Summarized Advantages: 1) Lightweight. Our Gmask(⋅)G_{mask}(\cdot) is extremely lightweight yet powerful, which consists of a OAFF and classification head with total 90K parameters and 3.6MB model size. 2) Fewer labels. We only need 522 images to train the Gmask(⋅)G_{mask}(\cdot) because our Gmask(⋅)G_{mask}(\cdot) is lightweight with only 90K parameters.

4. Quality-aware Discriminator

Our Dq(⋅)D_{q}(\cdot) can select high-quality synthesized image-mask pairs from noisy synthetic data pool, providing high-quality synthetic data to train the saliency network. We noticed that the synthetic data fails occasionally for non-rigid objects (e.g., dogs) due to their various poses, resulting in low-quality image-mask pairs. To alleviate this issue, we proposed a quality discriminator Dq(⋅)D_{q}(\cdot) adopting the lightweight MobileNetV3 (Howard et al., 2019) as backbone, which aims to select high-quality synthesized image-mask pair. During training, we feed two pairs to the quality discriminator Dq(⋅)D_{q}(\cdot), i.e., (xreal,yreal)(x_{real},y_{real}) and (xsyn,ysyn)(x_{syn},y_{syn}). Accordingly, the adversarial training loss for the Dq(⋅)D_{q}(\cdot) can be formulated as:

Note that our Dq(⋅)D_{q}(\cdot) is different from the typical discriminator, where the discriminator is designed for discriminating real or fake images, while our Dq(⋅)D_{q}(\cdot) performs image-mask quality control.

Summarized Advantages: 1) High-quality image-mask pairs. Our SODGAN can generate any desired number of high-quality image-mask pairs, which forms our synthetic dataset. The generated image-mask pairs can then be used to train any off-the-shelf SOD architecture just like real datasets are. 2) Strong generalization capabilities. Unlike previous works (Richter et al., 2016; Ros et al., 2016; Wang et al., 2019; Kar et al., 2019), which usually arises significant domain gap between the synthetic (from computer games) and real-world domains, the presented SODGAN can generate realistic images (see Fig. 4) and show strong generalization capabilities on the real test datasets (see Table 3).

Results and Analysis

In this section, we provide the detailed implementation regarding two aspects: convolutional neural networks and MLP.

CNN Architecture. We first use a linear embedding layer to reduce the input dimension from CC to 128, followed by 3 convolutional layers with kernel size of 3. The corresponding dimensions of the output channels are 128, 32, and 2 (the number of classes). All the layers are followed by a leaky ReLU activation function except for the last output layer. We call this standard version CNN-S. We also introduce CNN-M and CNN-L, where M/L denotes medium/large model size, and the architecture hyper-parameters of these model variants can be seen in the first 2 rows of Table 1.

MLP Architecture. We build our base model, called MLP-S, which consists of 3 fully-connected layers with 128, 32, and 2 hidden nodes, respectively. All layers except the output layer are followed by the BatchNorm layer and ReLU activation function. Similar to CNN-S, we also introduce its variants version MLP-M and MLP-L, and their hyper-parameters can be seen in the last 2 rows of Table 1.

2. Synthetic Data VS. Real DUTS-TR

As shown in Fig. 5, we provide analyses of our synthesized datasets compared to the real DUTS-TR datasets in terms of center bias, category distribution, color contrast, and salient object size.

Center bias. We visualize the salient object locations for the synthetic data and the DUTS-TR datasets in Fig. 5.a. Most objects are biased towards the image center for both datasets. Compared to the DUTS-TR, the synthetic data show lower center distributions. Category distribution. We use the SOTA classification model CoAtNet (Dai et al., 2021) to classify the filtered synthetic data and the DUTS-TR, which can be divided into 764 and 522 categories, respectively. As shown in Fig. 5.c, our synthetic data contains more object categories than the DUTS-TR. Color contrast & Object size. Since the DUTS-TR was designed for SOD tasks, the DUTS-TR’s images containing at least one salient object are higher color contrast than randomly generated synthetic data (see Fig. 5.d). Besides, we also statistics the object size of the DUTS-TR and our synthetic data in Fig. 5.e. As we can see, the synthetic data also contains smaller objects than the DUTS-TR. Additionally, BigGAN introduced the “truncation coefficient” λ\lambda, allowing explicit, fine-grained control of the trade-off between sample variety and complexity (see Fig. 5.b).

3. Ablation Study of Our Innovations

Eeffects of the proposed DEN(⋅)DEN(\cdot). To demonstrate the effects of our DEN(⋅)DEN(\cdot), we compared the proposed DEN(⋅)DEN(\cdot) with commonly used VAEs. As shown in Table 2, the proposed DEN(⋅)DEN(\cdot) improved by 1.8%1.8\% compared to the VAEs in terms of S-measure, which shows the effectiveness of the proposed diffusion model.

Eeffects of the proposed OAFF. In Table 2, we evaluate 3 settings of OAFF: 1) Gmask(⋅)G_{mask}(\cdot) without using the OAFF; 2) Gmask(⋅)G_{mask}(\cdot) only using the global attention GA(⋅)GA(\cdot); 3) the Gmask(⋅)G_{mask}(\cdot) with the OAFF. As we can see, the Gmask(⋅)G_{mask}(\cdot) with the GA(⋅)GA(\cdot) achieves better performance than the plain version, and the performance can be further improved by using OAFF, demonstrating the contribution of the OAFF to the final results.

The choice of classification head architecture. We evaluate 2 architectures on the proposed classification head network, i.e., CNN and MLP, with small (S), medium (M), and large (L) networks described in Sec. 4.1. As shown in Table 2, we notice that the MLP-S outperforms all the three CNN networks. Besides, we also notice that smaller networks obtain better performance due to the limited training data. Therefore, we take the MLP-S with channel dimension {128, 32, 2} as our classification head.

Eeffects of the proposed Dq(⋅)D_{q}(\cdot). To illustrate the effectiveness of the proposed Dq(⋅)D_{q}(\cdot), we implement 2 different settings, i.e., our SODGAN with/without using the Dq(⋅)D_{q}(\cdot). As shown in Table 2, the performance can be improved by 1.5%1.5\% in terms of F-measure on the DUTS-TE dataset by using the Dq(⋅)D_{q}(\cdot), verifying the contribution of our Dq(⋅)D_{q}(\cdot) to the final results.

Impacts of the amount of synthesized data. We further explore the number of synthesized data how to influence the saliency performance. As shown in left of Fig. 6, when the number of synthesized images is insufficient (¡ 12k), model performance can benefit substantially from the increased synthesized data. However, when the training set is large enough (¿ 12k), the application of more synthesized data does not necessarily lead to better performance. In this paper, unless otherwise specified, the reported SOD results were obtained by training on 12k synthetic image-mask pairs. Besides, to study the effects of λ\lambda, we vary the truncation coefficient λ∈{0.2,0.4,0.6,0.8,1}\lambda\in\{0.2,0.4,0.6,0.8,1\}. The results are shown in the right of Fig. 6. We observed that the saliency performance is inversely proportional to λ\lambda when λ>0.4\lambda>0.4, and the optimal setting is λ=0.4\lambda=0.4.

4. Synthetic Data for SOD

Setup. In this work, we do not focus on SOD network architecture design, so in our experiments, we adopt F3Net (Wei et al., 2020) as our saliency network by considering effectiveness and computational cost. Different from the previous works trained on the human wellannotated DUTS-TR (Wang et al., 2017) dataset (the detailed training data setting can be found in Table 4), we train our model on the SODGAN’s generated images-mask pairs (12k).

Datasets. We evaluate the performance of the proposed method on 5 commonly used benchmark datasets, including DUTS-TE (Wang et al., 2017), DUT-OMRON (Yang et al., 2013), ECSSD (Yan et al., 2013), HKU-IS (Zhao et al., 2015), and PASCAL-S (Li et al., 2014). Evaluation metrics. We adopt several widely-used metrics to evaluate our method, including the Precision-Recall (PR) curves, the F-measure curves, Mean Absolute Error (MAE), max and mean F-measure (Ran et al., 2014), S-measure (Fan et al., 2017) and Area Under Curve (AUC).

Competitors. We compare the proposed approach with 13 SOTA SOD models, including MWS (Zeng et al., 2019), EDNS (Zhang et al., 2020b), WS3A (Zhang et al., 2020c), SCWS (Yu et al., 2021), FCS (Zhang et al., 2020a), MFNet (Piao et al., 2021), DGRL (Wang et al., 2018), PAGR (Zhang et al., 2018), BAS (Qin et al., 2019), CPD (Wu et al., 2019), MINet (Pang et al., 2020), F3Net (Wei et al., 2020), PFSN (Ma et al., 2021), and SAMN (Liu et al., 2021a). For fair comparison, we evaluate these SOTA models by using the same metric code with the authors provided saliency maps.

Quantitative comparison. In Table 3, we compare our results with SOTA saliency methods. As indicated in Table 3, our method consistently achieves significant improvement compared with semi- and weakly- supervised methods in terms of 5 evaluation metrics. Concretely, our method improved by 1.13%1.13\%, 1.09%1.09\%, 2.32%2.32\%, 2.35%2.35\%, and 1.82%1.82\% on average compared to the second-best method in max F-measure on 5 datasets. Moreover, our saliency model even outperforms fully-supervised saliency models, such as CPD (Wu et al., 2019), BAS (Qin et al., 2019) and SAMN (Liu et al., 2021a), on ECSSD, HKU-IS and PASCAL-S datasets. Our approach trained on synthetic data achieves comparable or superior to the fully supervised F3Net (0.8422 vs. 0.8404 in terms of S-measure on the PASCAL-S) trained on more than 10k well-annotated image-label pairs. Besides, we also provide the PR and F-measure curves in Fig. 7, which also demonstrate the effectiveness of the synthesized high-quality image-mask pairs for saliency detection.

Qualitative comparison. As demonstrated in Fig. 8, our synthetic data supervised saliency model has better visual superiority than other SOTA models. Concretely, our model excels in dealing with various challenging scenarios, including cluttered backgrounds (the 1st row), low contrast objects (the 2nd row), inverted reflection in the water (the 3rd row), and small objects (the 4th row).

5. Conclusion

In this paper, we present a simple but powerful approach, namely SODGAN, to explore the potential of synthetic data for SOD. It opens up a new research paradigm for semi-supervised SOD, and shows that promising segmentation accuracy can be achieved by using controllable synthesized data. Our major novelty is to discover the interpretable direction that can disentangle the foreground object from the background in GANs feature space with only a few annotated images. Our work expands the application of the generative model to salient object detection tasks. We believe this is only the first step by utilizing synthetic data to train saliency deep networks.

References