Learning Noise-Aware Encoder-Decoder from Noisy Labels by Alternating Back-Propagation for Saliency Detection
Jing Zhang, Jianwen Xie, Nick Barnes
Introduction
Visual saliency detection aims to locate salient regions that attract human attention. Conventional saliency detection methods rely on human designed features to compute saliency for each pixel or superpixel. The deep learning revolution makes it possible to train end-to-end deep saliency detection models in a data-driven manner , outperforming handcrafted feature-based solutions by a wide margin. However, the success of deep models mainly depends on a large amount of accurate human labeling , which is typically expensive and time-consuming.
To relieve the burden of pixel-wise labeling, weakly supervised and unsupervised saliency detection models have been proposed. The former direction focuses on learning saliency from cheap but clean annotations, while the latter one studies learning saliency from noisy labels, which are typically obtained by conventional handcrafted feature-based methods. In this paper, we follow the second direction and propose a deep latent variable model that we call the noise-aware encoder-decoder to disentangle a clean saliency predictor from noisy labels. In general, a noisy label can be (1) a coarse saliency label generated by algorithmic pipelines using handcrafted features, (2) an imperfect human-annotated saliency label, or even (3) a clean label, which actually is a special case of noisy label, in which noise is none. Aiming at unsupervised saliency prediction, our paper assumes noisy labels to be produced by unsupervised handcrafted feature-based saliency methods, and places emphasis on disentangled representation of noisy labels by the noise-aware encoder-decoder.
Given a noisy dataset of examples, where and are image and its corresponding noisy saliency label, we intend to disentangle noise and clean saliency from each noisy label , and learn a clean saliency predictor . To achieve this, we propose a conditional latent variable model, which is a disentangled representation of noisy saliency . See Figure 1 for an illustration of the proposed model. In the context of the model, each noisy label is assumed to be generated by adding a specific noise or perturbation to its clean saliency map that is dependent on its image . Specifically, the model consists of two sub-models: (1) saliency predictor : an encoder-decoder network that maps an input image to a latent clean saliency map , and (2) noise generator : a top-down neural network that produces a noise or error from a low-dimensional Gaussian latent vector .
As a latent variable model, the rigorous maximum likelihood learning (MLE) typically requires to compute an intractable posterior distribution, which is an inference step. To learn the latent variable model, two algorithms can be adopted: variational auto-encoder (VAE) or alternating back-propagation (ABP) . VAE approximates MLE by minimizing the evidence lower bound with a separate inference model to approximate the true posterior, while ABP directly targets MLE and computes the posterior via Markov chain Monte Carlo (MCMC). In this paper, we generalize the ABP algorithm to learn the proposed model, which alternates the following two steps: (1) learning back-propagation for estimating the parameters of two sub-models, and (2) inferential back-propagation for inferring the latent vectors of training examples. As there may exist infinite combinations of and such that perfectly matches the provided noisy label , we further adopt the edge-aware smoothness loss to serve as a regularization to force each latent saliency map to have a similar structure as its input image . The learned disentangled saliency predictor is the desired model for testing.
Our solution is different from existing weak or noisy label-based saliency approaches in the following aspects: Firstly, unlike , we don’t assume the saliency noise distribution is a Gaussian distribution. Our noise generator parameterized by a neural network is flexible enough to approximate any forms of structural noises. Secondly, we design a trainable noise generator to explicitly represent each noise as a non-linear transformation of low-dimensional Gaussian noise , which is a latent variable that need to be inferred during training, while have no noise inference process. Thirdly, we have no constraints on the number of noisy labels generated from each image, while require multiple noisy labels per image for noise modeling or pseudo label generation. Lastly, our edge-aware smoothness loss serves as a regularization to force the produced latent saliency maps to be well aligned with their input images, which is different from , where object edges are used to produce pseudo saliency labels via multi-scale combinatorial grouping (MCG) .
Our main contributions can be summarized as follows:
We propose to learn a clean saliency predictor from noisy labels by a novel latent variable model that we call noise-aware encoder-decoder, in which each noisy label is represented as a sum of the clean saliency generated from the input image and a noise map generated from a latent vector.
We propose to train the proposed model by an alternating back-propagation (ABP) algorithm, which rigorously and efficiently maximizes the data likelihood without recruiting any other auxiliary model.
We propose to use an edge-aware smoothness loss as a regularization to prevent the model from converging to a trivial solution.
Experimental results on various benchmark datasets show the state-of-the-art performances of our framework in the task of unsupervised saliency detection, and also comparable performances with the existing fully-supervised saliency detection methods.
Related Work
Fully supervised saliency detection models mainly focus on designing networks that utilize image context information, multi-scale information, and image structure preservation. introduces feature polishing modules to update each level of features by incorporating all higher levels of context information. presents a cross feature module and a cascaded feedback decoder to effectively fuse different levels of features with a position-aware loss to penalize the boundary as well as pixel dissimilarity between saliency outputs and labels during training. proposes a saliency detection model that integrates both top-down and bottom-up saliency inferences in an iterative and cooperative manner. designs a pyramid attention structure with an edge detection module to perform edge-preserving salient object detection. uses a hybrid loss for boundary-aware saliency detection. proposes to use the stacked pyramid attention, which exploits multi-scale saliency information, along with an edge-related loss for saliency detection.
Learning saliency models without pixel-wise labeling can relieve the burden of costly pixel-level labeling. Those methods train saliency detection models with low-cost labels, such as image-level labels , noisy labels , object contours , scribble annotations , etc. introduces a foreground inference network to produce initial saliency maps with image-level labels, which are further refined and then treated as pseudo labels for iterative training. fuses saliency maps from unsupervised handcrafted feature-based methods with heuristics within a deep learning framework. collaboratively updates a saliency prediction module and a noise module to achieve learning saliency from multiple noisy labels. In , the initial noisy labels are refined by a self-supervised learning technique, and then treated as pseudo labels. creates a contour-to-saliency network, where saliency masks are generated by its contour detection branch via MCG and then those generated saliency masks are further used to train its saliency detection branch.
Learning from noisy labels techniques mainly focus on three main directions: (1) developing regularization ; (2) estimating the noise distribution by assuming that noisy labels are corrupted from clean labels by an unknown noise transition matrix and (3) training on selected samples . deals with noisy labeling by augmenting the prediction objective with a notion of perceptual consistency. proposes a framework to solve noisy label problem by updating both model parameters and labels. proposes to simultaneously learn the individual annotator model, which is represented by a confusion matrix, and the underlying true label distribution (i.e., classifier) from noisy observations. proposes to learn an extra network called MentorNet to generate a curriculum, which is a sample weighting scheme, for the base ConvNet called StudentNet. The generated curriculum helps the StudentNet to focus on those samples whose labels are likely to be correct.
Proposed Framework
The proposed model consists of two sub-models: (1) a saliency predictor, which is parameterized by an encoder-decoder network that maps the input image to the clean saliency ; (2) a noise generator, which is parameterized by a top-down generator network that produces a noise or error from a Gaussian latent vector . The resulting model is a sum of the two sub-models. Given training images with noisy labels, the MLE training of the model leads to an alternating back-propagation algorithm, which will be introduced in details in the following sections. The learned encoder-decoder network, which takes as input an image and outputs its clean saliency , is the disentangled model for saliency detection.
Let be the training dataset, where is the training image, is the noisy label of , is the size of the training dataset. Formally, the noise-aware encoder-decoder model can be formulated as follows:
where in Eq. (1) is an encoder-decoder structure parameterized by for saliency detection. It takes as input an image and predicts its clean saliency map . Eq. (2) defines a noise generator, where is a low-dimensional Gaussian noise vector following ( is the -dimensional identity matrix) and is a top-down deconvolutional neural network parametrized by that generates a saliency noise from the noise vector . In Eq. (3), we assume that the observed noisy label is a sum of the clean saliency map and the noise , plus a Gaussian residual , where we assume is given and is the -dimensional identity matrix. Although is a Gaussian noise, the generated noise is not necessarily Gaussian due to the non-linear transformation .
We call our network the noise-aware encoder-decoder network as it explicitly decomposes a noisy label into a noise and a clean label , and simultaneously learns a mapping from the image to the clean saliency map via an encoder-decoder network as shown in Fig. 1. Since the resulting model involves latent variables , training the model by maximum likelihood learning typically needs to learn the parameters and , and also infer the noise latent variable for each observed data pair . The noise and the saliency information are disentangled once the model is learned. The learned encoder-decoder sub-model is the desired saliency detection network.
2 Maximum Likelihood via Alternating Back-Propagation
For notation simplicity, let and . The proposed model is rewritten as a summarized form: , where and is the observation error. Given a dataset , each training example should have a corresponding , but all data shares the same model parameter . Intuitively, we should infer and learn to minimize the reconstruction error based on our formulation in Section 3.1. More formally, the model seeks to maximize the observed-data log-likelihood: . Specifically, let be the prior distribution of . Let be the conditional distribution of the noisy label given and . The conditional distribution of given is with the latent variable integrated out.
The gradient of can be calculated according to the following identity:
The expectation term is analytically intractable. The conventional way of training such a latent variable model is the variational inference, in which the intractable posterior distribution is approximated by an extra trainable tractable neural network . In this paper, we resort to Monte Carlo average through drawing samples from the posterior distribution . This step corresponds to inferring the latent vector of the generator for each training example. Specifically, we use Langevin Dynamics (a gradient-based Monte Carlo method) to sample . The Langevin Dynamics for sampling iterates:
where and are the time step and step size of the Langevin Dynamics respectively. In each training iteration, for a given data pair , we run steps of Langevin Dynamics to infer . The Langevin Dynamics is initialized with Gaussian white noise (i.e., cold start) or the result of obtained from the previous iteration (i.e., warm start). With the inferred along with , the gradient used to update the model parameters is:
To encourage the latent output of the encoder-decoder to be a meaningful saliency map, we add a negative edge-aware smoothness loss defined on to the log-likelihood objective . The smoothness loss serves as a regularization term to avoid a trivial decomposition of and given . Following , we use first-order derivatives (i.e., edge information) of both the latent clean saliency map and the input image to compute the smoothness loss
where is the Charbonnier penalty formula, defined as , represent pixel coordinates, and indexes over the partial derivative in and directions. We estimate by gradient ascent on . In practice, we set , and in Eq. (8).
The whole process of updating both and is summarized in Algorithm 1, which is implemented as alternating back-propagation, because both gradients in Eq. (5) and (7) can be computed via back-propagation.
3 Comparison with Variational Inference
The proposed model can also be learned in a variational inference framework, where the intractable in Eq. 4 is approximated by a tractable , such as , where both and are bottom-up networks that map to , with standing for all parameters of the bottom-up networks. The objective of variational inference is:
Recall that the maximum likelihood learning in our algorithm is equivalent to minimizing , where is the conditional training data distribution. The accuracy of variational inference in Eq. 9 depends on the accuracy of an approximation of the true posterior distribution by the inference model . Theoretically, the variational inference is equivalent to the maximum likelihood solution, when . However, in practice, there is always a gap between them due to the design of the inference model and the optimization difficulty. Therefore, without relying on an extra assisting model, our alternating back-propagation algorithm is more natural, straightforward and computationally efficient than variational inference. We refer readers to for a comprehensive tutorial on latent variable models.
4 Network Architectural Design
We now introduce the architectural designs of the encoder-decoder network ( in Eq. 1, or the green encoder-decoder in Fig. 1) and the noise generator network ( in Eq. 2, or the yellow decoder in Fig. 1) in this section.
Noise Generator: We construct the noise generator by using four cascaded deconvolutional layers, with a tanh activation function at the end to generate a noise map in the range of $d=8$.
Encoder-Decoder Network: Most existing deep saliency prediction networks are based on widely used backbone networks, including the VGG16-Net , ResNet , etc. Due to stride operations and multiple pooling layers used in these deep architectures, the saliency maps that are generated directly using the above backbone networks are low in spatial resolution, causing blurred edges. To overcome this, we propose an encoder-decoder-based framework with the VGG16-Net as the backbone as shown in Fig. 2. We denote the last convolutional layer of each convolutional group of VGG16-Net by (corresponding to “relu1_2”, “relu2_2”, “relu3_3”, “relu4_3”, and “relu5_3”, respectively). To reduce the channel dimension of , a convolutional layer is used to transform to of channel dimension . Then a Residual Channel Attention (RCA) module is adopted to effectively fuse the intermediate high- and low-level features. Specifically, given the high- and low-level feature maps and , we first upsample to , which has the same spatial resolution as , by bilinear interpolation. Then we concatenate and to form a new feature map . Similar to , we feed to the RCA block to achieve the discriminative feature extraction. Inside each channel attention block, we perform “squeeze and excitation” by first “squeezing” the input feature map to be half of the original channel size to obtain better nonlinear interactions across channels, and then “exciting” the squeezed feature map back to the original channel size. By adding a convolutional layer to the lowest level of the RCA module, we obtain a one-channel saliency map .
Experiments
Datasets: We evaluate our performance on five saliency benchmark datasets. We use 10,553 images from the DUTS dataset for training, and we generate noisy labels from images using handcrafted feature based-methods, such as RBD , MR and GS due to their high efficiencies. Testing datasets include the DUTS testing set, ECSSD , DUT , HKU-IS and THUR .
Evaluation Metrics: Four metrics are used to evaluate the performance of our method and the competing methods, including two widely used metrics, i.e., Mean Absolute Error () and mean F-measure (), and two newly released structure-aware metrics: mean E-measure () and S-measure () .
Training Details: Each input image is rescaled to pixels. The encoder part in Fig. 2 is initialized using the VGG16-Net weights pretrained for image classification . The weights of other layers are initialized using the “truncated Gaussian” policy, and the biases are initialized to be zeros. We use the Adam optimizer with a momentum equal to 0.9, and decrease the learning rate by 10% after running 80% of the maximum epochs . The learning rate is initialized to be 0.0001. The number of Langevin steps is 6. The Langevin step size is 0.3. The in Eq.(3) is 0.1. The whole training takes 8 hours with a batch size 10 on a PC with an NVIDIA GeForce RTX GPU. We use the PaddlePaddle deep learning platform.
2 Comparison with the State-of-the-art Methods
We compare our method with seven fully supervised deep saliency prediction models and five weakly supervised/unsupervised saliency prediction models, and their performances are shown in Table 1 and Fig. 3. Table 1 shows that compared with the weakly supervised/unsupervised models, the proposed method achieves the best performance, especially on DUTS and HKU-IS datasets, where our method achieves an approximately 2% performance improvement for S-measure, and a 4% improvement for mean F-measure. Further, the proposed method even achieves comparable performances with some newly released fully supervised models. For example, we achieve comparable performance with NLDF and DGRL on all the five benchmark datasets. Fig.3 shows the 256-dimensional F-measure and E-measure (where the x-axis represents threshold for saliency map binarization) of our method and the competing methods on four datasets, where the weakly supervised/unsupervised methods are represented by dotted curves. We can observe that the performances of the fully supervised models are better than those of the weakly supervised/unsupervised models. As shown in Fig.3, our performance shows stability with different thresholds relative to the existing methods, indicating the robustness of our model.
Figure 4 demonstrates a qualitative comparison on several challenging cases. For example, the salient object in the first row is large, and connects to the image border. Most competing methods fail to segment the border-connected region, while our method almost finds the whole salient region in this case. Also, salient object in the second row has a long and narrow shape, which is challenging to some competing methods. Our method performs very well and precisely detect the salient object.
3 Ablation Study
We conduct the following experiments for an ablation study.
(1) Encoder-decoder only: To study the effect of the noise generator, we evaluate the performance of the encoder-decoder (as shown in Fig. 2) directly learned from the noisy labels, without noise modeling or smoothness loss. The performance is shown in Table 2 with a label “”, which is clearly worse than ours. This result is also consistent with the conclusion that deep neural networks is not robust to noise .
(2) Encoder-decoder + smoothness loss : As an extension of method “”, one can add the smoothness loss in Eq. (8) as a regularization to better use the image prior information. We show the performance with a label “ & ” in Table 2. We observe a performance improvement compared with “”, which indicates the usefulness of the edge-aware smoothness loss.
(3) Noisy-aware encoder-decoder without edge-aware smoothness loss: To study the effect of the smoothness regularization, we try to remove the smoothness loss from our model. As a result, we find that it will lead to trivial solutions i.e., for all training images.
(4) Alternative smoothness loss: We also replace our smoothness loss by a cross-entropy loss that is also defined on the first-order derivative of the saliency map and that of the image . The performance is shown in Table 2 as “ & ”, which is better than or comparable with the existing weakly supervised/unsupervised methods shown in Table 1. By comparing the performance of “ & ” with that of the full model, we observe that the smoothness loss in Eq. 8 works better than the cross-entropy loss . The former puts a soft constraint on their boundaries, while the latter has a strong effect on forcing both boundaries of and to be the same. Although the saliency boundary are usually aligned with the image boundary, but they are not exactly the same. A soft and indirect penalty for edge dissimilarity seems to be more useful.
4 Model Analysis
We further explore our proposed model in this section.
(1) Learn the model from saliency labels generated by fully supervised pre-trained models: One way to use our method is treating it as a boosting strategy for the current fully-supervised models. To verify this, we first generate saliency maps by using a pre-trained fully-supervised saliency network, e.g., BASNet . We treat the outputs as noisy labels, on which we train our model. The performances are shown in Table 3 as -BAS. By comparing the performances of -BAS with those of BASNet in Table 1, we find that -BAS is comparable with or better than BASNet, which means that our method can further refine the outputs of the state-of-the-art pre-trained fully-supervised models if their performances are still far from perfect.
(2) Create one single noisy label for each image: In previous experiments, our noisy labels are generated by handcrafted feature-based saliency methods in the setting of multiple noisy labels per image. Specifically, we produce three noisy labels for each training image by methods RBD , MR and GS , respectively. As our method has no constraints on the number of generated noisy labels per image, we conduct experiments to test our models learned in the setting of one noisy label per image. In Table 3, we report the performances of the models learned from noisy labels generated by RBD , MR and GS , respectively. We use -RBD, -MR and -GS to represent their results, respectively. We observe comparable performances with those using the setting of multiple noisy labels per image, which means our method is robust to the number of noisy labels generated from each image and the quality of the generated noisy labels. (RBD ranks the among unsupervised saliency detection models in . RBD, MR and GS represent different levels of qualities of noisy labels). We also show in Table 3 the performances of the above handcrafted feature-based methods, which are denoted by RBD, MR and GS, respectively. The big gap between RBD/MR/GS and -RBD/-MR/-GS demonstrates the effectiveness of our model.
(3) Train the model from clean labels: The proposed noise-aware encoder-decoder can learn from clean labels, because clean label can be treated as a special case of noisy label, and the noise generator will learn to output zero noise maps in this scenario. We show experiments on training our model from clean labels obtained from the DUTS training dataset. The performances denoted by * are shown in Table 3. For comparison purpose, we also train the encoder-decoder component without the noise generator module from clean labels, whose results are displayed in Table 3 with a name *. We find that (1) our model can still work very well when clean labels are available, and (2) * achieves better performance than *, indicating that even though those clean labels are obtained from training dataset, they are still “noisy” because of imperfect human annotation. Our noise-handling strategy is still beneficial in this situation.
(4) Train the model by variational inference: In this paper, we train our model by alternating back-propagation algorithm that maximizes the observed-data log-likelihood, where we adopt Langevin Dynamics to draw samples from the posterior distribution , and use the empirical average to compute the gradient of the log-likelihood in Eq.(4). One can also train the model in a conditional variational inference framework as shown in Eq. (9). Following cVAE , we design an inference network , which consists of four cascade convolutional layers and a fully connected layer at the end, to map the image and the noisy label to the dimensional latent space . The resulting loss function includes a reconstruction loss , a KL-divergence loss and the edge-aware smoothness loss presented in Eq.(8). We present the cVAE results in Table 3. Our results learned by ABP outperforms those by cVAE. The main reason lies in the fact that the gap between the approximate inference model and the true inference model, i.e., , is hard to be zero in practise, especially when the capacity of is less than that of due to an inappropriate architectural design of . On the contrary, our Langevin Dynamics-based inference step, which is derived from the model, is more natural and accurate.
Conclusion
Although clean pixel-wise annotations can lead to better performances, the expensive and time-consuming labeling process limits the applications of those fully supervised models. Inspired by previous work , we propose a noise-aware encoder-decoder network for disentangled learning of a clean saliency predictor from noisy labels. The model represents each noisy saliency label as an addition of perturbation or noise from an unknown distribution to the clean saliency map predicted from the corresponding image. The clean saliency predictor is an encoder-decoder framework, while the noise generator is a non-linear transformation of a Gaussian noise vector, in which the transformation is parameterized by a neural network. Edge-aware smoothness loss is also utilized to prevent the model from converging to a trivial solution. We propose to train the model by a simple yet efficient alternating back-propagation algorithm , which is superior to variational inference. Extensive experiments conducted on different benchmark datasets demonstrate the effectiveness and robustness of our model and learning algorithm.
Acknowledgments. This research was supported in part by the Australia Research Council Centre of Excellence for Robotics Vision (CE140100016).