Learning Noise-Aware Encoder-Decoder from Noisy Labels by Alternating Back-Propagation for Saliency Detection

Jing Zhang, Jianwen Xie, Nick Barnes

Introduction

Visual saliency detection aims to locate salient regions that attract human attention. Conventional saliency detection methods rely on human designed features to compute saliency for each pixel or superpixel. The deep learning revolution makes it possible to train end-to-end deep saliency detection models in a data-driven manner , outperforming handcrafted feature-based solutions by a wide margin. However, the success of deep models mainly depends on a large amount of accurate human labeling , which is typically expensive and time-consuming.

To relieve the burden of pixel-wise labeling, weakly supervised and unsupervised saliency detection models have been proposed. The former direction focuses on learning saliency from cheap but clean annotations, while the latter one studies learning saliency from noisy labels, which are typically obtained by conventional handcrafted feature-based methods. In this paper, we follow the second direction and propose a deep latent variable model that we call the noise-aware encoder-decoder to disentangle a clean saliency predictor from noisy labels. In general, a noisy label can be (1) a coarse saliency label generated by algorithmic pipelines using handcrafted features, (2) an imperfect human-annotated saliency label, or even (3) a clean label, which actually is a special case of noisy label, in which noise is none. Aiming at unsupervised saliency prediction, our paper assumes noisy labels to be produced by unsupervised handcrafted feature-based saliency methods, and places emphasis on disentangled representation of noisy labels by the noise-aware encoder-decoder.

Given a noisy dataset D={(Xi,Yi)}i=1nD=\{(X_{i},Y_{i})\}_{i=1}^{n} of nn examples, where XX and YY are image and its corresponding noisy saliency label, we intend to disentangle noise Δ\Delta and clean saliency SS from each noisy label YY, and learn a clean saliency predictor f1:X→Sf_{1}:X\rightarrow S. To achieve this, we propose a conditional latent variable model, which is a disentangled representation of noisy saliency YY. See Figure 1 for an illustration of the proposed model. In the context of the model, each noisy label is assumed to be generated by adding a specific noise or perturbation Δ\Delta to its clean saliency map SS that is dependent on its image XX. Specifically, the model consists of two sub-models: (1) saliency predictor f1f_{1}: an encoder-decoder network that maps an input image XX to a latent clean saliency map SS, and (2) noise generator f2f_{2}: a top-down neural network that produces a noise or error Δ\Delta from a low-dimensional Gaussian latent vector ZZ.

As a latent variable model, the rigorous maximum likelihood learning (MLE) typically requires to compute an intractable posterior distribution, which is an inference step. To learn the latent variable model, two algorithms can be adopted: variational auto-encoder (VAE) or alternating back-propagation (ABP) . VAE approximates MLE by minimizing the evidence lower bound with a separate inference model to approximate the true posterior, while ABP directly targets MLE and computes the posterior via Markov chain Monte Carlo (MCMC). In this paper, we generalize the ABP algorithm to learn the proposed model, which alternates the following two steps: (1) learning back-propagation for estimating the parameters of two sub-models, and (2) inferential back-propagation for inferring the latent vectors of training examples. As there may exist infinite combinations of SS and Δ\Delta such that S+ΔS+\Delta perfectly matches the provided noisy label YY, we further adopt the edge-aware smoothness loss to serve as a regularization to force each latent saliency map SS to have a similar structure as its input image XX. The learned disentangled saliency predictor f1f_{1} is the desired model for testing.

Our solution is different from existing weak or noisy label-based saliency approaches in the following aspects: Firstly, unlike , we don’t assume the saliency noise distribution is a Gaussian distribution. Our noise generator parameterized by a neural network is flexible enough to approximate any forms of structural noises. Secondly, we design a trainable noise generator to explicitly represent each noise Δ\Delta as a non-linear transformation of low-dimensional Gaussian noise ZZ, which is a latent variable that need to be inferred during training, while have no noise inference process. Thirdly, we have no constraints on the number of noisy labels generated from each image, while require multiple noisy labels per image for noise modeling or pseudo label generation. Lastly, our edge-aware smoothness loss serves as a regularization to force the produced latent saliency maps to be well aligned with their input images, which is different from , where object edges are used to produce pseudo saliency labels via multi-scale combinatorial grouping (MCG) .

Our main contributions can be summarized as follows:

We propose to learn a clean saliency predictor from noisy labels by a novel latent variable model that we call noise-aware encoder-decoder, in which each noisy label is represented as a sum of the clean saliency generated from the input image and a noise map generated from a latent vector.

We propose to train the proposed model by an alternating back-propagation (ABP) algorithm, which rigorously and efficiently maximizes the data likelihood without recruiting any other auxiliary model.

We propose to use an edge-aware smoothness loss as a regularization to prevent the model from converging to a trivial solution.

Experimental results on various benchmark datasets show the state-of-the-art performances of our framework in the task of unsupervised saliency detection, and also comparable performances with the existing fully-supervised saliency detection methods.

Related Work

Fully supervised saliency detection models mainly focus on designing networks that utilize image context information, multi-scale information, and image structure preservation. introduces feature polishing modules to update each level of features by incorporating all higher levels of context information. presents a cross feature module and a cascaded feedback decoder to effectively fuse different levels of features with a position-aware loss to penalize the boundary as well as pixel dissimilarity between saliency outputs and labels during training. proposes a saliency detection model that integrates both top-down and bottom-up saliency inferences in an iterative and cooperative manner. designs a pyramid attention structure with an edge detection module to perform edge-preserving salient object detection. uses a hybrid loss for boundary-aware saliency detection. proposes to use the stacked pyramid attention, which exploits multi-scale saliency information, along with an edge-related loss for saliency detection.

Learning saliency models without pixel-wise labeling can relieve the burden of costly pixel-level labeling. Those methods train saliency detection models with low-cost labels, such as image-level labels , noisy labels , object contours , scribble annotations , etc. introduces a foreground inference network to produce initial saliency maps with image-level labels, which are further refined and then treated as pseudo labels for iterative training. fuses saliency maps from unsupervised handcrafted feature-based methods with heuristics within a deep learning framework. collaboratively updates a saliency prediction module and a noise module to achieve learning saliency from multiple noisy labels. In , the initial noisy labels are refined by a self-supervised learning technique, and then treated as pseudo labels. creates a contour-to-saliency network, where saliency masks are generated by its contour detection branch via MCG and then those generated saliency masks are further used to train its saliency detection branch.

Learning from noisy labels techniques mainly focus on three main directions: (1) developing regularization ; (2) estimating the noise distribution by assuming that noisy labels are corrupted from clean labels by an unknown noise transition matrix and (3) training on selected samples . deals with noisy labeling by augmenting the prediction objective with a notion of perceptual consistency. proposes a framework to solve noisy label problem by updating both model parameters and labels. proposes to simultaneously learn the individual annotator model, which is represented by a confusion matrix, and the underlying true label distribution (i.e., classifier) from noisy observations. proposes to learn an extra network called MentorNet to generate a curriculum, which is a sample weighting scheme, for the base ConvNet called StudentNet. The generated curriculum helps the StudentNet to focus on those samples whose labels are likely to be correct.

Proposed Framework

The proposed model consists of two sub-models: (1) a saliency predictor, which is parameterized by an encoder-decoder network that maps the input image XX to the clean saliency SS; (2) a noise generator, which is parameterized by a top-down generator network that produces a noise or error Δ\Delta from a Gaussian latent vector ZZ. The resulting model is a sum of the two sub-models. Given training images with noisy labels, the MLE training of the model leads to an alternating back-propagation algorithm, which will be introduced in details in the following sections. The learned encoder-decoder network, which takes as input an image XX and outputs its clean saliency SS, is the disentangled model for saliency detection.

Let D={(Xi,Yi)}i=1nD=\{(X_{i},Y_{i})\}_{i=1}^{n} be the training dataset, where XX is the training image, YY is the noisy label of XX, nn is the size of the training dataset. Formally, the noise-aware encoder-decoder model can be formulated as follows:

where f1f_{1} in Eq. (1) is an encoder-decoder structure parameterized by θ1\theta_{1} for saliency detection. It takes as input an image XX and predicts its clean saliency map SS. Eq. (2) defines a noise generator, where ZZ is a low-dimensional Gaussian noise vector following N(0,Id)\mathcal{N}(0,I_{d}) (IdI_{d} is the dd-dimensional identity matrix) and f2f_{2} is a top-down deconvolutional neural network parametrized by θ2\theta_{2} that generates a saliency noise Δ\Delta from the noise vector ZZ. In Eq. (3), we assume that the observed noisy label YY is a sum of the clean saliency map SS and the noise Δ\Delta, plus a Gaussian residual ϵ∼N(0,σ2ID)\epsilon\sim\mathcal{N}(0,\sigma^{2}I_{D}), where we assume σ\sigma is given and IDI_{D} is the DD-dimensional identity matrix. Although ZZ is a Gaussian noise, the generated noise Δ\Delta is not necessarily Gaussian due to the non-linear transformation f2f_{2}.

We call our network the noise-aware encoder-decoder network as it explicitly decomposes a noisy label YY into a noise Δ\Delta and a clean label SS, and simultaneously learns a mapping from the image XX to the clean saliency map SS via an encoder-decoder network as shown in Fig. 1. Since the resulting model involves latent variables ZZ, training the model by maximum likelihood learning typically needs to learn the parameters θ1\theta_{1} and θ2\theta_{2}, and also infer the noise latent variable ZiZ_{i} for each observed data pair (Xi,Yi)(X_{i},Y_{i}). The noise and the saliency information are disentangled once the model is learned. The learned encoder-decoder sub-model S=f1(X;θ1)S=f_{1}(X;\theta_{1}) is the desired saliency detection network.

2 Maximum Likelihood via Alternating Back-Propagation

For notation simplicity, let f={f1,f2}f=\{f_{1},f_{2}\} and θ={θ1,θ2}\theta=\{\theta_{1},\theta_{2}\}. The proposed model is rewritten as a summarized form: Y=f(X,Z;θ)+ϵY=f(X,Z;\theta)+\epsilon, where Z∼N(0,Id)Z\sim\mathcal{N}(0,I_{d}) and ϵ\epsilon is the observation error. Given a dataset D={(Xi,Yi)}i=1nD=\{(X_{i},Y_{i})\}_{i=1}^{n}, each training example (Xi,Yi)(X_{i},Y_{i}) should have a corresponding ZiZ_{i}, but all data shares the same model parameter θ\theta. Intuitively, we should infer ZiZ_{i} and learn θ\theta to minimize the reconstruction error ∑i=1n∥Yi−f(Xi,Zi;θ)∥2\sum_{i=1}^{n}\|Y_{i}-f(X_{i},Z_{i};\theta)\|^{2} based on our formulation in Section 3.1. More formally, the model seeks to maximize the observed-data log-likelihood: L(θ)=∑i=1nlog⁡pθ(Yi∣Xi)\mathcal{L}(\theta)=\sum_{i=1}^{n}\log p_{\theta}(Y_{i}|X_{i}). Specifically, let p(Z)p(Z) be the prior distribution of ZZ. Let pθ(Y∣X,Z)∼N(f(X,Z;θ),σ2I)p_{\theta}(Y|X,Z)\sim\mathcal{N}(f(X,Z;\theta),\sigma^{2}I) be the conditional distribution of the noisy label YY given ZZ and XX. The conditional distribution of YY given XX is pθ(Y∣X)=∫p(Z)pθ(Y∣X,Z)dZp_{\theta}(Y|X)=\int p(Z)p_{\theta}(Y|X,Z)dZ with the latent variable ZZ integrated out.

The gradient of L(θ)\mathcal{L}(\theta) can be calculated according to the following identity:

The expectation term Epθ(Z∣Y,X)\text{E}_{p_{\theta}(Z|Y,X)} is analytically intractable. The conventional way of training such a latent variable model is the variational inference, in which the intractable posterior distribution pθ(Z∣Y,X)p_{\theta}(Z|Y,X) is approximated by an extra trainable tractable neural network pϕ(Z∣Y,X)p_{\phi}(Z|Y,X). In this paper, we resort to Monte Carlo average through drawing samples from the posterior distribution pθ(Z∣Y,X)p_{\theta}(Z|Y,X). This step corresponds to inferring the latent vector ZZ of the generator for each training example. Specifically, we use Langevin Dynamics (a gradient-based Monte Carlo method) to sample ZZ. The Langevin Dynamics for sampling Z∼pθ(Z∣Y,X)Z\sim p_{\theta}(Z|Y,X) iterates:

where tt and ss are the time step and step size of the Langevin Dynamics respectively. In each training iteration, for a given data pair (Xi,Yi)(X_{i},Y_{i}), we run ll steps of Langevin Dynamics to infer ZiZ_{i}. The Langevin Dynamics is initialized with Gaussian white noise (i.e., cold start) or the result of ZiZ_{i} obtained from the previous iteration (i.e., warm start). With the inferred ZiZ_{i} along with (Xi,Yi)(X_{i},Y_{i}), the gradient used to update the model parameters θ\theta is:

To encourage the latent output SS of the encoder-decoder f1f_{1} to be a meaningful saliency map, we add a negative edge-aware smoothness loss defined on SS to the log-likelihood objective L(θ)\mathcal{L}(\theta). The smoothness loss serves as a regularization term to avoid a trivial decomposition of SS and Δ\Delta given YY. Following , we use first-order derivatives (i.e., edge information) of both the latent clean saliency map SS and the input image XX to compute the smoothness loss

where Ψ\Psi is the Charbonnier penalty formula, defined as Ψ(s)=s2+1e−6\Psi(s)=\sqrt{s^{2}+1e^{-6}}, (u,v)(u,v) represent pixel coordinates, and dd indexes over the partial derivative in xx and yy directions. We estimate θ\theta by gradient ascent on L(θ)−λls(X,S;θ)\mathcal{L}(\theta)-\lambda l_{s}(X,S;\theta). In practice, we set λ=0.7\lambda=0.7, and α=10\alpha=10 in Eq. (8).

The whole process of updating both {Zi}\{Z_{i}\} and θ={θ1,θ2}\theta=\{\theta_{1},\theta_{2}\} is summarized in Algorithm 1, which is implemented as alternating back-propagation, because both gradients in Eq. (5) and (7) can be computed via back-propagation.

3 Comparison with Variational Inference

The proposed model can also be learned in a variational inference framework, where the intractable pθ(Z∣Y,X)p_{\theta}(Z|Y,X) in Eq. 4 is approximated by a tractable qϕ(Z∣Y,X)q_{\phi}(Z|Y,X), such as qϕ(Z∣Y,X)∼N(μϕ(Y,X),diag(vϕ(Y,X)))q_{\phi}(Z|Y,X)\sim\mathcal{N}(\mu_{\phi}(Y,X),\text{diag}(v_{\phi}(Y,X))), where both μϕ\mu_{\phi} and vϕv_{\phi} are bottom-up networks that map (X,Y)(X,Y) to ZZ, with ϕ\phi standing for all parameters of the bottom-up networks. The objective of variational inference is:

Recall that the maximum likelihood learning in our algorithm is equivalent to minimizing KL(qdata(Y∣X)∥pθ(Y∣X))\text{KL}(q_{\text{data}}(Y|X)\|p_{\theta}(Y|X)), where qdata(Y∣X)q_{\text{data}}(Y|X) is the conditional training data distribution. The accuracy of variational inference in Eq. 9 depends on the accuracy of an approximation of the true posterior distribution pθ(Z∣Y,X)p_{\theta}(Z|Y,X) by the inference model pϕ(Z∣Y,X)p_{\phi}(Z|Y,X). Theoretically, the variational inference is equivalent to the maximum likelihood solution, when KL(pϕ(Z∣Y,X)∥pθ(Z∣Y,X))=0\text{KL}(p_{\phi}(Z|Y,X)\|p_{\theta}(Z|Y,X))=0. However, in practice, there is always a gap between them due to the design of the inference model and the optimization difficulty. Therefore, without relying on an extra assisting model, our alternating back-propagation algorithm is more natural, straightforward and computationally efficient than variational inference. We refer readers to for a comprehensive tutorial on latent variable models.

4 Network Architectural Design

We now introduce the architectural designs of the encoder-decoder network (f1f_{1} in Eq. 1, or the green encoder-decoder in Fig. 1) and the noise generator network (f2f_{2} in Eq. 2, or the yellow decoder in Fig. 1) in this section.

Noise Generator: We construct the noise generator by using four cascaded deconvolutional layers, with a tanh activation function at the end to generate a noise map Δ\Delta in the range of $.BatchnormalizationandReLUlayersareaddedbetweentwonearbydeconvolutionallayers.Thedimensionalityofthelatentvariable. Batch normalization and ReLU layers are added between two nearby deconvolutional layers. The dimensionality of the latent variabled=8$.

Encoder-Decoder Network: Most existing deep saliency prediction networks are based on widely used backbone networks, including the VGG16-Net , ResNet , etc. Due to stride operations and multiple pooling layers used in these deep architectures, the saliency maps that are generated directly using the above backbone networks are low in spatial resolution, causing blurred edges. To overcome this, we propose an encoder-decoder-based framework with the VGG16-Net as the backbone as shown in Fig. 2. We denote the last convolutional layer of each convolutional group of VGG16-Net by s1,s2,...,s5s_{1},s_{2},...,s_{5} (corresponding to “relu1_2”, “relu2_2”, “relu3_3”, “relu4_3”, and “relu5_3”, respectively). To reduce the channel dimension of sms_{m}, a 1×11\times 1 convolutional layer is used to transform sms_{m} to sm′s^{\prime}_{m} of channel dimension 3232. Then a Residual Channel Attention (RCA) module is adopted to effectively fuse the intermediate high- and low-level features. Specifically, given the high- and low-level feature maps sm′s^{\prime}_{m} and sm−1′s^{\prime}_{m-1}, we first upsample sm′s^{\prime}_{m} to sm′′s^{\prime\prime}_{m}, which has the same spatial resolution as sm−1′s^{\prime}_{m-1}, by bilinear interpolation. Then we concatenate sm′′s^{\prime\prime}_{m} and sm−1′s^{\prime}_{m-1} to form a new feature map FmF_{m}. Similar to , we feed FmF_{m} to the RCA block to achieve the discriminative feature extraction. Inside each channel attention block, we perform “squeeze and excitation” by first “squeezing” the input feature map FmF_{m} to be half of the original channel size to obtain better nonlinear interactions across channels, and then “exciting” the squeezed feature map back to the original channel size. By adding a 3×33\times 3 convolutional layer to the lowest level of the RCA module, we obtain a one-channel saliency map Si=f1(Xi;θ1)S_{i}=f_{1}(X_{i};\theta_{1}).

Experiments

Datasets: We evaluate our performance on five saliency benchmark datasets. We use 10,553 images from the DUTS dataset for training, and we generate noisy labels from images using handcrafted feature based-methods, such as RBD , MR and GS due to their high efficiencies. Testing datasets include the DUTS testing set, ECSSD , DUT , HKU-IS and THUR .

Evaluation Metrics: Four metrics are used to evaluate the performance of our method and the competing methods, including two widely used metrics, i.e., Mean Absolute Error (M\mathcal{M}) and mean F-measure (FβF_{\beta}), and two newly released structure-aware metrics: mean E-measure (EξE_{\xi}) and S-measure (SαS_{\alpha}) .

Training Details: Each input image is rescaled to 352×352352\times 352 pixels. The encoder part in Fig. 2 is initialized using the VGG16-Net weights pretrained for image classification . The weights of other layers are initialized using the “truncated Gaussian” policy, and the biases are initialized to be zeros. We use the Adam optimizer with a momentum equal to 0.9, and decrease the learning rate γ\gamma by 10% after running 80% of the maximum epochs K=20K=20. The learning rate is initialized to be 0.0001. The number of Langevin steps ll is 6. The Langevin step size ss is 0.3. The σ\sigma in Eq.(3) is 0.1. The whole training takes 8 hours with a batch size 10 on a PC with an NVIDIA GeForce RTX GPU. We use the PaddlePaddle deep learning platform.

2 Comparison with the State-of-the-art Methods

We compare our method with seven fully supervised deep saliency prediction models and five weakly supervised/unsupervised saliency prediction models, and their performances are shown in Table 1 and Fig. 3. Table 1 shows that compared with the weakly supervised/unsupervised models, the proposed method achieves the best performance, especially on DUTS and HKU-IS datasets, where our method achieves an approximately 2% performance improvement for S-measure, and a 4% improvement for mean F-measure. Further, the proposed method even achieves comparable performances with some newly released fully supervised models. For example, we achieve comparable performance with NLDF and DGRL on all the five benchmark datasets. Fig.3 shows the 256-dimensional F-measure and E-measure (where the x-axis represents threshold for saliency map binarization) of our method and the competing methods on four datasets, where the weakly supervised/unsupervised methods are represented by dotted curves. We can observe that the performances of the fully supervised models are better than those of the weakly supervised/unsupervised models. As shown in Fig.3, our performance shows stability with different thresholds relative to the existing methods, indicating the robustness of our model.

Figure 4 demonstrates a qualitative comparison on several challenging cases. For example, the salient object in the first row is large, and connects to the image border. Most competing methods fail to segment the border-connected region, while our method almost finds the whole salient region in this case. Also, salient object in the second row has a long and narrow shape, which is challenging to some competing methods. Our method performs very well and precisely detect the salient object.

3 Ablation Study

We conduct the following experiments for an ablation study.

(1) Encoder-decoder f1f_{1} only: To study the effect of the noise generator, we evaluate the performance of the encoder-decoder (as shown in Fig. 2) directly learned from the noisy labels, without noise modeling or smoothness loss. The performance is shown in Table 2 with a label “f1f_{1}”, which is clearly worse than ours. This result is also consistent with the conclusion that deep neural networks is not robust to noise .

(2) Encoder-decoder f1f_{1} + smoothness loss lsl_{s}: As an extension of method “f1f_{1}”, one can add the smoothness loss in Eq. (8) as a regularization to better use the image prior information. We show the performance with a label “f1f_{1} & lsl_{s}” in Table 2. We observe a performance improvement compared with “f1f_{1}”, which indicates the usefulness of the edge-aware smoothness loss.

(3) Noisy-aware encoder-decoder without edge-aware smoothness loss: To study the effect of the smoothness regularization, we try to remove the smoothness loss from our model. As a result, we find that it will lead to trivial solutions i.e., Si=0H×WS_{i}=\mathbf{0}_{H\times W} for all training images.

(4) Alternative smoothness loss: We also replace our smoothness loss lsl_{s} by a cross-entropy loss lc(S,X)l_{c}(S,X) that is also defined on the first-order derivative of the saliency map SS and that of the image XX. The performance is shown in Table 2 as “ff & lcl_{c}”, which is better than or comparable with the existing weakly supervised/unsupervised methods shown in Table 1. By comparing the performance of “ff & lcl_{c}” with that of the full model, we observe that the smoothness loss ls(S,X)l_{s}(S,X) in Eq. 8 works better than the cross-entropy loss lc(S,X)l_{c}(S,X). The former puts a soft constraint on their boundaries, while the latter has a strong effect on forcing both boundaries of SS and XX to be the same. Although the saliency boundary are usually aligned with the image boundary, but they are not exactly the same. A soft and indirect penalty for edge dissimilarity seems to be more useful.

4 Model Analysis

We further explore our proposed model in this section.

(1) Learn the model from saliency labels generated by fully supervised pre-trained models: One way to use our method is treating it as a boosting strategy for the current fully-supervised models. To verify this, we first generate saliency maps by using a pre-trained fully-supervised saliency network, e.g., BASNet . We treat the outputs as noisy labels, on which we train our model. The performances are shown in Table 3 as ff-BAS. By comparing the performances of ff-BAS with those of BASNet in Table 1, we find that ff-BAS is comparable with or better than BASNet, which means that our method can further refine the outputs of the state-of-the-art pre-trained fully-supervised models if their performances are still far from perfect.

(2) Create one single noisy label for each image: In previous experiments, our noisy labels are generated by handcrafted feature-based saliency methods in the setting of multiple noisy labels per image. Specifically, we produce three noisy labels for each training image by methods RBD , MR and GS , respectively. As our method has no constraints on the number of generated noisy labels per image, we conduct experiments to test our models learned in the setting of one noisy label per image. In Table 3, we report the performances of the models learned from noisy labels generated by RBD , MR and GS , respectively. We use ff-RBD, ff-MR and ff-GS to represent their results, respectively. We observe comparable performances with those using the setting of multiple noisy labels per image, which means our method is robust to the number of noisy labels generated from each image and the quality of the generated noisy labels. (RBD ranks the 1st1^{st} among unsupervised saliency detection models in . RBD, MR and GS represent different levels of qualities of noisy labels). We also show in Table 3 the performances of the above handcrafted feature-based methods, which are denoted by RBD, MR and GS, respectively. The big gap between RBD/MR/GS and ff-RBD/ff-MR/ff-GS demonstrates the effectiveness of our model.

(3) Train the model from clean labels: The proposed noise-aware encoder-decoder can learn from clean labels, because clean label can be treated as a special case of noisy label, and the noise generator will learn to output zero noise maps in this scenario. We show experiments on training our model from clean labels obtained from the DUTS training dataset. The performances denoted by ff* are shown in Table 3. For comparison purpose, we also train the encoder-decoder component without the noise generator module from clean labels, whose results are displayed in Table 3 with a name f1f_{1}*. We find that (1) our model can still work very well when clean labels are available, and (2) ff* achieves better performance than f1f_{1}*, indicating that even though those clean labels are obtained from training dataset, they are still “noisy” because of imperfect human annotation. Our noise-handling strategy is still beneficial in this situation.

(4) Train the model by variational inference: In this paper, we train our model by alternating back-propagation algorithm that maximizes the observed-data log-likelihood, where we adopt Langevin Dynamics to draw samples from the posterior distribution pθ(Z∣Y,X)p_{\theta}(Z|Y,X), and use the empirical average to compute the gradient of the log-likelihood in Eq.(4). One can also train the model in a conditional variational inference framework as shown in Eq. (9). Following cVAE , we design an inference network pϕ(Z∣Y,X)p_{\phi}(Z|Y,X), which consists of four cascade convolutional layers and a fully connected layer at the end, to map the image XX and the noisy label YY to the d=8d=8 dimensional latent space ZZ. The resulting loss function includes a reconstruction loss ∥Yi−f(Xi,Zi,θ)∥2\|Y_{i}-f(X_{i},Z_{i},\theta)\|^{2}, a KL-divergence loss KL(pϕ(Z∣Y,X)∥pθ(Z∣Y,X))\text{KL}(p_{\phi}(Z|Y,X)\|p_{\theta}(Z|Y,X)) and the edge-aware smoothness loss presented in Eq.(8). We present the cVAE results in Table 3. Our results learned by ABP outperforms those by cVAE. The main reason lies in the fact that the gap between the approximate inference model and the true inference model, i.e., KL(pϕ(Z∣Y,X)∥pθ(Z∣Y,X))\text{KL}(p_{\phi}(Z|Y,X)\|p_{\theta}(Z|Y,X)), is hard to be zero in practise, especially when the capacity of pϕ(Z∣Y,X)p_{\phi}(Z|Y,X) is less than that of pθ(Z∣Y,X)p_{\theta}(Z|Y,X) due to an inappropriate architectural design of pϕ(Z∣Y,X)p_{\phi}(Z|Y,X). On the contrary, our Langevin Dynamics-based inference step, which is derived from the model, is more natural and accurate.

Conclusion

Although clean pixel-wise annotations can lead to better performances, the expensive and time-consuming labeling process limits the applications of those fully supervised models. Inspired by previous work , we propose a noise-aware encoder-decoder network for disentangled learning of a clean saliency predictor from noisy labels. The model represents each noisy saliency label as an addition of perturbation or noise from an unknown distribution to the clean saliency map predicted from the corresponding image. The clean saliency predictor is an encoder-decoder framework, while the noise generator is a non-linear transformation of a Gaussian noise vector, in which the transformation is parameterized by a neural network. Edge-aware smoothness loss is also utilized to prevent the model from converging to a trivial solution. We propose to train the model by a simple yet efficient alternating back-propagation algorithm , which is superior to variational inference. Extensive experiments conducted on different benchmark datasets demonstrate the effectiveness and robustness of our model and learning algorithm.

Acknowledgments. This research was supported in part by the Australia Research Council Centre of Excellence for Robotics Vision (CE140100016).

References