Visual Saliency Transformer

Nian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao, Junwei Han

Introduction

SOD aims to detect objects that attract peoples’ eyes and can help many vision tasks, e.g., . Recently, RGB-D SOD has also gained growing interest with the extra spatial structure information from the depth data. Current state-of-the-art SOD methods are dominated by convolutional architectures , on both RGB and RGB-D data. They often adopt an encoder-decoder CNN architecture , where the encoder encodes the input image to multi-level features and the decoder integrates the extracted features to predict the final saliency map. Based on this simple architecture, most efforts have been made to build a powerful decoder for predicting better saliency results. To this end, they introduced various attention models , multi-scale feature integration methods , and multi-task learning frameworks . An additional demand for RGB-D SOD is to effectively fuse cross-modal information, i.e., the appearance information and the depth cues. Existing works propose various modality fusion methods, such as feature fusion , knowledge distillation , dynamic convolution , attention models , and graph neural networks . Hence, CNN-based methods have achieved impressive results .

However, all previous methods are limited in learning global long-range dependencies. Global contexts and global contrast have been proved crucial for saliency detection for a long time. Nevertheless, due to the intrinsic limitation of CNNs that they extract features in local sliding windows, previous methods can hardly exploit the crucial global cues. Although some methods utilized fully connected layers , global pooling layers , and non-local modules to incorporate the global context, they only did such in certain layers and the standard CNN-based architecture remains unchanged.

Recently, Transformer was proposed to model global long-range dependencies among word sequences for machine translation. The core idea is the self-attention mechanism, which leverages the query-key correlation to relate different positions in a sequence. Transformer stacks the self-attention layers multiple times in both encoder and decoder, thus can model long-range dependencies in every layer. Hence, it is natural to introduce the Transformer to SOD, leveraging the global cues in the model all the way.

In this paper, for the first time, we rethink SOD from a new sequence-to-sequence perspective and develop a novel unified model for both RGB and RGB-D SOD based on a pure transformer, which is named Visual Saliency Transformer. We follow the recently proposed ViT models to divide each image into patches and adopt the Transformer model on the patch sequence. Then, the Transformer propagates long-range dependencies between image patches, without any need of using convolution. However, it is not straightforward to apply ViT for SOD. On the one hand, how to perform dense prediction tasks based on pure transformer still remains an open question. On the other hand, ViT usually tokenizes the image to a very coarse scale. How to adapt ViT to the high-resolution prediction demand of SOD is also unclear.

To solve the first problem, we design a token-based transformer decoder by introducing task-related tokens to learn decision embeddings. Then, we propose a novel patch-task-attention mechanism to generate dense-prediction results, which provides a new paradigm for using transformer in dense prediction tasks. Motivated by previous SOD models that leveraged boundary detection to boost the SOD performance, we build a multi-task decoder to simultaneously conduct saliency and boundary detection by introducing a saliency token and a boundary token. This strategy simplifies the multitask prediction workflow by simply learning task-related tokens, thus largely reduces the computational costs while obtaining better results. To solve the second problem, inspired by the Tokens-to-Token (T2T) transformation , which reduces the length of tokens, we propose a new reverse T2T transformation to upsample tokens by expanding each token into multiple sub-tokens. Then, we upsample patch tokens progressively and fuse them with low-level tokens to obtain the final full-resolution saliency map. In addition, we also use a cross modality transformer to deeply explore the interaction between multi-modal information for RGB-D SOD. Finally, our VST outperforms existing state-of-the-art SOD methods with a comparable number of parameters and computational costs, on both RGB and RGB-D data.

Our main contributions can be summarized as follows:

For the first time, we design a novel unified model based on the pure transformer architecture for both RGB and RGB-D SOD, from a new perspective of sequence-to-sequence modeling.

We design a multi-task transformer decoder to jointly conduct saliency and boundary detection by introducing task-related tokens and patch-task-attention.

We propose a new token upsampling method for transformer-based framework.

Our proposed VST model achieves state-of-the-art results on both RGB and RGB-D SOD benchmark datasets, which demonstrates its effectiveness and the potential of transformer-based models for SOD.

Related Work

CNN-based approaches have become a mainstream trend in both RGB and RGB-D SOD and achieved promising performance. Most methods leveraged a multi-level feature fusion strategy by using UNet or HED-style network structures. Some works introduced the attention mechanism to learn more discriminative features, including spatial and channel attention or pixel-wise contextual attention . Other works tried to design recurrent networks to refine the saliency map step-by-step. In addition, some works introduced multi-task learning, e.g., fixation prediction , image caption , and edge detection to boost the SOD performance.

As for RGB-D SOD, many methods have designed various models to fuse RGB and depth features and obtained significant results. Some models adopted simple feature fusion methods, i.e., concatenation, summation, or multiplication. Some others leveraged the depth cues to generate spatial or channel attention to enhance the RGB features. Besides, dynamic convolution , graph neural networks , and knowledge distillation were also adopted to implement multi-modal feature fusion. In addition, adopted the cross-attention mechanism to propagate long-range cross-modal interactions between RGB and depth cues.

Different from previous CNN-based methods, we are the first to rethink SOD from a sequence-to-sequence perspective and propose a unified model based on pure transformer for both RGB and RGB-D SOD. In our model, we follow to leverage boundary detection to boost the SOD performance. However, different from these CNN-based models, we design a novel token-based multitask decoder to achieve this goal under the transformer framework.

2 Transformers in Computer Vision

Vaswani et al. first proposed a transformer encoder-decoder architecture for machine translation, where multi-head self-attention and point-wise feed-forward layers are stacked multiple times. Recently, more and more works have introduced the Transformer model to various computer vision tasks and achieved excellent results. Some works combined CNNs and transformers into hybrid architectures for object detection , panoptic segmentation , lane shape prediction , and so on. Typically, they first use CNNs to extract image features and then leverage the Transformer to incorporate long-range dependencies.

Other works design pure transformer models to process images from the sequence-to-sequence perspective. ViT divided each image into a sequence of flattened 2D patches and then adopted the Transformer for image classification. Touvron et al. introduced a teacher-student strategy to improve the data-efficiency of ViT and Wang et al. proposed a pyramid architecture to adapt ViT for dense prediction tasks. T2T-ViT adopted the T2T module to model local structures, thus generating multiscale token features. In this work, we adopt T2T-ViT as the backbone and propose a novel multitask decoder and a reverse T2T token upsampling method. It is noteworthy that our usage of task-related tokens is different from previous models. In , the class token is directly used for image classification via adopting a multilayer perceptron on the token embedding. However, we can not obtain dense prediction results directly from a single task token. Thus, we propose to perform patch-task-attention between patch tokens and the task tokens to predict saliency and boundary maps. We believe our strategy will also inspire future transformer models for other dense prediction tasks.

Another related work to ours is , which introduces transformer into the semantic segmentation task. The authors adopted a vision transformer as a backbone and then reshaped the token sequences to 2D image features. Then, they predicted full-resolution segmentation maps using convolution and bilinear upsampling. Their model still falls into the hybrid architecture category. In contrast, our model is a pure transformer architecture and does not rely on any convolution operation and bilinear upsampling.

Visual Saliency Transformer

Figure 1 shows the overall architecture of our proposed VST model. The main components include a transformer encoder based on T2T-ViT, a transformer convertor to convert patch tokens from the encoder space to the decoder space, and a multi-task transformer decoder.

Similar to other CNN-based SOD methods, which often utilize pretrained image classification models such as VGG and ResNet as the backbone of their encoders to extract image features, we adopt the pretrained T2T-ViT model as our backbone, as detailed below.

Given a sequence of patch tokens T′\bm{T}^{\prime} with length ll from the previous layer, T2T-ViT iteratively applies the T2T module, which is composed of a re-structurization step and a soft split step, to model the local structure information in T′\bm{T}^{\prime} and obtain a new sequence of tokens.

Different from ViT , the overlapped patch splitting adopted in T2T-ViT introduces local correspondence within neighbouring patches, thus bringing spatial priors.

The T2T transformation can be conducted iteratively multiple times. In each time, the re-structurization step first transforms previous token embeddings to new embeddings and also integrates long-range dependencies within all tokens. Then, the soft split operation aggregates the tokens in each k×kk\times k neighbour into a new token, which is ready to use for the next layer. Furthermore, when setting s<k−1s<k-1, the length of tokens can be reduced progressively.

1.2 Encoder with T2T-ViT Backbone

2 Transformer Convertor

We fuse TrE\bm{T}_{r}^{\mathcal{E}} and TdE\bm{T}_{d}^{\mathcal{E}} in the RGB-D converter to integrate the complementary information between the RGB and depth data. To this end, we design a Cross Modality Transformer (CMT), which consists of LCL^{\mathcal{C}} alternating cross-modality-attention layers and self-attention layers.

Under the pure transformer architecture, we modify the standard self-attention layer to propagate long-range cross-modal dependencies between the image and depth data, thus obtaining the cross-modality-attention, which is detailed as follows.

Next, we compute the “Scaled Dot-Product Attention” between the queries from one modality with the keys from the other modality. Then, the output is computed as a weighted sum of the values, formulated as:

We follow the standard Transformer architecture in and adopt the multi-head attention mechanism in the cross-modality-attention. The same positionwise feed-forward network, residual connections, and layer normalization are also used, forming our CMT layer.

After each adoption of the proposed CMT layer, we use one standard transformer layer on each RGB and depth patch token sequence, further enhancing their token embeddings. After alternately using CMT and transformer for LCL^{\mathcal{C}} times, we fuse the obtained RGB tokens and depth tokens by concatenation and then project them to the final converted tokens TC\bm{T}^{\mathcal{C}}, as shown in Figure 1.

2.2 RGB Convertor

To align with our RGB-D SOD model, for RGB SOD, we simply use LCL^{\mathcal{C}} standard transformer layers on TrE\bm{T}_{r}^{\mathcal{E}} to obtain the converted patch token sequence TC\bm{T}^{\mathcal{C}}.

3 Multi-task Transformer Decoder

Our decoder aims to decode the patch tokens TC\bm{T}^{\mathcal{C}} to saliency maps. Hence, we propose a novel token upsampling method with multi-level token fusion and a token-based multi-task decoder.

We argue that directly predicting saliency maps from TC\bm{T}^{\mathcal{C}} can not obtain high-quality results since the length of TC\bm{T}^{\mathcal{C}} is relatively small, i.e., l3=H16×W16l_{3}=\frac{H}{16}\times\frac{W}{16}, which is limited for dense prediction. Thus, we propose to upsample patch tokens first and then conduct dense prediction. Most CNN-based methods adopt bilinear upsampling to recover large scale feature maps. Alternatively, we propose a new token upsampling method under the transformer framework. Inspired by the T2T module that aggregates neighbour tokens to reduce the length of tokens progressively, we propose a reverse T2T (RT2T) transformation to upsample tokens by expanding each token into multiple sub-tokens, as shown in Figure 2(b).

Specifically, we first project the input patch tokens to reduce their embedding dimension from d=384d=384 to c=64c=64. Then, we use another linear projection to expand the embedding dimension from cc to ck2ck^{2}. Next, similar to the soft split step in T2T, each token is seen as a k×kk\times k image patch and neighbouring patches have ss overlapping. Then, we can fold the tokens as an image using pp zero-padding. The output image size can be computed using (2) reversely, i.e., given the length of the input patch tokens as ho×woh_{o}\times w_{o}, the spatial size of the out image is h×wh\times w. Finally, we reshape the image back to the upsampled tokens with size lo×cl_{o}\times c, where lo=h×wl_{o}=h\times w. By setting s<k−1s<k-1, the RT2T transformation can increase the length of the tokens. Motivated by T2T-ViT, we use RT2T three times and set k=k=, s=s=, and p=p=. Thus, the length of the patch tokens can be gradually upsampled to H×WH\times W, equaling to the original size of the input image.

Furthermore, motivated by the widely proved successes of multi-level feature fusion in existing SOD methods , we leverage low-level tokens with larger lengths from the T2T-ViT encoder, i.e., T1\bm{T}_{1} and T2\bm{T}_{2}, to provide accurate local structural information. For both RGB and RGB-D SOD, we only use the low-level tokens from the RGB transformer encoder. Concretely, we progressively fuse T2\bm{T}_{2} and T1\bm{T}_{1} with the upsampled patch tokens via concatenation and linear projection. Then, we adopt one transformer layer to obtain the decoder tokens TiD\bm{T}^{\mathcal{D}}_{i} at each level ii, where i=2,1i=2,1. The whole process is formulated as:

where $meansconcatenationalongthetokenembeddingdimension.“Linear”meanslinearprojectiontoreducetheembeddingdimensionaftertheconcatenationtomeans concatenation along the token embedding dimension. “Linear” means linear projection to reduce the embedding dimension after the concatenation toc.Finally,weuseanotherlinearprojectiontorecovertheembeddingdimensionof. Finally, we use another linear projection to recover the embedding dimension of\bm{T}^{\mathcal{D}}_{i}backtoback tod$.

3.2 Token Based Multi-task Prediction

Inspired by existing pure transformer methods , which add a class token on the patch token sequence for image classification, we also leverage task-related tokens to predict results. However, we can not obtain dense prediction results by directly using MLP on the task token embedding, as done in . Hence, we propose to perform patch-task-attention between the patch tokens and the task-related token to perform SOD.

In addition, motivated by the widely used boundary detection in SOD models , we also adopt the multi-task learning strategy to jointly perform saliency and boundary detection, thus using the latter to help boost the performance of the former.

Here we use the sigmoid activation for the attention computation since in each equation we only have one key.

Since TsD\bm{T}_{s}^{\mathcal{D}} and TbD\bm{T}_{b}^{\mathcal{D}} are at the 14\frac{1}{4} scale, we adopt the third RT2T transformation to upsample them to the full resolution. Finally, we apply two linear transformations with the sigmoid activation to project them to scalars in $$, and then reshape them to a 2D saliency map and a 2D boundary map, respectively. The whole process is given in Figure 1.

Experiments

For RGB SOD, we evaluate our VST model on six widely used benchmark datasets, including ECSSD (1,000 images), HKU-IS (4,447 images), PASCAL-S (850 images), DUT-O (5,168 images), SOD (300 images), and DUTS (10,553 training images and 5,019 testing images). For RGB-D SOD, we use nine widely used benchmark datasets: STERE (1,000 image pairs), LFSD (100 image pairs), RGBD135 (135 image pairs), SSD (80 image pairs), NJUD (1,985 image pairs), NLPR (1,000 image pairs), DUTLF-Depth (1,200 image pairs), SIP (929 image pairs), and ReDWeb-S (3,179 image pairs).

We adopt four widely used evaluation metrics to evaluate our model performance comprehensively. Specifically, Structure-measure SmS_{m} evaluates region-aware and object-aware structural similarity. Maximum F-measure (maxF) jointly considers precision and recall under the optimal threshold. Maximum enhanced-alignment measure EξmaxE_{\xi}^{\text{max}} simultaneously considers pixel-level errors and image-level errors. Mean Absolute Error (MAE) computes pixel-wise average absolute error. To evaluate the model complexity, we also report the multiply accumulate operations (MACs) and the number of parameters (Params).

2 Implementation Details

For fair comparisons, we follow most previous methods to use the training set of DUTS to train our VST for RGB SOD and use 1,485 images from NJUD, 700 images from NLPR, and 800 images from DUTLF-Depth to train our VST for RGB-D SOD. We follow to use a sober operator to generate the boundary ground truth from GT saliency maps. For depth data preprocessing, we normalize the depth maps to and duplicate them to three channels. Finally, we resize each image or depth map to 256×256256\times 256 pixels and then randomly crop 224×224224\times 224 image regions as the model input and use random flipping as data augmentation.

We use the pre-trained T2T-ViTt-14 model as our backbone since it has similar computational complexity as ResNet50 does. This model uses the efficient Performer and c=64c=64 in T2T modules, and sets LE=14L^{\mathcal{E}}=14. In our convertor and decoder, we set LC=L3D=4L^{\mathcal{C}}=L^{\mathcal{D}}_{3}=4 and L2D=L1D=2L^{\mathcal{D}}_{2}=L^{\mathcal{D}}_{1}=2 according to experimental results. We set the batchsizes as 11 and 8, and the total training steps as 40,000 and 60,000, for RGB and RGB-D SOD, respectively. For both of them, Adam is adopted as the optimizer and the binary cross entropy loss is used for both saliency and boundary prediction. The initial learning rate is set to 0.0001 and reduced by a factor of 10 at half and three-quarters of the total step, respectively. Deep supervision is also used to facilitate the model training, where we use the patch-task attention to predict saliency and boundary at each decoder level. We implemented our model using Pytorch and trained it on a GTX 1080 Ti GPU.

3 Ablation Study

Since our RGB-D VST is built by adding one more transformer encoder and additional CMT based on our RGB VST, while the other parts of the two models are the same, we conduct ablation studies based on our RGB-D VST to verify all of our proposed model components. The experimental results on four RGB-D SOD datasets, i.e., NJUD, DUTLF-Depth, STERE, and LFSD, are given in Table 1. We remove the transformer convertor and the decoder from our RGB-D VST as the baseline model. Specifically, it uses the two-stream transformer encoder to extract RGB encoder patch tokens TrE\bm{T}_{r}^{\mathcal{E}} and the depth encoder patch tokens TdE\bm{T}_{d}^{\mathcal{E}}, and then directly concatenate them and predict the saliency map with 1/16 scale by using MLP on each patch token.

For cross-modal information fusion, we deploy our proposed CMT right after the transformer encoder to substitute the concatenation fusion method in the baseline model, shown as “+CMT” in Table 1. Compared to the baseline, CMT brings performance gain especially on the NJUD and LFSD datasets, hence demonstrating its effectiveness.

Based on “+CMT” model, we further simply use bilinear upsampling (“+CMT+Bili”) to progressively upsample tokens to the full resolution and then predict the saliency map. The results show using bilinear upsampling to increase the resolution of the saliency map can largely improve the model performance. Then, we replace bilinear upsampling with our proposed RT2T token upsampling method (“+CMT+RT2T”). We find that RT2T leads to obvious performance improvement compared with using bilinear upsampling, which verifies its effectiveness.

We progressively fuse T1\bm{T}_{1} and T2\bm{T}_{2} in our decoder (“+CMT+RT2T+F”) to supply low-level fine-grained information. We find that this strategy further improves the model performance. Hence, leveraging low-level tokens in transformer is as important as fusing low-level features in CNN-based models.

Based on “+CMT+RT2T+F”, we further use our token-based multi-task decoder (TMD) to jointly perform saliency and boundary detection (“+CMT+RT2T+F+TMD”). It shows that using boundary detection can bring further performance gain for SOD on three out of four datasets. To very the effectiveness of our token-based prediction scheme, we try to directly use a conventional two-stream decoder (C2D) by using the “+RT2T+F” architecture twice to predict the saliency map and boundary map via MLP, without using task-related tokens. This model is denoted as “+CMT+RT2T+F+C2D” in Table 1. The parameters and MACs of TMD vs. C2D are 17.22 M vs. 20.35 M and 17.70 G vs. 28.27 G, respectively. The results show that using our TMD can achieve better results than using C2D on three out of four datasets, and also with much less computational costs. This clearly demonstrates the superiority of our proposed token-based transformer decoder.

4 Comparison with State-of-the-Art Methods

For RGB-D SOD, we compare our VST with 14 state-of-the-art RGB-D SOD methods, i.e., A2dele , JL-DCF , SSF-RGBD , UC-Net , S2S^{2}MA , PGAR , DANet , cmMS , ATSA , CMW , Cas-Gnn , HDFNet , CoNet , and BBS-Net . For RGB SOD, we compare our VST with 12 state-of-the-art RGB SOD models, including GateNet , CSF , LDF , MINet , ITSD , EGNet , TSPOANet , AFNet , PoolNet , CPD , BASNet , and PiCANet . Table 2 and Table 3 show the quantitative comparison results for RGB-D and RGB SOD, respectively. The results show that our VST outperforms all previous state-of-the-art CNN-based SOD models on both RGB and RGB-D benchmark datasets, with comparable number of parameters and relatively small MACs, hence demonstrating the great effectiveness of our VST. We also show visual comparison results among best-performed models in Figure 3. It shows our proposed VST can accurately detect salient objects in very challenging scenarios, e.g., big salient objects, cluttered backgrounds, foreground and background having similar appearances, etc.

Conclusion

In this paper, we are the first to rethink SOD from a sequence-to-sequence perspective and develop a novel unified model based on a pure transformer, for both RGB and RGB-D SOD. To handle the difficulty of applying transformers in dense prediction tasks, we propose a new token upsampling method under the transformer framework and fuse multi-level patch tokens. We also design a multi-task decoder by introducing task-related tokens and a novel patch-task-attention mechanism to jointly perform saliency and boundary detection. Our VST model achieves state-of-the-art results for both RGB and RGB-D SOD without relying on heavy computational costs, thus showing its great effectiveness. We also set a new paradigm for the open question of how to use transformer in dense prediction tasks.

This work was supported in part by the National Key R&D Program of China under Grant 2020AAA0105702, the National Science Foundation of China under Grant 62027813, 62036005, U20B2065, U20B2068.

References

Supplementary materials

We further report the results of ablation studies on four RGB SOD datasets, i.e., DUTS, HKU-IS, PASCAL-S, and SOD, in Table 4 to demonstrate the effectiveness of our VST model components.

The baseline model is using transformer encoder to extract patch tokens TrE\bm{T}_{r}^{\mathcal{E}} and then directly using TrE\bm{T}_{r}^{\mathcal{E}} to predict the saliency map with 1/16 scale by using MLP on each patch token. Based on the baseline, we insert RGB convertor right after the transformer encoder, shown as “+RC” in Table 4. Compared to the baseline, RC brings performance gains especially on the DUTS and PASCAL-S datasets, which demonstrates its effectiveness. For other components, i.e., RT2T, multi-level token fusion, and multi-task transformer decoder, we get consistent conclusions with the ablation studies on RGB-D SOD datasets as follows.

First, using bilinear upsampling (“+RC+Bili”) can significantly improve the model performance while using our proposed RT2T (“+RC+RT2T”) can further bring performance gains, hence demonstrating the effectiveness of our proposed RT2T. Second, based on “+RC+RT2T”, multi-level token fusion (“+RC+RT2T+F”) can lead to better performance on all four datasets, which verifies its effectiveness. Third, using multi-task transformer decoder (“+RC+RT2T+F+TMD”) can improve the model performance on all four datasets and it is also superior to the conventional two-stream decoder (“+RC+RT2T+F+C2D”).

To this end, the results of ablation studies on both RGB and RGB-D SOD datasets strongly demonstrate the effectiveness of our proposed VST components.

2 Layer Number Study

We conduct experiments to study the optimal numbers of different transformer layers, i.e., LCL^{\mathcal{C}} in the transformer convertor and LDL^{\mathcal{D}} in the multi-task transformer decoder, jointly considering computational costs and model performance. Note that there are three decoder modules at three scales in the multi-task transformer decoder, thus we set different transformer layer numbers for them, i.e., L3DL_{3}^{\mathcal{D}} for 1/16 scale, L2DL_{2}^{\mathcal{D}} for 1/8 scale, and L1DL_{1}^{\mathcal{D}} for 1/4 scale. The experimental results on four RGB-D SOD datasets, i.e., NJUD, DUTLF-Depth, STERE, and LFSD, are given in Table 5.

In our initial model setting, we set LC=L3D=8L^{\mathcal{C}}=L_{3}^{\mathcal{D}}=8. Since L2DL_{2}^{\mathcal{D}} and L1DL_{1}^{\mathcal{D}} are used at relatively large scales, we initially set both of them to 4, as shown in row I in Table 5. Then, we start to change the numbers of different layers.

We first reduce L2DL_{2}^{\mathcal{D}} and L1DL_{1}^{\mathcal{D}} from 4 to 2 to save computational costs. The experimental results on row II show that it can get comparable performance with less computational costs compared with row I. Hence, we set L2D=L1D=2L_{2}^{\mathcal{D}}=L_{1}^{\mathcal{D}}=2 and start to change L3DL_{3}^{\mathcal{D}} from 8 to 6, 4, 2, respectively, which are shown in row III, IV, V in Table 5. We find that as L3DL_{3}^{\mathcal{D}} decreases, the computation costs decrease gradually while the results are generally comparable. However, the model performance on row IV is better than that on row V on DUTLF-Depth and LFSD datasets. Thus, we set L3D=4L_{3}^{\mathcal{D}}=4 and start to change LCL^{\mathcal{C}} from 8 to 6, 4, 2, respectively, which are shown in row VI, VII, VIII. It can be seen that the performance on row VII is the best and the model has acceptable computational costs. Hence, we set LC=L3D=4L^{\mathcal{C}}=L_{3}^{\mathcal{D}}=4 and L2D=L1D=2L_{2}^{\mathcal{D}}=L_{1}^{\mathcal{D}}=2 as our final model setting.

3 More Visual Comparison with State-of-the-art Methods

We give more visual comparison results with the state-of-the-art RGB and RGB-D SOD methods in Figure 4 and Figure 5, respectively. It shows that our VST model can handle well in many challenging scenarios, i.e., big salient objects, cluttered backgrounds, foregrounds and backgrounds with very similar appearance, etc, while existing methods are heavily disturbed in these scenarios. Besides, we also show the boundary maps predicted by our RGB VST and RGB-D VST models in Figure 4 and Figure 5, respectively. It can be seen that our models can predict clear boundaries for salient objects.