Fast End-to-End Trainable Guided Filter

Huikai Wu, Shuai Zheng, Junge Zhang, Kaiqi Huang

I Introduction

Dense pixel-wise image prediction is a fundamental image processing and computer vision problem and has a wide range of applications. In image processing, dense pixel-wise image prediction enables smoothing an image while preserving the edges , enhancing the details of an image , transferring the style from a reference image , dehazing the photos , and retouching the images for global tonal adjustment . In computer vision, pixel-wise image prediction not only addresses the problem of segmenting an image into semantic parts , but also helps to estimate depth from a single image , and detect the most salient object in an image .

Recent methods usually employ Fully Convolutional Networks (FCNs) for these applications, achieving state-of-the-art performance. However, FCNs usually have a huge computational complexity and memory usage on high-resolution input images, which limits the deployment of pixel-wise image prediction algorithms in real-world applications. To accelerate FCNs, we present a general framework by following a coarse-to-fine fashion, which firstly downsamples the input image, executes the algorithm at low resolution, and then upsamples the result back to the original resolution. The main challenge is restoring the low-resolution output to the original resolution with rich details and sharp edges.

This challenge can be formulated as joint upsampling, which aims at generating a high-resolution output given the corresponding low-resolution one and a high-resolution guidance map. However, existing building blocks of FCNs have limited capability to handle such a problem. To enhance the ability of FCNs for joint upsampling, we propose to reformulate the widely used guided filter into a fully differentiable building block, which can be (1) jointly trained with FCNs, (2) adapted for different tasks by learnable parameters, and (3) directly supervised by high-resolution ground truth.

To this end, we propose a novel building block for FCNs named guided filtering layer. Concretely, the original guided filter is expressed as a computational graph consisting of dilated convolutions and pointwise convolutions with learnable parameters, which can adaptively evolve for different tasks. A trainable transformation function is introduced into the proposed layer, which can generate a task-specific guidance map. As a result, all the parameters of a guided filtering layer can be learned in a data-driven manner through end-to-end training. Moreover, such a layer can be easily integrated with a pre-defined FCN without extra efforts. By equipping FCNs with guided filtering layer, we present a general framework for pixel-wise image prediction tasks named Deep Guided Filtering Network (DGF), which can largely reduce the computational complexity and memory usage. The proposed framework can be widely employed for many image processing and computer vision tasks, as shown in Fig 1. Experiments show that DGF achieves the state-of-the-art performance in quality, speed, and memory usage.

In summary, the main contribution of this paper is that (1) we develop an end-to-end trainable guided filtering layer with learnable parameters and a trainable guidance map, which enhances the ability of FCNs for joint upsampling; (2) by combining with FCNs, the proposed layer significantly improves the state-of-the-art results in multiple image processing tasks, and runs 10-100×10\text{-}100\times faster than the alternatives; and (3) additional experiments show that our approach generalizes well to many computer vision tasks and achieves significant improvements over baseline methods.

An early version of this paper appeared in IEEE Conference on Computer Vision and Pattern Recognition (2018), to which we have made substantial extensions. The improvements are shown below: (1) formulate the original guided filter into a series of spatially varying linear transformation matrices without any learnable parameters. In this paper, we reformulate the original guided filter into a block of dilated convolutions and pointwise convolutions with learnable parameters. Such a formulation enables guided filtering layer to fit a specific task through end-to-end training. (2) Based on the improved guided filtering layer, we further boosted the performance of DGF in five image processing tasks. (3) We conduct a systematic ablation study on five image processing tasks to analysis the influence of each hyper-parameter in DGF. (4) We demonstrate the upper bound of the proposed layer’s ability in joint upsampling through a comprehensive experiment. (5) Both the training code and testing code are released for reproducing the experimental results in this paper and supporting further research as well as other applications.

II Related Work

The most related works to our method are along the direction of joint upsampling. Many algorithms have been developed to tackle this problem. Joint bilateral upsampling applies a bilateral filter to the high-resolution guidance map, resulting in a piecewise-smoothing high-resolution output. The underlying bilateral filter usually requires a large amount of computation resources. Thus, many methods are presented to reduce the computation complexity. Build on joint bilateral upsampling, Barron et al. present a new form of bilateral-space optimization that efficiently solves a regularized least-squares optimization problem to produce an output that is bilateral-smooth and close to the input. Gharbi et al. first compute a description of the transformation from a highly compressed input to output. Then a high-fidelity approximation of the output can be constructed by applying the recipe to the high-quality input. Similarly, bilateral guided upsampling fits an image operator with a grid of local affine models on the low-resolution input/output pair firstly. The high-resolution output is then generated by applying the local affine model to the high-resolution input image. This method serves as a post-processing operation, while our approach can be jointly trained with the entire FCN. Deep bilateral learning integrate bilateral filter with FCNs, which can be jointly learned through end-to-end training. However, this method requires producing affine coefficients before obtaining outputs, which lacks direct supervision from the ground truth. For computer vision tasks, the number of affine coefficients is usually very large, which becomes the bottleneck of performance and speed. Besides bilateral filter, guided filter is also widely used in joint upsampling, which derived from a local linear model and computes the filtering output by considering the content of a guidance image. Compared to it, our method is formulated as a fully differentiable building block with learnable parameters, which can be jointly trained with FCNs and adaptively adjusted according to a specific task. Similarly, Yuan et al. employ a locally-affine model to relate patches from low-resolution RAW images to high-resolution JPEG images.

The above methods are based on edge-preserving local filters. Differently, other methods produce high-resolution outputs by optimizing manually designed objective functions involving all or many pixels. The objective functions typically consist of data terms and regularization terms like total variation (TV) , weighted least squares (WLS) , and scale map scheme . Following these methods, Shen et al. propose mutual-structure to reserve the structural information that is contained in both images. Similarly, Ham et al. formulate the issue as a non-convex optimization problem, which is solved by the majorization-minimization algorithm. Compared to our method, the main drawbacks of these methods are (1) they rely on hand-designed objective functions, and (2) they are usually time-consuming.

II-B Deep Learning based Image Filter

Recently, deep learning based methods are proposed in image processing tasks, which largely advanced the state-of-the-art performance. Such tasks include image denoising , image demosaicking , image deblurring , image matting , rain drop removal , image dehazing , and image colorization .

The above methods mainly focus on solving one specific image processing task. Differently, some other works aim at approximating a general class of operators. Xu et al. employ deep neural networks to approximate a variety of edge-preserving filters with a gradient-domain training procedure, while Liu et al. combine a convolutional network and a set of recurrent networks to approximate various image filters.

Xu et al. and Liu et al. deploy neural networks to generate high-resolution output directly, accelerating the operation by dedicatedly designed network architectures. Similarly, Chen et al. propose context aggregation networks to accelerate a wide variety of image processing operators, which performs superior to the prior works , achieving the best results regarding speed and accuracy. Our approach is complementary to this method, which can deliver comparable or better results and runs 10-100×10\text{-}100\times faster.

Compared to all the related works, the proposed guided filtering layer can be end-to-end trained with the entire network and generalize well across different tasks ranging from image processing to computer vision, while achieving the state-of-the-art performance in both quality and speed.

III Guided Filtering Layer

Given a high-resolution image IhI_{h} and the corresponding low-resolution output OlO_{l}, joint upsampling aims at generating a high-resolution output OhO_{h} that is visually similar to OlO_{l} and preserves the edges and details from IhI_{h}. In the literature of joint upsampling, guided filter is one of the most widely used algorithms that has shown better performance regarding the trade-off between speed and accuracy.

III-B Guided Filter Revisited

To address joint upsampling, guided filter takes a low-resolution image IlI_{l}, the corresponding high-resolution image IhI_{h}, and a low-resolution output OlO_{l} as inputs, producing the high-resolution output OhO_{h}. Concretely, AlA_{l} and blb_{l} are firstly obtained by minimizing a reconstruction error between O^l\hat{O}_{l} and OlO_{l}, where O^l\hat{O}_{l} subjects to a local linear model:

ωk\omega_{k} is the kk-th local square window on IlI_{l}, and IliI_{l}^{i} is the ii-th pixel inside ωk\omega_{k}. AhA_{h} and bhb_{h} are then produced by upsampling AlA_{l} and blb_{l}. The high-resolution output OhO_{h} is finally generated by a linear transformation model:

where ∗* is element-wise multiplication.

III-C Fully Differentiable Guided Filter

The original guided filter can only be employed as a post-processing operation, which is not differentiable and cannot be end-to-end trained with FCNs. To enhance the ability of FCNs for joint upsampling, we propose a novel building block by reformulating guided filter into a fully differentiable layer. Such a layer, named guided filtering layer, can be jointly trained with FCNs from scratch, and directly supervised by high-resolution targets.

The computation graph of guided filtering layer is shown in Figure 2. AlA_{l} and blb_{l} are obtained by employing mean filter fμf_{\mu} and local linear model to IlI_{l} and OlO_{l}, where fμf_{\mu} is implemented as a box filter to reduce the computation complexity. AhA_{h} and bhb_{h} are then generated by bilinear upsampling f↑f_{\uparrow}. OhO_{h} is finally produced by a linear layer taking AhA_{h}, bhb_{h} and IhI_{h} as inputs. rr is the radius of fμf_{\mu} and ϵ\epsilon is the regularization term, which are set to be 11 and 1e-81e\text{-}8 by default.

The equations for propagating the gradients through guided filtering layer are shown in Algorithm 1. By formulating each operator into a differentiable function, the gradient of OhO_{h} back-propagates to OlO_{l}, IlI_{l}, and IhI_{h} through the computation graph, which enables both the joint training of FCNs and guided filtering layer with direct guidance from the high-resolution targets. As a result, FCNs can learn to generate a more suitable OlO_{l} for guided filtering layer to restore OhO_{h}.

III-D Learn to Generate Task-Specific Guidance Map

In Section III-C, IhI_{h}, IlI_{l} and OhO_{h}, OlO_{l} are assumed to have the same number of channels. When the channel sizes are different, a transformation function is required to transform IhI_{h} and IlI_{l} into a guidance map with the same number of channels as OhO_{h} and OlO_{l}. Even when the channel sizes are the same, a guidance map better than IhI_{h} and IlI_{l} is necessary for higher performance. Existing methods usually manually design the transformation function for different tasks, requiring lots of efforts and attempts. On the contrary, since the proposed guided filtering layer is fully differentiable, we can automatically learn a transformation function to generate more suitable, task-specific guidance maps by end-to-end training.

As shown in Figure 2, the transformation function F(I)F(I) transforms IhI_{h} and IlI_{l} into task-specific guidance maps GhG_{h} and GlG_{l}. F(I)F(I) is a FCN block composed of two convolution layers, between which are an adaptive normalization layer and a leaky ReLU layer. The kernel size of both convolution layers is set to be 1×11\times 1, and the channel size of the first convolution layer is set to be 16 by default.

III-E Convolutional Guided Filtering Layer

Except F(I)F(I), the proposed guided filtering layer is a parameter-free block, which behaves in the same manner for all different tasks. However, due to the huge differences between tasks, a single guided filtering layer without learnable parameters cannot perform well in all kinds of scenarios. To solve the problem, we introduce learnable parameters into guided filtering layer by replacing the non-parametric operations into convolution layers. As a result, the improved layer, convolutional guided filtering layer, becomes more powerful for processing various applications, which can adaptively fit a specific task through end-to-end training.

The architecture of convolutional guided filtering layer is shown in Figure 4. Compared to that in Figure 2, dilated convolutions are introduced to replace mean filter fμf_{\mu}, and a convolution block composed of pointwise convolutions takes the place of a local linear model. As for the hyper-parameters in Section III-C, ϵ\epsilon is removed, and rr represents the dilation rates in the dilated convolutions.

IV Deep Guided Filtering Network

Based on the proposed guided filtering layer, we present a general framework for pixel-wise image prediction tasks, named Deep Guided Filtering Network (DGF). By integrating the proposed layer with FCNs following a coarse-to-fine manner, DGF can generate high-resolution, edge-preserving outputs with a much lower computational cost and memory usage.

The architecture of DGF is shown in Figure 3. First, we downsample the original input image IhI_{h} to obtain the low-resolution input IlI_{l}. Then, a FCN Cl(Il)C_{l}(I_{l}) is applied to IlI_{l}, generating the corresponding low-resolution output OlO_{l}. Finally, the high-resolution output OhO_{h} is generated by guided filtering layer, taking IlI_{l}, IhI_{h} and OlO_{l} as inputs. The entire network is end-to-end trainable, which could be learned from scratch.

DGF is a general framework for pixel-wise image prediction tasks, which can remarkably reduce the computational complexity and memory usage of the underlying algorithms. Concretely, given a specific pixel-wise image prediction task, an FCN C(I)C(I) can be designed to achieve excellent performance without considering the speed and memory cost. In order to obtain significant optimization on speed and memory, we can simply drop C(I)C(I) into the proposed framework DGF to serve as Cl(Il)C_{l}(I_{l}) without any other modifications. Since C(I)C(I) processes the input images in low resolution rather than the original resolution, the speed, and memory usage can be largely improved. Moreover, the performance of our system is also comparable to the previous state-of-the-art one, thanks to the proposed guided filtering layer. This is due to that the proposed guided filtering layer significantly enhances the capabilities of FCNs in the task of joint upsampling.

IV-B Guided Filtering Layer

In this paper, there are four variants of DGF in total according to different configurations of the guided filtering layer.

DGFs\text{DGF}_{s}: The original guided filter is employed as post-processing operation without any training. Cl(Il)C_{l}(I_{l}) is trained with low-resolution input/output pairs before inserted into DGFs\text{DGF}_{s}.

DGFb\text{DGF}_{b}: Guided filtering layer in Figure 2 is employed in DGFb\text{DGF}_{b}. F(I)F(I) is an identity function when the inputs and outputs have the same number of channels. When the channel sizes are different, F(I)F(I) transforms the inputs into a grey image by averaging along the channel axis. Cl(Il)C_{l}(I_{l}) and guided filtering layer are jointly trained from scratch under the supervision directly from the high-resolution targets.

DGFbc\text{DGF}_{b}^{c}: Compared to DGFb\text{DGF}_{b}, the guided filtering layer is replaced by convolutional guided filtering layer in Figure 4.

DGFc\text{DGF}^{c}: Compared to DGFbc\text{DGF}_{b}^{c}, F(I)F(I) proposed in Section III-D is introduced, which can learn to generate task-oriented guidance maps without manually design. As a result, DGFc\text{DGF}^{c} is not only end-to-end trainable but can also fit different tasks better by adjusting the trainable convolution weights and the learnable F(I)F(I).

IV-C Objective Function

DGF is trained end-to-end under the supervision directly from the high-resolution targets. Concretely, given the high-resolution output OhO_{h} and the corresponding target ThT_{h}, the objective function is defined as L(Oh,Th)L(O_{h},T_{h}). The concrete formulation varies with different tasks. Usually, the objective function for training C(I)C(I) can be directly employed to train DGF without any adjustment.

V Experiments: Image Processing Tasks

To show the effectiveness of our method, we employ DGF to clone five widely-used image processing operators. Concretely, the ground truth images are first generated by applying L0L_{0} smoothing operator , detail manipulation operator , style transfer operator , non-local dehazing operator , and image retouching operator to the input images. Then, the input/ground-truth pairs are used to train DGF in a supervised way to clone the corresponding image processing operator.

L0L_{0} smoothing is effective for sharpening major edges while eliminating minor edges by the use of L0L_{0} gradient minimization. To generate the ground truth images, we use the official implementation with the default parametershttp://www.cse.cuhk.edu.hk/~leojia/projects/L0smoothing.

V-A2 Detail Manipulation

Multi-scale detail manipulation enhances an image by boosting features at multiple scales. Concretely, a three-level decomposition (coarse base level bb and two detail levels d1d^{1}, d2d^{2}) of the CIELAB lightness channel is first constructed given the input image. The resulting image is then obtained by a non-linear combination of bb, d1d^{1} and d2d^{2}. To generate the ground truth images, we first generate coarse-scale, medium-scale, and fine-scale images with the official implementation and the default parametershttp://www.cs.huji.ac.il/~danix/epd. The final output is then yielded by averaging the three images.

V-A3 Style Transfer

Photographic style transfer aims at transferring the photographic style of a reference image to the input image. To generate the ground truth images, we employ the official implementation with the default setting and the default reference imagehttp://www.di.ens.fr/~aubry/code/matlab_fast_llf_and_style_transfer.zip. The generated outputs are grey images, which are transformed into RGB images as the ground truth.

V-A4 Non-local Dehazing

Non-local dehazing employs a non-local prior to remove the effects of atmospheric absorption and scattering in the input image. We use the official implementation with default parameters to generate ground truth imageshttps://github.com/danaberman/non-local-dehazing.

V-A5 Image Retouching

Image retouching aims at automatically improving the aesthetic quality of the input image by global tonal adjustment. Human experts are employed to generate the ground truth.

V-B Details of DGF

We employ Context Aggregation Network (CAN) as Cl(Il)C_{l}(I_{l}) for all the five image processing operators. The detailed architectures of Cl(Il)C_{l}(I_{l}) and F(I)F(I) are shown in Table I. AdaptNorm represents adaptive normalization proposed by Chen et al. . Leaky ReLU is employed as the nonlinearity, of which the negative slope is set to 0.20.2. As for the objective function, we use L2L_{2} loss by following the convention of previous works .

V-C Experimental Setup

Our experiments are taken on MIT-Adobe FiveK Dataset , which contains 2,500/2,500 high-resolution photographs as the training/testing images. In the dataset, each photograph contains five annotations from five experts, which can be used as the ground truth for image retouching. Instead of all five annotations, we employ the annotation from expert A as the ground truth. As for the other four image processing operators, the ground truth images are generated following the instructions in Section V-A.

As for training, we first train the network for 150 epochs, with the input/target images resized to 512sxxs means the short side of an image is resized to xx without changing the aspect ratio.. To improve the generalization ability, we further train the network for 30 epochs, with training data randomly resized to a specific resolution between 512s and 1672s. As for IlI_{l}, the spatial resolution is 64s regardless of the resolution of IhI_{h}. Adam is employed as the optimizer, with learning rate set to 0.00010.0001 and batch size set to 11.

Our primary baseline is Deep Bilateral Learning (DBL) , which shares a similar architecture to ours and achieves a good trade-off between quality and speed. Another strong baseline is CAN , which achieves state-of-the-art performance while runs reasonably fast. To ensure a fair comparison, we train the models using the official implementations and training procedures for both methods.

V-D Experimental Results

The running time and memory usage are shown in Figure 5, which are measured on a workstation with Intel E5-2650 2.20GHz CPU and Nvidia Titan X (Pascal) GPU.

On GPU devices, both DGFb\text{DGF}_{\text{b}} and DGFbc\text{DGF}_{\text{b}}^{\text{c}} take less than 10ms to process an image with resolution ranging from 5122512^{2} to 307223072^{2}. DGFc\text{DGF}^{\text{c}} is slightly slower because of the usage of F(I)F(I), but it still runs in real-time on images with resolution 307223072^{2}. All the three variants of our method run much faster than CAN and DBL among all resolutions. Specifically, DGFb\text{DGF}_{\text{b}}, DGFbc\text{DGF}_{\text{b}}^{\text{c}}, and DGFc\text{DGF}^{\text{c}} take 6ms, 6ms, and 21ms respectively for an image in 204822048^{2}. CAN takes 160ms for an image in 204822048^{2}, which are more than 25×25\times, 25×25\times, and 7×7\times slower than our method. DBL takes 51ms in the same setting, which is slightly faster than CAN but more than 8×8\times slower than DGFb\text{DGF}_{\text{b}} and DGFbc\text{DGF}_{\text{b}}^{\text{c}}. The advantage of our method in speed is even more significant as the resolution grows.

For IhI_{h} with h×w×nIh\times w\times n_{I} and OhO_{h} with h×w×nOh\times w\times n_{O}, the theoretical computational complexities of DGFb\text{DGF}_{\text{b}}, DGFbc\text{DGF}_{\text{b}}^{\text{c}}, DGFc\text{DGF}^{\text{c}}, and DBL are O(nO×h×w)\mathcal{O}(n_{O}\times h\times w), O(nO×h×w)\mathcal{O}(n_{O}\times h\times w), O((nI+nO)×h×w)\mathcal{O}((n_{I}+n_{O})\times h\times w) and O(nI×nO×h×w)\mathcal{O}(n_{I}\times n_{O}\times h\times w) respectively.

As for memory usage, our method takes less GPU memory space than both baseline methods. CAN is the most memory inefficient method that takes nearly 10G GPU memory to process an image with resolution 204822048^{2}. DGFc\text{DGF}^{\text{c}} takes a similar amount of memory space to that of DBL but grows slower as the resolution increases. DGFb\text{DGF}_{\text{b}} and DGFbc\text{DGF}_{\text{b}}^{\text{c}} are the most memory efficient methods, which take less than 1G memory even on images with resolution 307223072^{2}.

V-D2 Quantitative and Qualitative Comparison

The performance of each method is evaluated on the test set of MIT-Adobe FiveK dataset with input/target images resized to 1024s. MSE, PSNR, and SSIM serve as the evaluation metrics.

As shown in Table II, our method achieves the state-of-the-art performance in style transfer, non-local dehazing, and image retouching; while obtaining comparable results in L0L_{0} smoothing and multi-scale detail manipulation. Concretely, DGFc\text{DGF}^{\text{c}} achieves 26.17 dB in PSNR for style transfer, which improves over CAN and DBL by 4.86 dB and 2.85 dB respectively. Compared to DBL, our method outperforms it across all the five tasks in all three metrics by a large margin.

The qualitative results are shown in Figure 8More qualitative results are shown in http://wuhuikai.me/DeepGuidedFilterProject/#visual..

V-D3 The Role of Guided Filtering Layer

To show the effect of convolutional guided filtering layer and F(I)F(I), we replace OlO_{l} with low-resolution ground truth to generate OhO_{h}. The obtained result represents the performance upper bound of each DGF variant. As shown in Table III, by reformulating guided filtering layer into learnable convolution layers, DGFbc\text{DGF}_{\text{b}}^{\text{c}} outperforms DGFb\text{DGF}_{\text{b}} in all five tasks. By further introducing F(I)F(I) into the convolutional guided filtering layer, DGFc\text{DGF}^{\text{c}} achieves the best performance.

Similar results can be observed in Table II. By jointly end-to-end training, DGFb\text{DGF}_{\text{b}} achieves better performance on most tasks than DGFs\text{DGF}_{\text{s}}. Concretely, DGFb\text{DGF}_{\text{b}} improves 1 dB and 0.83 dB (PSNR) for non-local dehazing and detail manipulation. By reformulating into convolutional guided filtering layer, the performance is further improved by comparing DGFb\text{DGF}_{\text{b}} and DGFbc\text{DGF}_{\text{b}}^{\text{c}}. By adding learnable F(I)F(I), we gain significant improvements in several tasks, especially in tasks that are resolution-dependent. Table II shows that DGFc\text{DGF}^{\text{c}} increases PSNR by 2.56 dB and 1.62 dB compared to DGFbc\text{DGF}_{\text{b}}^{\text{c}} for style transfer and detail manipulation.

DJF is the state-of-the-art method for joint upsampling. To verify the effectiveness of our method, we replace guided filtering layer in DGF with DJF. Results in Table II show that our method outperforms DJF in all tasks. Besides, our method also runs much faster than DJF, which takes 9×9\times less time than DJF on images with resolution 102421024^{2} (5ms v.s. 46ms).

V-D4 Cross Resolution Generalization

In the main experiment, our method is evaluated on 1024s images. To show the generalization ability of DGF for processing images in different resolutions, the pre-trained DGF is directed employed on images in 512s, 1024s, 1536s, and 2048s without finetuning. As shown in Figure 6, our method performs equally well across different resolutions on all tasks except style transfer. The reason is that style transfer is highly resolution-dependent. Concretely, given a reference image with a fixed resolution, the styles of the outputs are different for input images with different resolutions.

V-D5 Ablation Study

A series of experiments are taken in this section to validate the effect of each hyper-parameter in the proposed guided filtering layer.

The role of radius rr is shown in Figure 7a. The performance drops quickly as rr grows, and the default setting (r=1r=1) obtains the best PSNR score.

The effect of the resolution of IlI_{l} is shown in Figure 7b. For L0L_{0} smoothing, multi-scale detail manipulation, and non-local dehazing, the performance grows as the resolution of IlI_{l} increases. For style transfer and image retouching, higher resolution is not always better. The corresponding running time and memory usage are shown in Table IV. When the resolution of IlI_{l} is 128 or 256, our method can not only achieve an excellent performance but also run very fast.

The function of F(I)F(I) is also explored by varying the dilation rate. Figure 7c shows that increasing the dilation rate can improve the performance to a degree.

VI Experiments: Computer Vision Tasks

The proposed guided filtering layer can dramatically advance the performance of multiple image processing tasks in accuracy, speed, and memory usage. Moreover, our method can also be employed to replace the time-consuming conditional random field (CRF) in many computer vision applications. To evaluate the effectiveness of our method, we take an experiment on three computer vision tasks ranging from low-level vision to high level-vision, namely depth estimation , saliency object detection , and semantic segmentation .

Depth estimation is proposed by Saxena et al. , which aims at predicting the depth at each pixel of an image with monocular cues. For this task, KITTI is the most widely used dataset, which contains 42,382 rectified stereo pairs from 61 scenes. In this paper, 29,000/1,159 images from the official training set are used for training and evaluation, which covers 33 scenes. The remaining 28 scenes of the official training set contain 200 high-quality disparity images, which are used for testing in this paper.

VI-A2 Saliency Object Detection

Saliency object detection is used to detect the most salient object in an image, which is formulated as an image segmentation problem . MSRA-B and the official training/validation/test split is used in our experiment.

VI-A3 Semantic Segmentation

Semantic segmentation aims at assigning each pixel of an image to one of the pre-defined labels . To evaluate our method, PASCAL VOC 2012 benchmark is used in this paper, which involves 20 foreground object classes and one background class. The original dataset contains 1,464, 1,449, and 1,456 pixel-wise labeled images for training, validation, and testing, respectively. The training set is further augmented by extra annotations , resulting in 10,582 images. We use the 10,582 augmented images for training and the 1,449 validation images for testing.

VI-B Details of DGF

When applying DGF to computer vision tasks, the high-resolution input image is directly processed by Cl(Il)C_{l}(I_{l}) without downsampling, generating the low-resolution output OlO_{l}. As for the architecture of Cl(Il)C_{l}(I_{l}), MonoDepthhttps://github.com/mrharicot/monodepth , DSShttps://github.com/wuhuikai/DeepGuidedFilter , DeepLab-V2https://github.com/isht7/pytorch-deeplab-resnet are employed for depth estimation, saliency detection, and semantic segmentation respectively. The corresponding training and testing procedures and loss functions are also used to train our network. As for the hyper-parameters of guided filtering layer, rr and ϵ\epsilon are determined by grid search on the validation set, as shown in Table V. Notably, a second guided filtering layer is applied in the saliency detection task to achieve better performance.

VI-C Main Results

The performances of our method and baseline methods are shown in Table VI. For depth estimation, DGFs\text{DGF}_{\text{s}} obtain 0.177 improvements in rms over the baseline. By end-to-end training and adding the learnable guidance map, we achieve the best performance (5.887) in rms. Similar results are obtained in saliency detection and semantic segmentation. FβF_{\beta} increases from 90.61% to 91.29% by applying the guided filtering layer to saliency detection. By replacing DGFs\text{DGF}_{\text{s}} with DGF, FβF_{\beta} further improves to 91.75%. For segmentation, DGF obtains 73.58% in mean IOU, which has an improvement of 1.79% compared to the baseline method.

We also compare our method with DenseCRF , which is commonly used in saliency detection and semantic segmentation. Experiments show that our method is comparable to DenseCRF in saliency detection, and obtains better performance in semantic segmentation. Besides, the proposed layer performs at least 10×10\times faster than DenseCRF. Averagely, our approach takes 3434ms to process a 5122512^{2} image, while DenseCRF takes 432432ms.

Figure 9 shows the visual results of our method and baselines. The results obtained by our approach are better in preserving edges and detailsMore qualitative results are shown in http://wuhuikai.me/DeepGuidedFilterProject/#visual..

VII Conclusion

We present a novel building block for FCN, namely guided filtering layer, which aims at enhancing the ability of FCNs for joint upsampling. By formulating the guided filter into a fully differentiable module with learnable convolutional kernels, FCN-based pixel-wise image prediction approaches can benefit from end-to-end training and generate high-quality results. We further extend the proposed layer with a learnable transformation function, which makes it generalize well to different tasks by producing task-specific guidance maps. We integrate the guided filtering layer with FCNs and evaluate it on five image processing tasks and three computer vision tasks. Experiments show that the proposed layer could achieve state-of-the-art performance while taking 10-100×10\text{-}100\times less computational cost. We also conduct a comprehensive ablation study, which demonstrates the contribution of each component as well as the hyper-parameters.

Acknowledgement

This work is funded by the National Key Research and Development Program of China (Grant 2016YFB1001004 and Grant 2016YFB1001005), the National Natural Science Foundation of China (Grant 61673375, Grant 61721004 and Grant 61403383) and the Projects of Chinese Academy of Sciences (Grant QYZDB-SSW-JSC006 and Grant 173211KYS-B20160008). The authors would like to thank Patrick Pérez and Philip Torr for their helpful suggestions.

References