Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour Prediction
Dan Xu, Wanli Ouyang, Xavier Alameda-Pineda, Elisa Ricci, Xiaogang Wang, Nicu Sebe
Introduction
Considered as one of the fundamental tasks in low-level vision, contour detection has been deeply studied in the past decades. While early works mostly focused on low-level cues (e.g. colors, gradients, textures) and hand-crafted features , more recent methods benefit from the representational power of deep learning models . The ability to effectively exploit multi-scale feature representations is considered a crucial factor for achieving accurate predictions of contours in both traditional and CNN-based approaches. Restricting the attention on deep learning-based solutions, existing methods typically derive multi-scale representations by adopting standard CNN architectures and considering directly the feature maps associated to different inner layers. These maps are highly complementary: while the features from the first layers are responsible for predicting fine details, the ones from the higher layers are devoted to encode the basic structure of the objects. Traditionally, concatenation and weighted averaging are very popular strategies to combine multi-scale representations (see Fig. 1.a). While these strategies typically lead to an increased detection accuracy with respect to single-scale models, they severly simplify the complex relationship between multi-scale feature maps.
The motivational cornerstone of this study is the following research question: is it worth modeling and exploiting complex relationships between multiple scales of a deep representation for contour detection? In order to provide an answer and inspired by recent works exploiting graphical models within deep learning architectures , we introduce Attention-Gated Conditional Random Fields (AG-CRFs), which allow to learn robust feature map representations at each scale by exploiting the information available from other scales. This is achieved by incorporating an attention mechanism seamlessly integrated into the multi-scale learning process under the form of gates . Intuitively, the attention mechanism will further enhance the quality of the learned multi-scale representation, thus improving the overall performance of the model.
We integrated the proposed AG-CRFs into a two-level hierarchical CNN model, defining a novel Attention-guided Multi-scale Hierarchical deepNet (AMH-Net) for contour detection. The hierarchical network is able to learn richer multi-scale features than conventional CNNs, the representational power of which is further enhanced by the proposed AG-CRF model. We evaluate the effectiveness of the overall model on two publicly available datasets for the contour detection task, i.e. BSDS500 and NYU Depth v2 . The results demonstrate that our approach is able to learn rich and complementary features, thus outperforming state-of-the-art contour detection methods.
In the last few years several deep learning models have been proposed for detecting contours . Among these, some works explicitly focused on devising multi-scale CNN models in order to boost performance. For instance, the Holistically-Nested Edge Detection method employed multiple side outputs derived from the inner layers of a primary CNN and combine them for the final prediction. Liu et al. introduced a framework to learn rich deep representations by concatenating features derived from all convolutional layers of VGG16. Bertasius et al. considered skip-layer CNNs to jointly combine feature maps from multiple layers. Maninis et al. proposed Convolutional Oriented Boundaries (COB), where features from different layers are fused to compute oriented contours and region hierarchies. However, these works combine the multi-scale representations from different layers adopting concatenation and weighted averaging schemes while not considering the dependency between the features. Furthermore, these works do not focus on generating more rich and diverse representations at each CNN layer.
The combination of multi-scale representations has been also widely investigated for other pixel-level prediction tasks, such as semantic segmentation , visual saliency detection and monocular depth estimation , pedestrian detection and different deep architectures have been designed. For instance, to effectively aggregate the multi-scale information, Yu et al. introduced dilated convolutions. Yang et al. proposed DAG-CNNs where multi-scale feature outputs from different ReLU layers are combined through element-wise addition operator. However, none of these works incorporate an attention mechanism into a multi-scale structured feature learning framework.
Attention models have been successfully exploited in deep learning for various tasks such as image classification , speech recognition and image caption generation . However, to our knowledge, this work is the first to introduce an attention model for estimating contours. Furthermore, we are not aware of previous studies integrating the attention mechanism into a probabilistic (CRF) framework to control the message passing between hidden variables. We model the attention as gates , which have been used in previous deep models such as restricted Boltzman machine for unsupervised feature learning , LSTM for sequence learning and CNN for image classification . However, none of these works explore the possibility of jointly learning multi-scale deep representations and an attention model within a unified probabilistic graphical model.
Attention-Gated CRFs for Deep Structured Multi-Scale Feature Learning
Given an input image and a generic front-end CNN model with parameters , we consider a set of multi-scale feature maps . Being a generic framework, these feature maps can be the output of intermediate CNN layers or of another representation, thus is a virtual scale. The feature map at scale , can be interpreted as a set of feature vectors, , where is the number of pixels. Opposite to previous works adopting simple concatenation or weighted averaging schemes , we propose to combine the multi-scale feature maps by learning a set of latent feature maps with a novel Attention-Gated CRF model sketched in Fig.1. Intuitively, this allows a joint refinement of the features by flowing information between different scales. Moreover, since the information from one scale may or may not be relevant for the pixels at another scale, we utilise the concept of gate, previously introduced in the literature in the case of graphical models , in our CRF formulation. These gates are binary random hidden variables that permit or block the flow of information between scales at every pixel. Formally, is the gate at pixel of scale (receiver) from scale (emitter), and we also write . Precisely, when then the hidden variable is updated taking (also) into account the information from the -th layer, i.e. . As shown in the following, the joint inference of the hidden features and the gates leads to estimating the optimal features as well as the corresponding attention model, hence the name Attention-Gated CRFs.
2 Attention-Gated CRFs
The first term of the energy function is a classical unary term that relates the hidden features to the observed multi-scale CNN representations. The second term synthesizes the theoretical contribution of the present study because it conditions the effect of the pair-wise potential upon the gate hidden variable . Fig. 1c depicts the model formulated in Equ.(1). If we remove the attention gate variables, it becomes a general multi-scale CRFs as shown in Fig. 1b.
Given that formulation, and as it is typically the case in conditional random fields, we exploit the mean-field approximation in order to derive a tractable inference procedure. Under this generic form, the mean-field inference procedure writes:
where denotes the sigmoid function. This finding is specially relevant in the framework of CNN since many of the attention models are typically obtained after applying the sigmoid function to the features derived from a feed-forward network. Importantly, since the quantity depends on the expected values of the hidden features , the AG-CRF framework extends the unidirectional connection from the features to the attention model, to a bidirectional connection in which the expected value of the gate allows to refine the distribution of the hidden features as well.
3 AG-CRF Inference
In order to construct an operative model we need to define the unary and gated potentials and . In our case, the unary potential corresponds to an isotropic Gaussian:
where is a weighting factor.
The gated binary potential is specifically designed for a two-fold objective. On the one hand, we would like to learn and further exploit the relationships between hidden vectors at the same, as well as at different scales. On the other hand, we would like to exploit previous knowledge on attention models and include linear terms in the potential. Indeed, this would implicitly shape the gate variable to include a linear operator on the features. Therefore, we chose a bilinear potential:
Under these potentials, we can consequently update the mean-field inference equations to:
where is the expected a posteriori value of .
The previous expression implies that the a posteriori distribution for is a Gaussian. The mean vector of the Gaussian and the function write:
which concludes the inference procedure. Furthermore, the proposed framework can be simplified to obtain the traditional attention models. In most of the previous studies, the attention variables are computed directly from the multi-scale features instead of computing them from the hidden variables. Indeed, since many of these studies do not propose a probabilistic formulation, there are no hidden variables and the attention is computed sequentially through the scales. We can emulate the same behavior within the AG-CRF framework by modifying the gated potential as follows:
This means that we keep the pair-wise relationships between hidden variables (as in any CRF) and let the attention model be generated by a linear combination of the observed features from the CNN, as it is traditionally done. The changes in the inference procedure are straightforward and reported in the supplementary material due to space constraints. We refer to this model as partially-latent AG-CRFs (PLAG-CRFs), whereas the more general one is denoted as fully-latent AG-CRFs (FLAG-CRFs).
4 Implementation with neural network for joint learning
In order to infer the hidden variables and learn the parameters of the AG-CRFs together with those of the front-end CNN, we implement the AG-CRFs updates in neural network with several steps: (i) message passing from the -th scale to the current -th scale is performed with , where denotes the convolutional operation and denotes the corresponding convolution kernel, (ii) attention map estimation , where , and are convolution kernels and represents element-wise product operation, and (iii) attention-gated message passing from other scales and adding unary term: , where encodes the effect of the for weighting the message and can be implemented as a convolution. The symbol denotes element-wise addition. In order to simplify the overall inference procedure, and because the magnitude of the linear term of is in practice negligible compared to the quadratic term, we discard the message associated to the linear term. When the inference is complete, the final estimate is obtained by convolving all the scales.
Exploiting AG-CRFs with a Multi-scale Hierarchical Network
AMH-Net Architecture. The proposed Attention-guided Multi-scale Hierarchical Network (AMH-Net), as sketched in Figure 2, consists of a multi-scale hierarchical network (MH-Net) together with the AG-CRF model described above. The MH-Net is constructed from a front-end CNN architecture such as the widely used AlexNet , VGG and ResNet . One prominent feature of MH-Net is its ability to generate richer multi-scale representations. In order to do that, we perform distinct non-linear mappings (deconvolution , convolution and max-pooling ) upon , the CNN feature representation from an intermediate layer of the front-end CNN. This leads to a three-way representation: , and . Remarkably, while upsamples the feature map, maintains its original size and reduces it, and different kernel size is utilized for them to have different receptive fields, then naturally obtaining complementary inter- and multi-scale representations. The and are further aligned to the dimensions of the feature map by the deconvolutional operation. The hierarchy is implemented in two levels. The first level uses an AG-CRF model to fuse the three representations of each layer , thus refining the CNN features within the same scale. The second level of the hierarchy uses an AG-CRF model to fuse the information coming from multiple CNN layers. The proposed hierarchical multi-scale structure is general purpose and able to involve an arbitrary number of layers and of diverse intra-layer representations.
where , is the set of contour pixels of image and is the set of all parameters. The optimization is performed via the back-propagation algorithm with standard stochastic gradient descent.
Experiments
Datasets. To evaluate the proposed approach we employ two different benchmarks: the BSDS500 and the NYUDv2 datasets. The BSDS500 dataset is an extended dataset based on BSDS300 . It consists of 200 training, 100 validation and 200 testing images. The groundtruth pixel-level labels for each sample are derived considering multiple annotators. Following , we use all the training and validation images for learning the proposed model and perform data augmentation as described in . The NYUDv2 contains 1449 RGB-D images and it is split into three subsets, comprising 381 training, 414 validation and 654 testing images. Following in our experiments we employ images at full resolution (i.e. pixels) both in the training and in the testing phases.
Evaluation Metrics. During the test phase standard non-maximum suppression (NMS) is first applied to produce thinned contour maps. We then evaluate the detection performance of our approach according to different metrics, including the F-measure at Optimal Dataset Scale (ODS) and Optimal Image Scale (OIS) and the Average Precision (AP). The maximum tolerance allowed for correct matches of edge predictions to the ground truth is set to 0.0075 for the BSDS500 dataset, and to .011 for the NYUDv2 dataset as in previous works .
Implementation Details. The proposed AMH-Net is implemented under the deep learning framework Caffe . The implementation code is available on Githubhttps://github.com/danxuhk/AttentionGatedMulti-ScaleFeatureLearning. The training and testing phase are carried out on an Nvidia Titan X GPU with 12GB memory. The ResNet50 network pretrained on ImageNet is used to initialize the front-end CNN of AMH-Net. Due to memory constraints, our implementation only considers three scales, i.e. we generate multi-scale features from three different layers of the front-end CNN (i.e. res3d, res4f, res5c). In our CRF model we consider dependencies between all scales. Within the AG-CRFs, the kernel size for all convolutional operations is set to with stride and padding . To simplify the model optimization, the parameters are set as for all scales during training. We choose this value as it corresponds to the best performance after cross-validation in the range $$. The initial learning rate is set to 1e-7 in all our experiments, and decreases 10 times after every 10k iterations. The total number of iterations for BSDS500 and NYUD v2 is 40k and 30k, respectively. The momentum and weight decay parameters are set to 0.9 and 0.0002, as in . As the training images have different resolution, we need to set the batch size to 1, and for the sake of smooth convergence we updated the parameters only every 10 iterations.
2 Experimental Results
In this section, we present the results of our evaluation, comparing our approach with several state of the art methods. We further conduct an in-depth analysis of our method, to show the impact of different components on the detection performance.
Comparison with state of the art methods. We first consider the BSDS500 dataset and compare the performance of our approach with several traditional contour detection methods, including Felz-Hut , MeanShift , Normalized Cuts , ISCRA , gPb-ucm , SketchTokens , MCG , LEP , and more recent CNN-based methods, including DeepEdge , DeepContour , HED , CEDN , COB . We also report results of the RCF method , although they are not comparable because in an extra dataset (Pascal Context) was used during RCF training to improve the results on BSDS500. In this series of experiments we consider AMH-Net with FLAG-CRFs. The results of this comparison are shown in Table 2 and Fig. 4a. AMH-Net obtains an F-measure (ODS) of 0.798, thus outperforms all previous methods. The improvement over the second and third best approaches, i.e. COB and HED, is 0.5% and 1.0%, respectively, which is not trivial to achieve on this challenging dataset. Furthermore, when considering the OIS and AP metrics, our approach is also better, with a clear performance gap.
To perform experiments on NYUDv2, following previous works we consider three different types of input representations, i.e. RGB, HHA and RGB-HHA data. The results corresponding to the use of both RGB and HHA data (i.e. RGB+HHA) are obtained by performing a weighted average of the estimates obtained from two AMH-Net models trained separately on RGB and HHA representations. As baselines we consider gPb-ucm , OEF , the method in , SemiContour , SE , gPb+NG , SE+NG+ , HED and RCF . In this case the results are comparable to the RCF since the experimental protocol is exactly the same. All of them are reported in Table 2 and Fig. 4b. Again, our approach outperforms all previous methods. In particular, the increased performance with respect to HED and RCF confirms the benefit of the proposed multi-scale feature learning and fusion scheme. Examples of qualitative results on the BSDS500 and the NYUDv2 datasets are shown in Fig. 3. We show more examples of predictions from different multi-scale features on the BSDS500 dataset 5.
Ablation Study. To further demonstrate the effectiveness of the proposed model and analyze the impact of the different components of AMH-Net on the countour detection task, we conduct an ablation study considering the NYUDv2 dataset (RGB data). We tested the following models: (i) AMH-Net (baseline), which removes the first-level hierarchy and directly concatenates the feature maps for prediction, (ii) AMH-Net (w/o AG-CRFs), which employs the proposed multi-scale hierarchical structure but discards the AG-CRFs, (iii) AMH-Net (w/ CRFs), obtained by replacing our AG-CRFs with a multi-scale CRF model without attention gating, (iv) AMH-Net (w/o deep supervision) obtained removing intermediate loss functions in AMH-Net and (v) AMH-Net with the proposed two versions of the AG-CRFs model, i.e. PLAG-CRFs and FLAG-CRFs. The results of our comparison are shown in Table 3, where we also consider as reference traditional multi-scale deep learning models employing multi-scale representations, i.e. Hypercolumn and HED .
These results clearly show the advantages of our contributions. The ODS F-measure of AMH-Net (w/o AG-CRFs) is 1.1% higher than AMH-Net (baseline), clearly demonstrating the effectiveness of the proposed hierarchical network and confirming our intuition that exploiting more richer and diverse multi-scale representations is beneficial. Table 3 also shows that our AG-CRFs plays a fundamental role for accurate detection, as AMH-Net (w/ FLAG-CRFs) leads to an improvement of 1.9% over AMH-Net (w/o AG-CRFs) in terms of OSD. Finally, AMH-Net (w/ FLAG-CRFs) is 1.2% and 1.5% better than AMH-Net (w/ CRFs) in ODS and AP metrics respectively, confirming the effectiveness of embedding an attention mechanism in the multi-scale CRF model. AMH-Net (w/o deep supervision) decreases the overall performance of our method by 1.9% in ODS, showing the crucial importance of deep supervision for better optimization of the whole AMH-Net. Comparing the performance of the proposed two versions of the AG-CRF model, i.e. PLAG-CRFs and FLAG-CRFs, we can see that AMH-Net (FLAG-CRFs) slightly outperforms AMH-Net (PLAG-CRFs) in both ODS and OIS, while bringing a significant improvement (around 2%) in AP. Finally, considering HED and Hypercolumn , it is clear that our AMH-Net (FLAG-CRFs) is significantly better than these methods. Importantly, our approach utilizes only three scales while for HED and Hypercolumn we consider five scales. We believe that our accuracy could be further boosted by involving more scales.
Conclusions
We presented a novel multi-scale hierarchical convolutional neural network for contour detection. The proposed model introduces two main components, i.e. a hierarchical architecture for generating more rich and complementary multi-scale feature representations, and an Attention-Gated CRF model for robust feature refinement and fusion. The effectiveness of our approach is demonstrated through extensive experiments on two public available datasets and state of the art detection performance is achieved. The proposed approach addresses a general problem, i.e. how to generate rich multi-scale representations and optimally fuse them. Therefore, we believe it may be also useful for other pixel-level tasks.