EdgeStereo: A Context Integrated Residual Pyramid Network for Stereo Matching
Xiao Song, Xu Zhao, Hanwen Hu, Liangji Fang
Introduction
Stereo matching is a fundamental problem in computer vision. It has a wide range of applications, such as robotics and autonomous driving . Given a rectified image pair, the main goal is to find corresponding pixels from stereo images. Most traditional stereo algorithms follow the classical four-step pipeline , including matching cost computation, cost aggregation, disparity calculation and disparity refinement. However the hand-crafted features and multi-step regularized functions limit their improvements.
Since , CNN based stereo methods extract deep features to represent image patches and compute matching cost. Although the performance on several benchmarks is significantly promoted, there remains some difficulties, including the limited receptive fields and complicated regularized functions.
Recently end-to-end disparity estimation networks achieve state-of-the-art performance, however drawbacks still exist. Firstly, it is difficult to handle local ambiguities in ill-posed regions. Secondly, the cascade structures or 3D convolution based structures are computationally expensive. Lastly, disparity predictions of thin structures or near boundaries are not accurate.
Humans, on the other hand, can find stereo correspondences easily by utilizing edge cues. Accurate edge contours can help discriminating between different objects or regions. In addition, humans perform binocular alignment well in texture-less or occluded regions based on global perception at different scales.
Based on these observations, we design a multi-task network EdgeStereo that cooperates edge cues and edge regularization into disparity prediction pipeline. Firstly, we design a disparity network for EdgeStereo, called context pyramid based residual pyramid network (CP-RPN). Two modules are designed for CP-RPN: a context pyramid to encode multi-scale context information for ill-posed regions, and an one-stage residual pyramid to simplify the cascaded refinement structure. Secondly an edge detection sub-network is designed and employed in our unified model, to preserve subtle details with edge cues. Interactions between two tasks are threefold: (i) Edge features are embedded into disparity branch providing local and low-level representations. (ii) The edge map, acting as an implicit regularization term, is fed to residual pyramid. (iii) The edge map is also utilized in edge-aware smoothness loss, which further guides disparity learning.
In disparity branch of EdgeStereo, we use a siamese network with a correlation operation to extract image features and compute matching cost volumes, followed by a context pyramid. Based on different representations (unary features, edge features and matching cost volumes), context pyramid can encode contextual cues in multiple scales. Then they are aggregated as a hierarchical scene prior. Next we employ an hour-glass structure to regress the full-size disparity map. Different with the decoder in DispNetC or the cascade encoder-decoder in CRL, decoder in EdgeStereo is replaced by the proposed residual pyramid. We predict the disparity map on the smallest scale and learn disparity residuals on other scales. Hence learning and refining are conducted in a single decoder, making CP-RPN as an one-stage disparity estimation model. Based on experimental results in Table 1, our residual pyramid is better and faster than other cascade structures. In edge branch of EdgeStereo, the shallow part of backbone network is shared with CP-RPN. Edge feature and edge map are embedded into the disparity branch, under the guidance of edge-aware smoothness loss.
Both edge and disparity branches are fully-convolutional so that end-to-end training can be conducted for EdgeStereo. As there is no dataset providing both ground-truth disparities and edge labels, we propose a multi-phase training strategy. We adopt the supervised disparity regression loss and our adapted edge-aware smoothness loss to train the entire EdgeStereo, achieving a high accuracy on Scene Flow dataset . We further finetune our model on KITTI 2012 and 2015 datasets, achieving state-of-the-art performance on KITTI stereo benchmarks. After multi-task learning, both disparity estimation and edge detection tasks are improved in both quantitatively and qualitatively.
In summary, our main contribution is threefold.
(i) We propose EdgeStereo to support the joint learning of scene matching and edge detection, where edge cues and edge-aware smoothness loss serve as important guidance for disparity learning. The multi-task labels are not required during training due to our proposed multi-phase training strategy.
(ii) The effective context pyramid is designed to handle ill-posed regions, and the efficient residual pyramid is designd to replace cascade refinement structures.
(iii) Our unified model EdgeStereo achieves state-of-the-art performance on Scene Flow dataset, KITTI stereo 2012 and 2015 benchmarks.
Related Work
Among non-end-to-end deep stereo algorithms, each step in traditional stereo pipeline could be replaced by a network. For example, Luo et al. train a simple multi-label classification network for matching cost computation. Shaked and Wolf introduce an initial disparity prediction network pooling global information from cost volume. Gidaris et al. substitute hand-crafted disparity refinement functions with a three-stage refinement network.
For end-to-end deep stereo algorithms, all steps in traditional stereo pipeline are combined for joint optimization. To train end-to-end stereo networks, Mayer et al. create a large synthetic stereo dataset, meanwhile they also propose a baseline model called DispNet with an encoder-decoder structure. Based on DispNet, Pang et al. cascade a residual learning network for further refinement. Different from DispNet, Kendall et al. propose GC-Net that incorporates contextual information by means of 3D convolutions over a feature volume. Based on GC-Net, Yu et al. add an explicit cost aggregation structure. An unsupervised method is proposed in . Liang et al. formulate the disparity refinement task as Bayesian inference process for joint learning. PSMNet utilizes spatial pyramid pooling and 3D CNN to regularize cost volumes. Our CP-RPN is also an end-to-end network, but we explicitly encode context cues for disparity learning and our one-stage residual pyramid is efficient.
0.2 Combining Stereo Matching with Other Tasks.
Bleyer et al. first solve stereo and object segmentation problems together. Guney and Geiger propose Displets which utilizes foreground object recognition to help stereo matching. More tasks are fused through a slanted plane in . However, these hand-crafted multi-task methods are not robust.
0.3 Edge Detection.
To preserve details in disparity maps, we resort to edge detection task to supplement features and regularization. Inspired by FCN , Xie et al. first design an end-to-end edge detection network named holistically-nested edge detector (HED) based on VGG- network. Recently, Liu et al. modify the structure of HED, combining richer convolutional features from VGG backbone. These fully-convolutional edge networks can be easily incorporated with disparity estimation networks.
0.4 Deep Learning Based Multi-task Structure.
Cheng et al. propose an end-to-end network called SegFlow, which enables the joint learning of video object segmentation and optical flow. The segmentation branch and flow branch are iteratively trained offline. Our EdgeStereo is a different multi-task structure where multi-phase training is conducted rather than iterative training. Hence disparity branch can exploit more stable boundary information from pretrained edge branch. In addition, EdgeStereo does not require multi-task labels from a single dataset, hence it is easier to find proper datasets for training.
Approach
In this section, we describe our multi-task model EdgeStereo. We first present the basic network structure. Then we introduce two critical modules in disparity branch: context pyramid and residual pyramid. Next we detail the cooperation strategies of edge cues, including edge feature embedding, edge map feeding and the adapted edge-aware smoothness loss. Finally we show how to conduct multi-task learning via our multi-phase training strategy.
The overall architecture of our EdgeStereo is shown in Fig. 2. To combine two tasks efficiently, edge branch shares the shallow computation with disparity branch at the backbone network.
2 Context Pyramid
Context information is widely used in many tasks . For stereo matching, it can be regarded as the relationship between an object and its surroundings or its sub-regions, which can help inferring correspondences especially for ill-posed regions. Many stereo methods learn these relationships by stacking lots of convolution blocks. Differently we encode context cues explicitly through context pyramid, hence learning stereo geometry of the scene is easier. Moreover, single-scale context information is insufficient because objects with arbitrary sizes are existed. Over-focusing on global information may neglect small-size objects, while disparities of big stuff might be inconsistent or discontinuous if the receptive field is small. Hence the proposed context pyramid aims at capturing multi-scale context cues in an efficient way.
We use four parallel branches with similar structures in the context pyramid. As mentioned in , the size of receptive field roughly indicates how much we use context. Hence four branches own different receptive fields to capture context information at different scales. The largest context scale corresponds to the biggest receptive field. To our knowledge, convolution, pooling and dilation operations can enlarge the receptive field. Hence we design convolution context pyramid, pooling context pyramid and dilation context pyramid respectively. They are detailed in Section 4.1. The best one is embedded in EdgeStereo.
3 Residual Pyramid
Many stereo methods use a cascade structure for disparity estimation, where the first network generates initial disparity predictions and the second network produces residual signals to rectify initial disparities. However these residual signals are hard to learn (residuals are always close to zero), because initial disparity predictions are pretty good. Moreover these multi-stage structures are computationally expensive. In order to optimize the cascade structure, we design a residual pyramid so that initial disparity learning and disparity refining can be conducted in a single network.
To make multi-scale disparity estimation easier, we refer to the idea “From Easy to Tough” from curriculum learning . In other words, it is easier to regress disparity map on the smallest scale because searching range is narrow and few details is needed. To get larger disparity maps, we estimate residual signals relative to the disparity map at the smallest scale. The formulation of residual pyramid makes EdgeStereo an effective one-stage structure. Besides, the residual pyramid can be beneficial for overall training because it alleviates the problem of over-fitting.
Scale number in residual pyramid is consistent with the encoder structure. As shown in Fig. 3, the smallest scale in residual pyramid produces a disparity map ( of the full resolution), then it is continuously upsampled and refined with the residual map on a larger scale, until the full-size disparity map is obtained. The formulation is shown in Eq. (1), where denotes upsampling by a factor of and denotes the pyramid scale (e.g. represents the full-resolution).
For each scale, various information are aggregated to predict disparity or residual map, including the skip-connected feature map from encoder with higher frequency information, the edge feature and edge map (all interpolated to corresponding scale) to cooperate edge cues, and the geometrical constraints. For each scale except the smallest scale, we warp the resized right image according to disparity and obtain a synthesized left image . The error map is a heuristic cue which can help to learn residuals. Hence the concatenation of , , , and serves as the geometrical constraints.
4 Cooperation of Edge Cues
Basic disparity estimation network CP-RPN works well on ordinary and texture-less regions, where matching cues are clear or context cues can be easily captured through the context pyramid. However as shown in the second row of Fig. 1, details in disparity map are lost, due to too many convolution and down-sampling operations. Hence we utilize edge cues to help refining disparity maps.
Secondly we resize and feed the edge map to each scale in residual pyramid. The edge map acts as an implicit regularization term which can help smoothing disparities in non-edge regions and preserving edges in disparity map. Hence the edge sub-network does not behave like a black-box.
Finally we regularize the edge map into an edge-aware smoothness loss, which is an effective guidance for disparity estimation. For disparity smoothness loss , we encourage disparities to be locally smooth and the loss term penalizes depth changes in non-edge regions. To allow for depth discontinuities at object contours, previous methods weight this regularization term according to image gradients. Differently, we weight this term based on gradients of edge map, which is more semantically meaningful than intensity variation. As shown in Eq. (2), denotes the number of pixels, denotes disparity gradients and denotes gradients of edge probability map.
5 Multi-phase Training Strategy and Objective Function
In order to conduct multi-task learning for EdgeStereo, we propose a multi-phase training strategy where the training phase is split into three phases. Weights of the backbone network are fixed in all three phases.
In the first phase, edge sub-network is trained on a dataset for edge detection task, guided by a class-balanced cross-entropy loss proposed in .
In the second phase, we supervise the regressed disparities across scales on a stereo dataset. Deep supervision is adopted, forming the total loss as the sum where denotes the loss at scale . Besides the disparity smoothness loss, we adopt the disparity regression loss for supervised learning, as shown in Eq. (3).
where denotes the ground truth disparity map. Hence the overall loss at scale becomes , where is a loss weight for smoothness loss. In addition, weights of the edge sub-network are fixed.
In the third phase, all layers in EdgeStereo are optimized on the same stereo dataset used in the second phase. Similarly we also adopt the deep supervision across scales. However the edge-aware smoothness loss is not used in this phase, because edge contours in the second phase are more stable than those in the third phase. Hence the loss at scale is .
Experiments
Experiment settings and results are presented in this section. Firstly we evaluate key components of EdgeStereo on Scene Flow dataset. We also compare our approach with other state-of-the-art stereo matching methods on KITTI benchmarks. In addition, we demonstrate that better edge maps can be obtained after multi-task learning.
The encoder contains convolution layers with occasional strides of , resulting in a total down-sampling factor of . Correspondingly there are output scales in residual pyramid. For each scale, the estimation block consists of four convolution layers and the last convolution layer regresses disparity or residual map. Each convolution layer except the output is followed by a ReLU layer.
We modify the structure of HED and propose an edge sub-network called HEDβ, where low-level edge features are easier to obtain and the produced edge map is more semantic-meaningful. HEDβ uses the VGG-16 backbone from conv1_1 to conv5_3. In addition, we design five side branches from conv1_2, conv2_2, conv3_3, conv4_3 and conv5_3 respectively. Each side branch consists of two convolution layers, an upsampling layer and an convolution layer producing the edge probability map. In the end, feature maps from each upsampling layer in each side branch are concatenated as the final edge feature, meanwhile edge probability maps in each side branch are fused as the final edge map. The final edge feature and edge map are of full size.
Finally we describe the structure of each context pyramid.
Convolution context pyramid. Each branch consists of two convolution layers with a same kernel size. Kernel size for the largest context scale is biggest. For example, and for each branch, denoted as -.
Dilation context pyramid. Inspired by , each branch consists of a dilated convolution layer, followed by an convolution layer to reduce dimensions. Dilation rate for the largest context scale is biggest. For example, and for each branch respectively, denoted as -.
2 Datasets and Evaluation Metrics
Scene Flow dataset is a synthesised dataset containing training and test image pairs. Dense ground-truth disparities are provided and we perform the same screening operation as CRL . The real-world KITTI dataset includes two subsets with sparse ground-truth disparities. KITTI 2012 contains training and test image pairs while KITTI 2015 consists of training and test image pairs.
To pretrain the edge sub-network, we adopt the BSDS500 dataset containing training and test images. Consistent with , we combine the training data in BSDS500 with PASCAL VOC Context dataset .
To evaluate the stereo matching results, we apply the end-point-error (EPE) which measures the average Euclidean distance between estimated and ground-truth disparity. We also use the percentage of bad pixels whose disparity errors are greater than a threshold (), denoted as -pixel error.
3 Implementation Details
Our model is implemented based on Caffe . The model is optimized using the Adam method with , . In the first training phase, HEDβ is trained on BSDS500 dataset for iterations. The batch size is and the initial learning rate is which is divided by at the -th and -th iterations. The second and third training phases are all conducted on Scene Flow dataset with a batch size of . In the second phase, we train for iterations with a fixed learning rate of . The loss weight for edge-aware smoothness loss is set to . Afterwards in the third phase, we train for iterations with a learning rate of which is halved at the -th and -th iterations. When finetuning on KITTI datasets, the initial learning rate is set to which is halved at the -th and -th iterations. Since ground-truth disparities provided by the KITTI datasets are sparse, invalid pixels are neglected in .
4 Ablation Studies
In this section, we conduct several ablation studies on Scene Flow dataset to evaluate key components in the EdgeStereo model. The one-stage DispFulNet (a simple variant of DispNetC ) serves as the baseline model in our experiments. All results are shown in Table 1.
4.2 Context Pyramid.
Firstly we choose a context pyramid (-), then train a model consisting of the hybrid feature extraction part, the selected context pyramid and the encoder-decoder of DispFulNet. Compared with the model without context pyramid, -pixel error is reduced from to . Furthermore, as shown in the “Context Pyramid Comparisons” part in Table 1, adopting other context pyramids can also lower the -pixel error. Hence we argue that multi-scale context cues are beneficial for dense disparity estimation task.
4.3 Encoder-Decoder (Residual Pyramid).
We use the same encoder as DispFulNet and we adopt residual pyramid as the decoder. To prove its effectiveness, we train a model consisting of the hybrid feature extraction part and our encoder-decoder. Compared with the model containing the encoder-decoder in DispFulNet, -pixel error is reduced from to . Hence our multi-scale residual learning mechanism is superior to direct disparity regression.
Finally we train CP-RPN consisting of the hybrid feature extraction part, context pyramid - and our encoder-decoder. The -pixel error is and the EPE is , outperforming the baseline model by .
4.4 Context Pyramid Comparisons.
We train different CP-RPN models with different context pyramids. As shown in Table 1, convolution context pyramids don’t work well, reducing the -pixel error by only , and respectively. In addition, the large dilation rate is harmful for extracting context cues. The -pixel error of - is while for -. - has the best performance, achieving a -pixel error of . Hence pooling context pyramid - is embedded in the final model.
4.5 Comparisons with Multi-stage Refinement.
Firstly we compare with the three-stage refinement structure DRR . We replace our encoder-decoder with encoder-decoder in DispFulNet, then three additional networks are cascaded for refinement. CP-RPN outperforms this model by meanwhile being times faster. Next we compare with the two-stage cascade structure CRL . We replace our encoder-decoder with disparity prediction and disparity refinement networks in . As can be seen, performance is almost equal but our model is faster with less parameters, which proves the effectiveness of residual pyramid.
4.6 Benefits from Edge Cues.
We conduct several experiments where different disparity networks are cooperated with our edge sub-network. As can be seen, all stereo matching models are improved. We also present visual demonstrations as shown in Fig. 4. When edge cues are cooperated into the disparity estimation pipeline, subtle details are preserved hence the error rate is reduced.
5 Comparisons with Other Stereo Methods
In this section, we compare EdgeStereo with state-of-the-art stereo matching methods on Scene Flow dataset as well as KITTI 2012 and 2015 benchmarks.
Firstly we compare with several non-end-to-end methods, including SGM , SPS-St , MC-CNN-fst and DRR . We also compare with the most advanced end-to-end stereo networks, including DispNetC , DispFulNet , CRL , GC-Net and CA-Net . The comparisons are presented in Table 2, EdgeStereo achieves the best performance in terms of two evaluation metrics. As shown in Fig. 4, disparities predicted by EdgeStereo are very accurate, especially in thin structures and near boundaries.
5.2 KITTI Results.
For KITTI 2012, EdgeStereo is finetuned on all training image pairs, then test results are submitted to KITTI stereo 2012 benchmark. For evaluation, we use the percentage of erroneous pixels in non-occluded (Noc) and all (All) regions. We also conduct comparisons in challenging reflective (Refl) regions such as car windows. The results are shown in Table 3. By leveraging context and edge cues, our EdgeStereo model is able to handle challenging scenarios with large occlusion, texture-less regions and thin structures.
For KITTI 2015, we also finetune EdgeStereo on the whole training set. The test results are also submitted. For evaluation, we use the -pixel error of background (D1-bg), foreground (D1-fg) and all pixels (D1-all) in non-occluded and all regions. The results are shown in Table 4. EdgeStereo achieves state-of-the-art performance on KITTI 2015 benchmark and our one-stage structure is faster than most stereo models. Fig. 7 gives qualitative results on KITTI test sets. As can be seen, EdgeStereo produces high-quality disparity maps in terms of global scene and object details. We also provide visual demonstrations of “stereo benefits from edge” on KITTI datasets, as shown in Fig. 5.
6 Better Edge Map
We can’t evaluate on a stereo dataset whether the edge detection task is improved or not after multi-task learning, because the ground-truth edge map is not provided. Hence we first give visual demonstrations on Scene Flow dataset, as shown in Fig. 6. EdgeStereo produces edge maps with finer details, compared with HEDβ without multi-task learning. We argue that the learned geometrical knowledge from disparity branch can help highlighting image boundaries.
For quantitative demonstrations, we conduct further experiments on BSDS500 dataset. ODS F-measure (higher is better) is for original HEDβ, for HEDβ after multi-task learning and for the baseline model HED. All models are finetuned on BSDS500 dataset for same epochs.
Conclusion
In this paper, we present a multi-task architecture EdgeStereo where edge cues are incorporated into the disparity estimation pipeline. Also the proposed context pyramid and residual pyramid enable our unified model to handle challenging scenarios with an effective one-stage structure. Our method achieves state-of-the-art performance on Scene Flow dataset and KITTI stereo benchmarks, demonstrating the effectiveness of our design.