Depth Adaptive Deep Neural Network for Semantic Segmentation
Byeongkeun Kang, Yeejin Lee, Truong Q. Nguyen
I Introduction
Depth perception, which is one of the crucial abilities in the human visual system, allows human to perceive the distance to an object and to understand the world in three dimensions. The human visual system uses the perceived depth information to robustly estimate the size and shape of objects in three dimensions. The three-dimensional information helps to better understand the objects and scenes along with other cues such as color information. Thus, depth information plays a key role in understanding the visual world.
As depth information is crucial for understanding the visual world, many researches have been explored ways to acquire accurate depth information efficiently in both hardware systems and software systems. In hardware-based solutions, advanced depth sensors, such as Microsoft Kinect and light detection and ranging (LiDAR) sensors have been developed to capture better quality depth information with portability and low cost . In software-based solutions, disparity estimation algorithms using single or multiple cameras have been studied to estimate accurate depth cues in shorter processing time . Owing to these successes in both communities, depth information has been widely usable in many computer vision applications such as human pose estimation , indoor scene understanding , and autonomous driving .
After perceiving depth and/or color information, a machine processes the perceived information to understand the visual world. One of the recent popular frameworks for learning visual information is the deep neural network, which is loosely inspired by the neurons of a human brain. As computing capability of machines has increased drastically, deep neural networks have attained a huge improvement in understanding visual information and shown the state-of-the-art performance in many tasks such as image classification , object detection , and semantic segmentation .
Because of the importance of depth information and the improvements by using deep neural networks, it has been speculated that incorporating depth information with neural networks has the advantage in understanding visual information. In most researches on deep neural networks using depth information, the depth map has been treated as an image equivalent input to the networks . In such networks, the neurons share the predetermined receptive fields in a convolutional layer, which hinders the networks from learning common representations of an object. Considering that a pinhole camera captures an object at different distances, the camera captures the same object in different sizes, as demonstrated in Fig. 1. The illustration implies that a neural network can possibly learn/extract different features for the same object at various distances, yielding the confusions of recognizing objects. Hence, we propose the novel deep neural networks that learn common features of the same object by leveraging depth information (Section III-C). The proposed neural networks perceive the same region of the object regardless of the distance from the camera to each pixel as described in Fig. 2. This is achieved by the novel Depth-adaptive Multiscale or DaM convolution layer consisting of the adaptive perception neuron and the in-layer multiscale neuron in Section III-B. The adaptive perception neuron adjusts the size of the receptive field at each spatial location corresponding to the distance from the camera. The adjustment requires a coefficient to decide the ideal correlation between the size of the receptive field and the distance. Since the optimal coefficient differs depending on the objects, better performance can be achieved by utilizing multiple coefficients in a layer. This is implemented by the proposed in-layer multiscale neuron. The in-layer multiscale neuron learns/extracts diversely scaled representations in a layer by applying a different size of the receptive field at each feature representation. The adjustment of the receptive field is applied using the sparse convolution (dilated convolution) as demonstrated in Fig. 3. In Section IV, we verify the effectiveness of the proposed method on two tasks: indoor semantic segmentation and hand segmentation for hand-object interaction. We use publicly available NYUDv2 dataset for indoor semantic segmentation and collect a new challenging dataset including hand-object interaction for hand segmentation.
In summary, the contributions of our work are as follows:
We develop the depth-adaptive neural networks using the DaM convolution. The DaM convolution consists of the adaptive perception neurons and the in-layer multiscale neurons.
We propose the adaptive perception neuron. The neuron learns/extracts depth-adaptive representations.
We propose the in-layer multiscale neuron. The neuron learns/extracts variously scaled representations in a convolution layer.
We verify the effectiveness of the proposed networks on the task of semantic segmentation.
II Related Works
Deep neural networks using depth map. Most researches of deep neural networks using depth maps treated a raw depth map as an image equivalent. For instance, a raw depth map was given as a direct input to the networks in hand pose estimation , human pose estimation , and fingerspelling recognition .
Alternatively, Gupta et al. proposed the geocentric embedding for a depth map to learn better representations in convolutional neural networks . Specifically, the geocentric embedding encodes horizontal disparity, height above ground, and angle with gravity (HHA) for each pixel. The work showed that using the HHA encoded images, convolutional neural networks can learn better features for object detection and segmentation.
The networks we introduce are distinct from the works in . First, our proposed method utilizes depth information in convolution layers rather than converting a raw depth map into a better representation in a preprocessing stage. In other words, our method does not require any additional preprocessing to manipulate the raw data. Second, our proposed method can take any type of input (e.g. color image, depth map, etc.) to learn feature representations by giving the corresponding depth information as shown in Fig. 4.
Semantic segmentation. Long et al. proposed fully convolutional neural networks (FCN) for semantic segmentation by converting fully connected layers to convolution layers in the neural networks for image classification . The networks take an input of arbitrary size and produce the correspondingly-sized output.
Additional efforts have been made to improve the performance in . Zheng et al. proposed the convolutional neural networks that incorporate the strength of conditional random field (CRF)-based probabilistic graphical modeling. They formulated CRF as recurrent neural networks (RNN) and attached the RNN after FCN . Chen et al. improved semantic segmentation using convolution with up-sampled filters, atrous spatial pyramid pooling, and fully connected CRF . Yu et al. proposed an additional context module to aggregate multiscale information without losing resolution .
Unlike other methods, our approach improves the performance of neural networks using depth information without adding additional layers. In addition, the proposed networks can incorporate any aforementioned additional layers for further improvement.
Hand segmentation for hand-object interaction. Most algorithms for hand segmentation segment hands using skin color in color images. Oikonomidis et al. and Romero et al. segmented hands by thresholding skin color in the hue-saturation-value (HSV) color space . Wang et al. used the learned probabilistic model constructed from the color histogram of the first frame . The histogram was generated using super-Gaussian mixture model in . Tzionas et al. processed segmentation of hands using the Gaussian mixture model constructed for skin color .
However, skin color-based hand segmentation is sensitive to skin pigment difference and light condition variation. Similarly, in the segmentation, hands can be confused with other objects in skin color and other body parts (e.g. arm, face, etc.). To overcome these limitations, we decided to segment hands using only depth maps. For the experiment, we collected a new dataset with pixel-wise annotations because we were not able to find a publicly available dataset for hand-object interaction with annotations.
Convolution layer. Conventionally, most convolutional neural networks used typical, dense, and fixed convolution . Recently, dilated (atrous) convolution was applied for semantic segmentation to extract sparse features in higher resolution . The structure excluded pooling layers (which cause the decrement of spatial resolution) and replaced typical convolutions following the pooling layers by dilated convolutions. The dilated convolution was employed to increase the size of receptive fields and compensate the exclusion of pooling layers . The dilated convolution in these methods has different sparsity at each layer depending on the excluded pooling layers while it has the same sparsity for all spatial locations and all feature representations in a layer.
Contemporarily, active convolution and deformable convolution are presented in . The goal of both methods is to learn the shape of convolutions using a training dataset. Active convolution defines the learnable position parameters to represent various forms of the receptive fields in the task of image classification . The position parameters are shared across all kernels in a layer. Thus, the learned receptive field is the same at all spatial locations and for all feature representations. Deformable convolution uses the offset field similar to the position parameters . The offset field is computed using the input feature map and has different receptive fields at each spatial location. This deformable convolution was tested on semantic segmentation task and object detection task.
In the proposed networks, we apply dilated (sparse) convolution to adjust the size of receptive fields for two purposes. First, we adapt the sparsities in convolutions to learn/extract near depth-invariant representations using distance information. Thus, the sparsity is adjusted at each spatial location depending on the distance at the location. Second, we adapt the sparsity at each feature representation to learn variously scaled representations. That is, the proposed dilated convolution generates different sparsities at each spatial location and at each feature representation.
III Proposed Method
The goal of this work is to learn depth-invariant representations in deep neural networks using depth information. To achieve this goal, we propose the novel DaM convolution layer conceiving the adaptive perception neurons and the in-layer multiscale neurons as described in Fig. 4. The adaptive perception neuron is proposed to adjust the receptive field using the depth information at each spatial location. The in-layer multiscale neuron is designed to learn features in different scales at each feature space (or channel) in a layer.
In Section III-A, we introduce key notations for networks. We provide the detailed explanation of the DaM convolution consisting of the adaptive perception neuron and the in-layer multiscale neuron in Section III-B. The overall architecture of the proposed neural network is developed in Section III-C. In Section III-D, the details of the training procedure are derived for the proposed networks. Finally, we provide the mathematical proof of the depth invariant property of the proposed networks in Section III-E.
III-B Depth-adaptive Multiscale Convolution Layer
As observed in Fig. 1, an object appears to have different sizes in the image plane depending on its distance from the camera. The generalization performance of the trained networks using these depth-variant features may not be sufficiently good because learning a common representation is challenging from the features. As such, it is necessary to learn depth-invariant features for neural networks in order to achieve better generalization performance. To this end, we propose the DaM convolution layer containing the adaptive perception neurons and the in-layer multiscale neurons. The adaptive perception neuron in Section III-B1 adjusts its receptive field to offset the change of the spatial size of objects on captured images. The receptive field adjusted by the adaptive perception neuron is clearly sub-optimal because the ideal correlation between the size and the distance varies over objects (e.g. due to different sizes). Hence, we develop the in-layer multiscale neuron in Section III-B2 that effectively controls the size of receptive fields over individual objects. The in-layer multiscale neuron extracts the diversely scaled depth-invariant features by tuning a parameter that determines sparsity at each feature representation.
Given a depth map as an input of the networks, unlike color images, the intensity (value) of an object on the depth map is scaled by the distance from the camera. This implies that the networks may learn intensity-variant features for the same object. To avoid this misguiding, we propose to employ depth difference (relative depth) as an input for the feature extraction in Section III-B3.
The proposed adaptive perception neuron determines its size of receptive field based on the depth information at each spatial location while other methods used the predetermined receptive field in a convolution layer. Thus, the proposed networks having such adaptive perception neurons can apply different receptive field at each spatial location. Specifically, we increase the receptive field for objects at a close distance and decrease it for objects at a long distance to compensate for the variation of objects’ size on the captured images.
III-B2 In-layer multiscale neuron
Conventionally, learning/extracting features in various scales is advantageous in achieving higher segmentation accuracy by learning variant features. To learn features in multiple scales, the neural networks comprised of multiple neural networks were proposed in , known as the multiscale neural networks. In this type of neural networks, each constituting neural network takes an input in different resolution and learns features in various scales. However, these networks are structurally complex and require higher computational complexity. Thus, we propose the in-layer multiscale neuron that takes only an input and learns features with multiple scales in a network (see Fig. 6). The proposed in-layer multiscale neuron learns features at various scales by having a different parameter for the sparsity at each feature representation (channel).
where denominator is contributed by the adaptive perception neuron, and numerator is from the in-layer multiscale neuron.
III-B3 Depth difference
In practice, values on a depth map vary as the distance from the camera changes. For instance, objects at different distances are represented by different intensity levels. However, the relative distance between these objects is constant regardless of their distance from the camera . Consequently, we instead use the relative depth to measure distance-independent depth in the first convolution layer. The relative depth is computed as the difference between the depth at the receptive field and the depth at the center location of the receptive field. Replacing a depth by the relative depth, (3) is rewritten as
Although the input to the networks is replaced by the relative depth , the size of the receptive fields is computed using the raw depth map .
III-C Architecture
The proposed DaM convolution layer is applied to all convolution layers in two fully convolutional neural networks (Frontend module and DeepLab ). All original convolution layers are replaced by the proposed layers to achieve depth-invariance as demonstrated in Section III-E. Frontend module and DeepLab are selected as our baseline model since they are two of the state-of-the-art methods. For DeepLab, we employed the VGG-16 network-based architecture with large atrous spatial pyramid pooling (ASPP-L) and without conditional random field (CRF) .
III-D Back Propagation
and this is rewritten by the chain rule , as follows:
Recalling (3), since an output node has the input nodes determined by the depth-adaptive receptive field, is required to decode the connections from input nodes to output nodes (see Fig. 3). Considering this variation of receptive field, the second factor of the multinomial logistic loss is evaluated as
where and denote the momentum and the learning rate, respectively. The momentum was chosen as for Frontend module and for DeepLab, and the learning rate is explained in Section IV.
III-E Proof of Depth-Invariance
In this section, we present the mathematical proof of the depth-invariance property of the proposed networks. We first simplify the convolution in (3) by considering a single channel one-dimensional input and output. We, then, apply the proposed convolution to an input at different distances from the camera. By demonstrating that the outputs are equivalent regardless of the distances, we prove that the proposed DaM convolution is depth-invariant.
Considering a neural network having a single channel (feature space), (3) is substituted as follows:
For the one-dimensional input, (15) is further simplified as
Let’s first consider the example in Fig. 7, showing the proposed convolution layers for the input at distance in Fig. 7(a) and at distance in Fig. 7(b). In the example, the size of kernel is set to , and the size of receptive field is 1 at distance . Then, the output of the first convolution layer in Fig. 7(a) is
and the output of the second convolution layer is
and the output of the second convolution layer is
We conclude from this simple example that the proposed convolution extracts depth-invariant activations.
where is the ratio of distances between and . From the example of (19), (20) and the generalization of (21), we conclude that the proposed convolution extracts depth-invariant activations by adjusting the size of receptive field.
IV Experiments and Results
The proposed neural networks were tested on two applications: indoor semantic segmentation and hand segmentation for hand-object interaction. The experimental results verify that the proposed neural networks outperform original Frontend module and DeepLab without any additional layer or pre/post-processing.
For comparison, we report pixel-wise accuracy, mean accuracy, mean intersection over union (IoU), and frequency weighted (FW) IoU for both experiments. Additionally, for hand segmentation, we report precision, recall, and score. Let be the number of pixels which belong to the class and are predicted to the class , and be the total number of classes.
where for hand segmentation, class is hand, and class is others.
The NYUDv2 dataset consists of 1,449 pairs of RGB-D images including various indoor scenes with pixel-wise annotations . The pixel-wise annotations were coalesced into 40 dominant object categories by Gupta et al. . We experimented with this 40 classes problem with the standard separation of 795 training images and 654 testing images.
IV-A2 Experiments
All the models were initialized using the VGG-16 model trained using the ImageNet ILSVRC-2014 dataset except for the input of RGB-HHA. Then, the models were fine-tuned using the NYUDv2 training dataset . For the input of RGB-HHA, we initialized the model using the two fine-tuned models using NYUDv2 dataset (one model using RGB images and the other model using HHA images). Then, we fine-tuned the model using the pair of RGB images and HHA images similar to . The initial base learning rate was selected by trying several learning rates () with a factor of 10 such as . The decay factor () of the weight matrix is chosen as . The models used in the experiments were selected based on the mean IoU score. During training, we computed the mean IoU score at every 1,000 iterations for the input of RGB-HHA and at every 2,000 iterations for the other inputs.
For DeepLab, the initial base learning rate was selected as for HHA images and for other inputs. The learning rate was decreased using polynomial decay with the power of and the maximum iteration of . The scaling parameters for all layers and for all inputs were set to be linearly distributed in .
IV-A3 Results
We adopted the experimental settings in . We considered the inputs of an RGB image, the concatenated image of an RGB image and a depth map (early fusion), and an HHA encoded image . We also experimented with combining the scores from an RGB image and from an HHA encoded image at the last layer (late fusion). Table I and Fig. 8 show the quantitative results and the qualitative results. The proposed method achieves the improvements without any additional layers or pre/post-processing.
IV-A4 Analysis
We experimentally analyze the effects of multiscale parameters in Table II. The analysis shows that the proposed method outperforms other methods using the parameters in the reasonable ranges. We also analyze the effects of applying the different number of the DaM convolution in Table III. The experiments demonstrate that replacing all convolution layers outperforms other settings. The processing time is measured using a machine with Intel i7-4790K CPU and Nvidia Tesla K40c. Table IV shows that multi/random scale evaluation has the chance of further improving the segmentation performance. In the multiscale evaluation, the final results are combined with the results of original, twice enlarged, and half-scaled inputs. In the random scale evaluation, the final results are fused from the results of original and two randomly scaled inputs. The results using both scaling are combined with the results of the previously mentioned five inputs. Table V demonstrates that simply increasing receptive fields in DeepLab does not improve the accuracy. Lastly, we show the convergence curve for Frontend module and the network with the proposed DaM convolution in Fig. 9. The average loss is computed using the losses from 100 iterations. The graph shows that the proposed method converges slightly faster than Frontend module.
IV-B Hand-Object Interaction (HOI)
We collected a new datasethttps://github.com/byeongkeun-kang/HOI-dataset using Microsoft Kinect v2 since we were not able to find a publicly available dataset for hand-object interaction with pixel-wise annotation. The collected dataset consists of more than 9,175 pairs of depth maps and color images from 6 people (3 males and 3 females) interacting with 21 different objects. In addition, the dataset includes the cases of one hand and both hands in a scene. Ground truth was labeled by wearing a color glove during data collection and by finding the color of the glove on the color images.
To increase the variation of the dataset further (e.g. the distance from the camera to hands), 18,350 pairs of images were augmented by moving the camera closer/further to/from the scene as shown in Fig. 10. In total, the augmented dataset has pairs of depth maps and ground truth labels. Indeed, the standard deviation of the augmented data increases to relative to that of the collected dataset is , as evidenced in Fig. 11(a). The distances of the augmented dataset are distributed at more diverse distances as demonstrated in Fig. 11(b).
Among pairs, we used 19,470 pairs for training, 2,706 pairs for validation, and 5,349 pairs for testing.
IV-B2 Experiments
All the models were initialized using the VGG-16 model that were trained using the ImageNet ILSVRC-2014 dataset . Then, the initial models were fine-tuned using the HOI training dataset. The initial base learning rate was selected by trying several learning rates with the factor of 10 such as . In most cases, the initial learning rate was selected as . The decay factor of the weight matrix is chosen as .
The models used in the experiments were selected based on the score on the validation dataset. During training, we computed the score on the validation dataset at every 4,000 iterations. If the score stops improving, the base learning rate was decreased by a factor of 10. The training was terminated if the improvement of the score is negligible () or the score is not improved. The multiscale parameter was set to and for other layers was set to for each quarter of the feature spaces in each convolution layer.
IV-B3 Results
The performance of the proposed methods and the comparing methods is tabulated in Table VI for the inputs of the depth maps and the HHA encoded images . The visual segmentation results are displayed in Fig. 12. The proposed neural network improves about (depth maps) and (HHA) in score relative to the baseline Frontend model . Moreover, the proposed network with the input of depth map achieves higher score and mean IoU than Frontend module with the input of the HHA encoded image. These results verify that the proposed networks improve segmentation performance without any additional layer or pre/post-processing.
V Conclusion
In this paper, we presented the novel fully convolutional neural networks that adjust the receptive field using depth information to learn/extract depth-invariant feature representations. In the proposed neural networks, we introduced the DaM convolution layer consisting of the adaptive perception neuron and the in-layer multiscale neuron. The proposed neural networks were applied to indoor semantic segmentation and hand segmentation for hand-object interaction. The experimental results demonstrate that the proposed neural networks improve the accuracy of segmentation without any additional layers or pre/post-processing.
VI Acknowledgment
This work is supported in part by NSF grant IIS-1522125. We thank Nan Jiang for her contribution to data preprocessing and our colleagues for their participation in data collection. We also thank the anonymous reviewers for their insightful comments.