Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency

Zhenheng Yang, Peng Wang, Wei Xu, Liang Zhao, Ramakant Nevatia

Introduction

Depth and normal estimation has been explored in multiple

Related Work

2 Unsupervised learning for low-level vision

3 Spatial transformer network

Method

In this section, we describe the framework architecture and training procedure in detail. Our intuition is to train a CNN that is capable of modeling the geometry consistency of a mostly rigid scene. To facilitate the learning of the network, we explicitly propose to model the constraint between depth and normal. The training samples of the framework consist of frame sequences captured by a monocular moving camera.

To model a reasonable geometrical consistency, we propose the overall objective function as in Equation 1.

This objective function is a Lagrange fuction aiming to minimize the loss term Lwarp(D,I,Rt)+Lsmooth(D,N,I)+Lgrad(D,I,Rt)L_{warp}(D,I,Rt)+L_{smooth}(D,N,I)\\ +L_{grad}(D,I,Rt) subject to the constraint of geometrical constraint between depth map and normal map Ldn(D,N)=0L_{d}n(D,N)=0. The loss term consists of three components: photometric warping loss Lwarp(D,I,Rt)L_{warp}(D,I,Rt), smoothness loss Lsmooth(D,N,I)L_{smooth}(D,N,I), image gradient matching loss Lgrad(D,I,Rt)L_{grad}(D,I,Rt).

Photometric warping loss. One main supervision of our framework comes from novel view synthesis: given an input view of a scene and camera motion, synthesize an image of the scene seen from a different view. We can synthesize an image of the target view given the image of source view, camera motion from target view to source view and depth map of target view, using 3D inverse warping. The process of warping is shown in Figure 1.

Each pixel (point on the grid) of depth map is first mapped onto 3D space. The 3D point cloud is transformed based on camera motion and then mapped back to 2D plane. Each grid point in target view corresponds to one point in source view. Similar to (?), the bilinear interpolation is implemented to calculate each pixel value of warped image as a weighted sum of four nearest neighboring pixels in source image, weighted by the square area between the projected point and neighboring point, as shown in Figure 1 (b).

The warping loss is a photometric difference between the target image and warped image.

In which, ss iterates the number of source image, pp iterates each pixel in the image, ItI_{t} is the target image, IsI_{s} is the source image, Is^=τ(Is,Dt,Rt)\hat{I_{s}}=\tau(I_{s},D_{t},Rt) is the warped image, τ\tau is the warping function as introduced above.

Edge-aware smoothness loss. One issue with using only view synthesis as supervision is that the back-propagation gradients are solely derived from the pixel value difference between one point in target image and weighted sum of its four neighboring points in source image. The warping loss will not be useful for learning where the point falls on low-texture regions. The predicted depth on these regions can be of any value as long as the warped region has the similar pixel value. To overcome this issue, a prior knowledge of the scene geometry is incorporated for a smoothness loss:

This smoothness term penalizes the norm of second-order gradients of depth in order to encourage smoothly changing depth values. As depth discontinuity often happens at image gradients, the smoothness loss is weighted by the a function of image gradients to prevent smoothed depth at image gradients.

To further facilitate the macthing of target image and warped image, and to encourage the depth map to be sharp, the photometric difference of gradient maps of target image and warped image is cacluated as gradient matching loss.

2 Geometry consistency

As depth and surface normal are not independent under the same scene, thus we model the 3D geometry consistency by explicitly incorporating the relationship of depth and normal into the training procedure and use the relationship as a regularization in the objective function. The regularization term Ldn(D,N)=0L_{dn}(D,N)=0 is realized by two layers in our framework: depth2normal layer and normal2depth layer.

Depth2normal layer. The normal direction of each point is computed based on the neighboring points after projecting to 3D space. The process of calculating normal direction of point pp is shown in Figure 2. θ(p)\theta(p) is a set of neighboring (8) points of pp. Take point q∈θ(p)q\in\theta(p) for example. RqR_{q} is a set of points that satisfy the requirement: when projecting to 3D space, for q^∈θ^(p)\hat{q}\in\hat{\theta}(p) and for r^∈R^q\hat{r}\in\hat{R}_{q}, ((^q)−(^p)⋅(r^−(^p))≠0)(\hat{(}q)-\hat{(}p)\cdot(\hat{r}-\hat{(}p))\neq 0). Symbols with hat represent corresponding points in 3D space. Theoretically, the cross-product of any two non-collinear (in 3D space) vectors connecting p^\hat{p} and θ^(p)\hat{\theta}(p) is the normal direction N(p)N(p). To reduce the possiblity that the two vectors being collinear in 3D space, we require the vectors to be perpendicular when projected in 2D plane. The normal directions are averaged when iterating q∈θ(p)q\in\theta(p), and then l2l_{2} normalized to make it a unit vector. The normal direction is calculated as:

for each q∈θ(p)q\in\theta(p), RqR_{q} should satisfy (q−p)⋅(r−p)=0,r∈Rq(q-p)\cdot(r-p)=0,\quad r\in R_{q}.

Normal2depth layer. Normal2depth layer takes depth map and normal map as input and outputs a ”shifted“ depth map.

Evaluation

2 Ablation study

Image edge in depth2normal and normal2depth layers

3 Comparison with other methods

Conclusion