CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth

Jose M. Facil, Benjamin Ummenhofer, Huizhong Zhou, Luis Montesano, Thomas Brox, Javier Civera

Introduction

Recovering 3D information from 2D images is one of the fundamental problems in computer vision that, due to recent advances and applications, is receiving nowadays a renewed attention. Among others, there has been recent relevant results on problems such as 6D object pose detection , 3D model reconstruction , depth estimation from single and multiple views , 6D camera pose recovery or camera tracking and mapping . While traditional multi-view methods (e.g., ) are mostly based on geometry and optimization and, thus, are largely independent of the data, these recent deep learning approaches depend on training data that demonstrates the mapping from images to depth.

The common strategy to collect such data is by using an RGBD sensor, like the Kinect camera, which conveniently provides both the RGB image and what can be considered ground truth depth. It is implicitly assumed that training on this type of data will generalize to other RGB sensors that do not provide depth. However, the evaluation of recent learning-based methods relies largely on public benchmarks where images have been recorded with the same RGBD camera as the training data. Thus, evaluation on these benchmarks does not reveal whether a depth estimation method generalizes to RGB images from another camera.

Overfitting to a benchmark is a common problem in computer vision research. Other works have shown that datasets may have strong biases that make researchers over-confident regarding the performance of their method. In particular, train-test divisions of the same kind of data are not enough to prove generalization. In this work we show that, indeed, state-of-the-art single-view depth prediction networks do not generalize when the camera parameters of the test images are different from the training ones.

Moreover, we show that for single-view depth prediction the problem of missing generalization to images from different cameras is even more severe: it cannot be solved by training on images from a diverse set of cameras with different parameters. For present methods to adapt to a different camera model, they require changes in the architecture.

We present a deep neural network for single-view depth prediction that, for the first time, addresses the variability on the camera’s internal parameters. We show that this allows to use images from different cameras at train and test time without a performance degradation. This is of particular interest, as it enables the exploitation of images from any camera for training the data-hungry deep networks. Specifically, within our proposed network, our main contribution is a novel type of convolution, that we name as CAM-Convs (Camera-Aware Multi-scale Convolutions), that concatenates the camera internal parameters to the feature maps, and hence allows the network to learn the dependence of the depth from these parameters. Figure 1 shows an illustration of how CAM-Convs act in the typical encoder-decoder depth estimation pipeline. The network can be trained with a mixture of images from different cameras without overfitting to specific intrinsics. We show that the network generalizes also to images from cameras it has not been trained on. A comparison with the state of the art in single-image depth estimation demonstrates that the better generalization properties do not reduce the accuracy of the depth estimates.

Related Work

Estimating 3D structure and 6 degrees-of-freedom motion using deep learning has been addressed recently from several angles: Supervised and unsupervised , from single and multiple views , using end-to-end networks and fusing with multi-view geometry , completing depth maps , and estimating geolocation , relative motion , visual odometry , and simultaneous localization and mapping (SLAM) .

In this work we deal with single-view supervised depth learning, so we will focus our literature review in this case. Among the pioneering work we can reference , that similarly to pop-up illustrations, cut and fold a 2D image based on a segmentation into geometric classes and some geometric assumptions. is another seminal work that, with minimal assumptions on the scene, learned a model based on a MRF. was the first paper that used deep learning for single-view depth prediction, proposing a multi-scale depth network. Its results were improved later by .

Many methods focus on specific datasets which enable to train learning-based methods for specific tasks. For instance, Eigen and Fergus extend the multi-scale architecture in to the prediction of surface normals and semantic labels on the NYU dataset . Similarly, Wang et al. train a network that jointly predicts depth and segmentation on the same dataset. For depth, Laina et al. , Liu et al. and Eigen et al. show that their methods can be adapted to other datasets like Make3D or KITTI . However, they treat datasets like different tasks and require retraining for each dataset to achieve state-of-the-art performance.

Chen et al. , inspired by , introduce the Depth in the Wild dataset and train a CNN using ordinal relations between point pairs. While the images stem from internet photo collections taken with many different cameras, they do not make use of the camera parameters during training. Li and Snavely use a structure from motion pipeline to extract depth from internet photo collections and use this to train a CNN predicting depth up to a scale factor. Again, information about camera parameters is not exploited and generalization is solely driven by large diverse datasets. Extrinsic parameters have been considered for other tasks such as stereo estimates or synthesis of view point changes . Intrinsic parameters are usually left out in deep learning pipelines, with the exception of He et al. . They embed focal length information in a fully-connected approach, making it impossible to train and test in different image sizes, while our proposal is flexible and can deal with different image sizes.

In the next section we describe how to explicitly implement the internal camera parameters into the network and thereby improve generalization by CAM-Convs.

Camera-Aware Multi-scale Convolutions

CAM-Convs (standing for Camera-Aware Multi-scale Convolutions), is the variant of the convolution operation that we present in this paper. CAM-Convs include the camera intrinsics in the convolutions, allowing the network to learn and predict depth patterns that depend on the camera calibration. Specifically, we add CAM-Convs in the mapping from RGB features to 3D information–e.g. depth, normals–, that is, between the encoder and the decoder. As shown in Figure 2, we add them at every level, such that we include CAM-Convs on every skip-connection too. Notice that all the CAM-Convs are added after the encoder, allowing the use of pretrained models.

The basics of CAM-Convs are as follows: We pre-compute pixel-wise coordinates and field-of-view maps and feed them along with the input features to the convolution operation. CAM-Convs use the idea behind Coord-Convs , on adding normalized coordinates per pixel, but incorporating information on the camera calibration. An illustrative scheme of how CAM-Convs extra channels work is shown in Figure 3. The different maps included are computed using the camera intrinsic parameters (focal length ff and principal point coordinates (cx,cy)(c_{x},c_{y})) and the sensor size (width ww and height hh):

Centered Coordinates (cccc): To add the information of the principal point location to the convolutions, we include ccxcc_{x} and ccycc_{y} coordinate channels centered at the principal point–i.e. the principal point has coordinates (0,0)(0,0). Specifically, the channels are

We resize these maps to the input feature size using bilinear interpolation and concatenate them as new input channels. These channels are sensitive to the sensor size and resolution (pixel size) of the camera, as their values depend on it. We assume the sensor size is measured in pixels. In Figure 3 we represent cccc with a color gradient from red (for negative coordinates) to blue (for positive coordinates), white for 0. Notice in the figure how cccc values change when camera sensor size, principal point or pixel size change.

Field of View Maps (fovfov): The horizontal and vertical fovfov maps are calculated from the cccc maps and also depend on the camera focal length ff

where chch can be xx or yy (see Eq. 1 and 2). They give information about the captured context and the focal length. These maps are sensitive to sensor size and focal length. In Figure 3 we represent fovfov with a color gradient from green to pink; yellow represents an angle of 0 in the field of view map. Notice in the bottom part of the figure how the fovfov map values change when changing camera focal length, sensor size or principal point. Changes on the pixel size change the resolution of the map but the field of view and thus the available context in the image stays the same.

Normalized Coordinates (ncnc): We also include a Coord-Conv channel of normalized coordinates . The values of Normalized Coordinates vary linearly with the image coordinates between $.Thischanneldoesnotdependonthecamerasensor.However,itisveryusefultodescribethespatialextentofthecontext(infeaturespace)thatisleftineachdirection(e.g.,ifthevalueonthe. This channel does not depend on the camera sensor. However, it is very useful to describe the spatial extent of the context (in feature space) that is left in each direction (e.g., if the value on thexchannelisclosetochannel is close to-1$, it means the feature vector at this position is close to the left border and there is almost no context on the left side).

Notice that ncnc is not shown in Figure 3 as it remains constant.

where ξ=1d\xi=\frac{1}{d} is the inverse depth map. used a similar approach to correct depth values at test time, in this paper we propose for the first time to use it during training.

This normalization can be used together with our CAM-Convs. Although CAM-Convs allows the network to learn this normalization on its own, we found in our experiments that using this normalization accelerates the convergence. It should be remarked, though, that focal length normalization assumes a constant pixel size over the whole image set, and therefore can only be used in such cases. CAM-Convs are a more general model that overcomes this limitation.

Model and Training

The network we use in this work has an encoder-decoder architecture inspired by DispNet . Hence, we add skip-connections from the low-level feature maps of the encoder to the feature maps of the same size in the decoder, and concatenate them . Withal, we also estimate intermediate pyramid-resolution predictions, which converge faster and ensure that the network’s internal features are more aimed for the task. As it is common in the literature , our network’s backbone is ResNet-50, pretrained on the ImageNet Classification Dataset . As suggested in the literature and our experiments, pretraining the encoder on general image recognition tasks, as ImageNet, helps in both accuracy and convergence time reduction. A schematic of our network architecture can be seen in Figure 4.

ξ\xi: Inverse depth ξ=1d\xi=\frac{1}{d}. We chose inverse depth for its linear relationship with pixel variations.

c\mathit{c}: Depth confidence. As , we enforce the network to predict a confidence map for every depth prediction.

n\mathbf{n}: Surface normals. The normals are predicted only for small resolutions (all except the last two), as the ground-truth normals are too noisy at full resolution.

2 Losses

In this section we will present all the losses and their combination for the training.

Depth Loss: We minimize the L1 norm of the predicted inverse depth ξ\xi minus the ground truth inverse depth ξ^\hat{\xi}, that is

Note that for experiments with focal length normalization we scale depth values accordingly (see section 3.1).

Scale-Invariant Gradient Loss: We use the scale-invariant gradient loss proposed by , in order to favor smooth and edge preserving depth estimations. The loss based on the depths is

For the gradients, we use the same discrete scale-invariant finite differences operator g\mathbf{g} as defined in their work, which is

and we apply the scale-invariant loss to cover gradients at 5 different spacings hh.

Confidence Loss: The ground truth for the confidence map must be calculated online as it depends on the prediction. The confidence ground truth is calculated as

and its corresponding loss function is defined as

Normal Loss: For the normal loss, we use the L2 norm. The ground truth for the normals (n^\hat{\mathbf{n}}) is derived from the ground truth depth image. The loss for the normals is as follows:

Total Loss: The individual losses are weighted by factors obtained empirically, so the total loss L\mathcal{L} is

where λ1\lambda_{1}, λ2\lambda_{2}, λ3\lambda_{3} and λ4\lambda_{4} are 150150, 100100, 5050 and 2525 respectively.

Multi-Camera Experiments and Results

Most of the single-view depth prediction networks have been trained and tested using the same or very similar camera models. Generalizing to different camera models has several implications that are not straightforward. For this reason, we first present a thorough analysis on the generalization capabilities of current approaches. To this end we apply naïve generalization techniques (focal normalization and image resizing) during training on a network without our special convolutions (as Figure 4 but without CAM-Convs) and examine the limitations. Finally, we train and evaluate our network with CAM-Convs (as Figure 4) and show the improved generalization performance with respect to different camera parameters.

The major part of our experiments are done on the 2D-3D Semantics Dataset , that contains RGB-D equi-rectangular images. This dataset allows us to generate images with different camera intrinsics but the same content. We have observed that depth estimation networks overfit to the camera parameters and the image content distribution (the latter being different in indoors and outdoors datasets, for example). In this manner we eliminate the content distribution factor and isolate the effect of the camera parameters.

All the experiments were done using the 3-fold cross-validation suggested in . In this section we present median values for the most relevant experiments. To see the complete results, more details on the dataset and image generation process and additional experiments we refer the reader to the supplementary material.

The notation for sensor sizes and focal lengths used during the evaluation is in Table 1. As an example, if a network has been trained with sensor sizes 192×256192\times 256 and 224×224224\times 224, and focal length 72, we will denote this model as s2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>s</mi><mn>3</mn></msub></mrow><annotationencoding="application/x−tex">s3</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5806em;vertical−align:−0.15em;"></span><spanclass="mord"><spanclass="mordmathnormal">s</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:0em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">3</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f72s_{2}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>s</mi><mn>3</mn></msub></mrow><annotation encoding="application/x-tex">s_{3}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">3</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{72}. In some experiments we use a random distribution for the focal length. As an example, if the synthesized focal lengths are uniformly distributed between 72 and 128, the model will be denoted as U<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128\mathcal{U}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128}.

We evaluate the performance on both depth and inverse depth. All the error metrics we used in our experiments are standard from the literature. In addition we use relative metrics and the scale-invariant metric presented by , which are widely used in depth estimation.

2 Influence of context

Modifying the camera parameters affects the field of view, and hence the amount of context the image is capturing. We evaluate the influence of the context in the depth prediction of a standard U-Net encoder-decoder architecture (network in Figure 4 without CAM-Convs) with two different experiments. First, we compare two networks trained with images with sensor size s1s_{1} and two different focal lengths f128f_{128} and f64f_{64} (Table 2). Second, we compare two networks with images with the same focal length but different sensor sizes: s1s_{1} and s4s_{4} (Table 3).

As expected, context helps. The performance is better for the smallest focal f64f_{64}, which results in a wider FOV and hence more context. Also the performance is better for the bigger sensor size s1s_{1}, which also provides more context. To remove the context dependency in our analysis, for some of the experiments in next subsections we will generate images with uniformly distributed focal lengths.

3 Overfitting of standard networks

In this experiment we evaluate the performance of a standard U-Net architecture for variations of the camera parameters on the training and test sets. We will focus the study on two parameters: (a) focal length and (b) sensor size. First we will fix the sensor size to s1s_{1} and we will test on images with focal lengths f64f_{64}, f72f_{72} and f128f_{128} (first three test sets in Table 4). Second we will sample random focal lengths from a uniform distribution between f72f_{72} and f128f_{128} and we will evaluate on images with sensor sizes s1s_{1} and s2s_{2} (last two test sets in Table 4). For every test set there are 44 to 55 different train sets (referred in the \nth2 column of the table). For every test set we will refer to the case where the cameras from the training and test set are the same as the same-camera baseline. Training sets where we did not use focal length normalization are denoted with a ’*’. Networks trained on train sets with two sensor sizes have been trained either as Siamese networks with weight sharing or with image resizing to size s1s_{1} (denoted with a ’†’).

It is important to remark that, for all the experiments, the test and training data was generated from the exact same images and the networks have the same architecture and were trained for the same number of iterations. Any performance variation, then, should be attributed to the variations in the camera intrinsics and the naïve solutions we analyze. Notice in Table 4 that, in general, the same-camera baseline outperforms the rest, demonstrating the overfit to the camera parameters.

The conclusions of these experiments are as follows.

(a) Single-focal training overfits. The performance of a depth network degrades when trained on images from a particular camera and tested on images from different cameras. See, for example, the drop in performance between the \nth1 row (test: s1f64s_{1}f_{64}, train: s1f64∗{s_{1}f_{64}}^{*}) and the \nth2 (test: s1f64s_{1}f_{64}, train: s1f72{s_{1}f_{72}}) and \nth3 (test: s1f64s_{1}f_{64}, train: s1f128{s_{1}f_{128}}) rows in all metrics.

Multi-focal training with normalization helps. The results improve when the training set contains images with different focal lengths and is done with focal normalization. See, for example, that the results on test set s1s_{1}f64f_{64} with training set s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128} is close to the same-camera baseline. Notice, however, that the multi-focal train set does not reach the performance of the same-camera baseline. In section 5.4 we will show how CAM-Convs are able to outperform the same-camera baseline even when the training data does not contain the test focal length.

The performance degrades without focal normalization. Compare, for example, the error metrics of the train sets s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128}∗ and s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128}. Networks trained on f72f_{72}f128f_{128}∗, in fact, did not converge easily.

Limitations of focal normalization. Two things should be noticed regarding focal normalization: First, it does not model the changes on the sensor size and the resolution, and we will see now how changes on them degrade the performance. And second, Equation 4 only holds if the pixel size is the same for every camera in the training and test sets, which in general is not the case.

(b) Single-sensor size training overfits. Networks trained on a sensor size and tested on other sensor sizes do not perform as well as the same-camera baseline. This can be seen in Table 4 in the last two test sets s1Uf72f128s_{1}\mathcal{U}f_{72}f_{128} and s2Uf72f128s_{2}\mathcal{U}f_{72}f_{128}. Single-view depth estimation is a context-dependent task, and the network overfits to the amount of context in the training sensor size.

Multi-sensor size training with weight sharing does not generalize. Training with multiple sensor sizes works better than training with the wrong sensor size but cannot reach the same performance as same-camera baselines. Further, training a stack of weight sharing networks also does not scale to large numbers of different sensor sizes.

Resizing does not work. As a naïve approach, which scales to multiple sensor sizes, we use resizing (denoted with ’†’ in Table 4), which converts all the images to size (s1s_{1}) during training. Notice that resizing changes the aspect ratio. It also implies the recalculation of a new average focal length fr=frx+ry2f_{r}=f\frac{r_{x}+r_{y}}{2} for normalization. The performance degradation introduced by resizing is noticeable. Resizing creates inconsistent data in train and testing, which leads to learning and convergence difficulties.

Resizing helps only in a particular case (non-overlapping distributions of visual features). Table 5 shows an experiment, similar to the previous one, on two public datasets: KITTI , with sensor size sKs_{K}, and ScanNet , with sensor size sSs_{S}. In this case, training with both sensor sizes (by weight sharing) decreased the performance. However, resizing reduced the error to the level of the same-camera baselines. The reason for this is the completely different distribution of the two datasets, with null intersection of visual features (e.g. there are no chairs on KITTI and no cars on ScanNet). This is, however, a very particular case, resizing degrades significantly the accuracy in general.

4 Robust Generalization with CAM-Convs

In this experiment we show that CAM-Convs generalize to different camera models. In order to evaluate the influence of CAM-Convs we trained our model with two different sensor sizes (s1s_{1} and s2s_{2}) and weight sharing. Focal length during training is sampled randomly from a uniform distribution U<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128\mathcal{U}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128}. We evaluated the trained model in four different test sets, see Table 6. The first two include the camera model the network was trained with, the third has a sensor size unseen during training, and the last (s5f64s_{5}f_{64}) was generated from a camera completely different from the training ones with bigger sensor size and smaller focal length. This case augments considerably the context–e.g field of view–which proved to be the hardest case in previous experiments (see network trained with s1s_{1}f128f_{128} in Table 4).

CAM-Convs generalize over camera intrinsics, outperforming the same-camera baseline. Results on the test sets s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mimathvariant="script">U</mi></mrow><annotationencoding="application/x−tex">U</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6833em;"></span><spanclass="mordmathcal"style="margin−right:0.0993em;">U</span></span></span></span></span>f72s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="script">U</mi></mrow><annotation encoding="application/x-tex">\mathcal{U}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathcal" style="margin-right:0.0993em;">U</span></span></span></span></span>f_{72}f128f_{128} and s2<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mimathvariant="script">U</mi></mrow><annotationencoding="application/x−tex">U</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.6833em;"></span><spanclass="mordmathcal"style="margin−right:0.0993em;">U</span></span></span></span></span>f72s_{2}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="script">U</mi></mrow><annotation encoding="application/x-tex">\mathcal{U}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathcal" style="margin-right:0.0993em;">U</span></span></span></span></span>f_{72}f128f_{128} in Table 6 show that the network with CAM-Convs trained on images of two sizes clearly outperforms the baselines, which was trained on the exact test size. The addition of CAM-Convs allowed the network to learn the dependence of the image features from the calibration parameters.

CAM-Convs generalize to sensor sizes unseen during training. Remarkably, the network with CAM-Convs also outperforms the same-camera baseline on the test set with sensor size s3s_{3} (third test set in Table 6), which is not included in the training data. Further, it generalizes better than a network trained on the exact same conditions but without CAM-Convs (see s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>s</mi><mn>2</mn></msub></mrow><annotationencoding="application/x−tex">s2</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5806em;vertical−align:−0.15em;"></span><spanclass="mord"><spanclass="mordmathnormal">s</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:0em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">2</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>U<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>s</mi><mn>2</mn></msub></mrow><annotation encoding="application/x-tex">s_{2}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">2</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>\mathcal{U}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128} in the table).

CAM-Convs generalize to cameras unseen during training. With the last test set (s5s_{5}f64f_{64}) in Table 6 we evaluate our network on an extreme case of camera parameters with a very wide field of view and very different sensor size from the training ones. Table 6 shows that CAM-Convs improve considerably the generalization to new unseen cameras over the naïve approaches. Figure 5 shows a qualitative comparison between our network with CAM-Convs and the network without CAM-Convs (s1<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>s</mi><mn>2</mn></msub></mrow><annotationencoding="application/x−tex">s2</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5806em;vertical−align:−0.15em;"></span><spanclass="mord"><spanclass="mordmathnormal">s</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:0em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">2</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>U<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotationencoding="application/x−tex">f72</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8889em;vertical−align:−0.1944em;"></span><spanclass="mord"><spanclass="mordmathnormal"style="margin−right:0.1076em;">f</span><spanclass="msupsub"><spanclass="vlist−tvlist−t2"><spanclass="vlist−r"><spanclass="vlist"style="height:0.3011em;"><spanstyle="top:−2.55em;margin−left:−0.1076em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">72</span></span></span></span></span><spanclass="vlist−s">​</span></span><spanclass="vlist−r"><spanclass="vlist"style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f128s_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>s</mi><mn>2</mn></msub></mrow><annotation encoding="application/x-tex">s_{2}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">2</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>\mathcal{U}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>f</mi><mn>72</mn></msub></mrow><annotation encoding="application/x-tex">f_{72}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">72</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span></span>f_{128}) in the test set s5s_{5}f64f_{64}.

5 Experiments on Multiple Datasets

In our last experiment we demonstrate how CAM-Convs can generalize across datasests by training on four datasets with different cameras (KITTI , ScanNet, MegaDepth and Sun3D and testing on a different one (NYUv2 ).

Training: We trained our network for three different sensor sizes (320×320320\times 320, 256×256256\times 256 and 224×224224\times 224) using weight sharing. We augmented the training data by scaling the images and shifting the principal point to increase the variation of the camera parameters and then crop to image to one of the target sensor sizes. We did not use focal length normalization in this experiment, as we cannot ensure constant pixel size across datasets. As MegaDepth has only up-to-scale ground truth, we applied only scale-invariant losses and added the scale-invariant cost function of . The same network without CAM-Convs, and hence with no camera information, did not converge during training. The lack of calibration information creates inconsistencies (e.g. same-size objects may have different depths due to different focal lengths).

Testing: We evaluated our network on the official test set of NYUv2 and compared against the state of art (similar network without CAM-Convs) . Note that the network of was trained exclusively on NYUv2, while our network was trained on a set of datasets excluding NYUv2 with different cameras and data distributions (some of the datasets are outdoors, see Figure 8 and Figure 9). This is important since our model cannot benefit from the dataset bias . We predicted depths for images from 6 different cameras: the original camera of the NYUv2 dataset and 5 simulated ones by cropping (to shift principal point and reduce sensor size) and resizing (to change focal length).

Figure 6 shows the distribution of the mean error of the usual metrics obtained for the 6 different cameras. Since was trained on the NYUv2 dataset, it works slightly better when it predicts the images from the camera it was trained on (the point with the smallest error). However, performance degrades when the camera changes and CAM-Convs have always smaller error and variance. Figure 7 illustrate how CAM-Convs depth predictions are stable for different cameras, while predictions of vary significantly. Recall that CAM-Convs were not trained on NYUv2, which indicates that they are able to generalize over different camera models and outperform although they trained on the same dataset.

Figures 7, 8 and 9 show depth predictions for images (and cropped/resized versions) from the NYUv2, KITTI and MegaDepth test sets. Again, note the excellent performance across datasets with different data distributions and camera intrinsics. All predictions were done with the exact same network without further fine-tuning to a particular dataset or camera parameters.

Conclusions

This paper introduces CAM-Convs, a novel type of convolution that allows depth prediction networks to be camera-independent. Experimental results show that current networks overfit to the training camera model resulting on: 1) a lack of generalization to images from other cameras and 2) degraded performance when trained with images from different cameras. CAM-Convs learn how to use the camera intrinsics jointly with the image features to predict depth; solving both limitations. They maintain prediction accuracy for new cameras and better exploit training data from different cameras. The latter is an interesting direction to scale up systems that depend on camera parameters.

Acknowledgement: This project was in part funded by the Spanish government (DPI2015-67275), the EU Horizon 2020 project Trimbot2020, the Aragón government (DGA-T45_17R/FSE) and Fundación CAI-Ibercaja. We also thank Facebook for their P100 server donation and gift funding; and Nvidia for their Titan X and Xp donation.

References