Self-Calibrating Neural Radiance Fields

Yoonwoo Jeong, Seokjun Ahn, Christopher Choy, Animashree Anandkumar, Minsu Cho, Jaesik Park

Introduction

Camera calibration is one of the crucial steps in computer vision. Through this process, we learn how the incoming rays map to pixels and thus connect the images to the physical world. Thus, it is a fundamental step in many applications such as autonomous driving, robotics, augmented reality, and many more.

Camera calibration is typically done by placing calibration objects (e.g., a checkerboard pattern) in the scene and estimating the camera parameters using the known geometry of the calibration objects. However, in many cases, calibration objects are not readily available and can interfere with the perception tasks when cameras are deployed in the wild. Thus, calibrating without any external objects, or self-calibration, has been an important research topic; first proposed in Faugeras et al. . The paper has spurred many follow-ups, some of which propose to globally optimize or embed constraints into the self-calibration optimization process .

Although there has been much progress in developing self-calibration algorithms, all these methods have limitations: 1) the camera model used in self-calibration is a simple linear pinhole camera model. This camera-model design cannot incorporate generic non-linear camera noise that is prevalent in all commodity cameras resulting in less accurate camera calibration. 2) self-calibration algorithms use only a sparse set of image correspondences, and direct photometric consistency has not been used for self-calibration. 3) they use correspondences from a non-differentiable process and do not improve the 3D geometry of the objects, which could improve the camera model. Let us discuss each limitation in detail.

Second, conventional self-calibration methods solely rely on the geometric loss or constraints based on the epipolar geometry, such as Kruppa’s method that only uses a set of sparse correspondences extracted from a non-differentiable process. This could lead to diverging results with extreme sensitivity to noise when a scene does not have enough interest points. On the other hand, photometric consistency is a physically-based constraint that forces the same 3D point to have the same color in all valid viewpoints. It can create a large number of physically-based constraints to learn accurate camera parameters.

Lastly, conventional self-calibration methods use an off-the-shelf non-differentiable feature matching algorithm and do not improve or learn the geometry. It is well known that the better we know the geometry of the scene, the more accurate the camera model gets. This fact is essential since the geometry of the scene is the sole source of input for self-calibration.

In this work, we propose a self-calibration algorithm for generic camera models that end-to-end learn parameters for the basic pinhole model and radial distortion and non-linear camera noise. For this, our algorithm jointly learns geometry together with a unified end-to-end differentiable framework that allows better geometry to improve camera parameters. In particular, we use the implicit volumetric representation or Neural Radiance Fields for the differentiable scene geometry representation.

We also propose a geometric consistency designed for our camera model and train the system together with the photometric consistency for self-calibration, which provides a large set of constraints. The novel geometric consistency forces rays from corresponding points on images to be close to each other, which overcomes the pinhole camera assumption in the conventional geometric losses derived from Kruppa’s method for self-calibration .

Experimentally, we show that our models can learn camera parameters, including intrinsics and extrinsics, without the standard COLMAP initialization. Also, when the initialization values for these camera parameters are given, we fine-tune the camera parameters accurately, which improves the underlying geometry and novel view synthesis. We test our model on fish-eye images with COLMAP learned camera radial distortion parameters to analyze the distortion model and show that our model outperforms the baselines by a significant margin. In addition, we show that our non-linear camera model is modular and can be applied to NeRF variants such as NeRF and NeRF++ .

Related Work

Camera Distortion Model. Traditional 3D vision tasks often assume that the camera model is a simple pinhole model. With the development of camera models, various camera models have been introduced, including fish-eye models, per-pixel generic models. Although per-pixel generic models are more expressive, they are difficult to optimize. Schops et al. propose a model locating between 12 parameters and per-pixel generic models. They have shown the proposed model has less reprojection error than other camera models. Strum and Srikumar propose several methods that show how to calibrate a general imaging model, where structures are known, but viewpoints are unknown. The proposed methods allow learning central cameras without using any distortion model. Grossberg and Nayar propose a general imaging model that uses virtual sensing elements that describes the mapping between incoming ray and pixel. They also propose a calibration method that finds parameters of virtual sensing elements and shows that the method can be applied to any imaging system. Ramalingam and Sturm interpret the camera model as a function that maps pixel to a 3D ray. With this interpretation, they model various cameras, such as central cameras or axial cameras.

Camera auto-calibration is the process of estimating camera parameters from a set of uncalibrated images and cameras without using external calibration objects in the scene, such as checkerboard patterns. Zeller et al. propose a self-calibration method that adopts the Kruppa equation to self-calibrate the camera parameters in a video sequence. Pollefeys et al. propose a stratified method for calibration using modulus constraints. Chandraker et al. propose a self-calibration algorithm that incorporates the rank and positive semi-definite constraints into the optimization. Chandraker et al. incorporate the branch and bound method for the globally optimal stratified self-calibration algorithm. Ha et al. adopt a loss, which implicitly calibrates the camera models using the correspondences between image pairs to produce a high-quality depth map from uncalibrated small motion clips. Engel et al. propose a novel approach to calibrate the response function and the non-parametric vignetting function to generate a more accurate tracking model.

Novel View Synthesis. Neural Radiance Fields synthesize novel views by learning volumetric scene function with multi-layer perceptron. Several improvements on Neural Radiance Field have been proposed. Zhang et al. improve the original NeRF model by discriminating background and foreground. Liu et al. propose a sparse voxel field approach that skips ray marching of the voxels containing no relevant contents, enabling efficient and more precise rendering. Yariv et al. synthesize novel views by reconstructing the 3D surface as a level set of signed distance functions with a neural network. However, surface-based rendering requires a binary mask distinguishing background and foreground. Moreover, it is not suitable to reconstruct real scenes since the model also reconstructs the background surface. Yu et al. propose a learning framework to learn scene information using few images. Yen et al. address an inverse problem of NeRF, which estimates poses of observed images. They’ve used test images to predict the poses of the test images and re-trained the NeRF network with the predicted poses for better rendering quality.

Preliminary

We use the neural radiance fields to learn the 3D scene geometry, which is crucial for learning the photometric loss for self-calibration. In this section, we briefly cover the definitions of the neural radiance fields: NeRF and NeRF++ .

The color value C\mathbf{C} of a ray can be represented as an integral of all colors weighted by the opaqueness along a ray, or can be approximated as the weighted sum of colors at NN points along a ray.

where Δi=ti+1−ti\Delta_{i}=t_{i+1}-t_{i}. Thus, the accuracy of the method depends highly on the number of samples as well as how we sample points.

Background Representation with Inverse Depth. The volumetric rendering used in NeRF is effective and robust if the space the network to capture is bounded. However, on the outdoor scene, the volume of the space is unbounded, and the number of samples required to capture the space increases proportionally, often computationally prohibitive. Instead, Zhang et al. propose NeRF++ to model foreground and background with separate implicit networks while the background ray is reparametrized to have bounded volume. The network architecture of can be succinctly formulated as two implicit networks: one for foreground and one for background. In this paper, we will explore both NeRF and NeRF++ to analyze our camera self-calibration model.

Differentiable Self-Calibrating Cameras

In this section, we provide the definition of our differentiable camera model that combines the pinhole camera model, radial distortion, and a generic non-linear camera distortion for self-calibration . Mathematically, a camera model is a mapping p=π(r)\mathbf{p}=\pi(\mathbf{r}) that defines a 3D ray r\mathbf{r} to a 2D coordinate p\mathbf{p} in the image plane. In this work, we focus on the unprojection function, or a ray, r(p)=π−1(p)\mathbf{r}(\mathbf{p})=\pi^{-1}(\mathbf{p}) as the geometry learning and our projected ray distance only requires the unprojection of a pixel to a ray. Thus, we use the term camera model and camera unprojection interchangeably, and we represent a ray r(p)\mathbf{r}(\mathbf{p}) of a pixel p\mathbf{p} as a pair of 3-vectors: a direction vector rd\mathbf{r}_{d} and an offset or a ray origin vector ro\mathbf{r}_{o}.

Our camera unprojection process consists of two components: unprojection of pixels using a differentiable pinhole camera model and generic non-linear ray distortions. We first mathematically define each component.

The first component of our differentiable camera unprojection is based on the pinhole camera model, which maps a 4-vector homogeneous coordinate in 3D space to a 3-vector in the image plane.

Note that we will denote c=[cx,cy]\mathbf{c}=[c_{x},c_{y}] and f=[fx,fy]\mathbf{f}=[f_{x},f_{y}] for simplicity. Similarly, we use the extrinsics initial values R0R_{0} and t0\mathbf{t}_{0} and residual parameters to represent the camera rotation RR and translation t\mathbf{t}. However, directly learning the rotation offset for each element of a rotation matrix would break the orthogonality of the rotation matrix. Thus, we adopt the 6-vector representation which uses unnormalized first two columns of a rotation matrix to represent a 3D rotation:

Since these ray parameters (rd,ro\mathbf{r}_{d},\mathbf{r}_{o}) are functions of intrinsics and extrinsics residuals (Δf,Δc,Δa,Δt\Delta\mathbf{f},\Delta\mathbf{c},\Delta\mathbf{a},\Delta\mathbf{t}), we can pass gradients from the rays to the residuals to optimize the parameters. Note that we do not optimize K0,R0,t0K_{0},R_{0},\mathbf{t}_{0}.

Cameras are made of a set of circular lenses which warp rays to the center. Thus, distortions at the edge of the lenses create circular distortion patterns. We extend our model to incorporate such radial distortions. Following the radial fisheye model in COLMAP , we adopt the fourth order radial distortion model which drops rare higher order distortions, i.e. k=(k1+zk1,k2+zk2)\mathbf{k}=(k_{1}+z_{k_{1}},k_{2}+z_{k_{2}}).

Similar to other camera parameters, we learn these camera parameters using photometric errors.

2 Generic Non-Linear Ray Distortion

We model some distortions that are easy to express mathematically. However, complex optical abberations in real lenses cannot be modeled using a parametric camera. For such noise, we use non-linear model following Grossberg et al. to use local raxel parameters to capture generic non-linear aberration. Specifically, we use local ray parameter residuals zd=Δrd(p),zo=Δro(p)\mathbf{z}_{d}=\Delta\mathbf{r}_{d}(\mathbf{p}),\mathbf{z}_{o}=\Delta\mathbf{r}_{o}(\mathbf{p}) where p\mathbf{p} is the image coordinate.

We use bilinear interpolation to locally extract continuous ray distortion parameters

zd[x,y]\mathbf{z}_{d}[x,y] indicates the ray direction offset at a discrete 2D coordinate (x,y)(x,y). We learn the parameters of zd\mathbf{z}_{d} at discrete locations only. Similarly, we can define zo(p)\mathbf{z}_{o}(\mathbf{p}) as bilinear interpolation of zo[x,y]\mathbf{z}_{o}[x,y]. The final ray direction, ray offset generation can be summarized as Fig. 2.

Geometric and Photometric Consistency

Our camera model incorporates the generic non-linear distortions that increase the number of camera parameters drastically. In this work, we proposed using both geometric and photometric consistencies for self-calibration, which allows more accurate camera parameter calibration as these consistencies provide additional constraints. We discuss each of the constraints in this section.

The generic camera model poses a new challenge defining a geometric loss. In most traditional work, the geometric loss is defined as an epipolar constraint that measures the distance between an epipolar line and the corresponding point, or reprojection error where a 3D point for a correspondence is defined first which is then projected to an image plane to measure the distance between the projection and the correspondence. However, these methods have few limitations when we use our generic noise model.

First, the epipolar distance assumes a perfect pinhole camera, which breaks in our setup. Second, the 3D reprojection error requires creating a 3D point cloud reconstruction using a non-differentiable process, and the camera parameters are learned indirectly from the 3D reconstruction.

In this work, rather than requiring a 3D reconstruction to compute an indirect loss like the reprojection error, we propose the projected ray distance loss that directly measures the discrepancy between rays. Let (pA↔pB)(\mathbf{p}_{A}\leftrightarrow\mathbf{p}_{B}) be a correspondence on camera 1 and 2 respectively. When all the camera parameters are calibrated, the ray rA\mathbf{r}_{A} and rB\mathbf{r}_{B} should intersect at the 3D point that generated point pA\mathbf{p}_{A} and pB\mathbf{p}_{B}.

However, when there’s a misalignment due to an error in camera parameters, we can measure the deviation by computing the shortest distance between corresponding rays.

Let a point on line A be xA(tA)=ro,A+tArd,A\mathbf{x}_{A}(t_{A})=\mathbf{r}_{o,A}+t_{A}\mathbf{r}_{d,A} and a point on line B be xB(tB)=ro,B+tBrd,B\mathbf{x}_{B}(t_{B})=\mathbf{r}_{o,B}+t_{B}\mathbf{r}_{d,B}. A distance between the line A and a point on the line B is

If we solve for dd2dtB∣t^B=0\frac{\mathbf{d}d^{2}}{\mathbf{d}t_{B}}|_{\hat{t}_{B}}=0, we get

We substitute t^B\hat{t}_{B} to the line 2 and can get the x^B=xB(t^B)\hat{\mathbf{x}}_{B}=\mathbf{x}_{B}(\hat{t}_{B}). Similarly, we can get x^A\hat{\mathbf{x}}_{A}. For simplicity, we will denote x⋅\mathbf{x}_{\cdot} as x^⋅\hat{\mathbf{x}}_{\cdot} since we will focus primarily on the final solution. The distance between two points d^=xAxB‾\hat{d}=\overline{\mathbf{x}_{A}\mathbf{x}_{B}} is

However, this distance is not normalized for correspondences. Given the same camera distortions, a correspondence for a point farther from the cameras would have a larger deviation, while a correspondence for a point closer to the cameras would have a smaller deviation. Thus, we need to normalize the scale of the distance. Thus, we project the points xA,xB\mathbf{x}_{A},\mathbf{x}_{B} to image planes IA,IBI_{A},I_{B} and compute distance on the image planes, rather than directly using the distance in the 3D space.

where π(⋅)\pi(\cdot) is a projection function and equalizes the contribution from each correspondence irrespective of their distance from the cameras. We visualize the projected ray distance in Fig. 3.

This projected ray distance is a novel geometric loss different from the epipolar distance or the reprojection error. The epipolar distance is defined only for linear pinhole cameras and cannot model the non-linear camera distortions. On the other hand, the reprojection error requires extracting a 3D reconstruction in a non-differentiable preprocessing stage and optimizes the camera parameters via optimizing the 3D reconstruction. Our projected ray distance does not require the intermediate 3D reconstruction and can model the non-linear camera distortions.

2 Chirality Check

When the camera distortion is large and the baseline between cameras is small, the shortest line between rays from a correspondence might be located behind the cameras. Minimizing such invalid ray distance would result in suboptimal camera parameters. Thus, we check whether the points are behind a camera by computing the z-depth along with the camera rays. Mathematically,

where x[z]\mathbf{x}[z] indicates the z component of a vector. Finally, we only average valid projected ray distances for all correspondences to compute the geometric loss.

3 Photometric Consistency

Unlike geometric consistency, photometric consistency requires reconstructing the 3D geometry because the color of a 3D point is valid only if it is visible from the current perspective. In our work, we use a neural radiance field to reconstruct the 3D occupancy and color. This implicit representation is differentiable through both position and color value and allows us to capture the visible surface through volumetric rendering. Specifically, during the rendering process, a ray is parametrized using K0,R0,t0K_{0},R_{0},\mathbf{t}_{0} as well as ΔK,Δa,Δt\Delta K,\Delta a,\Delta t as well as zo[⋅],zd[⋅]\mathbf{z}_{o}[\cdot],\mathbf{z}_{d}[\cdot] as visualized in Fig. 2. We differentiate the following energy function with respect to the learnable camera parameters to optimize our self-calibration model.

Here, p\mathbf{p} is a pixel coordinate, and I\mathcal{I} is a set of pixel coordinates in an image. C^(r)\hat{C}(\mathbf{r}) is the output of the volumetric rendering using the ray r\mathbf{r}, which corresponds to the pixel p\mathbf{p}. C(p)C(\mathbf{p}) is the ground truth color. Thus, the gradient for the intrinsics is

Similarly, we can define gradients for the rest of the parameters Δa,Δt\Delta a,\Delta t as well as zo[⋅],zd[⋅]\mathbf{z}_{o}[\cdot],\mathbf{z}_{d}[\cdot] and calibrate cameras.

Optimizing Geometry and Camera

To optimize geometry and camera parameters, we learn the neural radiance field and the camera model jointly. However, it is impossible to learn accurate camera parameters when the geometry is unknown or too coarse for self-calibration. Thus, we sequentially learn parameters: geometry and a linear camera model first and complex camera model parameters.

The camera parameters determine the positions and directions of the rays for NeRF learning, and unstable values often result in divergence or sub-optimal results. Thus, we add a subset of learning parameters to the optimization process to jointly reduce the complexity of learning cameras and geometry. First, we learn the NeRF networks while initializing the camera focal lengths and focal centers to half the image width and height. Learning coarse geometry first is crucial since it initializes the networks to a more favorable local optimum for learning better camera parameters. Next, we sequentially add camera parameters for the linear camera model, radial distortion, and nonlinear noise of ray direction, ray origin to the learning. We learn simpler camera models first to reduce overfitting and faster training.

2 Joint Optimization

We present the final learning algorithm in Alg. 1. The get_paramsget\_params function returns a set of parameters for the curriculum learning which progressively adds complexity to the camera model. Next, we train the model with the projected ray distance by selecting a target image at random with sufficient correspondences. Heuristically, we found selecting images within maximum 30°from the source view gives an optimal result.

Experiment

We use three datasets to analyze different aspects of our model. Two outdoor scenes, Mildenhall et al. and Zhang et al. , are captured with a pinhole camera lens. LLFF and Tanks and Temples dataset are composed of 8 and 4 scenes, respectively, where their camera parameters are estimated using COLMAP .

Since these datasets are captured using professional cameras with small lens distortions, we collected a few scenes using a fish-eye camera to examine the end-to-end learning capacity of our model. We acquire the camera information with COLMAP.

2 Self-Calibration

We train our model from scratch to demonstrate that our model can self-calibrate the camera information. We initialize all the rotation matrices, the translation vectors, and focal lengths to an identity matrix, zero vector, and height and width of the captured images. Table 1 reports the qualities of the rendered images in the training set. Although our model does not adopt calibrated camera information, our model shows a reliable rendering performance. Moreover, for some scenes, our model outperforms NeRF, trained with COLMAP camera information. We have visualized the rendered images in Figure 7.

3 Improvement over NeRF

We have observed that our model shows better rendering qualities than NeRF when COLMAP initializes the camera information. We compare the rendering qualities of NeRF and our model in Table 2. Our model consistently shows better rendering qualities than the original NeRF. In addition, our model indicates much less projected ray distance, indicating that our model improves the camera information. We visualize the non-linear distortion that our camera model learned in Fig. 4.

4 Improvement over NeRF++

Since our model is designed to work on variants of NeRF, we have substituted the NeRF architecture to the NeRF++ architecture. We then compare NeRF++ and our model in tanks and temples dataset. Table 3 reports rendering qualities and projected ray distance loss in the training set. Our model results in better rendering qualities and much less train projected ray distance. The qualitative results are visualized in Figure 6.

5 Fish-eye Lens Reconstruction

We test our model on images with high distortion to contrast the importance of end-to-end learning of camera parameters. Conventional feature matching algorithms fail to acquire reliable correspondences for these scenes, so we skip the projected ray distance loss from our curriculum training. Table 4 reports the rendering qualities of our learned model and the baseline NeRF++. We trained the baseline and our model from the COLMAP initialization with a radial distortion model that provides fish-eye camera parameters. Since the NeRF++ camera model does not incorporate the radial distortion, we modify the implementation to incorporate the fish-eye distortion in ray computation.

6 Ablation Study

To check the effects of the proposed models, we conduct an ablation study. We check the performance for each phase in curriculum learning. We train 200K iterations for each phase. From this experiment, we have observed that extending our model is more potential in rendering clearer images. However, for some scenes, adopting projected ray distance increases the overall projected ray distance. Table 5 reports the results of the ablation study and Figure 8 visualizes the errors.

Conclusion

We propose a self-calibration algorithm that learns geometry and camera parameters jointly end-to-end. The camera model consists of a pinhole model, radial distortion, and non-linear distortion, which capture real noises in lenses. We also propose projected ray distance to improve accuracy, which allows our model to learn fine-grained correspondences. We show that our model learns geometry and camera parameters from scratch when the poses are not given, and our model improves both NeRF and NeRF++ to be more robust when camera poses are given.

Acknowledgements. This work was supported by the IITP grants (2019-0-01906: AI Grad. School Prog. - POSTECH and 2021-0-00537: visual common sense through self-supervised learning for restoration of invisible parts in images) funded by Ministry of Science and ICT, Korea.

References

Implementation Details

We use batch size of 1024 rays for NeRF , 512 rays for NeRF++ . We initially set the learning rate of NeRF to 0.0005. The learning rate decays exponentially to one-tenth for every 400000 steps. For NeRF++, we initially set the learning rate to 0.0005. The learning rate decays exponentially to one-tenth for every 7500000 steps. As explained in the main paper, we adopt curriculum learning for better stability of training. We extend our learnable camera parameters for every 200K iterations for NeRF experiments. For NeRF++, we have extended our learnable parameters in 500K iterations, 800K iterations, and 1.1M iterations. For NeRF, we use 64 samples for the coarse network and 128 samples for the fine network. In NeRF++ experiments, we have scaled extrinsic noises to 0.01. Especially for tanks and temples dataset, 64 points along a ray are sampled and fed to a coarse network. 128 points along the same ray are sampled and fed to a fine network. For FishEyeNeRF experiments, we have sampled 128 points for the coarse network and 256 for the fine network along the ray. Besides, 1024 rays are used for the FishEyeNeRF experiments for each iteration.

2 Projected Ray Distance Evaluation

The projected ray distance measures the deviation of a correspondence pair. However, as it only finds the shortest distance between rays in 3D space, a small change in the direction sometimes leads to a large change in the ray distance. Thus, we threshold ray distance with η\eta and remove pairs above this threshold. We set the η\eta to 5.0 for all the experiments.

Calibration with COLMAP initialization

We extend Table 2 in the main paper by conducting experiments in other scenes of the LLFF dataset . Table 1 reports the rendering qualities and projected ray distance of NeRF and our model. Our model shows a consistent improvement from NeRF when learnable camera parameters are initialized by COLMAP camera information.

Calibration without COLMAP initialization

We also extend Table 1 in the main paper by conducting the experiments in other scenes of the LLFF dataset . Table 2 reports the rendering qualities and projected ray distance metric. NeRF fails to render the scenes reliably; however, our model does. Qualitative results are shown in Figure 4.

Ablation Studies

We extend ablation study in our main paper by conducting experiments in the other scenes of LLFF dataset. Table 3 reports the quantitative results of the ablation study.

Qualitative Results

We report some qualitative results of experiments in the main paper. Figure 1 compares NeRF++ and our model in a tanks and temples dataset. Figure 2 compares NeRF++ and our model in two fish-eye scenes.

Figure 4 visualizes rendered images of our model when no calibrated camera information is provided. Lastly, Figure 5 visualizes the captured non-linear distortion in for all the scenes in LLFF dataset.