A Photometrically Calibrated Benchmark For Monocular Visual Odometry

Jakob Engel, Vladyslav Usenko, Daniel Cremers

Introduction

Structure from Motion or Simultaneous Localization and Mapping (SLAM) has become an increasingly important topic, since it is a fundamental building block for many emerging technologies – from autonomous cars and quadrocopters to virtual and augmented reality. In all these cases, sensors and cameras built into the hardware are designed to produce data well-suited for computer vision algorithms, instead of capturing images optimized for human viewing. In this paper we present a new monocular visual odometry (VO) / SLAM evaluation benchmark, that attempts to resolve two current issues:

Many existing methods are designed to operate on, and are evaluated with, data captured by commodity cameras without taking advantage of knowing – or even being able to influence – the full image formation pipeline. Specifically, methods are designed to be robust to (assumed unknown) automatic exposure changes, non-linear response functions (gamma correction), lens attenuation (vignetting), de-bayering artifacts, or even strong geometric distortions caused by a rolling shutter. This is particularly true for modern keypoint detectors and descriptors, which are robust or invariant to arbitrary monotonic brightness changes. However, recent direct methods as well attempt to compensate for automatic exposure changes, e.g., by optimizing an affine mapping between brightness values in different images . Most direct methods however simply assume constant exposure time .

While this is the only way to evaluate on existing datasets and with off-the-shelf commodity cameras (which often do not allow to either read or set parameters like the exposure time), we argue that for the above-mentioned use cases, this ultimately is the wrong approach: sensors – including cameras – can, and will be designed to fit the needs of the algorithms processing their data. In turn, algorithms should take full advantage of the sensor’s capabilities and incorporate knowledge about the sensor design. Simple examples are image exposure time and hardware gamma correction, which are intentionally built into the camera to produce better images. Instead of treating them as unknown noise factors and attempt to correct for them afterwards, they can be treated as feature that can be modelled by, and incorporated into the algorithm – rendering the obtained data more meaningful.

SLAM and VO are complex, very non-linear estimation problems and often minuscule changes can greatly affect the outcome. To obtain a meaningful comparison between different methods and to avoid manual overfitting to specific environments or motion patterns (except for cases where this is specifically desired), algorithms should be evaluated on large datasets in a wide variety of scenes. However, existing datasets often contain only a limited number of environments. The major reason for this is that accurate ground truth acquisition is challenging, in particular if a wide range of different environments is to be covered: GPS/INS is limited in accuracy and only possible in outdoor environments with adequate GPS reception. External motion capture systems on the other hand are costly and time-consuming to set up, and can only cover small (indoor) environments.

The dataset published in this paper attempts to tackle these two issues. First, it contains frame-wise exposure times as reported by the sensor, as well as accurate calibrations for the sensors response function and lens vignetting, which enhances the performance particularly of direct approaches. Second, it contains 50 sequences with a total duration of 105 minutes (see Figure 1), captured in dozens of different environments. To make this possible, we propose a new evaluation methodology which does not require ground truth from external sensors – instead, tracking accuracy is evaluated by measuring the accumulated drift that occurs after a large loop. We further propose a novel, straight-forward approach to calibrate a non-parametric response function and vignetting map with minimal set up required, and without imposing a parametric model which may not suit all lenses / sensors.

1 Related Work: Datasets

There exists a number of datasets that can be used for evaluating monocular SLAM or VO methods. We will here list the most commonly used ones.

21 stereo sequences recorded from a driving car, motion patterns and environments are limited to forward-motion and street-scenes. Images are pre-rectified, raw sensor measurements or calibration datasets are not available. The benchmark contains GPS-INS ground truth poses for all frames.

11 stereo-inertial sequences from a flying quadrocopter in three different indoor environments. The benchmark contains ground truth poses for all frames, as well as the raw sensor data and respective calibration datasets.

89 sequences in different categories (not all meant for SLAM) in various environments, recorded with a commodity RGB-D sensor. They contain strong motion blur and rolling-shutter artifacts, as well as degenerate (rotation-only) motion patterns that cannot be tracked well from monocular odometry alone. Sequences are pre-rectified, the raw sensor data is not available. The benchmark contains ground truth poses for all sequences.

8 ray-traced RGB-D sequences from 2 different environments. It provides a ground truth intrinsic calibration; a photometric calibration is not required, as the virtual exposure time is constant. Some of the sequences contain degenerate (rotation-only) motion patterns that cannot be tracked well from a monocular camera alone.

2 Related Work: Photometric Calibration

Many approaches exist to calibrate and remove vignetting artefacts and account for non-linear response functions. Early work focuses on image stitching and mosaicking, where the required calibration parameters need to be estimated from a small set of overlapping images . Since the available data is limited, such methods attempt to find low-dimensional (parametric) function representations, like radially symmetric polynomial representations for vignetting. More recent work has shown that such representations may not be sufficiently expressive to capture the complex nature of real-world lenses and hence advocate non-parametric – dense – vignetting calibration. In contrast to , our formulation however does not require a “uniformly lit white paper”, simplifying the required calibration set-up.

For response function estimation, a well-known and straight-forward method is that of Debevec and Malik , which – like our approach – recovers a 282^{8}-valued lookup table for the inverse response from two or more images of a static scene at different exposures.

3 Paper Outline

The paper is organized as follows: In Section 2, we first describe the hardware set-up, followed by both the used distortion model (geometric calibration) in 2.2, as well as photometric calibration (vignetting and response function) and the proposed calibration procedure in 2.3. Section 3 describes the proposed loop-closure evaluation methodology and respective error measures. Finally, in Section 4, we give extensive evaluation results of two state-of-the-art monocular SLAM / VO systems ORB-SLAM and Direct Sparse odometry (DSO) . We further show some exemplary image data and describe the dataset contents. In addition to the dataset, we publish all code and raw evaluation data as open-source.

Calibration

We provide both a standard camera intrinsic calibration using the FOV camera model, as well as a photometric calibration, including vignetting and camera response function.

The cameras used for recording the sequences are uEye UI-3241LE-M-GL monochrome, global shutter CMOS cameras from IDS. They are capable of recording 1280×10241280\times 1024 videos at up to 60fps. Sequences are recorded at different framerates ranging from 20fps to 50fps with jpeg-compression. For some sequences hardware gamma correction is enabled and for some it is disabled. We use two different lenses (Lensagon BM2420 with a field of view of 148∘×122∘148^{\circ}\times 122^{\circ}, as well as a Lensagon BM4018S118 with a field of view of 98∘×79∘98^{\circ}\times 79^{\circ}), as shown in Figure 2. Figure 1 shows a number of example images from the dataset.

2 Geometric Intrinsic Calibration

where ru:=(xz)2+(yz)2r_{u}:=\sqrt{(\frac{x}{z})^{2}+(\frac{y}{z})^{2}} is the radius of the point in normalized image coordinates.

A useful property of this model is the existence of a closed-form inverse: for a given point the image (ud,vd)(u_{d},v_{d}) and depth dd, the corresponding 3D point can be computed by first converting it back to normalized image coordinates

3 Photometric Calibration

We provide photometric calibrations for all sequences. We calibrate the camera response function GG, as well as pixel-wise attenuation factors V ⁣:Ω→V\colon\Omega\to (vignetting). Without known irradiance, both GG and VV are only observable up to a scalar factor. The combined image formation model is then given by

where tt is the exposure time, BB the irradiance image (up to a scalar factor), and II the observed pixel value. As a shorthand, we will use U:=G−1U:=G^{-1} for the inverse response function.

We first calibrate the camera response function from a sequence of images taken of a static scene with different (known) exposure time. The content of the scene is arbitrary – however to well-constrain the problem, it should contain a wide range of gray values. We first observe that in a static scene, the attenuation factors can be absorbed in the irradiance image, i.e.,

with B′(x):=V(x)B(x)B^{\prime}(\mathbf{x}):=V(\mathbf{x})B(\mathbf{x}). Given a number of images IiI_{i}, corresponding exposure times tit_{i} and a Gaussian white noise assumption on U(Ii(x))U(I_{i}(\mathbf{x})), this leads to the following Maximum-Likelihood energy formulation

For overexposed pixels, UU is not well defined, hence they are removed from the estimation. We now minimize (6) alternatingly for UU and B′B^{\prime}. Note that fixing either UU or B′B^{\prime} de-couples all remaining values, such that minimization becomes trivial:

where Ωk:={i,x∣Ii(x)=k}\Omega_{k}:=\{i,\mathbf{x}|I_{i}(\mathbf{x})=k\} is the set of all pixels in all images that have intensity kk. Note that the resulting UU may not be monotonic – which is a pre-requisite for invertibility. In this case it needs to be smoothed or perturbed; however for all our calibration datasets this does not happen. The value for U(255)U(255) is never observed since overexposed pixels are removed, and needs to be extrapolated from the adjacent values. After minimization, UU is scaled such that U(255)=255U(255)=255 to disambiguate the unknown scalar factor. Figure 3 shows the estimated values for one of the calibration sequences.

In contrast to , we do not employ a smoothness prior on UU – instead, we use large amounts of data (1000 images covering 120 different exposure times, ranging from 0.05ms to 20ms in multiplicative increments of 1.05). This is done by recording a video of a static scene while slowly changing the camera’s exposure. If only a small number of images or exposure times is available, a regularized approach will be required.

3.2 Non-parametric Vignette Calibration

Again, we assume Gaussian white noise on U(Ii(πi(x)))U(I_{i}(\pi_{i}(\mathbf{x}))), leading to the Maximum-Likelihood energy

The energy E(C,V)E(C,V) can be minimized alternatingly, again fixing one variable decouples all other unknowns, such that minimization becomes trivial:

Again, we do not impose any explicit smoothness prior or enforce certain properties (like radial symmetry) by finding a parametric representation for VV. Instead, we choose to solve the problem by using large amounts (several hundred images) of data. Since VV is only observable up to a scalar factor, we scale the result such that max⁡(V)=1\max(V)=1. Figure 4 shows the estimated attenuation factor map for both lenses. The high degree of radial symmetry and smoothness comes solely from the data term, without additional prior.

Without regularizer, the optimization problem (9) is well-constrained if and only if the corresponding bipartite graph between all optimization variables is fully connected. In practice, the probability that this is not the case is negligible when using sufficient input images: Let (A,B,E)(A,B,E) be a random bipartite graph with ∣A∣=∣B∣=n|A|=|B|=n nodes and ∣E∣=[n(log⁡n+c)]|E|=[n(\log n+c)] edges, where cc is a positive real number. We argue that the complex nature of 3D projection and perspective warping justifies the approximation of the resulting residual graph as random, provided the input images cover a wide range of viewpoints. Using the Erdős-Rényi theorem , it can be shown that for n→∞n\to\infty, the probability of the graph being connected is given by P=e−2e−cP=e^{-2e^{-c}} . In our case, n≈10002n\approx 1000^{2}, which is sufficiently large for this approximation to be good. This implies that 30n30n residuals (i.e., 30 images with the full plane visible) suffice for the problem to be almost certainly well-defined (P>0.999999P>0.999999). To obtain a good solution and fast convergence, a significantly larger number of input images is desirable (we use several hundred images) which are easily captured by taking a short (60s) video, that covers different image regions with different plane regions.

Evaluation Metrics

The dataset focuses on a large variety of real-world indoor and outdoor scenes, for which it is very difficult to obtain metric ground truth poses. Instead, all sequences contain exploring motion, and have one large loop-closure at the end: The first and the last 10-20 seconds of each sequence show the same, easy-to-track scene, with slow, loopy camera motion. We use LSD-SLAM to track only these segments, and – very precisely – align the start and end segment, generating a “ground truth” for their relative poseSome sequences begin and end in a room which is equipped with a motion capture system. For those sequences, we use the metric ground truth from the MoCap.. To provide better comparability between the sequences, the ground truth scale is normalized such that the full trajectory has a length of approximately 100.

The tracking accuracy of a VO method can then be evaluated in terms of the accumulated error (drift) over the full sequence. Note that this evaluation method is only valid if the VO/SLAM method does not perform loop-closure itself. To evaluate full SLAM systems like ORB-SLAM or LSD-SLAM, loop-closure detection needs to be disabled. We argue that even for full SLAM methods, the amount of drift accumulated before closing the loop is a good indicator for the accuracy. In particular, it is strongly correlated with the long-term accuracy after loop-closure.

It is important to mention that apart from accuracy, full SLAM includes a number of additional important challenges such as loop-closure detection and subsequent map correction, re-localization, and long-term map maintenance (life-long mapping) – all of which are not evaluated with the proposed set-up.

2 Error Metric

For this step it is important, that both EE and SS contain sufficient poses in a non-degenerate configuration to well-constrain the alignment – hence the loopy motion patterns at the beginning and end of each sequence. The accumulated drift can now be computed as Tdrift=Tegt(Tsgt)−1∈Sim(3)T_{\text{drift}}=T_{e}^{\text{gt}}(T_{s}^{\text{gt}})^{-1}\in\text{Sim}(3), from which we can explicitly compute (a) the scale-drift es:=scale(Tdrift)e_{s}:=\text{scale}(T_{\text{drift}}), (b) the rotation-drift er:=rotation(Tdrift)e_{r}:=\text{rotation}(T_{\text{drift}}) and (c) the translation-drift et:=∥translation(Tdrift)∥e_{t}:=\|\text{translation}(T_{\text{drift}})\|.

We further define a combined error measure, the alignment error, which equally takes into account the error caused by scale, rotation and translation drift over the full trajectory as

which is the translational RMSE between the tracked trajectory, when aligned (a) to the start segment and (b) to the end segment. Figure 7 shows an example. We choose this metric, since

it can be applied to other SLAM / VO modes with different observability modes (like visual-inertial or stereo),

it is equally affected by scale, rotation, and translation drift, implicitly weighted by their effect on the tracked position,

it can be applied for algorithms which compute only poses for a subset of frames (e.g. keyframes), as long as start- and end-segment contain sufficient frames for alignment, and

it better reflects the overall accuracy of the algorithm than the translational drift drift ete_{t} or the joint RMSE

as shown in Figure 6. In particular, ermsee_{\text{rmse}} becomes degenerate for sequences where the accumulated translational drift surpasses the standard deviation of p^\hat{p}, since the alignment (15) will simply optimize to scale(T)≈0\text{scale}(T)\approx 0.

Benchmark

When evaluating accuracy of SLAM or VO methods, a common issue is that not all methods work on all sequences. This is particularely true for monocular methods, as sequences with degenerate (rotation-only) motion or entirely texture-less scenes (white walls) cannot be tracked. All methods will then either produce arbitrarily bad (random) estimates or heuristically decide they are “lost”, and not provide an estimate at all. In both cases, averaging over a set of results containing such outliers is not meaningful, since the average will mostly reflect the (arbitrarily bad) errors when tracking fails, or the threshold when the algorithm decides to not provide an estimate at all.

A common approach hence is to show only results on a hand-picked subset of sequences on which the compared methods do not fail (encouraging manual overfitting), or to show large tables with error values, which is not practicable for a dataset containing 50 sequences. A better approach is to summarize tracking accuracy as cumulative distribution, visualizing on how many sequences the error is below a certain threshold – it shows both the accuracy on sequences where a method works well, as well as the method’s robustness, i.e., on how many sequences it does not fail.

Figure 8 shows such cumulative error-plots for two methods, DSO (Direct Sparse Odometry) and ORB-SLAM , evaluated on the presented dataset. Each of the 50 sequences is run 5 times forwards and 5 times backwards, giving a total of 500 runs for each line shown in the plots.

Since both methods do not support the FOV camera model, we run the evaluation on pinhole-rectified images with VGA (640×480640\times 480) resolution. Further, we disable explicit loop-closure detection and re-localization to allow application of our metric, and reduce the threshold where ORB-SLAM decides it is lost to 10 inlier observations. Note that we do not impose any restriction on implicit, “small” loop-closures, as long as these are found by ORB-SLAM’s local mapping component (i.e., are included in the co-visibility graph). Since DSO does not perform loop-closure or re-localization, we can use the default settings. We run both algorithms in a non-real-time setting (at roughly one quarter speed), allowing to use 20 dedicated workstations with different CPUs to obtain the results presented in this paper. Figure 8 additionally shows results obtained when hard-enforcing real-time execution (dashed lines), obtained on the same workstation which is equipped with an i7-4910MQ CPU.

A good way to further to analyse the performance of an algorithm is to vary the sequences in a number of ways, simulating different real-world scenarios:

Figure 9 shows the tracking accuracy when rectifying the images to different fields of view, while keeping the same resolution (640×480640\times 480). Since the raw data has a resolution of 1280×10241280\times 1024, the caused distortion is negligible.

Figure 10 shows the tracking accuracy when rectifying the images to different resolutions, while keeping the same field of view.

Figure 11 shows the tracking accuracy when playing sequences only forwards compared to the results obtained when playing them only backwards – switching between predominantly forward-motion and predominantly backward-motion.

In each of the three figures, the bold lines correspond to the default parameter settings, which are the same across all evaluations. We further give some example results and corresponding video snippets in Table 1.

Since for most sequences, the used loop-closure ground truth is computed with a SLAM-algorithm (LSD-SLAM) itself, it is not perfectly accurate. We can however validate it by looking at the RMSE when aligning the start- and end-segment, i.e., the minima of (12) and (13) respectively. They are summarized in Figure 12. Note the difference in order of magnitude: The RMSE within start- and end-segment is roughly 100 times smaller than the alignment RMSE, and very similar for both evaluated methods. This is a strong indicator that almost all of the alignment error originates from accumulated drift, and not from noise in the ground truth.

1 Dataset

The full dataset, as well as preview-videos for all sequences are available on

Raw camera images of all 50 sequences (43GB; 190’000 frames in total), with frame-wise exposure times and computed ground truth alignment of start- and end-segment.

Geometric (FOV distortion model) and photometric calibrations (vignetting and response function).

Calibration datasets (13GB) containing (1) checkerboard-images for geometric calibration, (2) several sequences suitable for the proposed vignette and response function calibration, (3) images of a uniformly lit white paper.

Minimal c++ code for reading, pinhole-rectifying, and photometrically undistorting the images, as well as for performing photometric calibration as proposed in Section 2.3.

Matlab scripts to compute the proposed error metrics, as well as the raw tracking data for all runs used to create the plots and figures in Section 4.

2 Known Issues

Even though we use industry-grade cameras, the SDK provided by the manufacturer only allows to asynchronously query the current exposure time. Thus, in some rare cases, the logged exposure time may be shifted by one frame.

References