Photorealistic Image Reconstruction from Hybrid Intensity and Event based Sensor

Prasan A Shedligeri, Kaushik Mitra

Introduction

Event-based sensors encode the local contrast changes in the scene as positive or negative events at the instant they occur. Event-based sensors provide a power efficient way of converting the megabytes of per-pixel intensity data into a stream of spatially sparse but temporally dense events. However, the event stream cannot be directly visualized like a normal video, with which we as human beings are familiar with. This calls for an algorithm that can convert this stream of event data to a more familiar version of image frames. These reconstructed intensity frames could also be used as an input for traditional frame-based computer vision algorithms like multi-view stereo, object detection etc. Previous attempts at converting the event stream into images have heavily relied on event data. Although these methods do a good job of recovering the intensity frames they suffer from two major disadvantages: a) The intensity frames don’t look photorealistic and b) some of the objects in the scene can go missing in the recovered frames because they are not producing any events (edges parallel to the sensor motion do not trigger any events).

In this paper, we propose a method to reconstruct photorealistic intensity images at a high frame rate. As the absolute intensity and fine texture information is lost during the encoding of events, we use the conventional image sensor to provide us with this information for reconstructing photorealistic images. There exists a commercially available hybrid sensor consisting of a co-located low-frame rate intensity sensor and an event-based sensor called DAVIS(Dynamic and Active pixel Vision Sensor). Fig. 2 summarizes our overall approach to reconstruct the temporally dense photorealistic intensity images using the hybrid sensor. We mainly have four steps. In the first step, we estimate the dense depth map from successive intensity frames. For this purpose, we use a traditional iterative optimization scheme, which we initialize by depth map obtained from a deep learning based optical flow estimation algorithm. In the second step, we map the event data between successive intensity frames to multiple pseudo-intensity frames using . Next, we use the pseudo-intensity frames and the dense depth maps obtained from the first step to estimate temporally dense camera ego-motion by direct visual odometry. And finally, in the fourth step, we warp the successive intensity frames to intermediate temporal locations of the pseudo-intensity frames to obtain photo-realistic reconstruction. Recently, have shown that it is possible to fuse temporally dense events with low frame-rate intensity frames to reconstruct intensity frames at a higher frame rate. However, due to lack of any regularization, the intensity frames reconstructed using tend to be noisy and blurry. With extensive experiments, we show that our proposed method is able to reconstruct photorealistic intensity images at a high frame rate and is also robust to noisy events in the event stream. To summarize, we make the following contributions:

We propose a pipeline using a hybrid event and low frame rate intensity sensor which can reconstruct temporally dense photorealistic intensity images. This would be difficult to obtain with only either the conventional image sensor or the event sensor.

We use the event sensor for estimating temporally dense sensor ego-motion and the low-frame rate intensity images for obtaining spatially dense depth map.

We demonstrate high quality temporally dense photorealistic reconstructions using the proposed method on real data captured from DAVIS.

We also demonstrate our algorithm’s robustness to abrupt camera motion and noisy events in the event sensor data.

Related work

Intensity image reconstruction from events: The proposed work is very closely related to other previous works which reconstruct intensity images from events . cannot recover the true intensity information of the scene as they use only the events to estimate the intensity images. Some works like reconstruct intensity images as a by-product of sensor tracking from event data over 3D scenes but are not able to recover the true intensity information. Recently, demonstrated that event data and the intensity image data can be used in a complementary filter to reconstruct intensity frames at a higher frame rate. Although makes use of the intensity images, the reconstructed images tend to be blurry and are adversely affected by noisy events due to lack of any regularization in their proposed method. The pre-print version of this work is also available on arxiv .

Visual odometry/SLAM with event sensors: The high temporal data acquisition of event sensors has made it extremely suitable for applications such as tracking which need low latency operation. Very recently, a dataset was also proposed to benchmark event based pose estimation, visual odometry, and SLAM algorithms . The dataset contains multiple video sequences captured with DAVIS and sub-millimeter accurate ground truth camera motion acquired using a motion-capture system. Previous works such as estimate ego-motion of the sensor directly from the event stream. Visual odometry/ SLAM with event sensors has also been a very popular topic of research. Although we estimate the scene depth and sensor ego-motion to warp intensity frames, visual odometry/SLAM is not the focus of this work.

Photorealistic image reconstruction

We propose to reconstruct photorealistic intensity images using the event stream obtained from an event sensor. The conventional image sensor will compensate for the fine texture and the absolute intensity information which is lost in the event stream. As can be seen from Fig. 2, we have four major steps to reconstruct the temporally dense photorealistic image reconstruction: (a) Estimate dense depth maps dkd_{k} and dk+1d_{k+1} corresponding to the successive intensity frames IkI_{k} and Ik+1I_{k+1} and the relative pose ξ\xi between them (§3.1); (b) Reconstruct pseudo-intensity frames EkjE_{k}^{j} at uniformly spaced temporally dense locations j=1,2,…Nj=1,2,\ldots N between every successive intensity frame IkI_{k} and Ik+1I_{k+1}; (c) Estimate temporally dense sensor ego-motion estimates ξkj\xi_{k}^{j} and ξk+1j\xi_{k+1}^{j} for each intermediate pseudo-intensity frame with respect to the intensity frames IkI_{k} and Ik+1I_{k+1}(§3.2) and (d)Forward warp the intensity frames IkI_{k} and Ik+1I_{k+1} to the intermediate location of each of the pseudo-intensity frames EkjE_{k}^{j} and blend them (§3.3).

One of the important steps in our proposed algorithm is forward warping the intensity images to multiple intermediate temporal locations between successive intensity frames. However, warping can introduce undesired holes in the final reconstructed images at regions of disocclusion. This can be solved by warping both the successive intensity frames, IkI_{k} and Ik+1I_{k+1}, to the intermediate locations. This requires us to estimate two dense depth maps dkd_{k} and dk+1d_{k+1} corresponding to the images IkI_{k} and Ik+1I_{k+1}, respectively. Fig. 4 shows the overall scheme of estimating dense depth maps from successive intensity frames. We initialize the depth estimates dkd_{k} and dk+1d_{k+1} from optical flow, and the 6-dof camera pose ξ\xi with zero rotation and translation. Here, ξ\xi is the 6-dof relative camera pose at Ik+1I_{k+1} with respect to IkI_{k}. We warp the intensity image Ik+1I_{k+1} to the location of IkI_{k} with the current estimate of dkd_{k} and ξ\xi, to obtain I^k\hat{I}_{k}. Similarly, we warp IkI_{k} to the location of Ik+1I_{k+1} to obtain I^k+1\hat{I}_{k+1}. We define the photometric reconstruction loss Lph\mathcal{L}_{ph} as,

By minimizing the above reconstruction loss, Lph\mathcal{L}_{ph}, it is possible to estimate the depth maps dkd_{k} and dk+1d_{k+1} and 6-dof relative pose ξ\xi. We also enforce an edge aware laplacian smoothness prior on the estimated depth maps dkd_{k} and dk+1d_{k+1}, by taking inspiration from . We define the smoothness loss Lsm\mathcal{L}_{sm} as,

where II is the intensity image, dd is the corresponding dense depth map and ∇x\nabla_{x} and ∇y\nabla_{y} are the x and y-gradient operators, respectively. Overall, we estimate the dense depth estimate dkd_{k}, dk+1d_{k+1} and the relative pose ξ\xi by,

Eq. (3) is a non-convex optimization problem and hence a good initialization of depth and pose is essential to avoid local minima. Here, we use optical flow between the successive intensity frames obtained from PWC-Net as an initial estimate of the depth. For a static scene, it is possible to estimate the scene depth and the 6-dof camera pose from the optical flow. However, in our experiments, we found that a simple inverse of the optical flow magnitude is good enough to initialize the depth for the optimization iterations in Eq. (3). We initialize pose with zero rotation and translation.

2 6-dof relative pose estimation by direct matching

To achieve the goal of photorealistic reconstruction we warp the successive intensity frames captured by the image sensor to the intermediate temporal location of an event frame. For warping, we need to determine the 6-dof camera pose between the temporal locations of the successive intensity frames and that of the intermediate event frames. We reconstruct pseudo-intensity images from events using at the temporal locations of the intermediate event frames as well as the successive intensity frames. As shown in Fig. 4, our goal here is to estimate the relative camera pose between Ek0E_{k}^{0}, Ek+10E_{k+1}^{0} and the pseudo-intensity images EkjE_{k}^{j} (j=1,2,…Nj=1,2,\ldots N). We use this relative pose, to warp the successive intensity frames to the intermediate locations specified by the event frames (EkjE_{k}^{j}) and hence reconstruct photorealistic intensity images.

Let ξkj\xi_{k}^{j} represent the 6-dof camera pose of the intermediate pseudo-intensity image EkjE^{j}_{k} with respect to Ek0E^{0}_{k} and ξk+1j\xi_{k+1}^{j} be the 6-dof camera pose of EkjE^{j}_{k} with respect to Ek+10E^{0}_{k+1}. We use the current estimate of relative camera pose ξkj\xi_{k}^{j} and the known depth estimate dkd_{k} to inverse warp the pseudo-intensity frame EkjE^{j}_{k} to the location of Ek0E^{0}_{k} to obtain E^k0\hat{E}^{0}_{k}. We similarly inverse warp the pseudo-intensity frame EkjE^{j}_{k} to the location of Ek+10E^{0}_{k+1} to obtain E^k+10\hat{E}^{0}_{k+1} using the current estimate of relative pose ξk+1j\xi_{k+1}^{j} and the known depth dk+1d_{k+1}. We define the photometric loss Lp\mathcal{L}_{p} as mean absolute error between the warped intensity frame and the ground truth frame.

By composing the relative pose estimates, ξkj\xi_{k}^{j} and (ξk+1j)−1(\xi_{k+1}^{j})^{-1} we obtain the overall pose between IkI_{k} and Ik+1I_{k+1}. We use this knowledge to regularize the relative camera pose estimates ξkj\xi_{k}^{j} and ξk+1j\xi_{k+1}^{j} with Lp(ξkj,ξk+1j)=∥Ik−I^k∥1\mathcal{L}_{p}(\xi_{k}^{j},\xi_{k+1}^{j})=\|I_{k}-\hat{I}_{k}\|_{1} . Overall,

where λr\lambda_{r} is the regularization parameter.

3 Forward Warping and Blending

At this stage, we have depth maps dkd_{k} and dk+1d_{k+1} corresponding to intensity images IkI_{k} and Ik+1I_{k+1} respectively. We do a source-target mapping (forward warping) from two images IkI_{k} and Ik+1I_{k+1} using the estimated relative pose ξkj\xi_{k}^{j} and ξk+1j\xi_{k+1}^{j}to the latent image IkjI_{k}^{j} and alpha-blend them. We splat the intensity values after forward warping to ensure that no holes are generated in the final image.

Experiments

For all our experiments we use DAVIS , which is commercially available and has a conventional image sensor and an event sensor bundled together. Since, we did not have access to DAVIS, we used the recently proposed dataset by and which consists of several video sequences captured using DAVIS. We obtain dense depth maps at the locations of low frame rate intensity frames and temporally dense sensor ego-motion using the event sensor data to warp the low frame-rate intensity frames to intermediate camera locations. For estimating depth, we initially enhance the edges of the depth obtained from optical flow estimate using a fast bilateral solver . The output of this bilateral solver is then used as an initialization for the iterative depth refinement scheme. We set β=10.0\beta=10.0 in Eq. (2) and λsm=1.0\lambda_{sm}=1.0 in Eq. (3). Using the event stream from each sequence in the dataset we generate pseudo-intensity estimates using the algorithm proposed in . We stack non-overlapping blocks of 2000 events into a frame and generate a corresponding pseudo-intensity frame using . These pseudo-intensity frames are then used for estimating the temporally dense sensor ego-motion. For pose estimation we use λr=0.01\lambda_{r}=0.01 in Eq. (6). We use the Adam optimizer to solve Eq. (3) and Eq. (6).

In Fig. 7 we demonstrate the effectiveness of our proposed method for estimating depth. We use an initial estimate of depth from a deep learning method and iteratively refine it. We empirically found that using PWC-Net to initialize the depth estimate for the iterative optimization scheme gave consistently good results. We also experiment with initializing the depth from . We provide comparisons in supplementary material.

2 Photorealistic intensity image reconstruction

In Fig. 1 and Fig. 7 we compare qualitatively the intensity images reconstructed using our proposed method to that proposed in . While MR utilizes only event sensor data, CF uses both event sensor data as well as information from intensity images. For fairness in comparison, we generate intensity images from for every 20002000 events in the sequence. In , we found that initializing the cut-off frequency to 6.28rad/s6.28rad/s and updating other parameters dynamically gave the best results.

Reinbacher et al. use only event information and are hence unable to recover the true intensity information present in the scene. Scheerlinck et al. do not use any kind of spatial regularization and hence the reconstructed images are noisy and blurry even though they have access to the intensity images. We acknowledge that run in real time, while our algorithm takes about two minutes to estimate the dense depth maps and about 40 seconds to render each intermediate frame. With recent advances in stereo depth estimation methods, we expect that in future we can completely eliminate the need for an iterative depth refinement scheme and directly use the output of a state-of-the-art stereo depth estimation algorithm. This will greatly reduce the computation time. It is possible to further reduce the computation time for estimating pose by using the Lucas-Kanade inverse compositional method.

3 Robustness to abrupt camera motion

In the case of abrupt motion of the sensor, the intensity images get blurred and the rate at which events are generated becomes high. We start with deblurring the intensity images using an existing deblurring technique(in our experiments we used ). These deblurred images are then used as an input to the reconstruction pipeline. Abrupt motion results in a high event rate and also produces many noisy events. These noisy events affect the reconstructions in as their trust on events increases exponentially with the rise in the event rate. As can be seen in Fig. 7, our method is robust to such abrupt motions as can be seen from the results shown in columns (b) and (c).

Conclusion

We combine the strength of texture-rich low frame rate intensity frames with high temporal rate event data to obtain temporally dense photo-realistic images. We achieve this by warping the low frame rate intensity frames from the conventional image sensor to intermediate locations. With extensive experiments, we have demonstrated that the images reconstructed from our algorithm are photorealistic compared to any of the previous methods. We also show the robustness of our algorithm to abrupt camera motion. Currently, our algorithm assumes a static scene. A future direction for us would be to build a generalized algorithm which can reconstruct photorealistic images for dynamic scenes as well.

References