Continuous-time Intensity Estimation Using Event Cameras

Cedric Scheerlinck, Nick Barnes, Robert Mahony

Introduction

Event cameras respond asynchronously to changes in scene-illumination at the pixel level, offering high-temporal-resolution information over a large dynamic range . Conventional cameras typically acquire intensity image frames at fixed time-intervals, generating temporally-sparse, low-dynamic-range image sequences. Fusing image frames with the output of event cameras offers the opportunity to create an image state infused with high-temporal-resolution, high-dynamic-range properties of event cameras. This image state can be queried locally or globally at any user-chosen time-instance(s) for computer vision tasks such motion estimation, object recognition and tracking.

Event cameras produce events; discrete packets of information containing the timestamp, pixel-location and polarity of a brightness change . An event is triggered each time the change in log intensity at a pixel exceeds a preset threshold. The result is a continuous, asynchronous stream of events that encodes non-redundant information about local brightness changes. The DAVIS camera also provides low-frequency, full-frame intensity images in parallel. Alternative event cameras output direct brightness measurements with every event , and may also allow user-triggered, full-frame acquisition.

Due to high availability of contrast event cameras that output polarity (and not absolute brightness) with each event, such as the DAVIS, many researchers have tackled the challenge of estimating image intensity from contrast events . Image reconstruction algorithms that operate directly on the event stream typically perform spatio-temporal filtering , or take a spatio-temporal window of events and convert them into a discrete image frame . Windowing incurs a trade-off between length of time-window and latency. SLAM-like algorithms maintain camera-pose and image gradient (or 3D) maps that can be upgraded to full intensity via Poisson integration , however, so far these methods only work well for static scenes. Another image reconstruction algorithmic approach is to combine image frames directly with events . Beginning with an image frame, events are integrated to produce inter-frame intensity estimates. The estimate is reset with every new frame to prevent growth of integration error.

In this paper, we present a continuous-time formulation of event-based intensity estimation using complementary filtering to combine image frames with events (Fig. 1). We choose an asynchronous, event-driven update scheme for the complementary filter to efficiently incorporate the latest event information, eliminating windowing latency. Our approach does not depend on a motion-model, and works well in highly dynamic, complex environments. Rather than reset the intensity estimate with arrival of a new frame, our formulation retains the high-dynamic-range information from events, maintaining an image state with greater temporal resolution and dynamic range than the image frames. Our method also works well on a pure event stream without requiring image frames. The result is a continuous-time estimate of intensity that can be queried locally or globally at any user-chosen time.

We demonstrate our approach on datasets containing image frames and an event stream available from the DAVIS camera , and show that the complementary filter also works on a pure event stream without image frames. If available, synthetic frames reconstructed from events via an alternative algorithm can also be used as input to the complementary filter. Thus, our method can be used to augment any intensity reconstruction algorithm. Our approach can also be applied in a setup with any conventional camera co-located with an event camera. Additionally, we show how an adaptive gain can be used to improve robustness against under/overexposed image frames.

In summary, the key contributions of the paper are;

a continuous-time formulation of event-based intensity estimation,

a computationally simple, asynchronous, event-driven filter algorithm,

a methodology for pixel-by-pixel adaptive gain tuning.

We also introduce a new ground truth dataset for reconstruction of intensities from combined image frame and event streams, and make it publicly available. Sequences of images taken on a high-speed camera form the ground truth. We retain full frames at 20Hz, and convert the inter-frame images to an event stream. We compare state-of-the-art approaches on this dataset.

The paper is organised as follows: Section 2 describes related works. Section 3 summarises the mathematical representation and notation used, and characterises the full continuous-time solution of the proposed filter. Section 4 describes asynchronous implementation of the complementary filter, and introduces adaptive gains. Section 5 shows experimental results including the new ground truth dataset, and high-speed, high-dynamic-range sequences from the DAVIS. Section 6 concludes the paper.

Related Works

Event cameras such as the DVS and DAVIS provide asynchronous, data-driven contrast events, and are widely popular due to their commercial availability. Alternative cameras such as ATIS and CeleX are capable of providing absolute brightness with each event, but are not commercially available at the time of writing. Estimating image intensity from contrast events is important because it grants computer vision researchers a readily available high-temporal-resolution, high-dynamic-range imaging platform that can be used for tasks such as face-detection , SLAM , or optical flow estimation .

Image reconstruction from events is typically done by processing a spatio-temporal window of events . Barua et al. learn a patch-based dictionary from simulated event data, then use the learned dictionary to reconstruct gradient images from groups of consecutive event-images. Intensity is obtained using Poisson integration . Optical flow , together with the brightness constancy equation can be used to estimate intensity. Bardow et al. simultaneously optimise optical flow and intensity estimates within a fixed-length, sliding spatio-temporal window using the primal-dual algorithm . Taking a spatio-temporal window of events imposes a latency cost at minimum equal to the length of the time window, and choosing a time-interval (or event batch size) that works robustly for all types of scenes is not trivial. Reinbacher et al. integrate events over time while periodically regularising the estimate on a manifold defined by the timestamps of the latest events at each pixel; the surface of active events . An alternative approach is to estimate camera pose and map in a SLAM-like framework . Intensity can be recovered from the map, for example via Poisson integration of image gradients. These approaches work well for static scenes, but are not designed for dynamic scenes. Belbachir et al. reconstruct intensity panoramas from a pair of 1D stereo event cameras on a single-axis rotational device by filtering events (e.g. temporal high-pass filter).

Combining different sensing modalities (e.g. a conventional camera) with event cameras can overcome limitations of pure events, including lack of information about static or texture-less regions of the scene that do not trigger many events. Brandli et al. combine image frames and event stream from the DAVIS camera to create inter-frame intensity estimates by dynamically estimating the contrast threshold (temporal contrast) of each event. Each new image frame resets the intensity estimate, preventing excessive growth of integration error, but also discarding important accumulated event information. Shedligeri et al. use events to estimate ego-motion, then warp low frame-rate images to intermediate locations. Liu et al. use affine motion models to reconstruct video of high-speed foreground/static background scenes.

We introduce the concept of a continuous-time image state that is asynchronously updated with every event. Our method is motion-model free and can be used with events only, or with image frames to complement events.

Approach

Let Y(p,t)Y(\boldsymbol{p},t) denote the intensity or irradiance of pixel p\boldsymbol{p} at time tt of a camera. We will assume that the same irradiance is observed at the same pixel in both a classical and event camera, such as is the case with the DAVIS camera . A classical image frame (for a global shutter camera) is an average of the received intensity over the exposure time

where tjt_{j} is the time-stamp of the image capture and ϵ\epsilon is the exposure time. In the sequel we will ignore the exposure time in the analysis and simply consider a classical image as representing image information available at time tjt_{j}. Although there will be image blur effects, especially for fast moving scenes in low light conditions (see the experimental results in §5), a full consideration of these effects is beyond the scope of the present paper.

The approach taken in this paper is to analyse image reconstruction for event cameras in the continuous-time domain. To this end, we define a continuous-time intensity signal YF(p,t)Y^{F}(\boldsymbol{p},t) as the zero-order hold (ZOH) reconstruction of the irradiance from the classical image frames:

Since event cameras operate with log intensity we convert the image intensity into log-intensity:

Note that converting the zero-hold signal into the log domain is not the same as integrating the log intensity of the irradiance over the shutter time. We believe the difference will be insignificant in the scenarios considered and we do not consider this further in the present paper.

Dynamic vision sensors (DVS), or event cameras, are biologically-inspired vision sensors that respond to changes in scene illumination. Each pixel is independently wired to continuously compare the current log intensity level to the last reset-level. When the difference in log intensity exceeds a predetermined threshold (contrast threshold), an event is transmitted and the pixel resets, storing the new illumination level. Each event contains the pixel coordinates, timestamp, and polarity (σ=±1\sigma=\pm 1 for increasing or decreasing intensity). An event can be modelled in the continuous-timeNote that events are continuous-time signals even though they are not continuous functions of time; the time variable tt on which they depend varies continuously. signal class as a Dirac-delta function δ(t)\delta(t). We define an event stream ei(p,t)e_{i}(\boldsymbol{p},t) at pixel p\boldsymbol{p} by

where σip\sigma_{i}^{\boldsymbol{p}} is the polarity and tipt_{i}^{\boldsymbol{p}} is the time-stamp of the ithi^{th} event at pixel p\boldsymbol{p}. The magnitude cc is the contrast threshold (brightness change encoded by one event). Define an event field E(p,t)E(\boldsymbol{p},t) by

The event field is a function of all pixels p\boldsymbol{p} and ranges over all time, capturing the full output of the event camera.

A quantised log intensity signal LE(p,t)L^{E}(\boldsymbol{p},t) can be reconstructed by integrating the event field

The result is a series of log intensity steps (corresponding to events) at each pixel. In the absence of noise, the relationship between the log-intensity L(p,t)L(\boldsymbol{p},t) and the quantised signal LE(p,t)L^{E}(\boldsymbol{p},t) is

where L(p,0)L(\boldsymbol{p},0) is the initial condition and μ(p,t;c)\mu(\boldsymbol{p},t;c) is the quantisation error. Unlike LF(p,t)L^{F}(\boldsymbol{p},t), the quantisation error associated with LE(p,t)L^{E}(\boldsymbol{p},t) is bounded by the contrast threshold; ∣μ(p,t;c)∣<c|\mu(\boldsymbol{p},t;c)|<c.

Events can be interpreted as the temporal derivative of LE(p,t)L^{E}(\boldsymbol{p},t)

2 Complementary Filter

We will use a complementary filter structure to fuse the event field E(p,t)E(\boldsymbol{p},t) with ZOH log-intensity frames LF(p,t)L^{F}(\boldsymbol{p},t). Complementary filtering is ideal for fusing signals that have complementary frequency noise characteristics; for example, where one signal is dominated by high-frequency noise and the other by low-frequency disturbance. Events are a temporal derivative measurement (10) and do not contain reference intensity L(p,0)L(\boldsymbol{p},0) information. Integrating events to obtain LE(p,t)L^{E}(\boldsymbol{p},t) amplifies low-frequency disturbance (drift), resulting in poor low-frequency information. However, due to their high-temporal-resolution, events provide reliable high-frequency information. Classical image frames LF(p,t)L^{F}(\boldsymbol{p},t) are derived from discrete, temporally-sparse measurements and have poor high-frequency fidelity. However, frames typically provide reliable low-frequency reference intensity information. The proposed complementary filter architecture combines a high-pass version of LE(p,t)L^{E}(\boldsymbol{p},t) with a low-pass version of LF(p,t)L^{F}(\boldsymbol{p},t) to reconstruct an (approximate) all-pass version of L(p,t)L(\boldsymbol{p},t).

The proposed filter is written as a continuous-time ordinary differential equation (ODE)

where L^(p,t)\hat{L}(\boldsymbol{p},t) is the continuous-time log-intensity state estimate and α\alpha is the complementary filter gain, or crossover frequency (Fig. 1).

The fact that the input signals in (11) are discontinuous poses some complexities in solving the filter equations, but does not invalidate the formulation. The filter can be understood as integration of the event field with an innovation term -\alpha\big{(}\hat{L}(\boldsymbol{p},t)-L^{F}(\boldsymbol{p},t)\big{)}, that acts to reduce the error between L^(p,t)\hat{L}(\boldsymbol{p},t) and LF(p,t)L^{F}(\boldsymbol{p},t).

The key property of the proposed filter (11) is that although it is posed as a continuous-time ODE, one can express the solution as a set of asynchronous-update equations. Each pixel acts independently, and in the sequel we will consider the action of the complementary filter on a single pixel p\boldsymbol{p}. Recall the sequence {tip}\{t_{i}^{\boldsymbol{p}}\} corresponding to the time-stamps of all events at p\boldsymbol{p}. In addition, there is the sequence of classical image frame time-stamps {tj}\{t_{j}\} that apply to all pixels equally. Consider a combined sequence of monotonically increasing unique time-stamps t^kp\hat{t}_{k}^{\boldsymbol{p}} corresponding to event {tip}\{t_{i}^{\boldsymbol{p}}\} or frame {tj}\{t_{j}\} time-stamps.

Within a time-interval t∈[t^kp,t^k+1p)t\in[\hat{t}_{k}^{\boldsymbol{p}},\hat{t}_{k+1}^{\boldsymbol{p}}) there are (by definition) no new events or frames, and the ODE (11) is a constant coefficient linear ordinary differential equation

It remains to paste together the piece-wise smooth solutions on the half-open intervals [t^kp,t^k+1p)[\hat{t}_{k}^{\boldsymbol{p}},\hat{t}_{k+1}^{\boldsymbol{p}}) by considering the boundary conditions. Let

denote the limits from below and above. There are two cases to consider:

New frame: When the index t^k+1p\hat{t}_{k+1}^{\boldsymbol{p}} corresponds to a new image frame then the right hand side (RHS) of (11) has bounded variation. It follows that the solution is continuous at t^k+1p\hat{t}_{k+1}^{\boldsymbol{p}} and

Event: When the index t^k+1p\hat{t}_{k+1}^{\boldsymbol{p}} corresponds to an event then the solution of (11) is not continuous at t^k+1p\hat{t}_{k+1}^{\boldsymbol{p}} and the Dirac delta function of the event must be integrated. Integrating the RHS and LHS of (11) over an event

yields a unit step scaled by the contrast threshold and sign of the event. Note the integral of the second term \int_{(\hat{t}_{k+1}^{\boldsymbol{p}})^{-}}^{(\hat{t}_{k+1}^{\boldsymbol{p}})^{+}}\alpha\big{(}\hat{L}(\boldsymbol{p},\tau)-L^{F}(\boldsymbol{p},\tau)\big{)}d\tau is zero since the integrand is bounded. We use the solution

as initial condition for the next time-interval. Eqns. (13), (16) and (19) characterise the full solution to the filter equation (11).

The filter can be run on events only without image frames by setting LF(p,t)=0L^{F}(\boldsymbol{p},t)=0 in (11), resulting in a high-pass filter with corner frequency α\alpha

This method can efficiently generate a good quality image state estimate from pure events. Furthermore, it is possible to use alternative pure event-based methods to reconstruct a temporally-sparse image sequence from events and fuse this with raw events using the proposed complementary filter. Thus, the proposed filter can be considered a method to augment any event-based image reconstruction method to obtain a high temporal-resolution image state.

Method

The complementary filter gain α\alpha is a parameter that controls the relative information contributed by image frames or events. Reducing the magnitude of α\alpha decreases the dependence on image frame data while increasing the dependence on events (α=0→L^(p,t)=LE(p,t)\alpha=0\to\hat{L}(\boldsymbol{p},t)=L^{E}(\boldsymbol{p},t)). A key observation is that the gain can be time-varying at pixel-level (α=α(p,t)\alpha=\alpha(\boldsymbol{p},t)). One can therefore use α(p,t)\alpha(\boldsymbol{p},t) to dynamically adjust the relative dependence on image frames or events, which can be useful when image frames are compromised, e.g. under- or overexposed.

We propose to reduce the influence of under/overexposed image frame pixels by decreasing α(p,t)\alpha(\boldsymbol{p},t) at those pixel locations. We use the heuristic that pixels reporting an intensity close to the minimum LminL_{\text{min}} or maximum LmaxL_{\text{max}} output of the camera may be compromised, and we decrease α(p,t)\alpha(\boldsymbol{p},t) based on the reported log intensity. We choose two bounds L1L_{1}, L2L_{2} close to LminL_{\text{min}} and LmaxL_{\text{max}}, then we set α(p,t)\alpha(\boldsymbol{p},t) to a constant (α1\alpha_{1}) for all pixels within the range [L1L_{1}, L2L_{2}], and linearly decrease α(p,t)\alpha(\boldsymbol{p},t) for pixels outside of this range:

where λ\lambda is a parameter determining the strength of our adaptive scheme (we set λ=0.1\lambda=0.1). For α1\alpha_{1}, typical suitable values are α1∈[0.1,10]\alpha_{1}\in[0.1,10] rad/s. For our experiments we choose α1=2π\alpha_{1}=2\pi rad/s.

2 Asynchronous Update Scheme

Given temporally sparse image frames and events, and using the continuous-time solution to the complementary filter ODE (11) outlined in §3 one may compute the intensity state estimate L^(p,t)\hat{L}(\boldsymbol{p},t) at any time. In practice it is sufficient to compute the image state L^(p,t)\hat{L}(\boldsymbol{p},t) at the asynchronous time instances t^kp\hat{t}_{k}^{\boldsymbol{p}} (event or frame timestamps). We propose an asynchronous update scheme whereby new events cause state updates (19) only at the event pixel-location. New frames cause a global update (16) (note this is not a reset as in ) The filter can also be updated (using (13)) at any user-chosen time instance (or rate). In our experiments we update the entire image state whenever we export the image for visualisation. . Algorithm 1 describes a per-pixel complementary filter implementation. At a given pixel p\boldsymbol{p}, let L^⋄\hat{L}_{\diamond} denote the latest estimate of L^(p,t)\hat{L}(\boldsymbol{p},t) stored in computer memory, and t^⋄\hat{t}_{\diamond} denote the time-stamp of the latest update at p\boldsymbol{p}. Let L⋄FL^{F}_{\diamond} and α⋄\alpha_{\diamond} denote the latest image frame and gain values at p\boldsymbol{p}. To run the filter in events only mode (high-pass filter (20)), simply let L⋄F=0L^{F}_{\diamond}=0.

Results

We compare the reconstruction performance of our complementary filter, both with frames (CFf\text{CF}_{\text{f}}) and without frames in events only mode (CFe\text{CF}_{\text{e}}), against three state-of-the-art methods: manifold regularization (MR) , direct integration (DI) and simultaneous optical flow and intensity estimation (SOFIE) . We introduce new datasets: two new ground truth sequences (Truck and Motorbike); and four new sequences taken with the DAVIS240C camera (Night drive, Sun, Bicycle, Night run). We evaluate our method, MR and DI against our ground truth dataset using quantitative image similarity metrics. Unfortunately, as code is not available for SOFIE we are unable to evaluate its performance on our new datasets. Hence, we compare it with our method on the jumping sequence made available by the authors.

Ground truth is obtained using a high-speed, global-shutter, frame-based camera (mvBlueFOX USB 2) running at 168Hz. We acquire image sequences of dynamic scenes (Truck and Motorbike), and convert them into events following the methodology of Mueggler et al. . To simulate event camera noise, a number of random noise events are generated (5% of total events), and distributed randomly throughout the event stream. To simulate low-dynamic-range, low-temporal-resolution input-frames, the upper and lower 25% of the maximum intensity range is truncated, and image frames are subsampled at 20Hz. In addition, a delay of 50ms is applied to the frame time-stamps to simulate the latency associated with capturing images using a frame-based camera.

The complementary filter gain α(p,t)\alpha(\boldsymbol{p},t) is set according to (21) and updated with every new image frame (Algorithm 1). We set α1=2π \alpha_{1}=2\pi\,rad/s for all sequences. The bounds [L1, L2][L_{1},\ L_{2}] in (21) are set to [Lmin+κ,[L_{\text{min}}+\kappa, Lmax−κ]L_{\text{max}}-\kappa], where κ=0.05(Lmax−Lmin)\kappa=0.05(L_{\text{max}}-L_{\text{min}}). The contrast threshold (cc) is not easy to determine and in practice varies across pixels, and with illumination, event-rate and other factors . Here we assume two constant contrast thresholds (ON and OFF) that are calibrated for each sequence using APS frames. We note that error arising from the variability of contrast thresholds appears as noise in the final estimate, and believe that more sophisticated contrast threshold models may benefit future works. For MR , the number of events per output image (events/image) is a parameter that impacts the quality of the reconstructed image. For each sequence we choose events/image to give qualitatively best performance. We set events/image to 1500 unless otherwise stated. All other parameters are set to defaults provided by .

Night drive (Fig. 2) investigates performance in high-speed, low light conditions where the conventional camera image frame (Raw frame) is blurry and underexposed, and dark details are lost. Data is recorded through the front wind shield of a car, driving down an unlit highway at dead of night. Our method (CF) is able to recover motion-blurred objects (e.g. roadside poles), trees that are lost in Raw frame, and road lines that are lost in MR. MR relies on spatial smoothing to reduce noise, hence faint features such as distant trees (Fig. 2; zoom) may be lost. DI loses features that require more time for events to accumulate (e.g. trees on the right), because the estimate is reset upon every new image frame. The APS (active pixel sensor in DAVIS) frame-rate was set to 7Hz.

Sun investigates extreme dynamic range scenes where conventional cameras become overexposed. CF recovers features such as leaves and twigs, even when the camera is pointed directly at the sun. Raw frame is largely over-saturated, and the black dot (Fig. 2; Sun) is a camera artifact caused by extreme brightness, and marks the position of the sun. MR produces a clean looking image, though some features (small leaves/twigs) are smoothed out (Fig. 2; zoom). DI is washed out in regions where the frame is overexposed, due to the latest frame reset. Because the sun generates so many events, MR requires more events/image to recover fine features (with less events the image looks oversmoothed), so we increase events/image to 2500. The APS frame-rate was set to 26Hz.

Bicycle explores the scenario of static background, moving foreground. Raw frame is underexposed in shady areas because of large intra-scene dynamic range. When the event camera is stationary, almost no events are generated by the static background and it cannot be recovered by pure event-based reconstruction methods such as MR and CFe\text{CF}_{\text{e}}. In contrast, CFf\text{CF}_{\text{f}} recovers both stationary and non-stationary features, as well as high-dynamic-range detail. The APS frame-rate was set to 26Hz.

Night run illustrates the benefit in challenging low-light pedestrian scenarios. Here a pedestrian runs across the headlights of a (stationary) car at dead of night. Raw frame is not only heavily motion-blurred, but also significantly delayed, since a large exposure duration is required for image acquisition in low-light conditions. DI is unreliable as an unfortunately timed new image frame could reset the image (Fig. 2). MR and CF manage to recover the pedestrian, and CFf\text{CF}_{\text{f}} also recovers the background without compromising clarity of the pedestrian. The APS frame-rate was set to 4.5Hz.

Ground truth evaluation. We evaluate our method with (CFf\text{CF}_{\text{f}}) and without (CFe\text{CF}_{\text{e}}) 20Hz input-frames, and compare against DI and MR . To assess similarity between ground truth and reconstructed images, each ground truth frame is matched with the corresponding reconstructed image with the closest time-stamp. Average absolute photometric error (%), structural similarity (SSIM) , and feature similarity (FSIM) are used to evaluate performance (Fig. 3 and Table 1). We initialise DI and CFf\text{CF}_{\text{f}} using the first input-frame, and MR and CFe\text{CF}_{\text{e}} to zero.

Fig. 3 plots the performance of each reconstruction method over time. Our method shows an initial improvement as useful information starts to accumulate, then maintains good performance over time as new events and frames are incorporated into the estimate. The oscillations apparent in DI arise from image frame resets. Table 1 summarises average performance for each sequence. Our method CFf\text{CF}_{\text{f}} achieves the lowest photometric error, and highest SSIM and FSIM scores for all sequences. Fig. 4 shows the reconstructed image halfway between two input-frames of Truck ground truth sequence. Pure event-based methods (MR and CFe\text{CF}_{\text{e}}) do not recover absolute intensity in some regions (truck body) due to sparsity of events. DI displays artifacts around edges, where many events are generated, because events are directly added to the latest input-frame. In CFf\text{CF}_{\text{f}}, event and frame information is continuously combined, reducing edge artifacts (Fig. 4) and producing a more consistent estimate over time (Fig. 3).

SOFIE. The code for SOFIE was not available at the time of writing, however the authors kindly share their pre-recorded dataset (using DVS128 ) and results. We use their dataset to compare our method to SOFIE (Fig. 5), and since no camera frames are available, we first demonstrate our method by setting input-frames to zero (CFe\text{CF}_{\text{e}}), then show that reconstructed image frames output from an alternative reconstruction algorithm such as SOFIE can be used as input-frames to the complementary filter (CFf\text{CF}_{\text{f}} (events + SOFIE)) to generate intensity estimates.

Conclusion

We have presented a continuous-time formulation for intensity estimation using an event-driven complementary filter. We compare complementary filtering with existing reconstruction methods on sequences recorded on a DAVIS camera, and show that the complementary filter outperforms current state-of-the-art on a newly presented ground truth dataset. Finally, we show that the complementary filter can estimate intensity based on a pure event stream, by either setting the input-frame signal to zero, or by fusing events with the output of a computationally intensive reconstruction method.

References