Focus Is All You Need: Loss Functions For Event-based Vision

Guillermo Gallego, Mathias Gehrig, Davide Scaramuzza

Introduction

Motion compensation approaches have been recently introduced for processing the visual information acquired by event cameras. They have proven successful for the estimation of motion (optical flow) , camera motion , depth (3D reconstruction) as well as segmentation . The main idea of such methods consists of searching for point trajectories on the image plane that maximize event alignment (Fig. Focus Is All You Need: Loss Functions For Event-based Vision, right), which is measured using some loss function of the events warped according to such trajectories. The best trajectories produce sharp, motion compensated images that reveal the brightness patterns causing the events (Fig. Focus Is All You Need: Loss Functions For Event-based Vision, middle).

In this work, we build upon the motion compensation framework and extend it to include twenty more loss functions for applications such as ego-motion, depth and optical flow estimation. We ask the question: What are good metrics of event alignment? In answering, we noticed strong connections between the proposed metrics (Table 1) and those used for shape-from-focus and autofocus in conventional, frame-based cameras , and so, we called the event alignment metrics “focus loss functions”. The extended framework allows mature computer vision tools to be used on event data while taking into account all the information of the events (asynchronous timestamps and polarity). Additionally, it sheds light on the event-alignment goal of functions used in existing motion-compensation works and provides a taxonomy of loss functions for event data.

The introduction and comparison of twenty two focus loss functions for event-based processing, many of which are developed from basic principles, such as the “area” of the image of warped events.

Connecting the topics of shape-from-focus, autofocus and event-processing by the similar set of functions used, thus allowing to bring mature analysis tools from the former topics into the realm of event cameras.

A thorough evaluation on a recent dataset , comparing the accuracy and computational effort of the proposed focus loss functions, and showing how they can be used for depth and optical flow estimation.

The rest of the paper is organized as follows. Section 2 reviews the working principle of event cameras. Section 3 summarizes the motion compensation method and extends it with the proposed focus loss functions. Experiments are carried out in Section 4 comparing the loss functions, and conclusions are drawn in Section 5.

Event-based Camera Working Principle

Event-based cameras, such as the DVS , have independent pixels that output “events” in response to brightness changes. Specifically, if L(x,t)≐log⁡I(x,t)L(\mathbf{x},t)\doteq\log I(\mathbf{x},t) is the logarithmic brightness at pixel x≐(x,y)⊤\mathbf{x}\doteq(x,y)^{\top} on the image plane, the DVS generates an event ek≐(xk,tk,pk)e_{k}\doteq(\mathbf{x}_{k},t_{k},p_{k}) if the change in logarithmic brightness at pixel xk\mathbf{x}_{k} reaches a threshold CC (e.g., 10-15% relative change):

where tkt_{k} is the timestamp of the event, Δtk\Delta t_{k} is the time since the previous event at the same pixel xk\mathbf{x}_{k} and pk∈{+1,−1}p_{k}\in\{+1,-1\} is the event polarity (i.e., sign of the brightness change).

Therefore, each pixel has its own sampling rate (which depends on the visual input) and outputs data proportionally to the amount of motion in the scene. An event camera does not produce images at a constant rate, but rather a stream of asynchronous, sparse events in space-time (Fig. Focus Is All You Need: Loss Functions For Event-based Vision, left).

Methodology

In short, the method in seeks to find the point-trajectories on the image plane that maximize the alignment of corresponding events (i.e., those triggered by the same scene edge). Event alignment is measured by the strength of the edges of an image of warped events (IWE), which is obtained by aggregating events along candidate point trajectories (Fig. Focus Is All You Need: Loss Functions For Event-based Vision). In particular, proposes to measure edge strength (which is directly related to image contrast ) using the variance of the IWE.

More specifically, the events in a set E={ek}k=1Ne\mathcal{E}=\{e_{k}\}_{k=1}^{N_{e}} are geometrically transformed

according to a point-trajectory model W\mathbf{W}, resulting in a set of warped events E′={ek′}k=1Ne\mathcal{E}^{\prime}=\{e^{\prime}_{k}\}_{k=1}^{N_{e}} “flattened” at a reference time treft_{\text{ref}}. The warp xk′=W(xk,tk;θ)\mathbf{x}^{\prime}_{k}=\mathbf{W}(\mathbf{x}_{k},t_{k};\boldsymbol{\theta}) transports each event along the point trajectory that passes through it (Fig. Focus Is All You Need: Loss Functions For Event-based Vision, left), until treft_{\text{ref}} is reached, thus, taking into account the space-time coordinates of the event. The vector θ\boldsymbol{\theta} parametrizes the point trajectories, and hence contains the motion or scene parameters.

The image (or histogram) of warped events (IWE) is given by accumulating events along the point trajectories:

where each pixel x\mathbf{x} sums the values bkb_{k} of the warped events xk′\mathbf{x}^{\prime}_{k} that fall within it (bk=pkb_{k}=p_{k} if polarity is used or bk=1b_{k}=1 if polarity is not used; see Fig. 1). In practice, the Dirac delta δ\delta is replaced by a smooth approximation, such as a Gaussian, δ(x−μ)≈N(x;μ,ϵ2Id)\delta(\mathbf{x}-\boldsymbol{\mu})\approx\mathcal{N}(\mathbf{x};\boldsymbol{\mu},\epsilon^{2}\mathtt{Id}), with typically ϵ=1\epsilon=1 pixel.

The contrast of the IWE (3) is given by its variance:

with mean μI≐1∣Ω∣∫ΩI(x;θ)dx\mu_{I}\doteq\frac{1}{|\Omega|}\int_{\Omega}I(\mathbf{x};\boldsymbol{\theta})d\mathbf{x}. Discretizing into pixels, it becomes Var⁡(I)=1Np∑i,j(hij−μI)2,\operatorname{Var}(I)=\frac{1}{N_{p}}\sum_{i,j}(h_{ij}-\mu_{I})^{2}, where NpN_{p} is the number of pixels of I=(hij)I=(h_{ij}) and μI=1Np∑i,jhi,j\mu_{I}=\frac{1}{N_{p}}\sum_{i,j}h_{i,j}.

2 Our Proposal: Focus Loss Functions

In (3) a group of events E\mathcal{E} has been effectively converted, using the spatio-temporal transformation, into an image representation (see Fig. 1), aggregating the little information from individual events into a larger, more descriptive piece of information. Here, we propose to exploit the advantages that such an image representation offers, namely bringing in tools from image processing (histograms, convolutions, Fourier transforms, etc.) to analyze event alignment.

We study different image-based event alignment metrics (i.e., loss functions), such as (4); that is, we study different objective functions that can be used in the second block of the diagram of Fig. Focus Is All You Need: Loss Functions For Event-based Vision (highlighted in red). To the best of our knowledge we are the first to address the following three related topics: (i) establishing connections between event alignment metrics and so-called “focus measures” in shape-from-focus (SFF) and autofocus (AF) in conventional frame-based imaging , (ii) comparing the performance of the different metrics on event-based vision problems, and (iii) showing that it is possible to design new focus metrics tailored to edge-like images like the IWE.

Table 1 lists the multiple focus loss functions studied, the majority of which are newly proposed, and categorizes them according to their nature. The functions are classified according to whether they are based on statistical or derivative operators (or their combination), according to whether they depend on the spatial arrangement of the pixel values or not, and according to whether they are maximized or minimized.

The next sections present the focus loss functions using two image characteristics related to edge strength: sharpness (Section 3.3) and dispersion (Section 3.4). But first, let us discuss the variance loss, since it motivates several focus losses.

An event alignment metric used in is the variance of the IWE (4), known as the RMS contrast in image processing . The variance is a statistical measure of dispersion, in this case, of the pixel values of the IWE, regardless of their spatial arrangement. Thus, event alignment (i.e., edge strength of the IWE) is here assessed using statistical principles (Table 1). The motion parameters θ\boldsymbol{\theta} that best fit the events are obtained by maximizing the variance (4). In the example of Fig. 1, it is clear that as event alignment increases, so does the visual contrast of the IWE.

Fourier Interpretation: Using (i) the formula relating the variance of a signal to its mean square (MS) and mean, Var⁡(I)=MS−μI2\operatorname{Var}(I)=\text{MS}-\mu_{I}^{2}, and (ii) the interpretations of MS and squared mean as the total “energy” and DC component of the signal, respectively, yields that the variance represents the AC component of the signal, i.e., the energy of the oscillating, high frequency, content of the IWE. Hence, the IWE comes into focus (Fig. 1) by increasing its AC component; the DC component does not change significantly during the process. This interpretation is further developed next.

3 Image Sharpness

In the Fourier domain, sharp images (i.e., those with strong edges, e.g., Fig. 1, right) have a significant amount of energy concentrated at high frequencies. The high-frequency content of an image can be assessed by measuring the magnitude of its derivative since derivative operators act as band-pass or high-pass filters. The magnitude is given by any norm; however, the L2L^{2} norm is preferred since it admits an inner product interpretation and simple derivatives. The following losses are built upon this idea, using first and second derivatives, respectively. Similar losses, based on the DCT or wavelet transforms instead of the Fourier transform are also possible. The main idea remains the same: measure the high frequency content of the IWE (i.e., edge strength) and find the motion parameters that maximize it, therefore maximizing event alignment.

Event alignment is achieved by seeking the parameters θ\boldsymbol{\theta} of the point-trajectories that maximize

Loss Function: Magnitude of Image Hessian.

For these loss functions, event alignment is attained by maximizing the magnitude of the second derivatives (i.e., Hessian) of the IWE, Hess⁡(I)\operatorname{Hess}(I). We use the (squared) norm of the Laplacian,

or the (squared) Frobenius norm of the Hessian,

Loss Functions: DoG and LoG.

The magnitude of the output of established band-pass filters, such as the Difference of Gaussians (DoG) and the Laplacian of the Gaussian (LoG), can also be used to assess the sharpness of the IWE.

Loss Function: Image Area.

Intuitively, sharp, motion-compensated IWEs, have thinner edges than uncompensated ones (Fig. 1). We now provide a definition of the “thickness” or “support” of edge-like images (see supplementary material) and use it for event alignment. We propose to minimize the support (i.e., area) of the IWE (3):

where F(λ)≐∫ρ(λ)dλF(\lambda)\doteq\int\rho(\lambda)d\lambda is the primitive of a decreasing weighting function ρ(λ)≥0\rho(\lambda)\geq 0, for λ≥0\lambda\geq 0. Fig. 2 shows an IWE and its corresponding support map, pseudo-colored as a heat map: red regions (large IWE values) contribute more to the support (8) than blue regions (small IWE values).

Our definition of the support of the IWE is very flexible, since it allows for different choices of the weighting function (Fig. 3). Specifically, we consider four choices:

Exponential: ρ(λ)=e−λ\rho(\lambda)=e^{-\lambda}, F(λ)=1−e−λF(\lambda)=1-e^{-\lambda}.

Gaussian: ρ(λ)=2πe−λ2\rho(\lambda)=\frac{2}{\sqrt{\pi}}e^{-\lambda^{2}}, F(λ)=erf(λ)F(\lambda)=\text{erf}(\lambda).

Lorentzian: ρ(λ)=2(1+λ2)π\rho(\lambda)=\frac{2}{(1+\lambda^{2})\pi}, F(λ)=2πarctan⁡(λ)F(\lambda)=\frac{2}{\pi}\arctan(\lambda).

Hyperbolic: ρ(λ)=sech2(λ)\rho(\lambda)=\text{sech}^{2}(\lambda), F(λ)=tanh⁡(λ)F(\lambda)=\tanh(\lambda).

4 Image Dispersion (Statistics)

Besides the variance (4), other ways to measure dispersion are possible. These may be categorized according to their spatial character: global (operating on the IWE pixels regardless of their arrangement) or local. Among the global metrics for dispersion we consider the mean square (MS), the mean absolute deviation (MAD), the mean absolute value (MAV), the entropy and the range. In the group of local dispersion losses we have: local versions of the above-mentioned global losses (variance, MS, MAD, MAV, etc.) and metrics of spatial autocorrelation, such as Moran’s I index and Geary’s Contiguity Ratio .

also measures the dispersion of the IWE (with respect to zero). As anticipated in the Fourier interpretation of the variance, the MS is the total energy of the image, which comprises an oscillation (i.e., dispersion) part and a constant part (DC component or squared mean).

Loss Function: Mean Absolute Deviation (MAD) and Mean Absolute Value (MAV).

The analogous of the variance and the MS using the L1L^{1} norm are the MAD,

respectively. The MAV provides a valid loss function to estimate θ\boldsymbol{\theta} if polarity is used. However, if polarity is not used, the MAV coincides with the IWE mean, μI\mu_{I} (which counts the warped events in Ω\Omega), and since this value is typically constant, the MAV does not provide enough information to estimate θ\boldsymbol{\theta} in this case (as will be noted in Table 2).

Loss Function: Image Entropy.

Information entropy measures the uncertainty or spread (i.e., dispersion) of a distribution, or, equivalently, the reciprocal of its concentration. This approach consists of maximizing Shannon’s entropy of the random variable given by the IWE pixels:

The PDF of the IWE is approximated by its histogram (normalized to unit area). Comparing the PDFs of the IWEs before and after motion compensation (last column of Fig. 1), we observe two effects of event alignment: (ii) the peak of the distribution at I=0I=0 increases since the image regions with almost no events grow (corresponding to homogeneous regions of the brightness signal), and (iiii) the distribution spreads out away from zero: larger values of ∣I(x)∣|I(\mathbf{x})| are achieved. Hence, the PDF of the motion-compensated IWE is more concentrated around zero and more spread out away from zero than the uncompensated one. Concentrating the PDF means decreasing the entropy, whereas spreading it out means increasing the entropy. To obtain a high contrast IWE, with sharper edges, the second approach must dominate. Hence, parameters θ\boldsymbol{\theta} are obtained by maximizing the entropy of the IWE (12).

Entropy can also be interpreted as a measure of diversity in information content, and since (ii) sharp images contain more information than blurred ones , and (iiii) our goal is to have sharp images for better event alignment, thus our goal is to maximize the information content of the IWE, which is done by maximizing its entropy.

Loss Function: Image Range.

As is well known, contrast is a measure of the oscillation (i.e., dispersion) of a signal with respect to its background (e.g., Michelson contrast), and the range of a signal, range⁡(I)=Imax⁡−Imin⁡\operatorname{range}(I)=I_{\max}-I_{\min}, measures its maximum oscillation. Hence, maximizing the range of the IWE provides an alternative way to achieve event alignment. However, the min and max statistics of an image are brittle, since they can drastically change by modifying two pixels. Using the tools that lead to (8) (suppl. material), we propose to measure image range more sensibly by means of the support of the image PDF,

where F(λ)F(\lambda) is a primitive of the weight function ρ(λ)≥0\rho(\lambda)\geq 0 (as in Fig. 3). Fig. 1 illustrates how the range of the IWE is related to its sharpness: as the IWE comes into focus, the image range (support of the histogram) increases.

Next, we present focus loss functions based on local versions of the above global statistics.

Loss Function: Local Variance, MS, MAD and MAV.

Mimicking (5), which aggregates local measures of image sharpness to produce a global score, we may aggregate the local variance, MS, MAD or MAV of the IWE to produce a global dispersion score that we seek to maximize. For example, the aggregated local variance (ALV) of the IWE is

where the local variance in a neighborhood B(x)⊂ΩB(\mathbf{x})\subset\Omega centered about the point x\mathbf{x} is given by

with local mean μ(x;I)≐∫B(x)I(v;θ) dv/∣B(x)∣\mu(\mathbf{x};I)\doteq\int_{B(\mathbf{x})}I(\mathbf{v};\boldsymbol{\theta})\,d\mathbf{v}/|B(\mathbf{x})| and ∣B(x)∣=∫B(x)du|B(\mathbf{x})|=\int_{B(\mathbf{x})}d\mathbf{u}. The local variance of an image (15) is an edge detector similar to the magnitude of the image gradient, ∥∇I∥\|\nabla I\|. It may be estimated using a weighted neighborhood B(x)B(\mathbf{x}) by means of convolutions with a smoothing kernel Gσ(x)G_{\sigma}(\mathbf{x}), such as a Gaussian:

Based on the above example, local versions of the MS, MAD and MAV can be derived (see suppl. material).

Loss Function: Spatial Autocorrelation.

Spatial autocorrelation of the IWE can also be used to assess event alignment. We present two focus loss functions based on spatial autocorrelation: Moran’s I and Geary’s CC indices.

Moran’s I index is a number between −1-1 and 11 that evaluates whether the variable under study (i.e., the pixels of the IWE) is clustered, dispersed, or random. Letting zi=I(xi)z_{i}=I(\mathbf{x}_{i}) be the value of the ii-th pixel of the IWE, Moran’s index is

where zˉ≡μI\bar{z}\equiv\mu_{I} is the mean of II,  Np\,N_{p} is the number of pixels of II,  (wij)\,(w_{ij}) is a matrix of spatial weights with zeros on the diagonal (i.e., wii=0w_{ii}=0) and W=∑i,jwijW=\sum_{i,j}w_{ij}. The weights wijw_{ij} encode the spatial relation between pixels ziz_{i} and zjz_{j}: we use weights that decrease with the distance between pixels, e.g., wij∝exp⁡(−∥xi−xj∥2/2σM2)w_{ij}\propto\exp(-\|\mathbf{x}_{i}-\mathbf{x}_{j}\|^{2}/2\sigma_{M}^{2}), with σM≈1\sigma_{M}\approx 1 pixel.

A positive Moran’s index value indicates tendency toward clustering pixel values while a negative Moran’s index value indicates tendency toward dispersion (dissimilar values are next to each other). Edges correspond to dissimilar values close to each other, and so, we seek to minimize Moran’s index of the IWE.

Geary’s Contiguity Ratio is a generalization of Von Neumann’s ratio of the mean square successive difference to the variance:

It is non-negative. Values around 1 indicate lack of spatial autocorrelation; values near zero are positively correlated, and values larger than 1 are negatively correlated. Moran’s I and Geary’s CC are inversely related, thus event alignment is achieved by maximizing Geary’s CC of the IWE.

5 Dispersion of Image Sharpness Values

We have also combined statistics-based loss functions with derivative-based ones, yielding more loss functions (Table 1), such as the variance of the Laplacian of the IWE Var⁡(ΔI(x;θ))\operatorname{Var}(\Delta I(\mathbf{x};\boldsymbol{\theta})), the variance of the magnitude of the IWE gradient Var⁡(∥∇I(x;θ)∥)\operatorname{Var}(\|\nabla I(\mathbf{x};\boldsymbol{\theta})\|), and the variance of the squared magnitude of the IWE gradient Var⁡(∥∇I(x;θ)∥2)≡Var⁡(Ix2+Iy2)\operatorname{Var}(\|\nabla I(\mathbf{x};\boldsymbol{\theta})\|^{2})\equiv\operatorname{Var}(I^{2}_{x}+I^{2}_{y}). They apply statistical principles to local focus metrics based on neighborhood operations (convolutions).

6 Discussion of the Focus Loss Functions

Connection with Shape-from-Focus and Autofocus: Several of the proposed loss functions have been proven successful in shape-from-focus (SFF) and autofocus (AF) with conventional cameras , showing that there is a strong connection between these topics and event-based motion estimation. The principle to solve the frame-based and event-based problems is the same: maximize a focus score of the considered image or histogram. In the case of conventional cameras, SFF and AF maximize edge strength at each pixel of a focal stack of images (in order to infer depth). In the case of event cameras, the IWE (an image representation of the events) plays the role of the images in the focal stack: varying the parameters θ\boldsymbol{\theta} produces a different “slice of the focal stack”, which may be used not only to estimate depth, but also to estimate other types of parameters θ\boldsymbol{\theta}, such as optical flow, camera velocities, etc.

Spatial Dependency: Focus loss functions based on derivatives or on local statistics imply neighborhood operations (e.g., convolutions with nearby pixels of the IWE), thus they depend on the spatial arrangement of the IWE pixels (Table 1). Instead, global statistics (e.g., variance, entropy, etc.) do not directly depend on such spatial arrangement The IWE consists of warped events, which depend on the location of the events; however, by “directly” we mean that the focus loss function, by itself, does not depend on the spatial arrangement of the IWE values.. The image area loss functions (8) are integrals of point-wise functions of the IWE, and so, they do not depend on the spatial arrangement of the IWE pixels. The PDF of the IWE also does not have spatial dependency, nor do related losses, such as entropy (12) or range (13). Composite focus losses (Section 3.5), however, have spatial dependency since they are computed from image derivatives.

Fourier Interpretation: Focus losses based on derivatives admit an intuitive interpretation in the Fourier domain: they measure the energy content in the high frequencies (i.e., edges) of the IWE (Section 3.3). Some of the statistical focus losses also admit a frequency interpretation. For example, image variance quantifies the energy of the AC portion (i.e., oscillation) of the IWE, and the MS measures the energy of both, the AC and DC components. Other focus functions, such as entropy, do not admit such a straightforward Fourier interpretation related to edge strength.

Experiments

In this section we compare the performance of the proposed focus loss functions, which depends on multiple factors, such as, the task for which it is used, the number of events processed, the data used, and, certainly, implementation. These factors produce a test space with an intractable combinatorial size, and so, we choose the task for best assessing accuracy on real data and then provide qualitative results on other tasks: depth and optical flow estimation.

Rotational camera motion estimation provides a good scenario for assessing accuracy since camera motion can be reliably obtained with precise motion capture systems. The acquisition of accurate per-event optical flow or depth is less reliable, since it depends on additional depth sensors, such as RGB-D or LiDAR, which are prone to noise.

2 Depth Estimation

3 Optical Flow Estimation

The profile of the loss function can also be visualized for 2D problems such as patch-based optical flow . Events in a small space-time window (e.g., 15×1515\times 15 pixels) are warped according to a feature flow vector θ≡v\boldsymbol{\theta}\equiv\mathbf{v}: xk′=xk−(tk−tref)v\mathbf{x}^{\prime}_{k}=\mathbf{x}_{k}-(t_{k}-t_{\text{ref}})\mathbf{v}. Fig. 6 shows the profiles of several competitive focus losses. They all have a clear extremal at the location of the visually correct ground truth flow (point 0, green arrow in Fig. 6a). The Laplacian magnitude and its variance show the narrowest peaks. Plots of more loss functions are provided in the supplementary material.

4 Unsupervised Learning of Optical Flow

The last row of Table 2 evaluates the accuracy and timing of a loss function inspired in the time-based IWE of : the variance of the per-pixel average timestamp of warped events. This loss function has been used in for unsupervised training of a neural network (NN) that predicts dense optical flow from events . However, this loss function is considerably less accurate than all other functions (Table 2). This suggests that (i) loss functions based on timestamps of warped events are not as accurate as those based on event count (the event timestamp is already taken into account during warping), (ii) there is considerable room for improvement in unsupervised learning of optical flow if better loss function is used.

To probe the applicability of the loss functions to estimate dense optical flow, we trained a network inspired in using (5) and a Charbonier prior on the flow derivative. The flow produced by the NN increases event alignment (Fig. 7). A deeper evaluation is left for future work.

Conclusion

We have extended motion compensation methods for event cameras with a library of twenty more loss functions that measure event alignment. Moreover, we have established a fundamental connection between the proposed loss functions and metrics commonly used in shape from focus with conventional cameras. This connection allows us to bring well-established analysis tools and concepts from image processing into the realm of event-based vision. The proposed functions act as focusing operators on the events, enabling us to estimate the point trajectories on the image plane followed by the objects causing the events. We have categorized the loss functions according to their capability to measure edge strength and dispersion, metrics of information content. Additionally, we have shown how to design new focus metrics tailored to edge-like images like the image of warped events. We have compared the performance of all focus metrics in terms of accuracy and time. Similarly to comparative studies in autofocus for digital photography applications , we conclude that the variance, the gradient magnitude and the Laplacian are among the best functions. Finally, we have shown the broad applicability of the functions to tackle essential problems in computer vision: ego-motion, depth, and optical flow estimation. We believe the proposed focus loss functions are key to taking advantage of the outstanding properties of event cameras, specially in unsupervised learning of structure and motion.

This work was supported by the Swiss National Center of Competence Research Robotics, the Swiss National Science Foundation and the SNSF-ERC Starting Grant.

Focus Is All You Need: Loss Functions For Event-based Vision —Supplementary Material—

Appendix A Notation

that is, the pp-th power of the norm of f\mathbf{f} is the sum of the pp-th power of the norms of its components.

The two most common cases are p=1p=1 and p=2p=2, which, for the gradient of an image, ∇I=(Ix,Iy)⊤\nabla I=(I_{x},I_{y})^{\top}, yield simple expressions:

A.2 Hessian Matrix

The Hessian matrix of a function, such as the IWE (used in (7)), is denoted by

The trace of the Hessian matrix is the Laplacian, which is used to define loss function (6).

Appendix B Area of the Image of Warped Events

To measure the “thickness” of the edges of the IWE (e.g., Fig. 2a), one could count the number of pixels with count of warped events above a threshold, e.g., one event. However, this is brittle since it depends on this arbitrary threshold. We propose to define the above-mentioned edge thickness or “area” of an edge-like image like the IWE (3) in a more sensible way as a weighted sum of the interior of the level sets of the image, as we show next.

Using a Gaussian function (kernel) as a smooth approximation to the Dirac delta, δ(x−μ)≈N(x;μ,σ2Id)\delta(\mathbf{x}-\boldsymbol{\mu})\approx\mathcal{N}(\mathbf{x};\boldsymbol{\mu},\sigma^{2}\mathtt{Id}), the image of warped events (3) has, strictly speaking, an unbounded support (area of pixels with non-zero value). To have a meaningful support measure, we instead count the number of pixels with value greater thanWe assume that I(x)≥0I(\mathbf{x})\geq 0 either because event polarity is not used (bk=1b_{k}=1 in (3)) or because the support here defined is applied to images of positive and negative events separately, and the results are added. λ\lambda,

where H(⋅)H(\cdot) is the Heaviside function. Fig. 8 shows several examples of it. This figure also illustrates the principle of area minimization, for a 1-D signal (3) with just two warped events. As observed, the area or thickness of II is minimized if the events are warped to the same location (Δx′=0  ⟺  xi′=xj′\Delta\mathbf{x}^{\prime}=\mathbf{0}\iff\mathbf{x}^{\prime}_{i}=\mathbf{x}^{\prime}_{j}), which is the desired event alignment condition of corresponding events.

To have a support metric that does not depend on the particular value of the threshold λ\lambda used (for fixed kernel width σ\sigma), we sum (24) over all threshold values,

Notice that it is not possible to use ρ=const\rho=\text{const} since this leads to supp⁡(I)=Ne\operatorname{supp}(I)=N_{e}, which does not depend on the motion parameters θ\boldsymbol{\theta} we wish to optimize for. Using weighting functions with unit area (i.e., ∫0∞ρ(λ)dλ=1\int_{0}^{\infty}\rho(\lambda)d\lambda=1) allows us to interpret (25) as a convex combination of supports (24), thus setting the correct scale so that (25) has the same units as (24).

B.2 Simplification of the Area of an Image

Substituting (24) in (25) and swapping the order of integration gives

where F(λ)≐∫ρ(λ)dλF(\lambda)\doteq\int\rho(\lambda)d\lambda is a primitive of ρ\rho, and F(0)F(0) is constant. This is an advantageous expression compared to (25), since it states that supp⁡(I)\operatorname{supp}(I) can be computed using the values of I(x)I(\mathbf{x}) directly, without having to compute (24) for every threshold λ\lambda and then sum up the results. By using a continuous image formulation, we have analytically integrated the partial sums (24).

Fig. 9 illustrates (26). It shows the warped events I(x;θ)I(\mathbf{x};\boldsymbol{\theta}) on a 31×3131\times 31 image patch for two different parameters θ1,θ2\boldsymbol{\theta}_{1},\boldsymbol{\theta}_{2} (depth values, in this example ). It also shows the corresponding integrands of (26), or “per-pixel support maps” F(I(x;θ)/λ0)−F(0)≡1−exp⁡(−(I(x;θ)/λ0))F(I(\mathbf{x};\boldsymbol{\theta})/\lambda_{0})-F(0)\equiv 1-\exp(-(I(\mathbf{x};\boldsymbol{\theta})/\lambda_{0})), with λ0=10\lambda_{0}=10 warped events. Pixels with I(x)≳λ0I(\mathbf{x})\gtrsim\lambda_{0} events contribute more to the support (26) than pixels with I(x)≲λ0I(\mathbf{x})\lesssim\lambda_{0} events, as shown in the support maps (right column of Fig. 9), which are color-coded from blue (low contribution) to red (high contribution).

Basically, the red regions of the support maps approximately indicate the area of the IWE, whereas the blue regions indicate the pixels where few warped events accumulate and therefore do not effectively contribute to the area of the IWE. Clearly, the bottom patch has a smaller area (i.e., thinner edges) than the top patch, as indicated by the smaller area of the red regions. The image area (26) is used to define the focus loss function (8).

Appendix C Loss Function: Image Entropy

As anticipated in (12), event alignment may be achieved by maximizing the entropy of the IWE, where

is Shannon’s (differential) entropy for a continuous random variable whose density function (PDF) is p(z)p(z).

The PDF of an image is approximated by its histogram, normalized to have unitary area. In a continuous formulation, this is written as (see )

using the Dirac delta. This equation intuitively says that pI(z)p_{I}(z) is computed as a ratio of areas: the “number of pixels” of the IWE with value zz, divided by the total “number of pixels”, ∣Ω∣=∫Ωdx=Np|\Omega|=\int_{\Omega}d\mathbf{x}=N_{p}.

Substituting (28) into (27), the entropy of the IWE becomes

Observe that the entropy is maximized by favoring large values of log⁡(1/pI(I(x)))\log(1/p_{I}(I(\mathbf{x}))) over smaller ones. Since log⁡\log is concave, it means that large values of 1/pI(I(x))1/p_{I}(I(\mathbf{x})) are favored, i.e., small values of pI(I(x))p_{I}(I(\mathbf{x})) are favored. For a PDF that is concentrated around I=0I=0 (large pI(0)p_{I}(0), as shown on the last column of Fig. 1), favoring small density values implies that they must be achieved away from I=0I=0, i.e., for large ∣I∣|I| values (which are caused by the aggregation of aligned events). Thus, maximizing the entropy increases the range of I(x)I(\mathbf{x}), producing a higher contrast image.

Appendix D Loss Function: Image Range

We measure the image range by means of the support of its PDF (28),

where the weight function ρ(λ)≥0\rho(\lambda)\geq 0 emphasizes the contributions of small ∣λ∣|\lambda| over those of large ∣λ∣|\lambda|, according to the typical shape of the PDF of the event image (concentrated around λ=0\lambda=0).

Mimicking the steps in Section B, Eq. (30) can be rewritten as

where F(λ)F(\lambda) is a primitive of ρ(λ)\rho(\lambda), and F(0)F(0) is constant.

The motion parameters are found by maximizing (31), i.e., (13). The same weighting functions and primitives as for the image area (Section 3.3) may be used (with even symmetry if event polarity is used in the IWE (3)). This approach is inspired by the maximization of the entropy of the PDF of the image of warped events, as explained in Section C.

Appendix E Loss Function: Spatial Autocorrelation

Moran’s I index (17) (or “serial correlation coefficient” ) is a measure of spatial autocorrelation, i.e., it measures how similar is one object with respect to its neighbors. It is a concept that applies to variables whose values are known in unstructured grids (spatial units), in general (see Fig. 10).

If the variable of interest zz consists of the intensity values of an image, zi=I(xi)z_{i}=I(\mathbf{x}_{i}), which is defined on a regular (pixel) lattice {xi}\{\mathbf{x}_{i}\}, and the weights wijw_{ij} are shift-invariant (they do not depend on the particular location of pixels ii and jj, only on their relative spatial arrangement) and symmetric wij=wjiw_{ij}=w_{ji}, then it is possible to write Moran’s I index using a convolution. In the formalism of continuous images z(x)z(\mathbf{x}) over a domain Ω\Omega, Moran’s I index becomes

E.2 Geary’s Contiguity Ratio

Geary’s contiguity ratio is a generalization of Von Neumann’s ratio of the mean square successive difference (numerator of (18)) to the variance (denominator of (18)). Geary’s contiguity ratio is non-negative, and its mean is 1 for random images. Values of CC significantly lower than 1 demonstrate positive spatial autocorrelation (the variable of interest is regarded as contiguous), while values significantly higher than 1 illustrate negative autocorrelation.

In the formalism of continuous images, Geary’s contiguity ratio can be written as

with local score efficiently computed using convolutions:

Notice that the last term in (35) also appears in Moran’s I index (32). Thus, Geary’s CC is inversely related to Moran’s I, but they are not identical.

Homogeneous regions of an image z(x)z(\mathbf{x}) have a positive spatial autocorrelation, indicated by c(x)<1c(\mathbf{x})<1. The regions with large values of (35), i.e., negative spatial autocorrelation, are those corresponding to the edges of the objects (dissimilar intensity values on each side of the edge). Thus, Geary’s local statistic (35) acts as an edge detector (see Figs. 12 or 14).

Appendix F Loss Function: Aggregation of Local Statistics

Similarly to (14), aggregating other local statistics of the IWE also yield focus measures. For example, the aggregation of the local mean absolute deviation (ALMAD),

is closely related to (14) since both aggregate local measures that are edge-detectors of II ((15) uses the the L2L^{2} norm, whereas (37) uses the L1L^{1} norm). Using a weighted neighborhood (e.g., Gaussian kernel GσG_{\sigma}), (37) can be efficiently approximated by the formula with two convolutions:

where the inner convolution approximates the local mean, μ(x;I)≈I(x)∗Gσ(x)\mu(\mathbf{x};I)\approx I(\mathbf{x})*G_{\sigma}(\mathbf{x}), and the outer convolution averages the magnitude of the local, centered IWE (integrand of (37)) over the neighborhood around x\mathbf{x}.

Omitting the local mean in (15) and (37) leads to local versions of the MS and the MAV, respectively. These operators, however, are not edge detectors; nevertheless, they also work as focus loss functions since the images on which they are applied, the IWEs, are edge-like images (the events are brightness changes, i.e., they are related to the temporal derivative of the brightness signal). Similarly to the (global) MAV, the local MAV does not provide enough information to estimate the warp parameters θ\boldsymbol{\theta} if polarity is not used (as indicated in Table 2).

Appendix G Plots of the Local Loss Maps

Most of the loss functions considered can be written as integrals over the IWE domain. Figs. 11, 12, 13 and 14 visualize the integrands (i.e., “local loss”) of most focus loss functions, for two scenes: dynamic and boxes from the dataset . Images are given in pairs (with a common caption below the images): local loss before optimization (without motion compensation, on the left), and after optimization of the corresponding focus loss function (motion-compensated, on the right). Each image pair shares the same color scale for proper visualization of how the local loss changes before and after optimization. For reference, since the local loss are transformations of the IWE, Figs. 11 and 13 also provide, on the top right, the IWE before and after optimization with one of the loss functions (the variance).

The local loss of area-based loss functions is the support map, as in Figs. 2b and 9. The focus loss given by the IWE range (13) is not expressed as an integral over the image domain, therefore, no image integrand is visualized in the above-mentioned figures. MAV and local MAV are not displayed either since they cannot be optimized with respect to the parameters.

Notice that all local loss maps are represented using the same color scheme, from blue (small) to large (yellow). For objective functions formulated as maximization problems (variance, gradient magnitude, etc.), visually good maps are those that are almost “blue” for IWEs with bad event-alignment parameters, and that clearly show “yellow” regions where events align (due to good parameters θ\boldsymbol{\theta}). For objective functions formulated as minimization problems (e.g., area-based, loss based on the mean timestamp per pixel), the situation is the opposite: good local loss maps become less yellowish and more blueish as event alignment improves due to good parameters θ\boldsymbol{\theta}.

Appendix H Additional Experiments on Accuracy Evaluation

Appendix I Additional Plots of Focus Loss Functions in Optical Flow Space

Rows 2 and 3 of the figures show the variance, MS, MAD, MAV and their aggregation of local versions. They are all visualized in the range $,bydividingbythemaximumvalueofthefocusfunction.Withoutusingpolarity(Fig.15),thevariancepresentsanicepeakatthecorrectopticalflow, by dividing by the maximum value of the focus function. Without using polarity (Fig. 15), the variance presents a nice peak at the correct optical flow\boldsymbol{\theta}_{2},thepeaksoftheMSandMADfunctionsarenotaspronounced,andtheMAVfunctiondoesnothavethemaximumatthegroundtruthlocation(weexplainedthat,withoutpolarity,theMAVcannotbeusedtoestimate, the peaks of the MS and MAD functions are not as pronounced, and the MAV function does not have the maximum at the ground truth location (we explained that, without polarity, the MAV cannot be used to estimate\boldsymbol{\theta}).Thelocalversions(thirdrowofFig.15)areslightlynarrowerthantheglobalversions.Thefourthrowpresentsthefourarea−basedfocuslosses(SectionB),whosegoalistobeminimized,andindeed,theypresentalocalminimumat). The local versions (third row of Fig. 15) are slightly narrower than the global versions. The fourth row presents the four area-based focus losses (Section B), whose goal is to be minimized, and indeed, they present a local minimum at\boldsymbol{\theta}_{2}.Therearenotbigdifferencesinthesefourarea−basedlosses.Thefifthrowshowsmorestatistics−basedlosses.TherangeandGeary’sCshowalocalmaximumatthecorrectflow.Moran’sindexshowsalocalminimum,asexpected,atthecorrectflow.Theentropy,withoutusingpolarity,doesnothavealocalmaximumatthecorrectflow.Instead,usingpolarity(Fig.16),itdoeshavealocalmaximumat. There are not big differences in these four area-based losses. The fifth row shows more statistics-based losses. The range and Geary’s C show a local maximum at the correct flow. Moran’s index shows a local minimum, as expected, at the correct flow. The entropy, without using polarity, does not have a local maximum at the correct flow. Instead, using polarity (Fig. 16), it does have a local maximum at\boldsymbol{\theta}_{2}.ThelasttworowsofFigs.15and16showfocuslossfunctionsbasedontheIWEderivativesandtheirvariances(compositelosses).Theyallpresentaclearpeakatthecorrectdepth(asthecaseofthevarianceandlocalvariance);someofthemaremorenarrowthanothers(allarevisualizedintherange. The last two rows of Figs. 15 and 16 show focus loss functions based on the IWE derivatives and their variances (composite losses). They all present a clear peak at the correct depth (as the case of the variance and local variance); some of them are more narrow than others (all are visualized in the range$, for ease of comparison). The gradient magnitude (based on Sobel operator), the DoG magnitude and the LoG magnitude seem to be the smoothest of these two rows.

With Polarity.

Fig. 16 shows the results on the same experiment as Fig. 15, but using event polarity in the IWE. In a scene with approximately equal number of dark-to-bright and bright-to-dark transitions, the number of positive and negative events is approximately balanced, and so, the mean of the IWE is approximately zero. Thus, in this case, the MS is approximately equal to the variance, and the MAV approximates the MAD. This is noticeable in the second row of Fig. 16. A similar trend is observe in the local versions of the four above statistics, albeit the local variance and MAD present narrower peaks than the local MS and MAV, respectively. The area-based focus functions are computed by splitting the events according to polarity, computing the areas of the two resulting IWEs and adding their area values. The corresponding plots, in the fourth row of Fig. 16 are similar to those without polarity (Fig. 15, except for the vertical scale). The entropy and range improve if event polarity is used, basically because the PDF becomes double-sided and it allows us to distinguish positive and negative IWE edges/values. Moran’s index and Geary’s CC ratio are not good focus losses if polarity is used, since they present brittle local minimum/maximum, respectively. The derivative-based losses present a clear peak at the correct flow, and slightly more pronounced than their counterparts in Fig. 15.

Appendix J Additional Plots on Depth Estimation

Fig. 18 shows depth estimation for every pixel of a reference view along the trajectory of the event camera. For every pixel, we compute focus curves, as those in Fig. 17b, and select the depth at the peak. To capture fine spatial details, the focus functions are computed on patches of 3×33\times 3 pixels in the reference view, weighted by a Gaussian kernel to emphasize the contribution of the center pixel. We also record the value of the focus function at the peak for every pixel of the reference view. These values are displayed as a “focus confidence map” in Fig. 18. For better visualization, the focus values are represented in negative form, from bright (low focus value) to dark (high focus value). The confidence map is used to select the pixels in the reference view with largest focus, i.e., the pixels for which depth is most reliably estimated. The above selection yields a semi-dense depth map, which is displayed color coded, overlaid on the intensity frame from the DAVIS camera at the reference view. We used adaptive thresholding [46, p.780] on the focus confidence map, and a median filter to remove spike noise from the depth map. As it is seen, depth is most reliably estimated at strong brightness edges of the scene. Finally, the depth map is also visualized as a point cloud, color-coded according to depth (Fig. 18).

The figure compares some representative focus functions. In general, we obtain good depth 3D reconstructions with the methods tested. Some methods produce slightly noisier 3D reconstructions than others, and some recover more edges than others. This is due to both the shape of the focus confidence maps and the adaptive thresholding parameters. We observe that focus functions as simple as the local mean square (MS) or the local mean absolute deviation (MAD) produce good results. These semi-dense 3D reconstruction methods may be used as the mapping module of an event-based visual odometry system, such as , to enable camera pose estimation from the 3D reconstructed scene.

Appendix K Analytical Derivatives of Focus Loss Functions

In this section, we provide the analytical derivatives of some of the focus loss functions used. A significant advantage of the proposed focus loss functions is that they are defined in terms of well-known analytical operations, and so, we can use powerful tools from Calculus to compute and simplify their derivatives. This becomes useful in the optimization framework of Fig. Focus Is All You Need: Loss Functions For Event-based Vision, both, for speed-up and increased accuracy over numerical derivatives.

The derivative of the IWE (3) with respect to the warp parameters θ\boldsymbol{\theta} is, replacing the Dirac delta with an approximation, such as a Gaussian, δ(x)≈N(x;0,ϵ2Id)\delta(\mathbf{x})\approx\mathcal{N}(\mathbf{x};\mathbf{0},\epsilon^{2}\mathtt{Id}), and using the chain rule,

where ∇N\nabla\mathcal{N} is the gradient of the 2D Gaussian PDF, and the derivative of the warp is a purely geometric term: ∂xk′(θ)∂θ=W′(xk′,tk;θ).\frac{\partial\mathbf{x}^{\prime}_{k}(\boldsymbol{\theta})}{\partial\boldsymbol{\theta}}=\mathbf{W}^{\prime}(\mathbf{x}^{\prime}_{k},t_{k};\boldsymbol{\theta}).

Loss Function: Mean Square (MS).

The derivative of the MS, (9), is, by the chain rule,

Loss Function: Variance.

The derivative of the variance of the IWE (4) is given in . Letting

be the centered IWE, the derivative of the variance is, by the chain rule,

(formally, the same formula as (40), but with the centered IWE playing the role of the IWE in (40)) with

since the mean μ(⋅)\mu(\cdot) and the derivative are linear operators, and therefore, commute. The previous result (39) may be substituted in (43).

Loss Function: Mean Absolute Value (MAV).

From a numerical point of view, it is sensible to replace sign(x)\text{sign}(x) by a smooth approximation, e.g., sign(x)≈tanh⁡(kx)\text{sign}(x)\approx\tanh(kx), with parameter k≫1k\gg 1 controlling the width of the transition around x=0x=0 (see ).

Loss Function: Mean Absolute Deviation (MAD).

The derivative of the MAD, (10), is, by the chain rule and using (43),

Loss Function: Entropy.

The derivative of entropy (12) is, stemming from (29),

where pI′(z)=dpI/dzp_{I}^{\prime}(z)=dp_{I}/dz. This formula can be obtained by differentiating (29) and applying the chain rule,

Implementation Details: We approximate the density function pI(z)p_{I}(z) by a smooth histogram of the IWE (3). First, a high-resolution histogram (e.g., 200 bins) is computed and normalized to unit area (like a PDF), and then it is smoothed by a Gaussian filter, e.g., of standard deviation σ=5\sigma=5 bins. The smoothed PDF is the convolution

and it is straightforward to show, using the convolution properties and interchanging the order of integration, that it is equivalent to the function obtained by replacing the Dirac delta in (28) with a smooth approximation (such as the Gaussian: δ←N\delta\leftarrow\mathcal{N}):

Smoothing (i.e., filtering) mitigates the noise due to bin discretization and improves robustness (size of the basin of attraction) of the entropy-based focus function (12). Thus, the term in the integrand of (46) is approximated by

where the derivative (pIσ)′(p_{I}^{\sigma})^{\prime} is computed using central, finite-differences on the samples of pIσ(z)p_{I}^{\sigma}(z). Linear interpolation is used to interpolate the samples of pIσp_{I}^{\sigma} and (pIσ)′(p_{I}^{\sigma})^{\prime}.

Loss Function: Image Area.

with weighting functions ρ(λ)\rho(\lambda). See Section 3.3 for different choices (exponential, Gaussian, Lorentzian and Hyperbolic).

Loss Function: Image Range.

The derivative of the support of the PDF of the IWE (13) is

where ρ′(λ)=dρ/dλ\rho^{\prime}(\lambda)=d\rho/d\lambda is the derivative of the weighting function. This can be shown using the chain rule and the derivative of the PDF of the IWE (28), expressed in terms of the derivative of the Dirac delta.

Result: Derivative of a Convolution.

The derivative of a convolution of the IWE with a kernel K(x)K(\mathbf{x}) is computed component-wise:

Loss Function: Local MS.

The derivative of the aggregated local MS of the IWE is, by the chain rule and (53),

Loss Function: Local Variance.

The derivative of the aggregated local variance of the IWE is, by the chain rule on (14),

where, using (16) and result (53), the integrand becomes

Loss Function: Local MAV.

The derivative of the aggregated local MAV of the IWE is, using the chain rule and the compact notation (53),

Loss Function: Local MAD.

Letting Ic(x)≐I(x)−μ(I)(x)I^{c}(\mathbf{x})\doteq I(\mathbf{x})-\mu(I)(\mathbf{x}) be the locally-centered IWE, with local mean μ(I)(x)≐I(x)∗Gσ(x)\mu(I)(\mathbf{x})\doteq I(\mathbf{x})*G_{\sigma}(\mathbf{x}), the derivative of the aggregated local MAD of the IWE is, using the chain rule and the compact notation (53),

with derivative of the locally-centered IWE

since the local mean and the derivative are linear operators, and hence, commute.

Loss Function: Gradient Magnitude.

The derivative of the squared magnitude of the gradient of the IWE (5) is

with 2×M2\times M matrix (assuming equality of mixed derivatives by Schwarz’s theorem),

Loss Function: Laplacian Magnitude.

Following similar steps as for the gradient magnitude, the derivative of the magnitude of the Laplacian of the IWE (6) is

Loss Function: Hessian Magnitude.

The derivative of Hessian magnitude (7) is

where xi,xjx_{i},x_{j} are variables x,yx,y of the image plane. Using once more Schwarz’s theorem to swap the differentiation order, each of the four terms in the sum (65) is

Loss Function: Difference of Gaussians (DoG).

The derivative of the squared difference of Gaussians applied to the IWE can be written in a compact way, using (53) on the DoG(x)≐(Gσ1−Gσ2)(x)\text{DoG}(\mathbf{x})\doteq(G_{\sigma_{1}}-G_{\sigma_{2}})(\mathbf{x}) filter, with σ1>σ2\sigma_{1}>\sigma_{2}, as

Loss Function: Laplacian of the Gaussian (LoG).

The derivative of the squared difference of the Laplacian of the Gaussian of the IWE can be computed from the formula for the DoG, using the fact that the DoG approximates the LoG if σ2≈1.6σ1\sigma_{2}\approx 1.6\sigma_{1}.

Loss Function: Variance of the Laplacian.

Derivative of the variance of the Laplacian:

where the mean is μΔI≐1∣Ω∣∫ΩΔI(x)dx\mu_{\Delta I}\doteq\frac{1}{|\Omega|}\int_{\Omega}\Delta I(\mathbf{x})d\mathbf{x}. The derivative of (69) with respect to the parameters θ\boldsymbol{\theta} is, using (64),

Loss Function: Variance of the Squared Gradient Magnitude.

The derivative of the variance of the squared gradient magnitude

with mean μ∥∇I∥2≐1∣Ω∣∫Ω∥∇I(x)∥2dx\mu_{\|\nabla I\|^{2}}\doteq\frac{1}{|\Omega|}\int_{\Omega}\|\nabla I(\mathbf{x})\|^{2}d\mathbf{x}, is, using (61),

Loss Function: Variance of the Gradient Magnitude.

Derivative of the variance of the gradient magnitude

with mean μ∥∇I∥≐1∣Ω∣∫Ω∥∇I(x)∥dx\mu_{\|\nabla I\|}\doteq\frac{1}{|\Omega|}\int_{\Omega}\|\nabla I(\mathbf{x})\|d\mathbf{x}, is, also using (61),

References