Optical Flow in Mostly Rigid Scenes

Jonas Wulff, Laura Sevilla-Lara, Michael J. Black

Introduction

The world is composed of things that move and things that do not. The 2D motion field, which is the projection of the 3D scene motion onto the image plane, arises from observer motion relative to the static scene and the independent motion of objects. A large body of work exists on estimating camera motion and scene structure in purely static scenes, generally referred to as Structure-from-Motion (SfM). On the other hand, methods that estimate general 2D image motion, or optical flow, make much weaker assumptions about the scene. Neither approach fully exploits the mixed structure of natural scenes. Most of what we see in such scenes is static - houses, roads, desks, etc.In KITTI-2015 and MPI-Sintel, independently moving regions make up only 15% and 28% of the pixels, respectively. Here, we refer to these static parts of the scene as the rigid scene, or rigid regions. At the same time, moving objects like people, cars, and animals make up a small but often important part of natural scenes. Despite the long history of both SfM and optical flow, no state-of-the art optical flow method synthesizes both into an algorithm that works on general scenes like those in the MPI-Sintel dataset (Fig. LABEL:fig:teaser). In this work, we propose such a method to estimate optical flow in video sequences of generic scenes that contain moving objects within a rigid scene.

For the rigid scene, the camera motion and depth structure fully determine the motion, which forms the basis of SfM methods. Modern optical flow benchmarks, however, are full of moving objects such as cars or bicycles in KITTI, or humans and dragons in Sintel. Assuming a fully static scene or treating these moving objects as outliers is hence not viable for optical flow algorithms; we want to reconstruct flow everywhere.

Independent motion in a scene typically arises from well defined objects with the ability to move. This points to a possible solution. Recently, convolutional neural networks (CNN) have achieved good performance on detecting and segmenting objects in images, and have been successfully incorporated into optical flow methods . Here we take a slightly different approach. We modify a common CNN and train it on novel data to obtain a rigidity score from the labels, taking into account that some objects (e.g. humans) are more likely to move than others (e.g. houses). This score is combined with additional motion cues to obtain an estimate of rigid and independently moving regions.

After partitioning the scene into rigid and moving regions, we can deal with each appropriately. Since the motion of moving objects can be almost arbitrary, it is best computed using a classical unconstrained flow method. The flow of the rigid scene, on the other hand, is extremely restricted, and only depends on the depth structure and the camera motion and calibration. In theory, one could use an existing SfM algorithm to reconstruct the camera motion and the 3D structure of the scene, and project this structure back to obtain the motion of the rigid scene regions. Two factors make this hard in practice. First, the number of frames usually considered in optical flow is small; most methods only work on two or three consecutive frames. SfM algorithms, on the other hand, require tens or hundreds of frames to work reliably. Second, SfM algorithms require large camera baselines in order to reliably estimate the fundamental matrices. In video sequences, large baselines are rare, since the camera usually translates very little between frames. An exception to this are automotive scenarios such as the KITTI benchmark, where the recording car often moves rapidly and the frame rate is low.

Since full SfM is unreliable in general flow scenarios, we adopt the Plane+Parallax (P+P) framework In this framework, frames are registered to a common plane, which is aligned in all images after the registration. This removes the motion caused by camera rotation and simple intrinsic camera parameter changes, leaving parallax as the sole source of motion. Since all parallax is oriented towards or away from a common focus of expansion in the frame, computing the parallax is reduced to a 1D search problem and therefore easier than computing the full optical flow.

Here we show that using the P+P framework brings an additional advantage: the parallax can be factored into a structure component, which is independent of the camera motion and constant across time, and a temporally varying camera component, which is a single number per frame. We integrate the structure information across time; by definition, the structure of the rigid scene does not change. By combining the structure information from multiple frames, our algorithm generates a better structure component for all frames, and fills in areas that are unmatched in a single pair of frames due to occlusion.

Additionally, the relationship between the structure component and the parallax (and thus, the optical flow) enables us to regularize the flow in a physically meaningful way, since regularizing the structure implicitly regularizes the flow. We use a robust second-order regularizer, which corresponds to a locally planar prior.

We integrate the regularization into a novel objective function measuring the photometric error across three frames as a function of the structure and camera motion. This allows us to optimize the structure and also to recover from poor initializations. We call the method MR-Flow for Mostly-Rigid Flow and show an overview in Fig. 1.

We test MR-Flow on MPI-Sintel and KITTI 2015 (Fig. LABEL:fig:teaser). Among published monocular methods, at time of writing, we achieve the lowest error on MPI-Sintel on both passes; on KITTI-2015, our accuracy is second only to , a method specifically designed for automotive scenarios. Our code, the trained CNN, and all data is available at .

In summary, we present three main contributions. First, we show how to segment the scene into rigid regions and independently moving objects, allowing us to estimate the motion of each type of region appropriately. Second, we extend previous plane+parallax methods to express the flow in the rigid regions via its depth structure. This allows us to regularize this structure instead of the flow field and to combine information across more than two frames. Third, we formulate the motion of the rigid regions as a single model. This allows us to iterate between estimating the structure and to recover from unstable initializations.

Previous work

SfM and optical flow have both made significant, but mostly independent, progress. Roughly speaking, SfM methods require purely rigid scenes and use sparse point matches, wide baselines between frames, solve for accurate camera intrinsics and extrinsics, and exploit bundle adjustment to optimize over many views at once. In contrast, optical flow is applied to scenes containing generic motion, exploits continuous optimization, makes weak assumptions about the scene (e.g. that it is piecewise smooth), and typically processes only pairs of video frames at a time.

Combining optical flow and SfM. There have been many attempts to combine SfM and flow methods, dating to the 80’s . For video sequences from narrow-focal-length lenses, the estimation of the camera motion is challenging, as it is easy to confuse translation with rotation and difficult to estimate the camera intrinsics .

More recently there have been attempts to combine SfM and optical flow . The top monocular optical flow method on the KITTI-2012 benchmark estimates the fundamental matrix and computes flow along the epipolar lines . This approach is limited to fully rigid scenes. Wedel et al. compute the fundamental matrix and regularize optical flow to lie along the epipolar lines. If they detect independent motion, they revert to standard optical flow for the entire frame. In contrast, we segment static from moving regions and use appropriate constraints within each type of region. Roussos et al. assume a known calibrated camera and solve for depth, motion and segmentation of a scene with moving objects. They perform batch processing on sequences of about 30 frames in length, making this more akin to SfM methods. While they have impressive results, they consider relatively simple scenes and do not evaluate flow accuracy on standard benchmarks.

Plane+Parallax. P+P methods were developed in the mid-90’s . The main idea is that stabilizing two frames with a planar motion (homography) removes the camera rotation and simplifies the geometric reasoning about structure . In the stabilized pair, motion is always oriented towards or away from the epipole and corresponds to parallax, which is related to the distance of the point from the plane in the 3D scene.

Estimating a planar homography can be done robustly and with more stability than estimating the fundamental matrix . While one is not able to estimate metric depth, the planar stabilization simplifies the matching process, turning the 2D optical flow estimation problem into a 1D problem that is equivalent to stereo estimation. Given the practical benefits, one may ask why P+P methods are not more prevalent in the leader boards of optical flow benchmarks. The problem is that such methods work only for rigid scenes. Making the P+P approach usable in general natural scenes is one of our main contributions.

Moving region segmentation. There have been several attempts to segment moving scenes into regions corresponding to independently moving objects by exploiting 3D motion cues and epipolar motion . Several methods use the P+P framework to detect independent motions, but those methods typically only do detection and not flow estimation, and are often applied to simple scenes where there is a dominant motion like the ground plane and small moving objects . Irani et al. develop mosaic representations that include independently moving objects but do not explicitly compute their flow. Given two frames as input, Ranftl et al. segment a general moving scene into piecewise-rigid components and reason about the depth and occlusion relationships. While they produce impressive depth estimates, they rely on accurate flow estimates between the frames and do not refine the flow itself.

Combining multiple flow methods. There is also existing work on combining motion estimates from different algorithms into a single estimate , but these do not attempt to fuse rigid and general motion. Bergen et al. define a framework for describing optical flow problems using different constraints from rigid motion to generic flow, but do not combine these models into a single method.

Recent work combines segmentation and flow. Sevilla et al. perform semantic segmentation and use different models for different semantic classes. Unlike them, we use semantic segmentation to estimate the rigid scene and then impose stronger geometric constraints in these regions. Hur and Roth integrate semantic segmentation over time, leading to more accurate flow estimation for objects and better segmentation performance.

Most similar to our approach is , which first segments the scene into objects using a CNN. A fundamental matrix is then computed and used to constrain the flow within each object. Our work is different in a number of important ways. (i) Their approach is sequential and cannot recover from an incorrect fundamental matrix estimate. We propose a unified objective function where the parts of the solution inform and improve each other. (ii) relies exclusively on the CNN to segment moving regions. While this works in specific scenarios such as automotive, it may not generalize to new scenes. We combine semantic segmentation and motion to classify rigid regions and thus require less accurate semantic rigidity estimates. This makes our algorithm both more robust and more general, as demonstrated by the fact that in contrast to we evaluate on the challenging MPI-Sintel benchmark. (iii) requires moving objects to be rigid (i.e., rigidly moving vehicles) and assumes a small rotational component of the egomotion. This works for KITTI-2015 but does not apply to more general scenes. (iv) uses only two frames at a time and extrapolates into occlusions. Our model combines information across time, and thus it is able to compute accurate flow in occlusions.

Plane + Parallax background

The P+P paradigm has been used in rigid scene analysis for a long time. Since it forms the foundation of our algorithm, we briefly review the parts that are important for this work and refer the reader to for more details.

The core idea of P+P is to align two or more images to a common plane Π\mathbf{\Pi}, so that

where x\mathbf{x} and x′\mathbf{x^{\prime}} represent a point in the reference frame and the corresponding point in another frame of the sequence, x⁡h\operatorname{\mathbf{x}}_{h} denotes x⁡\operatorname{\mathbf{x}} in homogeneous coordinates, HH is the homography mapping the image of Π\mathbf{\Pi} between frames, and ⟨a⟩=(a1/a3,a2/a3)\left\langle\mathbf{a}\right\rangle=\left(a_{1}/a_{3},a_{2}/a_{3}\right) is the perspective normalization.

This alignment removes the effects of camera rotation and the effect of camera calibration change (such as a zoom) between the pair of frames . Getting rid of rotation is especially convenient, since the ambiguity between rotation and translation in case of small displacements is a major source of numerical instabilities in the estimation of the structure of the scene.

When computing optical flow between aligned images, the flow of the pixels corresponding to points on the plane is zeroNote that the plane does not have to correspond to a physical surface, but merely to a rigid, “virtual” plane.. For an image point x\mathbf{x} corresponding to a 3D point X\mathbf{X} off the plane, the residual motion is given as

where d(C2)d(C_{2}) is the distance of the second camera center to Π\mathbf{\Pi}, zz is the distance of point X\mathbf{X} to the first camera, TzT_{z} is the depth displacement of the second camera, d(X)d(\mathbf{X}) is the distance from point X\mathbf{X} to Π\mathbf{\Pi}, and e\mathbf{e} is the common focus of expansion that coincides with the epipole corresponding to the second camera. This representation has two main advantages. First, instead of an arbitrary 2D vector, each flow is confined to a line; therefore computing the optical flow is reduced to a 1D search problem. Second, when considering the flow of a pixel to different frames tt which are registered to the same plane, Eq. (2) can be written as

where A(x)=d(X)/zA(x)=d(\mathbf{X})/z is the structural component of the flow field, which is independent of tt. It is hence convenient to accumulate structure over time via AA. bt=Tz/d(C2)b_{t}=T_{z}/d(C_{2}), on the other hand, encodes the camera motion to frame tt, and is a single number per frame. To simplify notation, we express the residual flow in terms of the parallax field w(x,t)w(\mathbf{x},t), so that

with q=(e−x)\mathbf{q}=\left(\mathbf{e}-\mathbf{x}\right). Here, ww denotes the flow in pixels along the line towards e\mathbf{e}.

We can thus parametrize the motion across multiple frames as a common structure component AA and per-frame parameters θt={Ht,bt,et}\theta_{t}=\left\{H_{t},b_{t},\mathbf{e}_{t}\right\}. Since we use the center frame of a triplet of frames as the reference and compute the motion to the two adjacent frames, from here on we denote the two parameter sets as θ+={H+,b+,e+}\theta^{+}=\left\{H^{+},b^{+},\mathbf{e}^{+}\right\} for the forward direction and θ−\theta^{-} for the backward direction.

Initialization

Given a triplet of images and a coarse, image-based rigidity estimation (described in Sec. 5.1), the goal of our algorithm is to compute (i) a segmentation into rigid regions and moving objects and (ii) optical flow for the full frame. We start by computing initial motion estimates using an existing optical flow method . For a triplet of images {I−,I,I+}\{I^{-},I,I^{+}\}, we compute four initial flow fields, u0+\mathbf{u_{0}^{+}} from II to I+I^{+} and u0−\mathbf{u_{0}^{-}} from II to I−I^{-}, and their respective backwards flows uˉ0+\mathbf{\bar{u}_{0}^{+}} and uˉ0−\mathbf{\bar{u}_{0}^{-}}. Due to the non-convex nature of our model (see Sec. 6) we need to compute good initial estimates for the P+P parameters θ^+,θ^−\hat{\theta}^{+},\hat{\theta}^{-}, visibility maps V+,V−V^{+},V^{-} denoting which pixels are visible and which are occluded in forward and backward directions, and an initial structure estimate A^\hat{A}.

Initial alignment and epipole detection. First we compute the planar alignments (homographies) between frames. Since P+P only holds in the rigid scene, in this section we only consider points that are marked as rigid by the initial semantic rigidity estimation. While computing a homography between two frames is usually easy, two factors make it challenging in our case: (i) when aligning multiple frames, the plane to which the frames are aligned has to be equivalent for each frame for P+P to work, and (ii) the 3D points corresponding to the four points used to estimate the homographies have to be coplanar for Eq. (3) to hold.

The second step is to ensure the coplanarity of the points inducing the homographies. For this, we can turn around Eq. (3), and simultaneously refine the homographies and estimate the epipoles e{+,−}\mathbf{e}^{\{+,-\}} so that Eq. (3) holds. Let ur=⟨H(x⁡+u0)h⟩−x⁡\mathbf{u_{r}}=\left\langle H(\operatorname{\mathbf{x}}+\mathbf{u_{0}})_{h}\right\rangle-\operatorname{\mathbf{x}} be the residual flow after registration with HH. Each pair x,ur\mathbf{x},\mathbf{u_{r}} defines a residual flow line, and in the noise-free case, the epipole e\mathbf{e} is simply the intersection of these lines. Since the computed optical flow contains noise, we compute the epipole using the method described in , which we found to be sufficiently robust to noise. Therefore, e\mathbf{e} is a function of the optical flow and of the computed homography. Enforcing coplanarity of the homographies is now equivalent to enforcing that the residual flow lines in both directions each pass through a common point as well as possible. The refined homographies are thus computed as

To initialize b+,b−b^{+},b^{-}, we first compute the parallax fields by projecting ur\mathbf{u}_{r} onto the parallax flow lines,

Inserting (6) into (4) and solving for AA, we get

Note that Eq. (3) contains a scale ambiguity between the structure AA and the camera motion parameter bb. Therefore, we can freely choose one of b+,b−b^{+},b^{-}, which only affects the scaling of AA; we choose b^+\hat{b}^{+} so that the initial forward structure A+A^{+} defined by Eq. (7) has a MAD of 1. Since A−A^{-} is a function of b−b^{-} and should be as close as possible to A+A^{+}, we obtain the estimate b^−\hat{b}^{-} by solving

Using b^−\hat{b}^{-}, we compute the initial backward structure A^−\hat{A}^{-} using Eq. (7), and set the full sets of P+P parameters to θ+={H^+,b^+,e^+}\theta^{+}=\{\hat{H}^{+},\hat{b}^{+},\mathbf{\hat{e}}^{+}\}, and θ−\theta^{-} accordingly.

Occlusion estimation. Pixels can become occluded in both directions. In occluded regions, we expect the flow to be wrong, since it can at best be extrapolated. Given the initial flow fields, we compute the visibility masks V+(x)V^{+}(\mathbf{x}), V−(x)V^{-}(\mathbf{x}) using a forward-backward check .

Initial structure estimation. Using the computed structure maps A^{+,−}\hat{A}^{\{+,-\}} and visibility maps V{+,−}V^{\{+,-\}}, the initial estimate for the full structure is

Rigidity estimation

Different cues provide different, complementary information about the rigidity of a region. The semantic category of an object tells us whether it is capable of independent motion, rigid scene parts have to obey the parallax constraint (3), and the 3D structure of rigid parts cannot change over time. We integrate all of them in a probabilistic framework to estimate a rigidity map of the scene, marking each pixel as belonging to the rigid scene or to a moving object.

We leverage the recent progress of CNNs for semantic segmentation to predict rigid and independently moving regions in the scene. In short, we model the relationship between an object’s appearance and its ability to move.

Obviously object appearance alone does not fully determine whether something is moving independently. A car may be moving, if driving, or static, if parked. However, for the purpose of motion estimation, not all errors are the same. Assuming an object is static when in reality it is not imposes false constraints that hurt the estimation of the global motion, while assuming a rigid region is independently moving does little harm. Thus, when in doubt, we predict a region to be independently moving.

The main optical flow benchmarks, KITTI-2015 and MPI-Sintel, provide different training data. While the essence of our model is the same for both, our training process varies to adapt to the available data. In both cases we start with the DeepLab architecture , pre-trained on the 21 classes of Pascal VOC , substitute all fully connected layers with convolutional layers, and densify the predictions . Both networks produce a rigidity score between 0 and 1 which we call the semantic rigidity probability psp_{s}.

MPI-Sintel contains many objects that are not contained in Pascal VOC, such as dragons. Thus using the CNN to predict a semantic segmentation is not possible. Also, no ground truth semantic segmentation is provided, so training a CNN to recognize these categories is not possible. However, the dataset provides ground truth camera calibration, depth and optical flow for the training set. With these we estimate rigidity maps that we take as ground truth. We do this by computing a fully rigid motion field, using the depth and camera calibration, and comparing it with the ground truth flow field. Pixels are classified as independently moving if these two fields differ by more than a small amount. We make this data publicly available .

We modify the last layer of the CNN to predict 2 classes, rigid and independently moving, instead of the original 21. We train using the last 30 frames of each sequence in the training set, and validate on the first 5 frames of each sequence. Sequences shorter than 50 frames are only included in the validation set. At test time, the probability of being rigid is computed at each pixel and then thresholded. Examples of the estimated rigidity maps can be seen in Fig. 2.

In KITTI 2015, some independently moving objects (e.g. people) are masked out from the depth and flow ground truth. Therefore, the approach we followed for MPI-Sintel cannot be used. The objects in KITTI, however, appear in standard datasets like the enriched Pascal VOC. We modify the last layer of the network to predict the 22 classes that may be present in KITTI (e.g. person or road) similar to . We then classify an object as moving if it has the ability to move independently (e.g. cars, or buses) and as rigid otherwise. Training details appear in the Sup. Mat. .

Note that the same approach we use for KITTI can be used for general video sequences by using a generic pre-trained semantic segmentation network together with a definition of which semantic classes can move and which are static. This allows our method to directly benefit from advances in semantic segmentation and novel, fine-grained semantic segmentation datasets.

2 Physical rigidity estimation

For objects that have not been seen previously or that exhibit phenomena like motion blur, the semantic rigidity may be wrong. Hence, we use two additional cues, motion direction and temporal consistency of the structure.

Moving regions from motion direction. A simple approach to classify a pixel as rigid or independently moving is to test whether its parallax flow points to the epipole . Here, we employ a probabilistic framework for this classification. Due to space limitations, we just present the final result here; for the derivation, please see the Sup. Mat. .

For a given point x\mathbf{x}, our model assumes the measured corresponding point x′=x+ur\mathbf{x}^{\prime}=\mathbf{x}+\mathbf{u_{r}} to have a Gaussian error distribution around the true correspondence with covariance matrix Σ=σd2I\Sigma=\sigma_{d}^{2}\mathbf{I}. Let c=∥ur∥c=\|\mathbf{u_{r}}\| and α\alpha be the angle between ur\mathbf{u_{r}} and the line connecting x⁡\operatorname{\mathbf{x}} to e\mathbf{e}. Assuming a uniform distribution of motion directions for moving objects, the likelihood of a point being rigid is then given as

Moving regions from structure consistency. Another cue for rigidity is the temporal consistency of the structure. This is particularly helpful where semantics and motion direction cannot disambiguate the rigidity, for example when an object such as a car moves parallel to the observer’s motion.

Recall that according to the P+P framework the structure of the rigid scene is independent of time. In rigid regions that are visible in all frames, we assume the forward and backward structure A+A^{+} and A−A^{-} to be close to each other. A structure based rigidity estimate psp_{s} can thus be computed as

Combined rigidity probability from motion. The motion-based probabilities pdp_{d}, psp_{s} can be seen as orthogonal. Surfaces that move independently along the parallax direction are considered to be rigid according to pdp_{d}, while surfaces that move by small amounts orthogonal to the parallax direction are considered to be rigid according to psp_{s}. Hence, for a region to be considered actually rigid, we require both pdp_{d} and psp_{s} to be high. The final motion-based rigidity probability pmp_{m} is

3 Combining rigidity estimates

The previously computed rigidity probabilities pcp_{c}, pmp_{m} yield per-pixel rigidity probabilities. To combine those into a coherent estimate, we first compute a rigidity unary

with R(x⁡)=1R(\operatorname{\mathbf{x}})=1 if x⁡\operatorname{\mathbf{x}} is rigid, and 0 otherwise. Since we expect the rigidity to be spatially coherent, we estimate the full labelling by solving R^=\hat{R}=

where wx,yw_{\mathbf{x},\mathbf{y}} is the image-based Potts modulation from and N(x)\mathcal{N}(\mathbf{x}) is the 8-connected neighborhood of x\mathbf{x}. Eq. (16) is solved using TRWS .

Figure 2 (top) shows the importance of combining different cues to recover from errors and accurately estimate the rigidity. The semantic estimation (b) misses a large part of the dragon’s head, while both the direction-based (b) and structure-based estimations misclassify different segments of the scene. Combining cues yields a good estimate (e).

Model and optimization

Model. The final structure should fulfill a number of criteria. First, as in the classical flow approach, warping the images using the flow induced by the structure should result in a low photometric error. Second, we assume that our initial flow fields are reasonable, hence, the final structure should be similar to the structures defined by the initial forward and backward flow. Third, the structure directly corresponds to the surface structure of the world, and thus we can regularize it using a locally planar model. This implicitly regularizes the flow in a more geometrically meaningful way than traditional priors on the flow.

Under these considerations, the full model for the motion of the rigid parts of the scene is defined as E(A,θ+,θ−)=E(A,\theta^{+},\theta^{-})=

EdE_{d} is the photometric error, modulated by the estimated visibilities in forward and backward directions:

where Ia−,Ia,Ia+I_{a}^{-},I_{a},I_{a}^{+} are augmented versions of I−,I,I+I^{-},I,I^{+}, i.e. stacked images containing the respective grayscale images and the gradients in xx and yy directions. The warping function s(x⁡,A,θ)s(\operatorname{\mathbf{x}},A,\theta) defines the correspondence of x according to the structure AA and the P+P parameters θ\theta,

The consistency term EcE_{c} encourages similarity between AA and A{+,−}A^{\{+,-\}}.

To ensure a constant error for all A∈[A−,A+]A\in\left[A^{-},A^{+}\right], we use the Charbonnier function as the robust penalty ρc\rho_{c}.

The locally-planar regularization uses a 2nd order prior,

Here, wx,wyw_{x},w_{y} are again the modulation terms from , and, using a slight abuse of notation, ∇xx,∇xy,∇yy\nabla_{xx},\nabla_{xy},\nabla_{yy} are the second derivative operators. Since the second order prior by itself is highly sensitive to noise, we add a first order prior

where ∇x,∇y\nabla_{x},\nabla_{y} are the first derivative operators in the horizontal and vertical direction respectively.

Optimization. To minimize the energy (17) we employ an iterative scheme, and alternate between optimizing for AA with θ{+,−}\theta^{\{+,-\}} fixed, and for θ{+,−}\theta^{\{+,-\}} with AA fixed. When optimizing AA, we use a standard warping-based variational optimization with 1 inner and 5 outer iterations and no downscaling. To optimize for θ\theta, we first optimize for H,bH,b using L-BFGS and then recompute e\mathbf{e} as described in Sec. 4. We use two iterations, since we found that more do not decrease the error significantly. This yields the final estimates Aˉ,θˉ+,θˉ−\bar{A},\bar{\theta}^{+},\bar{\theta}^{-} for the structure and the P+P parameters.

Due to the non-convex nature of (17), a global optimum is not guaranteed. However, in practice we found that our initializations are close to a good optimum, and hence our optimization procedure works well.

Final flow estimation. Finally, we convert the estimated structure Aˉ\bar{A} into an optical flow field

In the moving regions, we use the initial forward flow u0+\mathbf{u_{0}^{+}}, and compose the full flow field as

Experiments

To quantify our method, we evaluate on the MPI-Sintel and KITTI-2015 flow benchmarks. The parameters are chosen to minimize errors on the training sets, and are set to {σd,σs,λr,c,λr,p,λc,λ1st,λ2nd}={0.75,2.5,0.1,1.1,0,0.1,5e3}\{\sigma_{d},\sigma_{s},\lambda_{r,c},\lambda_{r,p},\lambda_{c},\lambda_{1st},\lambda_{2nd}\}=\{0.75,2.5,0.1,1.1,0,0.1,5e3\} for Sintel and {1.0,0.25,\{1.0,0.25, 0.5,1.1,0.01,1,5e4}0.5,1.1,0.01,1,5e4\} for KITTI. Table 1 shows the errors for our method, our initialization (DF), and for top performing methods on MPI-Sintel (FF+) and KITTI-2015 (SDF) . Both evaluate only on one dataset; in contrast, our method achieves high accuracy on both datasets. Figure 3 visualizes results; for more results see .

On MPI-Sintel, our method currently outperforms all published works. In particular, the structure estimation gives flow in occluded regions, producing the lowest errors in the unmatched regions of any published or unpublished work. On a 2.2 GHz i7 CPU, our method takes on average 2 minutes per triplet of frames without the initial flow computation, 74s for the initialization and rigidity estimation, and 46s for the optimization.

In KITTI-2015 the scenes are simpler and contain only automotive situations; however, the images suffer from artifacts such as noise and overexposures. Among published monocular methods, MR-Flow is second after , which is designed for automotive scenarios and not tested on Sintel.

Conclusion

We have demonstrated an optical flow method that segments the scene and improves accuracy by exploiting rigid scene structure. We combine semantic and motion information to detect independently moving regions, and use an existing flow method to compute the motion of these regions. In rigid regions of the scene, the flow is directly constrained by the 3D structure of the world. This allows us to implicitly regularize the flow by constraining the underlying structure to a locally planar model. Furthermore, since the structure is temporally coherent, we combine information from multiple frames. We argue that this uses the right constraints in the right place and produces accurate flow in challenging situations and competitive results on Sintel and KITTI.

This opens several directions for future work. First, the rigidity estimation could be improved using better inference algorithms and training data. Jointly refining the foreground flow with the rigid flow estimation could improve performance. Our method could also use longer sequences, and enforce temporal consistency of the rigidity maps.

Acknowledgements. JW and LS were supported by the Max Planck ETH Center for Learning Systems.

References

Appendix A Supplemental Material

As described in our paper, each of the two datasets provide different data, and thus we used slightly different procedures to estimate rigidity on each of the datasets. Here we describe the details. Additionally, Fig. 4 and Fig. 5 show some more examples of the outputs of the networks on images that were not seen during test time.

KITTI 2015: We followed the same procedure as previous work on semantic segmentation on KITTI , since it was shown to be successful. We used the DeepLab architecture modifying the output layer to classify 22 classes: aeroplane, bicycle, bird, boat, building, bus, car, cat, cow, dog, floor, grass, horse, motorbike, road, sheep, sidewalk, sky, train, water, person and background, where background includes anything that is not one of the other 21 classes. We initialized the weights with the VGG model trained on Pascal VOC and then fine-tuned it on our categories using a fixed momentum of 0.9, weight decay of 0.0005 and learning rate of 0.0001 for the first 100K iterations, reduced by 0.1 after every 50K steps, during 200K iterations. We used a dense CRF on top, where the unaries are the CNN output at each pixel and the pairwise potentials are position and a bilateral kernel with both position and RGB values. The inference in the dense CRF model is performed using 10 steps of mean-field approximate inference. At test time we obtain a probability over the classes. We estimate rigidity by choosing the class with the highest probability, and classifying the pixel as rigid or non-rigid based on whether an object in the class is capable of moving independently (for example, car) or not (for example, building). The accuracy of rigidity classification on the training set is 96.09%, where rigid parts are correctly classified 96.93% of the time, and independent moving parts are correctly classified 91.51% of the time.

MPI-Sintel: In this dataset there is no previous work on estimating rigidity. Thus, we used one of the latest released versions of the DeepLab architecture, called DeepLab-Coco-LargeFov, which is pretrained including extra annotations from the MS-COCO datasetFurther details can be found on their webpage http://ccvl.stat.ucla.edu/software/deeplab/deeplab-coco-largefov/. We modified the output layer to classify pixels as rigid or nonrigid. We fine-tuned all layers using the same parameters as before for 1.4K iterations. This small number of iterations was selected to avoid overfitting. At test time, we obtain a probability of rigidity, and we compute the final estimate of rigidity by thresholding at 0.5. The accuracy of the estimation on a validation set is 94.2%.

A.2 Derivation of direction-based rigidity likelihood

While the accuracy of the CNN-based rigidity estimation is surprisingly high, there are still occasions when it fails. This may happen for example when there is strong motion blur in the rigid regions, causing these to be classified as independently moving, which is a false positive. On the other hand, a noisy or over-saturated appearance of objects can cause the semantic segmentation to fail and result in a false negative. Therefore, we combine the CNN-based rigidity estimation with a motion-based estimation, described in this section.

For a given point x\mathbf{x}, our model assumes the measured corresponding point x′=x+u\mathbf{x}^{\prime}=\mathbf{x}+\mathbf{u} to have a Gaussian error distribution around the true correspondence with covariance matrix Σ=σ2I\Sigma=\sigma^{2}\mathbf{I}. For a given rigid point (denoted by the conditioning on r=1r=1) and a given focus of expansion e\mathbf{e}, the probability of the true correspondence pointing towards e\mathbf{e} is then given as

where l(x,e)l(\mathbf{x},\mathbf{e}) denotes the line that goes through x\mathbf{x} and e\mathbf{e}. Figure 6 shows an illustration.

Since the error distribution is Gaussian, every marginal is also Gaussian. Therefore, the line integral in (25) is given as

where dd denotes the distance of x′\mathbf{x^{\prime}} to l(x,e)l(\mathbf{x},\mathbf{e}), as shown in Figure 6.

In the following, it will be convenient to express the correspondence point x′\mathbf{x^{\prime}} in (26) in terms of the angle α\alpha and the displacement magnitude c=∥u∥c=\|\mathbf{u}\|. Then,

We can thus drop e\mathbf{e} from the equation.

The likelihood of a point being rigid given the measurements α,c\alpha,c is now

Abbreviating the prior for rigidity p(r=1)p(r=1) as p1p_{1} and setting p(α∣r=0,c)=12πp(\alpha|r=0,c)=\frac{1}{2\pi} (the motion of independently moving regions is supposed to be uniformly distributed), we get

What remains is to compute ZZ. Using Eq. (27), it can be computed as

Appendix B Ablation study

This section provides an additional ablation study that had to be omitted from the paper due to space limitations. Here, we test how different subcomponents of our algorithm impact the end resultThe results in this section were obtained using a reduced version of the MPI-Sintel training set containing every 4th frame, and only the final pass.. To assess the impact in different regions of the frame, we provide the errors both on the full frames and only in the ground truth rigid regions.

For the ablation study, we successively switch on four steps: occlusion reasoning, coplanarity refinement, nonlinear initialization of b^−\hat{b}^{-}, and spatial priors.

Occlusion reasoning refers to the estimation of the visibility maps V+,V−V^{+},V^{-} using the forward-backward consistency, as described in Section 4 in the paper. If this step is switched off, we set V−=V+=1V^{-}=V^{+}=1 everywhere, and therefore do not explicitly exclude occluded pixels from subsequent computations.

Nonlinear initialization of b^−\hat{b}^{-} refers to initializing b^−\hat{b}^{-} using Eq.(8), i.e. choosing b^−\hat{b}^{-} so that the resulting backward A^−\hat{A}^{-} structure is as similar as possible under a robust error norm to the initial forward structure A^+\hat{A}^{+}. Note that, without loss of generality, b^+\hat{b}^{+} is always chosen so that the MAD of A^+\hat{A}^{+} is 1.

If this step is switched off, use Eq. (7) to equate A^+\hat{A}^{+} and A^−\hat{A}^{-}. This produces an estimate of b^−\hat{b}^{-} per pixel, of which we take the median to arrive at the global estimate of b^−\hat{b}^{-}.

Spatial priors refer to the 1st- and 2nd-order spatial smoothness regularizers in our objective function (17). To disable those, we set λ1st=λ2nd=0\lambda_{1st}=\lambda_{2nd}=0.

Table 2 shows the improvement when successively switching on more parts of the algorithm. The occlusion reasoning has the largest positive impact on the error, since it allows the algorithm to properly merge the flow in both directions from the reference frame. Following this, the most important parts are the nonlinear initialization and ensuring the coplanarity. The spatial priors serve mostly to remove flow noise near boundaries. This improves the result visually, but has a fairly small numerical impact.

Table 3 shows the impact when turning off individual components, but leaving all others intact. Again, we can observe that disabling the occlusion reasoning has the largest negative impact, followed by the nonlinear initialization and ensuring the coplanarity.

Table 4 shows the impact of the different terms of the variational refinement (Eq. (17)). The cases are as described above. In addition, table 4 includes a complete omission of the variational refinement (no-opt), which uses only the intial structure estimate as described by Eq. (9) in the paper, and selective disabling of the individual regularizers (λ1st=0\lambda_{1st}=0 for no-1st and λ2nd=0\lambda_{2nd}=0 for no-2nd).

When using the data term only (no-spatial-priors), the error is higher than when not using any optimization. Due to effects in the Sintel final pass such as motion blur, fog, vignetting etc. this is to be expected. Using the 2nd order regularization improves the results; interestingly, however, the impact of the 1st order regularization is negligible.

Appendix C Failure cases

This section gives examples of when our algorithm fails. We consider our algorithm to fail when the flow computed by our algorithm is worse than the initial flow . For each of Sintel clean, Sintel final, and KITTI we give one failure example. All examples are among the worst overall in the respective training sets. In our training set, we observe two primary sources of error, segmentation failures and alignment failures.

Figures 7 and Figure 8 show examples for the first type of error, segmentation failures. In these cases, moving regions are mistaken as parts of the rigid background, such as the car in Fig. 7 or the girl’s head in Fig. 8. These failures occur if the CNN does not pick up a region strongly enough and if, at the same time, the motion of the object is consistent with the motion of the rigid scene. In Fig. 7, the CNN picks up only the frontal part of the car. Since in this example the camera is not moving, the focus of expansion is mistakenly determined by the few parts of the frame that move (i.e. the car), and the motion-based rigidity estimation cannot correct the mistake made by the CNN.

In Fig. 8, the camera pans to the left, and at the same time, the head moves to the right. Since both directions are approximately parallel, the head is considered to be rigid. Note how in this case the estimated flow has a very similar hue to the ground truth flow, even in most of the regions that are misclassified. This confirms that the direction of the flow is approximately consistent with the motion of the rigid parts of the scene; however, since the head still moves slightly out of the rigidity constraints, our method increases the error over the initialization.

Figure 9 shows the second type of error, a failure to align the images. As can be seen in Fig. 9(a), the background in this sequence contains heavy motion blur and a slight vignetting. Together, these two effects cause a high uncertainty of the initial optical flow in the background regions, which in turn causes our initial alignment procedure to fail.