Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects

Martin Rünz, Lourdes Agapito

I INTRODUCTION

The wide availability of affordable structured light and time of flight depth sensors has had enormous impact both on the democratization of the acquisition of 3D models in real time from hand-held cameras and on providing robots with powerful but low-cost 3D sensing capabilities. Tracking the motion of a camera while maintaining a dense representation of the 3D geometry of its environment in real time has become more important than ever .

While solid progress has been made towards solving this problem in the case of static environments, where the only motion is that of the camera, dealing with dynamic scenes where an unknown number of objects might be moving independently is significantly harder. The typical strategy adopted by most systems is to track only the motion of the camera relative to the static background and treat moving objects as outliers whose 3D geometry and motion is not modeled over time. However, in robotics applications often it is precisely the objects moving in the foreground that are of most interest to the robot. If we want to design robots that can interact with dynamic scenes it is crucial to equip them with the capability to (i) discover objects in the scene via segmentation (ii) track and estimate the 3D geometry of each object independently. These high level object-based representations of the scene would greatly enhance the perception and physical interaction capabilities of a robot.

Consider for instance a SLAM system on-board a self-driving car – tracking and maintaining 3D models of all the moving cars around it and not just the static parts of the scene could be critical to avoid collisions. Or think of a robot that arrives at a scene without a priori 3D knowledge about the objects it must interact with – the ability to segment, track and fuse different objects would allow it actively to discover and learn accurate 3D models of them on the fly through motion, by picking them up, pushing them around or simply observing how they move. An object level scene description of this kind, has the potential to enable the robot to interact physically with the scene.

In this paper we introduce Co-FUSION a new RGB-D based SLAM system that can segment a scene into the background and different foreground objects, using either motion or semantic cues, while simultaneously tracking and reconstructing their 3D geometry over time. Our underlying assumption is that objects of interest can be detected and segmented in real-time using efficient segmentation algorithms and then tracked independently over time. Our system offers two alternative grouping strategies – motion segmentation that groups together points that move consistently in 3D and object instance segmentation that both detects and segments individual objects of interest (at the pixel level) in an RGB image given a semantic label. These two forms of segmentation allow us not only to detect objects due to their motion but also objects that might be static but are semantically of interest to the robot.

Once detected and segmented, objects are added to the list of active models and are subsequently tracked and their 3D shape model updated by fusing only the data labeled as belonging to that object. The tracking and fusion threads for each object are based on recent surfel-based approaches . The main contribution of this paper is a system that would allow a robot not only to reconstruct its surrounding environment but also to acquire the detailed 3D geometry of unknown objects that move in the scene. Moreover, our system would equip a robot with the capability to discover new objects in the scene and learn accurate 3D models of them through active motion. We demonstrate Co-Fusion on different scenarios – placing different previously unseen objects on a table and learning their geometry (see Figure Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects), handing over an object from one person to another (see Figure 2), hand-held 3D capture of a moving object with a moving camera (see Figure 8) and on a car driving scenario (see Figure 4a). We also demonstrate quantitatively the robustness of the tracking and the reconstruction on some synthetic and ground truth sequences of dynamic scenes.

II RELATED WORK

The arrival of the Microsoft Kinect device and the sudden availability of inexpensive depth sensors to consumers, triggered a flurry of research aimed at real-time 3D scanning. Systems such as KinectFusion first made it possible to map the 3D geometry of arbitrary indoor scenes accurately and in real time, by fusing the images acquired by the depth camera simply by moving the sensor around the environment. Access to accurate and dense 3D geometry in real time opens up applications to rapid scanning or prototyping, augmented/virtual reality and mobile robotics that were previously not possible with offline or sparse techniques. Successors to KinectFusion have quickly addressed some of its shortcomings. While some have focused on extending its capabilities to handle very large scenes or to include loop closure others have robustified the tracking or improved memory and scale efficiency by using point-based instead of volumetric representations that lead to increased 3D reconstruction quality . Achieving higher level semantic scene descriptions by using a dense planar representation or real-time 3D object recognition further improved tracking performance while opening the door to virtual or even real interaction with the scene. More recent approaches such as incorporate semantic segmentation and even recognition within a SLAM system in real time. While they show impressive performance, they are still limited to static scenes.

The core underlying assumption behind many traditional SLAM and dense reconstruction systems is that the scene is largely static. How can these dense systems be extended to track and reconstruct more than one model without compromising real time performance? The SLAMMOT project represented an important step towards extending the SLAM framework to dynamic environments by incorporating the detection and tracking of moving objects into the SLAM operation. It was mostly demonstrated on driving scenarios and limited to sparse reconstructions. It is only very recently that the problem of reconstruction of dense dynamic scenes in real time has been addressed. Most of the work has been devoted to capturing non-rigid geometry in real time with RGB-D sensors. The assumption here is that the camera is observing a single object that deforms freely over time. DynamicFusion is a prime example of a monocular real time system that can fuse together scans of deformable objects captured from depth sensors without the need for any pre-trained model or shape template. With the use of a sophisticated multi-camera rig of RGB-D sensors 4DFusion can capture live deformable shapes with an exceptional level of detail and can deal with large deformations and changes in topology. On the other hand template based techniques can also obtain high levels of realism but are limited by their need to add a preliminary step to capture the template or are dedicated to tracking specific objects by their use of hand-crafted or pre-trained models . These include general articulated tracking methods that either require a geometric template of the object in a rest pose , or prior knowledge of the skeletal structure .

In contrast, capturing the full geometry of dynamic scenes that might contain more than one moving object has received more limited attention. Ren et al. propose a method to track and reconstruct 3D objects simultaneously by refining an initial simple shape primitive. However, in contrast to our approach, it can only track one moving object and requires a manual initialization. propose a combined approach for estimating pose, shape, and the kinematic structure of articulated objects based on motion segmentation. While it is also based on joint tracking and segmentation, the focus is on discovering the articulated structure, only foreground objects are reconstructed and its performance is not real time. Stückler and Behnke propose a dense rigid-body motion segmentation algorithm for RGB-D sequences. They only segment the RGB-D images and estimate the motion but do not simultaneously reconstruct the objects. Finally build a model of the environment and consider as new objects parts of the scene that become inconsistent with this model using change detection. However, this approach requires a human in the loop to acquire known-correct segmentation and does not provide real time operation.

Several recent RGB-only methods have also addressed the problem of monocular 3D reconstruction of dynamic scenes. Works such as are similar in spirit to our simultaneous segmentation, tracking and reconstruction approach. Russell et al. perform multiple model fitting to decompose a scene into piecewise rigid parts that are then grouped to form distinct objects. The strength of their approach is the flexibility to deal with a mixture of non-rigid, articulated or rigid objects. Fragkiadaki et al. follow a pipeline approach that first performs clustering of long term tracks into different objects followed by non-rigid reconstruction. However, both of these approaches act on sparse tracks and are batch methods that require all the frames to have been captured in advance. Our method also shares commonality with the dense RGB multi-body reconstruction approach of , who also perform simultaneous segmentation, tracking and 3D reconstruction of multiple rigid models, with the notable difference that our approach is online and real time while theirs is batch and takes several seconds per frame.

III OVERVIEW OF OUR METHOD

Co-Fusion is a live RGB-D SLAM system that processes each new frame in real time. As well as maintaining a global model of the detailed geometry of the background our system stores models for each object segmented in the scene and is capable of tracking their motions independently. Each model is stored simply as a set of 3D points. Our system maintains two sets of object models: while active models are objects that are currently visible in the live frame, inactive models are objects that were once visible, therefore their geometry is known, but are currently out of view.

Figure 1 illustrates the frame-to-frame operation of our system. At the start of live capture, the scene is initialized to contain a single active model – the background. Once the fused 3D model of the background and the camera pose are stable after a few frames our system follows the pipeline approach described below. For each new frame acquired by the camera the following steps are performed:

Tracking First, we track the 6DOF rigid pose of each active model in the current frame. This is achieved by minimizing an objective function independently for each model that combines a geometric error based on dense iterative closest point (ICP) alignment and a photometric cost based on the difference in color between points in the current live frame and the stored 3D model.

Segmentation In this step we segment the current live frame associating each of its pixels with one of the active models/objects. Our system can perform segmentation based on two different cues: (i) motion and (ii) semantic labels. We now describe each of these two grouping strategies.

(i) Motion segmentation We formulate motion segmentation as a labeling problem using a fully connected Conditional Random Field and optimize it in real time on the CPU with the efficient approach of . The unary potentials encode the geometric ICP cost incurred when associating a pixel with a rigid motion model. The optimization is followed by the extraction of connected components in the segmented image. If the connected region occupied by outliers has sufficient support an object is assumed to have entered the scene and a new model is spawned and added to the list.

(ii) Multi-class image segmentation As an alternative to motion segmentation our system can segment object instances at the pixel level given a class label using an efficient state of the art approach based on deep learning. This allows us to segment objects based on semantic cues. For instance, in an autonomous driving application our system could segment not just moving but also stationary cars.

Fusion Using the newly estimated 6-DOF pose, the dense 3D geometry of each active model is updated by fusing the points labeled as belonging to that model. We used a surfel-based fusion approach related to the methods of and .

While the tracking and fusion steps of our pipeline run on the GPU, the segmentation step runs on the CPU. The result is an RGB-D SLAM system that can maintain an up-to-date 3D map of the static background as well as detailed 3D models for up to 5 different objects at 12 frames per second.

IV NOTATION AND PRELIMINARIES

V TRACKING ACTIVE MODELS

For each input frame at time tt and for each active model Mm\mathcal{M}_{m} we track its global pose Ttm{\bf T_{tm}} by registering the current live depth map with the predicted depth map in the previous frame, obtained by projecting the stored 3D model using the estimated pose for t−1t-1. We track each active model independently by running the optimization described below selecting only the 3D map points that are labeled as belonging to that specific model.

For each active model Mm\mathcal{M}_{m}, we minimize a cost function that combines a geometric term based on point-to-plane ICP alignment and a photometric color term that minimizes differences in brightness between the predicted color image resulting from projecting the stored 3D model in the previous frame and the current live color frame.

This cost function is closely related to the tracking threads of other RGB-D based SLAM systems . However, the most notable difference is that while assume that the scene is static and only track a single model, Co-Fusion can track various models while maintaining real-time performance.

V-B Geometry Term

For each active model mm in the current frame tt we seek to minimize the cost of the point-to-plane ICP registration error between (i) the 3D back-projected vertices of the current live depth map and (ii) the predicted depth map of model mm from the previous frame t−1t-1:

where vti{\bf v_{t}^{i}} is the back-projection of the ii-th vertex in the current depth-map Dt\mathcal{D}_{t}; and vi{\bf v^{i}} and ni{\bf n^{i}} are respectively the back-projection of the ii-th vertex of the predicted depth-map of model mm from the previous frame t−1t-1 and its normal. Tm{\bf T_{m}} describes the transformation that aligns model mm in the previous frame t−1t-1 with the current frame tt.

V-C Photometric Color Term

Given (i) the current depth image; (ii) the current estimate of the 3D geometry of each active model; and (iii) the estimated rigid motion parameters that align each model with respect to the previous frame t−1t-1, it is possible to synthesize projections of the scene onto a virtual camera aligned with the previous frame.

The tracking problem then becomes one of photometric image registration where we minimize the brightness constancy between the live frame and the synthesized view of the 3D models in frame t−1t-1.

where Tm{\bf T_{m}} is the rigid transformation that aligns active model Mm\mathcal{M}_{m} between the previous frame t−1t-1 and the current frame and It−1(⋅){\bf I_{t-1}(\cdot)} is a function that provides the color attached to a vertex on the model in the previous frame t−1t-1.

For reasons of robustness and efficiency this optimization is embedded in a coarse-to-fine approach using a 4-layer spatial pyramid. Our GPU implementation builds on the open source code release of .

VI MOTION SEGMENTATION

Following the tracking step we have new estimates for the MtM_{t} rigid transformations {Ttm}\{{\bf T_{tm}}\} that describe the absolute pose of each active model with respect to the global reference frame at time tt.

In practice, to allow the motion segmentation to run in real time on the CPU, we first over segment the current frame into SLIC super-pixels using the fast implementation of and apply the labeling algorithm at the super-pixel level. The position, color and depth of each super-pixel is estimated by averaging those of the pixels inside it.

We follow the energy minimization approach of that optimizes the following cost function with respect to the labeling xt∈LS{\bf x_{t}}\in{\mathcal{L}}^{S}

where ii and jj are indices over the image super-pixels ranging from 11 to SS (the total number of super-pixels).

The pairwise potentials ψp(xi,xj)\psi_{p}(x_{i},x_{j}) can be expressed as

where μ(xi,xj)\mu(x_{i},x_{j}) encapsulates the classic Potts model that penalizes nearby pixels taking different labels, and km(fi,fj)k_{m}(f_{i},f_{j}) are contrast-sensitive potentials that measure the similarity between the appearance of pixels. This results in a cost that encourages super-pixels ii and jj to take the same label if the distance between their feature vectors fif_{i} and fjf_{j} is small. In practice we characterize each super-pixel ii with the 6D feature vector fif_{i} that encodes its 2D location, RGB color and depth value. We set kmk_{m} to be Gaussian kernels

km(fi,fj)=exp⁡(−12(fi−fj)TΛm(fi−fj))k_{m}(f_{i},f_{j})=\exp(-{{1}\over{2}}(f_{i}-f_{j})^{T}\Lambda_{m}(f_{i}-f_{j})) with Λm\Lambda_{m} the inverse covariance matrix In practice we set K=2K=2 and the inverse covariance matrices to Λ1=\Lambda_{1}=diag(1/θα2,1/θα2,1/θβ2,1/θβ2,1/θβ2,1/θγ2)(1/\theta_{\alpha}^{2},1/\theta_{\alpha}^{2},1/\theta_{\beta}^{2},1/\theta_{\beta}^{2},1/\theta_{\beta}^{2},1/\theta_{\gamma}^{2}) and Λ2=\Lambda_{2}=diag(1/θδ2,1/θδ2,0,0,0,0)(1/\theta_{\delta}^{2},1/\theta_{\delta}^{2},0,0,0,0).

We use the efficient inference method of to optimize the labeling, which can be computed in real time on the CPU. The output of this optimization is a soft assignment of labels to each super-pixel ii. To convert this into a hard assignment we simply take the maximum of all the label assignments and associate each super-pixel with the motion of a single active model.

Post-processing. Following the segmentation we perform a series of post-processing steps to obtain more robust results. First we perform connected components for all the labels and we merge models that have similar rigid transformations. Secondly we ensure that disconnected regions are modeled separately by suppressing all except the largest component with the same label. In a similar way, components whose size falls below a threshold τ\tau are removed.

If the connected region occupied by outliers is larger than 3%3\% of the total number of pixels, an object is assumed to have entered the scene and a new label/object is spawned. If part of the geometry of this new object was already in the map (for instance, if an object started moving after having been part of the background map for a while) we attempt to remove the duplicate reconstruction. In practice we found that a good strategy is to remove areas with a high ICP error from the background. This is illustrated in Figure 5.

On the other hand, if a label disappeared and does not reappear within a certain number of frames, it is assumed that the respective model left the scene. In this case the model will be added to the inactive list, if it contains enough surfels with a high confidence and is deleted otherwise.

VII OBJECT INSTANCE SEGMENTATION

In this section we investigate the use of semantic cues to segment objects in the scene which allows to deal both with moving and static objects. We use the top performing state of the art method for object instance segmentation to segment objects of interest. SharpMask is an augmented feed-forward network able to predict object proposals and object masks simultaneously. The architecture has 3 elements: A pre-trained network for feature map extraction, a segmentation branch and a branch that scores the ‘objectness’ of an image patch. The results of SharpMask (an example segmentation can be seen in Figure 4a) can be given directly to Co-Fusion after temporal consistency is imposed between consecutive frames. The segmentation can be run on a limited set of labels to segment only objects of a chosen class, for instance all the tools lying on a table. We used the publicly available models pre-trained on the COCO dataset .

VIII FUSION

During the tracking stage, active models Mm\mathcal{M}_{m} are projected to the camera view using splat rendering in order to align individual model poses. In the subsequent fusion stage the surfel maps are updated by merging the newly available RGB-D frame into the existing models. After projectively associating image coordinates u{\bf u} with corresponding surfels in the model Mm\mathcal{M}_{m}, an update scheme similar to is used.

IX Evaluation

We carried out a quantitative evaluation both on synthetic and real sequences with ground truth data. Appropriate synthetic sequences with Kinect-like noise were specifically created for this work (ToyCar3 and Room4) and have been made publicly available, along with evaluation tools. For the ground truth experiments on real data we attached markers to a set of objects, as shown in Figure 9, and accurately reconstructed them using a NextEngine 3D-scanner. The scenes were recorded with a motion-capture system (OptiTrack) to obtain ground-truth data for the trajectories. An Asus Xtion was used to acquire the real sequences. Although the quality of each stage in our pipeline depends on the performance of every other stage, i.e. a poor segmentation might be accountable for a poor reconstruction, it is valuable to evaluate the different elements.

Pose estimation We compared the estimated and ground-truth trajectories by computing the absolute trajectory (AT) root-mean-square errors (RMSE) for each of the objects in the scene. Results on synthetic sequences are shown in table II and Figure 6. Results on the real GT sequences comparing estimated and GT trajectories (given by OptiTrack) can be found in supplementary material Please see http://visual.cs.ucl.ac.uk/pubs/cofusion/index.html for additional experimental evaluation and video..

Motion segmentation As the result of the segmentation stage is purely 2D, conventional metrics for segmentation quality can be used. We calculated the intersection-over-union measure per label for each frame of the synthetic sequences (we did not have ground truth segmentation for the real sequence). Figure 6 shows the IoU for each frame in the ToyCar3 and Room4 sequences.

Fusion To assess the quality of the fusion, one could either inspect the 3D reconstruction errors of each object separately or jointly, by exporting the geometry in a unified coordinate system. We used the latter on the synthetic sequences. This error is strongly conditioned on the tracking, but nicely highlights the quality of the overall system. For each surfel in the unified map of active models, we compute the distance to the closest point on the ground-truth meshes, after aligning the two representations. Figure 7 visualizes the reconstruction error as a heat-map and highlights differences to Elastic-Fusion. For the real scene Esone1 we computed the 3D reconstruction errors of each object independently. The results are shown in Table I and Figure 9.

Qualitative results We performed a set of qualitative experiments to demonstrate the capabilities of Co-Fusion. One of its advantages is that it eases the 3D scanning process, since we do not need to rely on the static-world assumption. In particular, a user can hold and rotate an object in one hand while using the other to move a depth-sensor around the object. This mode of operation offers more flexibility, when compared to methods that require a turntable, for instance. Figure 8 shows the result of such an experiment.

Our final demonstration shows Co-Fusion continuously tracking and refining objects as they are placed on a table one after the other, as depicted in Figure Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects. This functionality can be useful in robotics applications, where objects have to be moved by an actuator. The result of the successful segmentation is shown in Figure Co-Fusion: Real-time Segmentation, Tracking and Fusion of Multiple Objects(b).

X CONCLUSIONS

We have presented Co-Fusion, a real time RGB-D SLAM system capable of segmenting a scene into multiple objects using motion or semantic cues, tracking and modeling them accurately while also maintaining a model of the environment. We have demonstrated its use in robotics and 3D scanning applications. The resulting system could enable a robot to maintain a scene description at the object; even in the case of dynamic scenes.

ACKNOWLEDGMENT

This work has been supported by the SeconHands project, funded from the EU Horizon 2020 Research and Innovation programme under grant agreement No 643950.

References