MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects

Martin Rünz, Maud Buffier, Lourdes Agapito

RELATED WORK

The field of Visual SLAM has a long history of offering solutions to the problem of jointly tracking the pose of a moving camera (see for a recent survey) while reconstructing a map of the environment. The advent of inexpensive, consumer-grade RGB-D cameras – such as the Microsoft Kinect – stimulated further research, and enabled the leap to dense real-time methods .

Dense RGB-D SLAM: Resulting methods are capable of accurately mapping indoor environments and gained popularity in augmented reality and robotics. KinectFusion proved that a truncated signed distance function (TSDF) based map representation can achieve fast and robust mapping and tracking in small environments. Subsequent work showed that the same principles are applicable to large scale environments by choosing appropriate data structures.

Surface elements (surfels) have a long history in computer graphics and have found many applications in computer vision . More recently, surfel-based map representations were also introduced to the domain of RGBD-SLAM. A map of surfels is similar to a point cloud with the difference that each element encodes local surface properties – typically a radius and normal – in addition to its location. In contrast to a TSDF-based map, surfel clouds are naturally memory efficient and avoid the overhead due to switching representations between mapping and tracking that is typical of TSDF-based fusion methods. Whelan et al. presented a surfel-based RGBD-SLAM system for large environments with local and global loop-closure.

Scene segmentation: The computer graphics and vision communities have devoted substantial effort to object and scene segmentation. Segmented data can broaden the functionality of visual tracking and mapping systems, for instance, by enabling robots to detect objects. Some methods have proposed to segment RGBD data based on geometric properties of surface normals , mainly by assuming that objects are convex. While the clear strength of geometry-based segmentation systems is that they produce accurate object boundaries, their weakness is that they typically result in over-segmentations and they do not convey any semantic information.

Semantic scene segmentation: Another line of work aims at segmenting 3D scenes semantically, using Markov Random Fields (MRFs). These methods require labelled 3D data, however, which in contrast to labelled 2D image data is not readily available. This is exemplified by the fact that all three works involved manual annotation of training data. Datasets containing isolated RGBD frames, such as NYUv2 , are not applicable here and it requires significant effort to build consistent reconstructed datasets for segmentation, as recently shown by Dai et al. .

Semantic SLAM: Motivated by the success of convolutional neural networks, Tateno et al. and McCormac et al. integrate deep neural networks in real-time SLAM systems. As inference is solely based on 2D information, the need for 3D annotated data is circumvented. The resulting systems offer strategies to fuse labelled image data into segmented 3D maps. Earlier work by Hermans et al. implements a similar scheme, using a randomised decision forest classifier. However, since the systems are not considering object instances, tracking multiple models independently is unattainable.

Dynamic SLAM: There are two main scenarios in dynamic SLAM: non-rigid surface reconstruction and multibody formulations for independently moving rigid objects. In the first case, a deformable world is assumed and as-rigid-as-possible registration is performed, while in the second, rigid object instances are identified and tracked sparsely or densely . Both categories use template- or descriptor-based formulations , which require pre-observing objects of interest, and template-free methods. In the case when the dynamic parts of the scene are not of interest, it is valuable to recognise them as outliers to avoid errors in the optimisation back-end. Methods for the explicit detection of dynamic regions for static fusion were proposed by Jaimez et al. and Scona et al. .

Table 1 provides an overview of related real-time capable methods comparing them under five important properties.

Only two dynamic SLAM system, to the best of our knowledge, have previously attempted incorporates semantic knowledge, but both fall short of the functionality of MaskFusion. Co-Fusion demonstrated the ability to track, segment and reconstruct objects based on their semantic labels, but the overall system was not real-time capable and limited functionality was shown. DynSLAM developed a mapping system for autonomous driving applications capable of separately reconstructing both the static environment and the moving vehicles. However, the overall system was not real-time (this is the reason it does not appear in Table 1) and vehicle was the only dynamic object-class it reconstructed, so its functionality was limited to road scenes.

System Overview

MaskFusion enables real-time dense dynamic RGBD SLAM at the level of objects. In essence, MaskFusion is a multi-model SLAM system that maintains a 3D representation for each object that it recognises in the scene (in addition to the background model). Each model is tracked and fused independently. Figure 1 illustrates its frame-to-frame operation. Each time a new frame is acquired by the camera, the following steps are performed: Tracking: The 3D geometry of each object is represented as a set of surfels. The six degree of freedom pose of each model is tracked by minimizing an energy that combines a geometric iterative closest point (ICP) error with a photometric cost based on brightness constancy between corresponding points in the current frame and the stored 3D model, aligned with the pose in the previous frame. In order to lower computational demand and increase robustness, only non-static objects are tracked separately. Two different strategies were tested to decide whether an object is static or not: one based on motion inconsistency, similar to , and another that treats objects which are being touched by a person as dynamic. Segmentation: MaskFusion combines two types of cues for segmentation: semantic and geometric cues. Mask-RCNN is used to provide object masks with semantic labels. While this algorithm is impressive and provides good object masks, it suffers from two drawbacks. First, the algorithm does not run in real time and can only operate at a maximum of 5 Hz. Second, the object boundaries are not perfect – they tend to leak into the background. To overcome both of these limitations, we run a geometric segmentation algorithm, based on an analysis of depth discontinuities and surface normals. In contrast to the semantic instance segmentation, the geometric segmentation runs in real time and produces very accurate object boundaries (see Figures 2(d) and (e) for an example visualisation of the geometric edge map and the geometric components returned by the algorithm). On the negative side, geometry-based segmentation tends to oversegment objects. The combination of these two segmentation strategies – geometric segmentation on a per-frame basis and semantic segmentation as often as possible – provides the best of both worlds, allowing us to (1) run an overall system in real time (geometric segmentation is used for frames without semantic object masks, while the combination of both is used for frames with object masks) and (2) obtain semantic object masks with improved object boundaries, thanks to the geometric segmentation. Fusion: The geometry of each object is fused over time by using the object labels to associate surfels with the correct model. Our fusion follows the same strategy as .

The rest of the paper is organised as follows. We first describe the principles of our dynamic RGBD-SLAM method in Section 3; further details regarding the integration of the semantic and geometric segmentation results are provided in Section 4. A quantitative and qualitative evaluation of the proposed approach is presented in Section 5.

MULTI-OBJECT SLAM

The alignment is performed by minimising a joint geometric and photometric error function :

where EmicpE^{icp}_{m} and EmrgbE^{rgb}_{m} are the geometric and photometric error terms respectively and ξm\mathbf{\xi}_{m} is the unknown rigid transformation, expressed in a minimal 6D Lie algebra representation se3\mathfrak{se}_{3}, which is subject to optimisation.

The first term in equation (1) is a sum of projective ICP residuals. Given a vertex vti\mathbf{v}_{t}^{i}, which is the back-projection of the ii-th vertex in Dt\mathcal{D}_{t}; and vi\mathbf{v}^{i} and ni\mathbf{n}^{i}, the corresponding vertex and normal in Dt−1a\mathcal{D}^{a}_{t-1} (the geometry expressed in the camera coordinate frame at time t−1t-1), EmicpE^{icp}_{m} is written as:

The photometric term, on the other hand, is a sum of photo-consistency residuals between It\mathcal{I}_{t} and It−1a\mathcal{I}^{a}_{t-1}, and reads as follows:

2 Fusion

Given Rtm\mathbf{R}_{tm} and ttm\mathbf{t}_{tm}, surfels for each model Mm\mathcal{M}_{m} are updated by performing a projective data association with the current RGBD frame. This step is inspired by but a stencilling based on the segmentation discussed in Section 4 is used to adhere to object boundaries. As a result, each newly created surfel is part of exactly one model. Further, we introduce a confidence penalty for surfels outside the stencil, which is required due to imperfect segmentations.

SEGMENTATION

MaskFusion reconstructs and tracks multiple objects simultaneously, maintaining separate models. As a consequence, new data has to be associated with the correct model before fusion is performed. Inspired by Co-Fusion , instead of associating data in 3D, segmentation is carried out in 2D and model-to-segment correspondences are established. Given these correspondences, new frames are masked and only subsets of the data are fused with existing models. Masking is based on the semantic instance segmentation labels proposed by a DNN , in conjunction with geometric segmentation, which improves the quality of object boundaries. Our semantic segmentation pipeline provides masks at 30Hz or more.

The design of the pipeline is based on the following observations: (i) Current semantic segmentation methods are good at detecting objects, but tend to provide imperfect object boundaries. (ii) The current state-of-the-art approach, Mask-RCNN , cannot be executed at frame rate. (iii) The information contained in RGBD frames enables fast over-segmentation of the image, for instance by assuming object convexity.

The second observation directly implies that to achieve overall real-time performance our system must execute instance level semantic segmentation in a parallel thread concurrently to the tracking and fusion threads. However, executing two programs at different frequencies concurrently requires a synchronisation strategy. We buffer new frames in a queue QfQ_{f} and refer the SLAM system to the head of the queue, while the semantic segmentation operates on the back of the queue, as illustrated in Figure 1a. This way, the execution of the SLAM pipeline is delayed by the worst-case processing time of the semantic segmentation. In our experiments we picked a queue length of 12 frames, which involves a delay of approx. 400ms. Whether this delay can be neglected or not, depends on the use-case of the system. Even though a latency exists, the system runs at a frame-rate of 30fps. Furthermore, a semantic segmentation is not available for most frames due to the lower execution frequency of the masking component, yet each frame requires a labelling in order to fuse new data. This issue is solved by associating regions of mask-less frames with existing models only, as discussed in Section 4.3.

To compensate for inexact boundaries, as mentioned in observation 1, we make use of observation 3 and map components from a geometric over-segmentation to semantic masks. This results in improved masks, due to higher-quality boundaries of the geometric segmentation.

Mask-RCNN achieves this by extending the Faster-RCNN architecture. Faster-RCNN is a two-stage approach that proposes regions of interest first and then predicts an object class and bounding box per region and in parallel. He et al. added a third branch to the second stage, which generates masks independently of class IDs and bounding boxes. Both stages rely on a feature map, which is extracted by a ResNet-based backbone network, and apply convolutional layers for inference.

Figure 2c visualises the output of Mask-RCNN. Note that instances of the same class are highlighted with different colours, and also that masks are not perfectly aligned with object boundaries.

2 GEOMETRIC SEGMENTATION

Assuming that objects – especially man-made objects – are largely convex, it is possible to build fast segmentation methods that place edges in concave areas and depth discontinuities. In practice, such methods tend to oversegment data, due to the simplified premise. Moosmann et al. successfully segment 3D laser data based on this assumption. The same principle is also used by other authors to segment objects in RGBD frames .

Our geometric segmentation method follows this approach and, similarly to , generates an edginess-map based on a depth discontinuity term ϕd\phi_{d} and concavity term ϕc\phi_{c}. Specifically, a pixel is defined as an edge pixel if ϕd+λ^ϕc>τ\phi_{d}+\hat{\lambda}\phi_{c}>\tau, where τ\tau is a threshold and λ^\hat{\lambda} a relative weight. Given a local neighbourhood N\mathcal{N}, ϕd\phi_{d} and ϕc\phi_{c} are computed as follows:

Here, v\mathbf{v} and vi\mathbf{v}_{i} indicate vertex positions, while n\mathbf{n} and ni\mathbf{n}_{i} represent normals, obtained by back-projecting Dt\mathcal{D}_{t}. Since ϕd+λ^ϕc\phi_{d}+\hat{\lambda}\phi_{c} depends on a local neighbourhood only, the edginess of a pixel can be evaluated quickly on a GPU. Figure 2d shows the edge map for a frame that was captured with an Asus Xtion RGBD-camera. Edge maps are converted to a geometric labelling Ltg:Ω→{0..Ntg}\mathcal{L}_{t}^{g}:\Omega\rightarrow\{0..N_{t}^{g}\}, where NtgN_{t}^{g} is the number of extracted components excluding the background, by running an out-of-the-box connected components algorithm, as illustrated in Figure 2e.

3 MERGED SEGMENTATION

For each frame that is processed by the SLAM system, the pipeline illustrated in Figure 4 is executed. While the geometric segmentation, shown on the left-hand-side, is performed for all frames, geometric labels are mapped to semantic masks only if these are available. In the absence of semantic masks, geometric labels are associated with existing models directly and the following steps are skipped:

After over-segmenting input frames geometrically, the resulting components Cti  ∀i∈{1..Ntg}C_{ti}\;\forall i\in\{1..N_{t}^{g}\} are mapped to masks Ltns\mathcal{L}_{tn}^{s} by identifying the one with maximal overlap. Only if this overlap is greater than a threshold – in our experiments 65%⋅∣Cti∣65\%\cdot|C_{ti}|, where ∣Cti∣|C_{ti}| denotes the number of pixels belonging to component CtiC_{ti} – a mapping is assigned. Note that multiple components can be mapped to the same mask, but no more than a single mask is linked to a component. An updated labelling Ltc:Ω→1..Nts\mathcal{L}_{t}^{c}:\Omega\rightarrow{1..N_{t}^{s}} is computed, which replaces component with mask IDs, if an assignment was made.

3.2 Mapping masks to models

Next, a similar overlap between grouped components CtjC_{tj} in Ltc\mathcal{L}_{t}^{c} and projected object labels La\mathcal{L}^{a}, as shown in Figure 2f, is evaluated. Requiring that the camera and objects are tracked correctly, La\mathcal{L}^{a} is generated by rendering all models using the OpenGL pipeline. Besides testing an analogous threshold to before (5%⋅∣Ctj∣5\%\cdot|C_{tj}|), it is verified that the object class IDs of model and mask coincide.

Components that are not yet assigned to a model are now considered to be assigned directly. This is necessary because Mask-RCNN can fail to recognise objects, and most frames are expected to not exhibit any masks. Once again, an overlap of 65%⋅∣Ci∣65\%\cdot|C_{i}| between remaining components and labels in La\mathcal{L}^{a} is evaluated.

The final segmentation Lt:Ω→{0..N}\mathcal{L}_{t}:\Omega\rightarrow\{0..N\} contains the object ids of the models associated with relevant components. A special pre-defined valueWe use the value 255255, as we represent labels as unsigned bytes and assume a number of models less than that. is used to specify areas that ought to be ignored during fusion. This is especially useful to explicitly prevent the reconstruction of certain object classes, such as the arm of the person in Figure 2g, highlighted in white.

EVALUATION

Since the mapping and tracking components of MaskFusion are based on the work of , we focus on the ability to tackle challenging problems that are not solvable by traditional SLAM systems and refer the reader to the corresponding publications for additional details.

To objectively compare MaskFusion with other methods, we evaluate its performance on an established RGBD benchmark dataset . This dataset offers sequences of colour and depth frames and includes ground-truth camera poses to compare with. Measures commonly used for the analysis of visual SLAM or visual odometry methods are the absolute trajectory error (ATE) and the relative pose error (RPE). While the ATE evaluates the overall quality of a trajectory by summing positional offsets of ground-truth and reconstructed locations, the RPE considers local motion errors and therefore surrogates drift. To provide scene-length independent measures, both entities are usually expressed as root-mean-square-error (RMSE). Since MaskFusion is designed to work in dynamic environments, we chose according sequences from the dataset.

First, we estimate camera motion on scenes that involve rapid movement of persons. As our method – as with the methods to which we compare – is not capable of reconstructing deformable parts, we exploit the contextual knowledge of MaskFusion to neglect data associated with persons. Table 2 lists AT-RMSE and RP-RMSE measurements of five methods, including MaskFusion (MF):

VO-SF : A close to real-time method that computes piecewise-rigid scene flow to segment dynamic objects.

ElasticFusion EF : A visual SLAM system that assumes a static environment.

Co-Fusion (CF) : A visual SLAM system that separates objects by motion.

StaticFusion (SF) : A 3D reconstruction system that segments and ignores dynamic parts.

Note that Co-Fusion and MaskFusion are the only systems that maintain multiple object models. The sequences in Table 2 are roughly ordered by difficulty and latter rows exhibit an increasing amount of dynamic motion. While f3s abbreviates freiburg3_sitting, f3w stands for freiburg3_walking.

Interestingly, ElasticFusion performs best in the presence of slight motion, even though it assumes static scenes. Our interpretation of this is that other methods label points as dynamic / outlier that would still be beneficial for tracking, and hence show inferior performance.

Making use of context information proves to be especially useful in highly dynamic scenes, or when the beginning of a scene is difficult. These cases can be hard to tackle by energy minimisation, whereas semantic segmentation results are shown to be robust.

Further, we reconstruct and track the teddy bear in sequence f3_long_office independently from the background motion. This way it is possible to compare the estimated object trajectory with the ground-truth camera trajectory, as highlighted in Figure 5. The trajectory of the bear is only available for a subsection of the sequence as it is out-of-view otherwise.

1.2 Reconstruction

We conducted a quantitative evaluation of the quality of the 3D reconstruction achieved by MaskFusion using objects from the YCB Object and Model Set , a benchmark designed to facilitate progress in robotic manipulation applications. The YCB set provides physical daily life objects of different categories, which are supplied to research teams, as well as a database with mesh models and high-resolution RGB-D scans of the objects. We selected a ground truth model from the dataset (a bleach bottle), and acquired a dynamic sequence to quantitatively evaluate the errors in the 3D reconstruction. Figure 8 shows an image of the object, the ground truth 3D model, our reconstruction and a heatmap showing the 3D error per surfel. The average 3D error for the bleach bottle was 7.07.0mm with a standard deviation of 5.85.8mm (where the GT bottle is 250mm tall and 100mm across).

1.3 Segmentation

To assess the quality of the segmentation quantitatively we acquired a 600600 frame long sequence and provided ground truth 2D annotations for the masks of one of the objects (teddy). Figure 7 shows the intersection over union (IoU) graphs for three different runs. The IoU of the per-frame segmentation masks obtained with MaskRCNN only and MaskRCNN combined with the geometric segmentation are shown in red and blue respectively. The blue curve shows the IoU obtained using our full method, where the object masks are obtained by reprojecting the reconstructed 3D model. This graph shows how combining semantic and geometric cues results in more accurate segmentations, but even better results are achieved when maintaining temporally consistent 3D models over the sequence through tracking and fusion.

2 Qualitative results

We tested MaskFusion on a variety of dynamic sequences, which show that it presents an effective toolbox for different use cases.

A common but challenging task in robotics is to grasp objects. Aside from requiring sophisticated actuators, a robot needs to identify grasping points on the correct object. MaskFusion is well suited to provide the relevant data, as it detects and reconstructs objects densely. Further, and in contrast to most other systems, it continues the tracking during interaction. If the appearance of the actuator is known in advance or if a person interacts with objects, the neural network can be trained to exclude these parts from the reconstruction. Figure 11 shows a timeline of frames that illustrate a grasping performance. In this example, the first 600 frames were used to detect and model 5 objects in the scene, while tracking the camera. We implemented a simple hand-detector that is used to recognise when an object is touched, and as soon as the person interacts with the spray-bottle, the object is tracked reliably until it is placed back on the table at frame 1100.

2.2 Augmented reality

Visual SLAM is a building block of many augmented reality systems and we believe that adding semantic information enables new kinds of applications. To illustrate that MaskFusion can be used for augmented reality applications, we implemented demos that rely and geometric as well as semantic data in dynamic scenes:

Calories demo This prototype aims at estimating the calories of an object-based on its class and shape. By estimating body volumes, using simple primitive fitting, and providing a database with calories per volume unit ratios for different classes, it is straightforward to augment footage with the desired information. Experiments based on this prototype are shown in Figure 10.

Skateboard demo Another demo program presents a virtual character that actively reacts to its environment. As soon as the skateboard appears in the scene the character jumps and remains on it, as depicted in Figure 9. Note that the character stays attached to the board even after a person kicks it and sets it into motion. This requires accurate tracking of the skateboard and camera at the same time.

3 Performance

The convolutional masking component runs asynchronously to the rest of MaskFusion and requires a dedicated GPU. It operates at 5Hz, and since it is blocking the GPU for long periods of time, we use another GPU for the SLAM pipeline, which operates at >>30Hz if a single model is tracked. In the presence of multiple non-static objects, the performance declines and results in a frame-rate of 20Hz for 3 models. Our test system is equipped with two Nvidia GTX Titan X and an Intel Core i7, 3.5GHz.

CONCLUSIONS

This paper introduced MaskFusion, a real-time visual SLAM system that utilises semantic scene understanding to map and track multiple objects. While inferring semantic labels from 2D image data, the system maintains independent 3D models for each object instance and for the background. We showed that MaskFusion can be used to implement novel augmented reality applications or perform common robotics tasks.

While MaskFusion makes meaningful progress towards achieving an accurate, robust and general dynamic and semantic SLAM system, it comes with limitations in the three main problems it addresses: recognition, reconstruction and tracking. Regarding the recognition, MaskFusion can only recognise objects from classes on which MaskRCNN has been trained (currently the 80 classes of the MS-COCO dataset) and does not account for miss-classification of object labels. Secondly, although MaskFusion can cope with the presence of some non-rigid objects, such as humans, by removing them from the map, tracking and reconstruction is limited to rigid objects. Thirdly, tracking small objects with little geometric information when no 3D model is available can result in errors. Solving these limitations opens up opportunities for future work.

References