Fusion++: Volumetric Object-Level SLAM

John McCormac, Ronald Clark, Michael Bloesch, Andrew J. Davison, Stefan Leutenegger

Introduction

Indoor scene understanding and 3D mapping is a foundational technology that can enable autonomous real-world robotic task completion and also provide a common interface for more intelligent and intuitive human-map and human-robot interactions. To enable this requires a careful choice of map representation. One particularly useful representation is to build an object-oriented map. We argue this is a natural and efficient way to represent the things that are most important for robotic scene understanding, planning and interaction; and it is also highly suitable as the basis for human-robot communication.

In an object level map, the geometric elements which make up an object are grouped together as instances and can be labelled and reasoned about as units, in contrast to approaches which independently label dense geometry such as surfels or points. This approach also naturally paves the way towards interaction and dynamic object reasoning, although our system currently assumes a static environment and does not yet aim to track individual dynamic objects.

In this work we demonstrate an object-oriented online SLAM system with a focus on indoor scene understanding using RGB-D data. We aim to produce semantically labelled TSDF reconstructions of object instances without strong a priori knowledge of the object types present in a scene. We use Mask R-CNN to provide 2D instance mask predictions and fuse these masks online into the TSDF reconstruction (see Figure 1) along with a 3D ‘voxel mask’ to fuse the instance foreground (see Figure 3).

Unlike many dense reconstruction systems we make no attempt to keep a dense representation of the entire scene. Our persistent map consists of only reconstructed object instances. This allows the use of rigid TSDF volumes for high-quality reconstructions to be combined with the flexibility of a pose-graph system without the complication of performing intra-TSDF deformations. Each object is contained within a separate volume, allowing each one to have a different, suitable, resolution with larger objects integrated into lower fidelity TSDF volumes than their smaller counterparts. It also enables tracking large scenes with relatively small memory usage and high-fidelity reconstructions by excluding large volumes of free-space. A throw-away local TSDF of unidentified structure is used to assist tracking and model occlusions.

We capture a repeated loop of an indoor office scene to evaluate the system under conditions of occasional poorly constrained ICP tracking. The scene also contains a large number and variety of objects which not only exhibit the generality of the approach but is useful for evaluating the memory and run-time scaling of the method with many objects. While not optimised for real-time operation, we achieve ∼\sim4-8Hz operating performance (excluding relocalisation/graph optimisation) on our office sequence and are confident that with sufficient optimisation true real-time operation is possible. We also quantitatively evaluate the trajectory error improvement of our system over a baseline approach on the RGB-D SLAM Benchmark .

In this work we make the following contributions:

A generic object-oriented SLAM system which performs mapping as variable resolution 3D instance reconstruction.

Per-frame instance detections are robustly fused using voxel foreground masks and missing detections are accounted for with an “existence” probability.

We show high quality object reconstruction within globally consistent loop-closed object SLAM maps.

Related work

For reconstruction, we follow the TSDF formulation of Curless and Levoy and the KinectFusion approach of Newcombe et al. for local tracking. Our approach to object-level reconstruction is related to the work of Zhou and Koltun , where “points of interest” were detected and the aim was to reconstruct the scene so as to preserve detail in these areas while distributing drift and registration errors throughout the rest of the environment. In our work we analogously aim to optimise the quality of object reconstructions and allow residual error to be absorbed in the edges of the pose graph.

SLAM++ by Salas-Moreno et al. was an early RGB-D object-oriented mapping system. They used point pair features for object detection and a pose graph for global optimisation. The drawback was the requirement that the full set of object instances, with their very detailed geometric shapes, had to be known beforehand and pre-processed in an offline stage before running. Stückler and Behnke also previously tracked object models learned beforehand by registering them to a multi-resolution surfel map. Tateno et al. used a pre-trained database of objects to generate descriptors, but they used a KinectFusion TSDF to incrementally segment regions of a reconstructed TSDF volume and match 3D descriptors directly against those of other objects in the database.

A number of approaches to object discovery exist . Most related to ours is the work of Choudhary et al. where they localised the camera in an online manner using discovered objects as landmarks in a pose-graph formulation similar to ours, although they used the point cloud centroid only whereas our pose-graph object landmark edges are full 6 DoF SE(3)SE(3) constraints provided from ICP on dense volumes. They showed that the approach improves SLAM results by detecting loop closures. However, unlike our work they use point-clouds rather than TSDFs and do not train an object detector but instead they use the unsupervised segmentation approach of Trevor et al. .

Another approach to object discovery is through dense change detection between successive mappings of the same scene . Unlike these systems, our system is designed for online use and does not require changes to occur in a scene before objects are detected. These approaches are complementary to our proposed approach, providing supervisory signals for CNN fine-tuning, and enabling additional object database filtering mechanisms.

In RGB-only SLAM for object detection, Pillai and Leonard use ORB-SLAM to assist object recognition. They use a semi-dense map to produce object proposals and aggregate detection evidence across multiple views for object detection and classification. MO-SLAM by Dharmasiri et al. focused on object discovery through duplicates. They use ORB descriptors to search for sets of landmarks which can be grouped by a single rigid body transformation. This approach is similar to our relocalisation method, which uses BRISK features but augmented with depth.

Very closely related to ours is work by Sünderhauf et al. , who proposed an object-oriented mapping system composed of instances using bounding box detections from a CNN and an unsupervised geometric segmentation algorithm using RGB-D data. Although the premise is closely related, there are a number of differences when compared to our system. They use a separate SLAM system, ORB-SLAM2 , whereas in our system the discovered object instances are tightly integrated into the SLAM system itself. We also fuse instances into separate TSDF volumes with a foreground mask from 2D instance mask detection rather than using point cloud segments.

A number of very recent related works have also been announced. Pham et al. fuse a TSDF of the entire scene and semantically label voxels using a CNN followed by a progressive CRF. To segment instances, instead of fusing native instance detections, they opt to cluster semantically labelled voxels in 3D. This approach, although a natural next-step from dense 3D semantic mapping, is not suitable for object-level pose graph optimisation and reconstruction as the instances are embedded within a shared TSDF. It also requires semantic recognition as a pre-requisite for object discovery which could prove problematic for similar or unrecognised objects in close proximity (Figure 3).

Rünz and Agapito , as in our method, use Mask R-CNN predictions to detect object instances. They aim to densely reconstruct and track moving instances using an ElasticFusion surfel model for each object, as well as for the background static map. Although using the same prediction model, the approach and goals of these two systems differ substantially. Unlike the present work, they do not aim to reconstruct high-quality objects as pose-graph landmarks in room-scale SLAM. We on the other hand do not currently tackle dynamic scenes and assume all objects to be static during an observation. Clearly there is the long-term potential to combine these two approaches.

Method

Our pipeline is visualised in Figure 2. From RGB-D input, a coarse background TSDF is initialised for local tracking and occlusion handling (Section 3.3). If the pose changes sufficiently or the system appears lost, relocalisation (Section 3.4) and graph optimisation (Section 3.5) are performed to arrive at a new camera location, and the coarse TSDF is reset. In a separate thread RGB frames are processed by Mask R-CNN and the detections are filtered and matched to the existing map (Section 3.2). When no match occurs, new TSDF object instances are created, sized, and added to the map for local tracking, global graph optimisation, and relocalisation. On future frames, associated foreground detections are fused into the object’s 3D ‘foreground’ mask alongside semantic and existence probabilities (Section 3.1).

Initialisation and resizing: Detections not matched by the procedure described in 3.2 are used to initialize an appropriately sized and positioned instance TSDF. In the kkth frame each detection ii produces a binary mask MikM_{i}^{k}. We project all the masked image coordinates u=(u1,u2)\mathbf{u}=(u_{1},u_{2}) into F→W{\smash{\underrightarrow{\mathcal{F}}_{W}}} using the depth map Dk(u)D_{k}(\mathbf{u}),

To robustly size the TSDF in the presence of masks which can occasionally include far-away background surfaces, we do not directly accept the maximum and minimum of this point cloud. Instead we use the 10th10^{th} and 90th90^{th} percentiles of this point cloud (separately for each axis) to define points p10\mathbf{p}_{10} and p90\mathbf{p}_{90} respectively, which are used to calculate the volume centre po=p90+p102\mathbf{p}_{o}=\frac{\mathbf{p}_{90}+\mathbf{p}_{10}}{2} and volume size so=m∥(p90−p10)∥∞s_{o}=m\lVert(\mathbf{p}_{90}-\mathbf{p}_{10})\rVert_{\infty}. We use an mm of 1.5 to account for erosion and provide additional padding.

Each instance TSDF has an initial fixed resolution along a given axis of ror_{o}, which we choose to be 6464, and sos_{o} is used to calculate the physical size of a voxel vo=sorov_{o}=\frac{s_{o}}{r_{o}}. Therefore, small objects will be reconstructed with fine details and large objects more coarsely, making the map as useful as possible for a given memory footprint.

During operation matched objects may need to be re-sized as new detections include additional areas. To do this, the point cloud of the current mask described above is combined with a similarly eroded point cloud generated from the current TSDF reconstruction. The 3D volume encompassing them both is used to calculate the new volume centre and size as before. To avoid aliasing when re-sizing, we translate the volume centre by discrete multiples of vov_{o}, and maintain the same vov_{o} but increase ror_{o}, while maintaining an even parity. We also limit the maximum voxel resolution to 128128, by re-initialising the volume as though new if ro>128r_{o}>128, and limit the maximum object size to be 3m.

Before initialising an instance we require the volume centre to be within 5m of the camera, and a 3D axis-aligned bounding box Intersection over Union (IoU) <0.5<0.5 with any other volume already in the map. When an object centre is moved, the pose-graph node and associated measurements are also updated as described in Section 3.5.

Integration: For integrating surface measurements from a depth map DkD^{k} into Vo\mathcal{V}^{o} we take an approach similar to Newcombe et al. Code based on https://github.com/GerhardR/kfusion.. Vo\mathcal{V}^{o} stores at each discrete voxel location v=(vx,vy,vz)\mathbf{v}=(v_{x},v_{y},v_{z}) both the current normalised truncated signed distance value Sk−1o(v)S^{o}_{k-1}(\mathbf{v}) and its associated weight Wk−1o(v)W^{o}_{k-1}(\mathbf{v}). If v\mathbf{v} projects into a camera frame pixel with a depth value less than the depth measurement plus the truncation distance, μ\mu (here chosen as 4vo4v_{o}), then that measurement is fused into the volume in a weighted average fashion. Integration is performed on every frame where the TSDF volume is visible, when 50% of TSDF pixels are validly tracked and the ICP RMSE <0.03<0.03 (these error metrics are described in more detail in Section 3.3). This is to maintain the reconstruction quality of instances when the camera frame may have drifted.

It is also important to note that the above surface integration is performed throughout the entire volume, regardless of whether it is a masked region or not. To store which voxels correspond to this instance’s ‘foreground’ we also fuse instance mask detections. We view each positive or negative detection as the result of a binomial trial sampled from a latent foreground probability, po(v∈foreground)p^{o}(\mathbf{v}\in\text{foreground}). We store foreground Fk−1o(v)F^{o}_{k-1}(\mathbf{v}) and not foreground Nk−1o(v)N^{o}_{k-1}(\mathbf{v}) detection counts as the (α,β)(\alpha,\beta) shape parameters in a beta distribution conjugate prior which are initialised with (1,1)(1,1). When a new detection is matched and the depth measurement is within the truncation distance as above, then we also update the detection counts using the corresponding mask ii:

with π([x,y,z]⊺)=[x/z,y/z,1]⊺\boldsymbol{\pi}([x,y,z]^{\intercal})=[x/z,y/z,1]^{\intercal} denoting the projection. Finally, to compute whether a voxel is part of the foreground we calculate the expectation,

and use a decision threshold of E[po(v)]>0.5E[p^{o}(\mathbf{v})]>0.5. A visualisation of this is shown in Figure 3.

Raycasting: For tracking, data association, and visualisation we render depth, normals, vertices, RGB, and object indices. Within each object volume Vo\mathcal{V}^{o} we step along the ray with a stepsize of vsov^{o}_{s} (and 0.5vso0.5v^{o}_{s} when Sko(v)<0.8S_{k}^{o}(\mathbf{v})<0.8, where Sko(v)S_{k}^{o}(\mathbf{v}) is the SDF normalised by μ\mu) and search for the zero-crossing point in Sko(v)S_{k}^{o}(\mathbf{v}) where E[po(v)]>0.5E[p^{o}(\mathbf{v})]>0.5 (both values are trilinearly interpolated from neighbouring voxels to smooth the representation). We store the ray length of the nearest of these intersections to avoid searching past that point in another volume.

This alone results in occluding surfaces which are not part of the foreground failing to occlude the ray. If a background TSDF is available, and either no intersection with a foreground object occurs or the intersection is farther than 5cm behind the background TSDF intersection, then the background TSDF ray intersection is used instead.

Existence Probability: To prevent spurious instances from building up over time, we also model the probability of each instance’s existence as p(o)p(o) using the Beta distribution, in a manner identical to the foreground mask. For any frame where a predicted instance should be clearly visible (i.e. our raycasted image has more than 50250^{2} pixels of that instance), then if the instance has been associated to a detection its existence count eoe_{o} is incremented, and if not its non-existence count, dod_{o}, is incremented. If E[p(o)]E[p(o)] falls below 0.10.1, the instance is deleted and the object node with all associated edges are removed from the pose graph (described in Section 3.5).

Semantic Labels: Each TSDF also stores a probability distribution over potential class labels lol_{o}. Mask R-CNN provides a probability distribution p(lo∣Ik)p(l_{o}|I_{k}) over the classes given the image, IkI_{k}. We found that the standard multiplicative Bayesian update scheme :

where ZZ is a normalising constant, often leads to an overly confident class probability distribution, with scores unsuitable for ranking in object detection. Instead here we fuse multiple associated detections by simple averaging:

which produces a more even class probability distribution.

2 Detection and Data Association

Detections from the Mask R-CNN model for a given frame kk contain instances ii with a binary mask MkiM^{i}_{k} and class probability distribution p(li∣Ik)p(l_{i}|I_{k}). A forward pass takes ∼\sim250ms, and although our system is not real-time, this still represents a significant bottleneck and so can be performed in a parallel thread. For GPU memory efficiency, we take only the top 100 detections (scored according to the region proposal network ‘object’ score ) and filter for masks not near the image border (within 20 pixels) and where both max(p(li∣Ik))>0.5\text{max}(p(l_{i}|I_{k}))>0.5 and ∑Mki>502\sum M^{i}_{k}>50^{2}.

3 Layered Local Tracking

We maintain an instance-agnostic coarse background TSDF, aa, to assist local frame-to-model tracking where/when there are no instances and to handle occlusions. It has a resolution of 2563256^{3} with a voxel size of 2cm. Its initialisation point Wpa=TWCk[002.56]⊺{}_{W}\mathbf{p}_{a}=\mathbf{T}^{k}_{WC}[0\quad 0\quad 2.56]^{\intercal}, is 2.56m along the zz-axis in the camera frame F→C{\smash{\underrightarrow{\mathcal{F}}_{C}}} to prevent wasted volume as in . The volume is reset when its new initialisation point exits a spherical threshold (1.28m) around the previous volume centre, i.e. ∥Wpa−TWCk[002.56]⊺∥2>1.28\lVert{}_{W}\mathbf{p}_{a}-\mathbf{T}^{k}_{WC}[0\quad 0\quad 2.56]^{\intercal}\rVert_{2}>1.28.

The Gauss-Newton iteration can then be implemented as follows (with iteration index tt):

4 Relocalisation

If the system is lost or we reset the coarse TSDF, we perform relocalisation to align the current frame to the current set of instances (if there are any). We found direct dense ICP methods using only the volume reconstructions did not produce accurate results for wide baseline relocalisation as they are sensitive to the initial pose and small objects were often ambiguous without texture constraints. Although alternative dense methods may also prove useful here, we took the approach of using snapshots of sparse BRISK featuresBRISK v.2 with homogeneous Harris scale space corner detection on only the highest image resolution. (with a detection threshold of 10) projected to 3D using the depth map. For a given detection of an object if there is no existing snapshot of the object within 15∘15^{\circ} view angle difference, we then add a new snapshot of the object from that pose (see Figure 4).

To re-localise we perform 3D-3D RANSAC against each instance where the dot product with the predicted class distribution is greater than 0.6. We use OpenGV with a minimum of 5 inlier features (within 2cm) to match each object individually. If we find one or more matching objects in the scene, we run a final 3D-3D RANSAC on every point in the scene (from all objects and the background jointly) with a minimum of 50 inlier features (within 5cm) to arrive at a final camera pose. This pose is used to render a new reference image of the map to produce the constraints required for the pose graph optimisation described below.

5 Object-Level Pose Graph

Our pose-graph formulation is similar to that of . For every frame with a Mask R-CNN detection (including coarse TSDF resets), we add a new camera pose node to our graph. When a new instance, index oo, is initialised, a corresponding landmark node is added to the graph, defined by the coordinate frame attached to the centre of the object’s volume, po\mathbf{p}_{o}. The first camera pose node is fixed and defined to be the origin of the world frame, F→W{\smash{\underrightarrow{\mathcal{F}}_{W}}}. Each node consists of a full SE(3)SE(3) transformation from object to World, TWO\mathbf{T}_{WO}, or camera to world, TWC\mathbf{T}_{WC}, and the measurements are SE(3)SE(3) relative pose constraints between nodes.

The final error to be minimised in the pose graph is the sum over all the edges from the camera to objects, O\mathcal{O}, and camera to camera, C\mathcal{C}, given their state, the measurement, and the information matrix,

where LσL_{\sigma} denotes a robust Huber kernel. We solve this graph in the g2o framework using sparse Cholesky decomposition and Levenberg-Marquart. After optimisation we update the pose of the instance TSDFs and the camera before initialising the new coarse TSDF to that pose and continuing with local tracking.

Experiments

We evaluate the performance and memory usage of our system on a Linux system with an Intel Core i7-5820K CPU at 3.30GHz, and an nVidia GeForce GTX1080 Ti GPU with 11.175GB of memory. Our core pipeline is implemented in Python and uses Tensorflow for instance predictions, and Python wrappers around other core components which are developed in C++ and/or CUDA, such as KFusion, g2o, BRISK, and OpenGV. Our input is standard 640×480640\times 480 resolution RGB-D video. To allow for reproducibility, instead of running an asynchronous CNN thread we here perform predictions synchronously every 30 frames.

Our Mask-RCNN uses the ResNet-101 base model (up to the conv4_x block) and is finetuned from the publicly available tensorpack implementation and weights .http://models.tensorpack.com For finetuning on indoor scenes we use the NYUv2 dataset. We lock the ResNet-101 weights from the COCO pre-training and fine-tune the remaining layers. As the COCO dataset consists of 80 classes we re-size and reinitialise the class-specific upper layers of Mask R-CNN and Faster R-CNN. We train using stochastic gradient descent with momentum of 0.90.9 for 3030 epochs with a learning rate of 0.001.

To evaluate the performance of our system while repeatedly viewing a scene of instances we captured a 3,685 frame sequence of an indoor office scene. We tailored this sequence to evaluate the consistency of our map in the presence of poorly constrained (planar floor) geometry and ICP drift, after which we loop over the same scene again. The pose-graph and loop closure is shown in Figure 5, it can be seen that despite the accumulated drift, the system re-localises and corrects the pose graph, this allows the previously reconstructed objects to be correctly associated in future frames. On the entirety of the trajectory our system reconstructed 105 landmark object instances, however, it must be noted that despite our filtering mechanisms, a build up of noisy partially reconstructed sub-objects still occurs.

2 Reconstruction Quality

To evaluate the reconstruction quality we use objects from the YCB dataset which provides ground truth models and reconstruct discovered objects from sequence 0001 of the public YCB video dataset . Figure 6 shows a qualitative comparison against the ground truth. The missing portion of the cracker box was caused by an occlusion by another object, and a missed foreground detection on one of the few frames where the cracker box was unoccluded.

3 RGB-D SLAM Benchmark

We evaluate the trajectory error of our system against the baseline approach of simple coarse TSDF odometry, i.e. using the same coarse resetting background without instances layered on top, and without loop-closure pose graph optimisation. Table 1 shows the results. It can be seen that in all but one of the sequences evaluated our Fusion++ system improved upon the baseline approach (while providing an inventory of objects as Figure 1 visualises for the fr2_desk sequence). It is also worth noting that our system does not achieve state-of-the-art performance on these sequences such as , and would require additional work, such as including joint depth and photometric tracking, to become competitive. We focused on a usable object map here and leave accuracy of motion tracking for future work.

4 Memory and Run-time Analysis

Memory usage: We use the office sequence to evaluate the run-time performance and memory usage of our system. As memory usage scales cubically with the size of a TSDF, it is significantly more efficient to compose a map of many relatively small, highly detailed, volumes in dense areas of interest than to use one large one with a resolution equal to the smallest. After loading the CNN and image buffers, our remaining ∼\sim7GB GPU memory budget (and 10 bytes per voxel) would allow a single 9003900^{3} volume or, as here, a 2563256^{3} background volume and up to 2.5K object volumes with dimension 64364^{3}, 2MB. Our object volumes dynamically vary up to 1283128^{3} and on our office sequence used 377MB for 105 objects (∼\sim4MB/object), as shown in Figure 7. Of course, more efficient alternatives such as an octree or voxel hashing can also be used to directly eliminate wasted free-space voxels, and are also directly applicable to our approach.

Runtime performance: Our system, although not real-time, scales well with the number of objects. Excluding re-localisation on the office sequence the average frame rate was 4-8Hz (shown in Figure 7), with an average additional computational cost of 1ms per object. A more detailed breakdown of the runtime performance of different components and their scaling factors is given in Table 2.

Conclusions

We have shown consistent instance mapping and classification of numerous objects of previously unknown shape in real, cluttered indoor scenes. Our online and near real-time system, which is built from modules for image-based instance segmentation, TSDF fusion and tracking, and pose graph optimisation, makes a long-term map which focuses on the most important object elements of a scene with variable, object size-dependent resolution.

A number of shortcomings of the current approach remain to be addressed in future work. There is a balance to be struck between filtering detections and providing good coverage of a scene, and even with the existence probability and deletion mechanism detailed here, over time spurious detections result in a growing clutter of partial object reconstructions. More thorough object detection precision/recall evaluations as well as semantic accuracy metrics will assist in this. A learned mechanism for filtering and reconstructing these objects, such as may prove useful in this regard, or combining view-based segmentation and classification with 3D methods which take advantage of object databases such as ShapeNet .

There is also significant scope in future to better combine information from multiple duplicate objects seen from different views to reconstruct a single better model, rather than maintaining separate TSDFs for each. Our object-oriented representation can also naturally be extended to model moving objects with individually changing poses. This attribute would be particularly useful when reasoning about dynamic applications in robotics or augmented reality.

Acknowledgements

This research was supported by Dyson Technology Ltd.

References