RIO: 3D Object Instance Re-Localization in Changing Indoor Environments

Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, Matthias Nießner

Introduction

3D scanning and understanding of indoor environments is a fundamental research direction in computer vision laying the foundation for a large variety of applications ranging from indoor robotics to augmented and virtual reality. In particular, the rapid progress in RGB-D scanning systems allows to obtain 3D reconstructions of indoor scenes using only low-cost scanning devices such as the Microsoft Kinect, Intel Real Sense, or Google Tango. Along with the ability to capture 3D maps, researchers have shown significant interest in using these representations to perform 3D scene understanding and developed a rapidly-emerging line of research focusing on tasks such as 3D semantic segmentation or 3D instance segmentation . However, the shared commonality between these works is that they only consider static scene environments. In this work, we focus on environments that change over time. Specifically, we introduce the task of object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a different point in time. Therefore, the captured reconstructions naturally cover a variety of temporal changes; see Fig. 1. We believe this is a critical task for many indoor applications, for instance, for a robot or virtual assistant to find a specific object in its surrounding environment.

The main challenge in RIO – finding the 6DoF of each object – lies in establishing good correspondences between re-scans, which is non-trivial due to different scanning patterns and changing geometric context. These make the use of hand-crafted geometric descriptors, such as FPFH or SHOT , less effective. Similarly, learned 3D feature matching approaches, such as 3DMatch , cannot be easily leveraged since they are trained on self-supervised correspondences from static 3D scenes, and are hence very susceptible to geometry changes. One of the major limitations in using data-driven approaches for object instance localization is the scarce availability of supervised training data. While existing RGB-D datasets, such as ScanNet or SUN RGB-D , provide semantic segmentations for hundreds of scenes, they lack temporal annotations across scene changes. In order to address this shortcoming, we introduce 3RScan a new dataset that is composed of 1482 RGB-D sequences. An essential novelty of the proposed dataset is that several re-scans are provided for every environment. The dataset includes not only dense ground truth semantic instance annotations (for every scan), but also associates objects that have changed in appearance and/or location between re-scans. In addition to using 3RScan for training feature descriptors, we also introduce a new benchmark for object instance localization.

In order to learn from this data, we propose a fully-convolutional multi-scale network capable of learning geometric features in dynamic environments. The network is trained with corresponding TSDF (truncated signed distance function) patches on moved objects extracted at two different spatial scales in a self-supervised fashion. As a result, we obtain change-invariant local features that outperform state-of-the-art baselines in correspondence matching and on our newly created benchmark for re-localization of object instances. In summary, we explore the task of 3D Object Instance Re-Localization in changing environments and contribute:

3RScan, a large indoor RGB-D dataset of changing environments that are scanned multiple times. We provide ground truth annotations for dense semantic instance labels and changed object associations.

a new data-driven object instance re-localization approach that learns robust features in changing 3D contexts based on a geometric multi-scale neural network.

Related Work

3D object localization and pose estimation via keypoint matching are long standing areas of interest in computer vision. Until very recently, 3D hand-crafted descriptors where prominently used to localize objects under occlusion and clutter by determining 3D point-to-point correspondences. However with the success of machine learning, the interest shifted to deep learned 3D feature descriptors capable of embedding 3D data, such as meshes or point clouds, in a discriminative latent space . Even though these approaches show impressive results on tasks such as correspondences matching and registration, they are restricted to static environments. In this work, we go one step further by focusing on dynamic tasks; specifically, we aim to localize given 3D objects from a source scan in a cluttered target scan which contains common geometric and appearance changes.

Scene understanding methods based on RGB-D data generally rely on volumetric or surfel-based SLAM to reconstruct the 3D geometry of the scene while fusing semantic segments extracted via Random Forests or CNNs . Other works such as SLAM++ or Fusion++ operate on an object level and create semantic scene graphs for SLAM and loop closure. Non-incremental scene understanding methods, in contrast, process a 3D scan directly to obtain semantic, instance or part segmentation . Independently from the approach, all these methods rely on the assumption that objects are static and the scene structure does not change over time.

Driven by the great interest in the development of scene understanding applications, several large-scale datasets based on RGB-D data have been recently proposed . We have summarized the most prominent efforts in Table 1, together with their main features (e.g., number of scenes, mean of acquisition). The majority of datasets do not include changes in the scene layout and objects therein, and assume each scene is static over time. This is the case of ScanNet , currently the largest real dataset for indoor scene understanding consisting of 15001500 scans of approx. 750750 unique scenes. Notably, only a few recent proposals started exploring the idea of collecting scene changes to allow long-term scene understanding. InteriorNet is a large-scale synthetic dataset, in which random physics-based furniture shuffles and illumination changes are applied to generate appearance and geometry variations which indoor scenes typically undergo. Several state-of-the-art sparse and dense SLAM approaches are compared on this benchmark. Despite the impressive size and indisputable usefulness, we argue that, due to the domain gap between real and synthetic imagery, the availability of real sequences remain crucial for the development of long-term scene understanding. To the best of our knowledge, the only real dataset encompassing scene changes is the one released by Fehr et al. , which includes 23 sequences of 3 different rooms used to segment the scene structure from the movable furniture, though lacking the annotations and necessary size to train and test current learned approaches.

3RScan-Dataset

We propose 3RScan, a large scale, Real-world dataset which contains multiple (2−122-12) 3D snapshots (Re-scans) of naturally changing indoor environments, designed for benchmarking emerging tasks such as long-term SLAM, scene change detection and camera or object instance Re-Localization. In this section, we describe the acquisition of the scene scans under dynamic layout and moving objects, as well as that of annotation in terms of object pose and semantic segmentation.

The recorded sequences are either a) controlled, where pairs are acquired within a time frame of only a few minutes under known scene changes or b) uncontrolled, where unknown changes naturally occurred over time (up to a few months) via scene-user interaction. All 1482 sequences were recorded with a Tango mobile application to enable easy usage for untrained users. Each sequence was processed offline to get bundle-adjusted camera poses with loop-closure and texture mapped 3D reconstructions. To ensure high variability, 45+45+ different people recorded data in more than 1313 different countries. Each sequence comes with aligned semantically annotated 3D data and corresponding 2D frames (approximately 363k in total), containing in detail:

calibrated RGB-D sequences with variable nn RGB Ri,...Rn\mathbf{R_{i}},...\mathbf{R_{n}} and depth images Di,...Dn\mathbf{D_{i}},...\mathbf{D_{n}}.

camera poses Pi,...Pn\mathbf{P_{i}},...\mathbf{P_{n}} and calibration parameters K\mathbf{K}.

global alignment among scans from the same scene as a global transformation T\mathbf{T}.

dense instance-level semantic segmentation where each instance has a fixed ID that is kept consistent across different sequences of the same environment.

object alignment, i.e. a ground truth transformation TGT=RGT+tGT\mathbf{T_{GT}}=\mathbf{R_{GT}}+\mathbf{t_{GT}} for each changed object together with its symmetry property.

intra-class transformations A\mathbf{A} of ambiguous instances in the reference to recover all valid object poses in the re-scans (see Figure 3).

2 Scene Changes

Due to the repetitive recording of interactive indoor environments, our data naturally captures a large variety of temporal scene changes. Those changes are mostly rigid and include a) objects being moved (from a few centimeters up to a few meters) or b) objects being removed or added to the scene. Additionally, non-rigid objects such as curtains or blankets and the presence of lighting changes create additional challenging scenarios.

3 Annotation

The dataset comes with rich annotations which include scan-to-scene-mappings and 3D transformations (section 3.3.2) together with dense instance segmentation (section 3.3.1). More details and statistics regarding the annotations are given in the supplementary material.

Similarly to ScanNet , instance-level semantic annotations are obtained by labeling on a segmented 3D surface directly. For this, each reference scan was annotated with a modified version of ScanNet’s publicly available annotation framework. To reduce annotation time, we propagate the annotations in a segment-based fashion from the reference scan to each re-scan using the global alignment T\mathbf{T} with the scan-to-scene mappings. This gives us very good annotation estimates for the re-scans, with the assumption that most parts of the scene remain static. Figure 4 gives an example of automatic label propagation from a hand-annotated scene in the presence of noise and scene changes. Semantic segments were annotated by human experts using a web-based crowd-sourcing interface and verified by the authors. The average annotation coverage of the semantic segmentation for the entire dataset is 98.5%98.5\%.

3.2 Instance Changes

To obtain instance-level 3D transformations, a keypoint-based 3D annotation and verification interface was developed based on the CAD alignment tool used in . A 3D transformation is obtained by applying Procrustes on manually annotated 3D keypoint correspondences on the object from the reference and its counterpart in the re-scan (see Figure 5). Additionally to this 3D transformation, a symmetry property was assigned to each instance.

4 Benchmark

3D Object Instance Re-Localization

In order to address the task of RIO, we propose a new data-driven approach that finds matching features in changing 3D scans using a 3D correspondence network. Our network operates on multiple spatial scales to encode change-invariant neighborhood information around the object and the scene. Object instances are re-localized by combining the learned correspondences with RANSAC and a 6DoF object pose optimization.

2 Network Architecture

The network architecture of RIO is visualized in Figure 6. Due to non-padded convolutions and two pooling layers the input volumes are reduced to a 512-dimensional feature vector. It consists of two separate single scale encoders (SSE) and a subsequent multi-scale encoder (MSE). The two different input resolutions capture different neighborhoods with a different level of detail. Since both single scale encoder branches are identical, their network responses are concatenated before being fed into the MSE, as visualized in Figure 6. This multi-scale architecture helps to simultaneously capture fine geometric details as well as higher-level semantics of the surroundings. We show that our multi-resolution network produces richer features and therefore outperforms single scale architectures that process each scale independently by a large margin. Please also note that the two network branches do not share weights since they process the geometry of different context. To achieve a strong gradient near the object surface the raw TSDF is inverted in the first layer of the network such that

3 Training

During training, a triplet network architecture together with a triplet loss (equation 2) is used. It maximizes the L2L_{2} distance of negative patches and minimizes the L2L_{2} distance of positive patches. We choose the margin α\alpha to be 11. For optimization, Adam optimizer with an initial learning rate of 0.0010.001 is used.

4 Training Data: From Static to Dynamic

We initially train our network fully self-supervised with static TSDF patches extracted from RGB-D sequences of our dataset. To be able to deal with partial reconstructions induced by different scanning patterns, two small sets of non-overlapping frames are processed to produce two different TSDF volumes of the same scene. Then, first Harris 3D keypoints are extracted on one volume, then these same locations are refined on the other volume via non-maxima suppression of the Harris responses within a small radius around each extracted keypoint. If corresponding keypoints on the two volumes are above a certain threshold, we consider them a suitable patch pair and use it for pre-training of our network.

The goal of our method is to produce a local feature encoding that maps the local neighborhood of an object around a 3D keypoint on a 3D surface to a vector while being invariant to local changes around the object of interest. We learn this change-invariant descriptor by using the object alignments and sampling dynamic patches from our proposed 3RScan-dataset. So, once converged we fine-tune the static network with dynamic 3D patches specifically generated around points of interest on moving objects. To learn only higher level features, during fine-tuning, we freeze the first layers and only train the multi-scale encoder branch of our network. Correspondence pairs are generated in a self-supervised fashion while using the ground truth pose annotations of our training set to find high keypoint responses in the same small radius around each source 3D keypoint. The negative counterpart of each triplet is randomly selected from another training scene but also includes TSDF patches on removed objects. Random rotation augmentation is applied to enlarge our training data.

5 6DoF Pose Alignment

To re-localize object instances, we first compute features for keypoints on the source objects and the whole target scene. Correspondences for the model keypoints are then found via k-nearest neighbour search in the latent space of the feature encoding of the points in the scene. After outliers are filtered with RANSAC, remaining correspondences serve as an input of a 6DoF pose optimization. Given the remaining two sets of correspondences on the source object O=p1,p2,...pn\mathbf{O}={p_{1},p_{2},...p_{n}} ∈R3\in R^{3} and the target scene S=q1,q2,...qn\mathbf{S}={q_{1},q_{2},...q_{n}} ∈R3\in R^{3} we then want to find a optimal rigid transformation that aligns the two sets. Specifically, we want to find a rotation R\mathbf{R} and a translation t\mathbf{t} such that

We solve this optimization using Singular Value Decomposition (SVD). The resulting 6DoF transformation gives us a pose that aligns the model to the scene. Qualitative results of our alignment method with corresponding ground truth alignments on some scans of our 3RScan-dataset are shown in Figure 7.

Evaluation

In the following, we show quantitative experimental results of our method by evaluating it on our newly created 3RScan-dataset. In the first section, we compare the ability of different methods to match dynamic patches around keypoints on annotated changed objects. Our proposed multi-scale network is then evaluated on the newly-created benchmark for re-localization of object instances.

For accurate 6D pose estimation in changing environments, robust correspondence matching is crucial. The feature matching accuracy of different network architectures is reported in Table 3. Each network is pre-trained with static samples (marked as static) and then fine-tuned on dynamic patches from the training (marked as dynamic). The F1 score, accuracy, precision, false positive rate (FPR) and error rate (ER) at 95% recall are listed and visualized with their respective PRC graphs in Figure 8. Additionally to the 1:1 matching accuracy, we also use a top-1 metric: the percentage of top-1 placements of a positive patch given 50 randomly chosen negative patches. Such a metric better represents the real test case of object instance re-localization where several negative samples are compared against a positive keypoint. In can be seen that our multiscale network architecture – even if only trained with static data – outperforms all single scale architectures by a large margin and improves further to an F1F1 score of 94.37 if additionally trained with dynamic data.

2 Object Instance Re-localization

Evaluation results for all object instances are listed in Table 4 and Table 5. While classical hand-crafted methods still perform reasonable well – especially for more descriptive objects such as sofas and beds – our method outperforms them with a large margin. Qualitative results are shown in Figure 7.

Conclusion

In this work, we release the first large-scale dataset of real-world sequences with temporal discontinuity that consists of multiple scans of the same environment. We believe that the new task of object instance re-localization (RIO) in changing indoor environments is a very challenging and particularly important task, yet to be further explored. Besides 6D object instance alignments in those changing environments, 3RScan comes with a large variety of annotations designed for multiple benchmark tasks including – but not limited to – persistent dense and sparse SLAM, change detection or camera re-localization. We believe that 3RScan helps the development and evaluation of these new algorithms and we are excited to see more work in this domain to, in the end, accomplish persistent, long-term understanding of indoor environments.

Acknowledgment

We would like to thank the volunteers who helped with 3D scanning, all expert annotators, as well as Jürgen Sturm, Tom Funkhouser and Maciej Halber for fruitful discussions. This work was funded by the Bavarian State Ministry of Education, Science and the Arts in the framework of the Centre Digitisation.Bavaria (ZD.B), the ERC Starting Grant Scan2CAD (804724), TUM-IAS for the Rudolf Mößbauer Fellowship, and a Google Research and Faculty award.

References

Supplemental Material

In this supplemental document, we provide additional information about the proposed dataset such as statistics, scene examples and a detailed description about the annotation process.

Dataset

We tailored a mobile app running on Google Tango with pre-annotation functionality as a scanning interface (see Figure 9). Some users gave lightweight instructions on the scene changes; these instructions served as guidelines later in the annotation process.

For each uploaded scan, scene candidates are computed. Since a 3D scene matching is expensive scan pairs are found in 2D instead by conducing a similarity search in the texture uv-map of the mesh. These matches are then to be manually adjusted. Once the reference for each scene is assigned, the IMU normalized scans are globally registered via a coarse to fine correspondence-based 2D ICP together with RANSAC and refined with a global 3D ICP. An additional verification as well as an optional manual, keypoint based alignment ensures high quality.

Additionally to a server-side offline processing of the RGB-D sequences that results in texture mapped 3D reconstructions, an offline 3D segmentation is triggered. This 3D segmentation is utilized by the semantic segmentation interface proposed by Dai et al. . Further, this 3D segmentation – together with the aforementioned 3D alignment – serves as the basis for the propagation of semantic labels from the references to the re-scans and after a manual clean-up procedure results in the final instance segmentation shown in Figure 17. In the current snapshot of the dataset almost all of our scans have an instance segmentation coverage of above 90% (see Figure 10) with an average scene coverage >> 98%. In total, 48k instances are annotated with 534 unique labels.

Our dataset consists of around 363k calibrated RGB-D and depth images. Since raw RGB and depth sequences from Tango are of varying frame rates and spatial resolution, a spatial and temporal calibration procedure of the raw images is applied. Further, to remove rectification lines present in Google Tango depth images, a median filter is used before the spatial calibration of the images. Camera trajectories for SLAM are visualized in Figure 18 and since global scene-to-scene mappings are provided these can easily be transferred into the same coordinate systems (see last two rows of Figure 18). Further, we also show 2D projections of our textured 3D models with aligned camera poses in Figure 14.

Further, instead of assigning a 1−n−1-n-relationship of different room types per scene, our 3D reconstructions are annotated with mm corresponding scene functionalities (sleeping, eating, working, etc.) in an n−mn-m fashion. This shows the high variety of scenes in 3RScan.

The annotation interface for annotating instance changes is a web-based tool (Figure 16) where each scene is rendered next to its corresponding reference. When an object is selected in the re-scan (see green dot) its instance segmentation from the reference scan is automatically segmented. Please note, that this requires the instance IDs to be consistent across scans of the same environments. Hovering over the objects gives shows the label and the ID of the instance and allows to potentially fix the underlying semantic segmentation. In the alignment view this instance is then shown next to the re-scan such that corresponding keypoints can easily be selected. Once enough keypoints are set, a Procrustes based alignment (Kabsch algorithm) is triggered that computes a transformation that aligns the object to the scene. For non-rigid changes and removed as well as added objects the instances IDs are tracked.

A subset of the changed objects in the dataset are symmetric. We follow the symmetry annotation described in Avetisyan et al. and categorize each object’s rotational symmetry around the canonical axis to the classes C2C_{2}, C4C_{4} and C∞C_{\infty}. 22% of the objects have a symmetry as listed in Table 6. We take this into account when evaluating the predictions against ground truth poses.

A focus during data acquisition was the capturing of a variety of realistic scene changes in controlled and uncontrolled environments over a time span of more than 1212 months. The number of references scenes with re-scanning frequencies are plotted in Figure 11. Further, 3289 instance transformations – of 1947 different objects – are provided with the data. But since the transformations give the object pose from the reference to one re-scan the alignment for another re-scan can easily be computed. For evaluation these changed object categories are mapped to 9 different classes as listed in 7.

These changed object instances are labelled with 187 different categories. The majority of instances include movements of objects and more portable furniture items such as chairs, pillows, boxes or smaller tables. Naturally, these objects involve most human interaction. Figure 13 gives an overview of the motion of these annotated objects. However, we also annotated objects that slightly change their appearance over time such as toilets. Detailed statistics are given in Figure 12.