ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Gilad Baruch, Zhuoyuan Chen, Afshin Dehghan, Tal Dimry, Yuri Feigin, Peter Fu, Thomas Gebauer, Brandon Joffe, Daniel Kurz, Arik Schwartz, Elad Shulman

Introduction

Indoor 3D scene understanding is becoming key for many applications in the domains of augmented reality, robotics, photography, games, and real estate. More recently, modern machine learning techniques have fueled many state-of-the-art scene understanding algorithms. A variety of methods are addressing different parts of the challenge, like depth estimation, 3D reconstruction, instance segmentation , object detection and more. Most of these research works are enabled through a variety of real and synthetic RGB-D datasets that have been made available over the past few years . Even though commercially available RGB-D sensors, like Microsoft Kinect, have made collection of such datasets possible, it is still not a trivial task to capture data at large scale with ground truth. In addition, almost all previous devices used for data collection is datasets such as SunRGBD or ScanNet , are different from the hardware used by people nowadays. The lack of diversity in data and the gap in the depth sensing technology brings challenges in making the innovative research of the last decade practical for day-to-day use.

Recently, Apple released iPads and iPhones equipped with the LiDAR scanner . It unleashed a new era in availability and accessibility of depth sensors. This work provides the first large-scale dataset that is captured with Apple’s LiDAR scanner using handheld devices. It helps bridge the domain gap between existing datasets and widely available mobile depth sensors, and is the largest RGB-D dataset in terms of number of sequences and scene diversity collected in people’s homes.

Our dataset, which we named ARKitScenes, consist of 5,048 RGB-D sequences which is more than three times the size of the current largest available indoor dataset . These sequences include 1,661 unique scenes. Additionally we provide estimated ARKit camera poses as well as the LiDAR scanner-based ARKit scene reconstruction for all the sequences. A comparison of ARKitScenes with some of the existing datasets is shown in Table 1. In addition to the raw and processed data above, we provide high quality ground truth and demonstrate its usability in two downstream supervised learning tasks: 3D object detection and color-guided depth upsampling. For the 3D object detection task, ARKitScenes provides the largest RGB-D dataset labeled with oriented 3D bounding boxes for 17 room-defining furniture categories. ARKitScenes further uses high-resolution ground truth scene geometry that is captured with a professional stationary laser scanner (Faro Focus S70). We describe a technique used to register the high quality laser scans with mobile RGB-D frames captured with an iPad Pro. To our best knowledge, this is the first dataset that provides high quality ground truth depth data registered to frames from a widely available depth sensor. Finally, we evaluate the performance of state-of-the-art methods when trained and evaluated on ARKitScenes and highlight the challenges of existing methods in generalizing to real-world scenarios. In summary the contributions of this paper are as follows:

We present ARKitScenes, the first RGB-D dataset captured with the widely available Apple LiDAR scanner. Along with the per-frame raw data (Wide camera RGB, Ultra Wide camera RGB, LiDAR scanner depth, IMU) we provide the estimated ARKit camera pose and ARKit scene reconstruction for each iPad Pro sequence.

ARKitScenes is the largest indoor 3D dataset consisting of 5,048 captures of 1,661 unique scenes.

We provide high quality ground truth of (a) depth registered with RGB-D frames and (b) oriented 3D bounding boxes of room-defining objects registered with the scene reconstructions.

We demonstrate the effectiveness of the dataset in advancing state-of-the-art methods while highlighting the limitations of current methods and datasets in generalizing to realistic scenarios.

We expect ARKitScenes to stimulate the development of novel algorithms. Furthermore, we call for an evaluation and comparison of future work on ARKitScenes as it represents a diverse set of homes in the wild. And finally, we hope ARKitScenes bridges the gap between innovation and usability by the general public as it provides data captured with an RGB-D sensor which many people carry in their pockets.

Related work

Availability of large-scale datasets such as ImageNet stimulated research, especially since supervised deep learning techniques re-gained popularity. In the context of scene understanding and 3D point clouds, several datasets were released in the past few years . These datasets have enabled a series of research investigations in various areas including semantic segmentation, object detection, room layout estimation, depth estimation and more. For outdoor scene understanding, several large scale datasets with a variety of scenes in real scenarios were released that have powered deep learning algorithms . However, when it comes to indoor scene understanding we are only limited to a few datasets . Outdoor and indoor datasets have very different characteristics, both caused by the size of the space and the type of sensors that are used to collect those datasets. Because of that, the methods designed for one will not necessarily perform the same in the other.

Indoor scene understanding. NYU v2 is one of the earliest RGB-D datasets focusing on indoor scenes. It is composed of 464 scenes from three cities captured with a Kinect device. It also includes 1,449 densely labeled pairs of aligned RGB and depth maps annotated with 2D polygons. Sun RGB-D further expands on previous work by introducing a dataset with over 10,000 RGB-D frames along with 2D polygon and 3D bounding box labels. The labels are provided at frame level and not scene level, lacking view-point diversity. There are several other datasets which have focused on long indoor RGB-D captures to address the view-point diversity. Among these datasets ScanNet is the largest and closest to ours in terms of scene diversity and assets. ScanNet provides 1,513 scans of 707 unique scenes along with dense 3D labels and CAD models. Our dataset on the other hand is three times the size of ScanNet, captured with Apple’s LiDAR scanner instead of Kinect, and it provides 3D oriented bounding boxes and ground truth depth, which are not provided in .

3D object detection is a computer vision task that has gained a lot of popularity in recent years . The majority of the published techniques focus on outdoor environments mainly in the context of autonomous driving, where diverse datasets such as exists. Most of these techniques make assumptions in the algorithm (e.g Bird’s Eye View projection) that do not generalize well to indoor scenes. For indoor 3D object detection the number of datasets is limited. SunRGB-D and ScanNet are the two most commonly used ones. The former lacks scene level labels and the latter lacks oriented 3D bounding box labels. More recent datasets such as provide a variety of objects with 3D labels but do not provide depth sensor data. ARKitScenes provides the largest set of 3D oriented bounding boxes for a set of 17 room-defining object categories that addresses the gaps of previous indoor datasets.

Color-guided depth upsampling is the task of generating a high resolution (HR) depth map by using a high-resolution color image as guidance for the upsampling of a low resolution (LR) depth map . HR depth maps are essential for many depth use cases which require high frequency depth information. Prior works are using datasets of a few images with high resolution ground truth, such as 34 images from Middlebury , or 58 synthetic images from MPI-Sintel, or using the low resolution depth sensors as HR ground truth , and downscaling it even further in order to obtain the LR image. Table 1 compares those datasets and their properties. ARKitScenes is unique since it does not use the ground truth image as the source to generate the LR image by simple downscaling, instead it is providing LR depth maps captured with a consumer grade handheld LiDAR scanner and registered ground truth high-resolution depth maps captured with a professional stationary laser scanner. As a result upsampling methods trained with ARKitScenes are expected to generalize better to real-world scenarios as will be demonstrated below.

The rest of this paper is organized as follows. First, in Section 3, we introduce our data collection protocol as well as details around hardware and software used to capture the data. Moreover we cover the details around how we gather ARKitScenes ground truth. Next, in Section 4, we explore ARKitScenes for two downstream tasks of 3D object detection and depth upsampling. Finally, in Section 5, we summarize our findings and propose some future work.

ARKitScenes dataset

In this section, we describe the steps we pursued to acquire this dataset from collecting raw data in real-world homes, our data collection app, to fully-automatic spatial registration of the collected video sequences with highly accurate high-resolution stationary laser scans as well as manual annotation of 3D bounding box labels in the dataset.

We used two main devices for data collection: The 2020 iPad Pro and Faro Focus S70. The 2020 iPad Pro is used to collect various sensor outputs such as IMU, RGB (for both Wide and Ultra Wide cameras) as well as the dense depth map from the LiDAR scanner via ARKit. We use the official ARKit SDKhttps://developer.apple.com/documentation/arkit to collect such information. Our data collection app runs ARKit world tracking and scene reconstruction during the capture. This is to provide live feedback to the operators, who are not computer vision experts, on tracking robustness and reconstruction quality. In addition to the handheld iPad Pro we utilized a Faro Focus S70 stationary laser scanner on a tripod to collect high-resolution XYZRGB point clouds of the environment.

For data collection locations, we use real-world homes which we rent for a full day. The home owners consented to this data being released publicly to facilitate research and development of indoor 3D scene understanding. The operator was instructed to remove any personally identifiable information prior to starting the captures. Data is collected in three major cities in Europe, London, Newcastle, and Warsaw. To increase the indoor scene diversity and coverage we took two criteria into account when selecting homes for data collection: the socioeconomic status (SES) of the household as well as the location of the house in the city. The houses in our dataset are selected from rural, suburban, and urban location in each of the aforementioned cities. Additionally we included houses from all three tiers of low, medium, and high SES levels.

After selecting the house for data collection, we divide each house into multiple scenes (in most cases, each scene covers one room), and perform the following steps. First, we use a Faro Focus S70 stationary laser scanner on a tripod to collect highly accurate XYZRGB point clouds of the environment. Tripod locations are chosen to maximize surface coverage, and on average we collected four laser scans per room to ensure good coverage. Second, we record up to three video sequences attempting to capture all surfaces in each room using the iPad Pro. Each sequence follows a different motion pattern and captures the ceiling, floors, walls, and room-defining objects. The on-device ARKit world tracking poses as well as the scene reconstruction are stored and provided with the dataset, and they are also overlaid on the camera stream in the data collection app, to ensure the objects in the room are well covered. An example of such scan patterns and our app UI is shown in our supplementary materials.

Throughout the collection of all data, we attempt to keep the environment completely static, i.e. we make sure no objects move or change their appearance. However, since data collection of a venue takes an average of six hours and many venues are lit by sunlight, the lighting situation can change during that time resulting in potentially inconsistent illumination between the different sequences and scans.

2 Ground truth generation

Ground truth poses and depth maps. After data collection, in a one-time offline step, we first spatially register all XYZRGB point clouds from the stationary laser scanner into a common coordinate system using the proprietary software Faro Scene, which for most scenes fully automatically estimates a 6DoF rigid body transformation for each scan transforming it into a common venue coordinate system. Note that a venue (usually a house or apartment) can comprise multiple unique scenes. After this step, and throughout the rest of this paper, the XYZRGB point clouds are always assumed to be expressed in common venue coordinates.

Our approach to estimate the ground truth 6DoF pose of the iPad Pro’s RGB cameras with respect to the venue coordinate system requires the generation of synthetic views of our laser scan of the venue. Rendering these XYZRGB point clouds from novel viewpoints poses unique challenges. In particular, we require that far geometry is correctly occluded by near geometry and that geometry is discarded for which a direct line-of-sight from the novel view point cannot be guaranteed (e.g. they might be occluded by surfaces that were not captured in the scan). Naively rasterizing the scans as unstructured point clouds would violate both of these requirements.

Instead, we first find for each scanned 3D point cloud a triangulation by reducing it to two dimensions via stereographic projection with respect to the laser scanner’s nodal point and computing a 2D Delaunay triangulation. When applied to the 3D point cloud this triangulation is forming a watertight mesh, for which we compute texture coordinates referring to an equirectangular texture into which we write the RGB color information from the XYZRGB point cloud.

The triangles of this mesh are then split into two sets by applying a threshold on the angle between the triangle normal and the ray from the triangle center to the laser scanner nodal point. When this angle exceeds a threshold the triangle is considered to manifest a discontinuity and will be used as occlusion geometry; otherwise it will be used as foreground geometry. This separation enables reasoning about unobstructed line-of-sight. See Figure 3 (b) for an example visualization.

Finally, the two sets of triangles are rendered in separate passes using OpenGL, both writing to the depth buffer, while only front-facing triangles of the foreground geometry write a nonzero value to the stencil buffer. In all other cases the stencil buffer is cleared. As a result only fragments with unobstructed line-of-sight towards front-facing foreground geometry will have a nonzero stencil value, hence the stencil buffer can be used to mask pixels in the rasterized output. To create a joint rendering of multiple scans each scan is rendered separately and the renderings are merged in screen space using the individual depth and stencil buffers. After repeating this process for every scan the rasterization of a synthetic view is complete. This strategy is independent of the order in which individual rasterization results are merged and enables rasterizing depth maps and RGB images of one or multiple stationary laser scans from arbitrary viewpoints.

To be used as ground truth, the next objective is to determine the 6DoF pose of the handheld iPad Pro’s cameras with respect to the venue coordinate system. To this end we extract local image features and descriptors in keyframes of the camera sequence and store them as query features. Using the presented rasterization method we then create renderings of the laser scans, in which we detect and describe the same kind of local image features and store them as reference features together with their 3D locations, which can be looked up in the rasterized depth maps. We then match each query feature with the most similar reference feature leading to a set of 2D-3D correspondences for each keyframe. We use RANSAC and PnP to estimate initial camera poses using these sparse correspondences. To improve over these initial estimates we jointly solve for the camera poses of all keyframes that minimize a dense photometric error metric. For a single keyframe the metric optimizes photometric consistency between (1) the keyframe image and rendered views of the laser scans and (2) the keyframe image and projected views of neighboring keyframes that share visibility of the same parts of the laser scan surface geometry. This registration is performed once offline per video sequence.

When the refined ground truth poses are determined we render camera frame-aligned ground truth depth maps encoding per-pixel orthographic depth. This dense ground truth depth map, along with the per-keyframe ground truth pose are provided as part of the dataset. Figure 3 visualizes the results for a set of keyframes from the dataset.

3D object bounding boxes. We use a custom tool to manually annotate 3D oriented bounding boxes for 17 categories of room-defining furniture. The annotation happens on the ARKit scene reconstruction, which leads to a colored mesh of the scene. Additionally, our labeling tool allows annotators to see real-time projections of 3D bounding boxes onto video frames to facilitate accurate annotation.

Train/Validation/Test split. We split the venues of ARKitScenes into 80% for training, 10% for validation and the remaining 10% are a held-out test set that we do not release. The training and validations set include the 5,0485,048 sequences which we release. Since the split is determined on a per-venue basis, all laser scans and iPad sequences of a given venue fall into the same bin. The split is common across all downstream tasks including those discussed below.

Tasks and benchmarks

To evaluate the performance of different algorithms on ARKitScenes, we chose two computer vision tasks and trained state-of-the-art machine learning models using our dataset.

Problem details. As a fundamental task in computer vision, the goal of 3D object detection is to localize and recognize objects in a 3D scene. In ARKitScenes, we focus on two setups for our 3D object detection: single-frame based and whole-scene based. Given an RGB-D video sequence, the former targets detecting objects in each single RGB-D frame, while the latter detects objects from the whole reconstruction of the 3D scene.

Ground Truth Processing. Given that our ground truth bounding boxes are labeled on each scene reconstruction, we can directly use them for whole-scene 3D object detection scenario. However, for single-frame, we need to pre-process them as follows. Similar to outdoor benchmarks such as KITTI , we only keep boxes that at leasat have five corners in the camera frustum. Additionally a bounding box is removed if it contains fewer than 10 points.

After the bounding box exclusion, 65% of frames will be empty with no bounding boxes remaining. Moreover, we keep at most 300 frames for each video scan to down-weight really long videos. After filtering and long video sampling, we are left with over one million frames. Training current state-of-the-art detection algorithms off the shelf with this many frames will take a very long time. To further speedup, we subsampled one third of the data for our training, ending up with 365,007 frames: 323,868 for training and 41,139 for testing.

Models and training details. Building on the PointNet++ backbone and Hough voting modules, VoteNet achieves state-of-the-art performance on indoor scenarios. Along this line, MLCVNet and H3DNet further improves the VoteNet model by leveraging an attention model and extra geometric primitive prediction.

We followed the original design of these three approaches: the backbone network is a PointNet++ with several set-abstraction layers and feature propagation (upsampling) layers with skip connections, which outputs a subset of the input points with XYZ and an enriched CC-dimensional feature vector. The results are MM seed points of dimension (3+C)(3+C). Each seed point generates one vote. Each seed goes through a Hough voting module with supervision to guide each foreground point to vote to its bounding box center. Finally a last proposal module aggregates the votes with a shared PointNet to predict center, size and category of each bounding box. We use ADAM optimizer to train our model for 200 epochs, with learning rate 0.0010.001 and decay rate 0.10.1 at the 80-th and 120-th epoch. We augment our data with rotation, scaling and translation.

For single frame evaluation, training the vanilla version of all these models will take a very long time to converge. To speedup, we train all three models with constrained computational budget (at most two weeks) by using a lighter network backbone and fewer training epochs. First, as all three models are based on PointNet++ backbone with four set abstraction (SA) layers and two feature propagation/upsamplng (FP) layers, we reduced the output dimension of four SA from 256 to 128 and the depth of each MLP module from three layers to two layers. Second, we reduced the training epoch number from 180, 360, 360 to 100, 80, 60 for VoteNet , MLCVNet and H3DNet respectively. Each of the three models training take about approximately two weeks after our speedup.

Benchmark. As a baseline evaluation, we first show the performance of object detection on whole-scene in Table2. VoteNet is able to achieve mAP (mean average precision) of 0.358, while extra primitive supervision and attention model can further improve the overall performance to 0.383 and 0.419 respectively. We observe that these models perform better on large objects, such as refrigerator and bathtub, and struggles on small objects such as stove, dishwasher and TV/monitor. In Table 3, we show performance of single-frame detection. We observe a similar trend for different categories in both task settings. Finally, Fig 4 shows qualitative results of VoteNet on ARKitScenes for both single-frame and whole-scene. For more results and additional experiments please refer to the supplementary material.

2 Color-guided depth upsampling

Problem details. Depth upsampling is a common approach used to enhance low resolution (LR) depth maps to a high resolution (HR), higher fidelity depth map using an HR color image as guidance. HR accurate depth maps are imperative for downstream tasks such as 3D reconstruction, augmented reality, and photography. All of these require high frequency depth information which is often lost at low resolution.

Initial classical approaches to color-guided depth upsampling include both optimization and filtering-based methods . These approaches performed relatively well but are often hand-crafted and lack the ability to capture global structure and context. More recently, a data-driven approach using deep neural networks helps overcome some of these challenges. One of the prominent works in this area is Multi-Scale Guided Networks (MSG) , an encoder-decoder network extracting features at different resolutions from the guiding image in the encoder branch, and concatenating it in the corresponding resolution of the decoder branch of the depth map. The network is trained to learn the differential correction of the naïve upsampling. The current state-of-the-art in the field of guided depth upsampling is Multi-Scale Progressive Fusion (MSPF) , where the authors suggest the use of two different encoder branches, one for depth and one for color, along with a reconstruction branch that applies fusion blocks to restore the HR depth map.

ARKitScenes adaptations. As mentioned in Section 1, prior works on depth upsampling were demonstrated over LR depth maps that were generated by down-sampling ground truth HR depth maps, sometimes with the addition of artificial noise. This inherently makes these datasets easier for processing but limits the evaluation to non-realistic scenarios in which the low resolution sensor is only suffering from synthetic artifacts not necessarily representative of real-world scenarios. However, ARKitScenes brings a more realistic challenge of upsampling a low resolution depth map captured with the LiDAR scanner on a mobile device that has artifacts inherent to active sensing. Hence the challenge becomes twofold: depth upsampling and depth artifacts correction.

Another topic is that the HR ground truth depth map in ARKitScenes is a projection of laser scans that were taken from different viewpoints, therefore occlusions may cause some parts of the image to lack depth information. Hence, some of the methods in the existing research need to be adapted to handle the special value of no-depth pixels in the ground truth. Specifically, the Structural Similarity Index Measure (SSIM) loss used by MSPF cannot be used over ARKitScenes, as it is a full-reference method, requiring the information in all pixels without masking. In addition, the edge loss used by MSPF operates on the entire image and therefore needed to be changed to a more robust loss. We opted to use the edge loss suggested by . More details about these adaptations and experiments can be found in the supplementary material.

Experimental results. We would like to compare the results of existing methods on ARKitScenes. For these experiments we reproduced three classical approaches for depth upsampling - naïve Bilinear interpolation, Joint Bilateral Upsampling (JBU) and Fast Guided Global Interpolation (FGI) , as well as the two mentioned modern DNN-based solutions - MSG and MSPF .

In order to have cleaner data for training, we removed frames where regions of missing depth information took more than 40% of the HR depth map. Also, in order to avoid issues originating from the ground truth registration process, we ignore frames at which the Root Mean Square Error (RMSE) between the LR and a downscaled HR depth map is more than 7cm or when comparing per pixel, more than 20% of the pixels in the frame differ by more than 5cm.

Conclusions

We presented ARKitScenes, it is not only the first dataset that is captured with Apple’s LiDAR scanner, but also to the best of our knowledge the largest indoor RGB-D dataset ever collected with a mobile device. We showed how our dataset can be used for two downstream computer vision tasks of 3D object detection and color-guided depth upsampling. We believe ARKitScenes will enable the research community to push the boundaries of existing state of the art and develop technologies that better generalizes to real-world scenarios.

References