Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Santhosh K. Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Undersander, Wojciech Galuba, Andrew Westbury, Angel X. Chang, Manolis Savva, Yili Zhao, Dhruv Batra

Introduction

As we seek to develop intelligent AI agents that can assist us in our daily activities, good models of indoor 3D environments are becoming increasingly important. Consequently, recent years have seen growing demand for datasets of 3D interiors, whether acquired from the real world, or authored by artists using 3D design tools. Scene datasets based on real-world interiors can be used to develop and evaluate computer vision systems (e.g., on object detection and semantic segmentation tasks), or to train AI agents to navigate and follow instructions in an embodied setting. The latter research agenda in particular has been accelerated by the availability of realistic 3D datasets and high-performance simulators that dramatically reduce the time and logistical complexity for developing AI agents.

Unfortunately, there are only a handful of datasets of indoor 3D environments captured from the real world. Early efforts on 3D scene datasets such as SceneNN and ScanNet collected reconstructions of regions of rooms, and individual rooms. Other datasets that provide 3D reconstructions of entire buildings such as the BuildingParser , Matterport3D and Gibson efforts are either limited in total size or suffer from incomplete reconstructions.

Three key characteristics distinguish HM3D relative to prior work on real-world scanned indoor 3D datasets: scale, completeness, and visual fidelity. Unlike prior datasets, each scene in HM3D typically represents a complete building such as a multi-floor private residence. Therefore, HM3D has significantly higher total navigable area (1.4 \mathchar45 3.7×1.4~{}\mathchar 45\relax~{}3.7\times larger), which is particularly important for embodied AI tasks such as navigation. The completeness of HM3D is reflected in 34 \mathchar45 91%34~{}\mathchar 45\relax~{}91\% reduction in reconstruction artifacts due to missing surfaces, holes, or untextured surface regions when compared to prior photorealistic 3D datasets. This increased surface completeness leads to lower incidence of highly unrealistic ‘seeing through a hole in the wall’ issues that can be detrimental to embodied AI agent training. Finally, the visual fidelity of images rendered from HM3D is 20 \mathchar45 85%20~{}\mathchar 45\relax~{}85\% higher than prior large-scale datasets, which can help to train better embodied AI agents that generalize to real-world settings. As the name suggests, HM3D is ‘Habitat-ready’, meaning that it comes prepacked with meta-data and support necessary to be used with the Habitat simulator for training embodied AI agents to understand and navigate 3D spaces.

We carry out a number quantitative analyses and experiments to understand the characteristics of HM3D. First, we compare rendered images from HM3D and other 3D scan datasets to camera-captured images from the counterpart real-world interiors, and find that HM3D has significantly higher visual fidelity than other datasets. Second, we find that HM3D has fewer artifacts leading to incompleteness and ‘holes’ in surface reconstruction. Finally, we train agents for the task of PointGoal navigation using HM3D and other datasets, and find that agents trained in HM3D generalize well across environments. In particular, HM3D is pareto-optimal in the sense that HM3D-trained agents achieve the best performance across Gibson, MP3D, and HM3D test sets. HM3D-trained agents also achieve perfect success on Gibson test scenes, and obtain 3 \mathchar45 43~{}\mathchar 45\relax~{}4 points higher success and SPL on MP3D test scenes when compared to the next-best agent. These results strongly suggest that embodied agents benefit from the increased scale and diversity of HM3D.

Related Work

3D datasets can broadly be categorized into synthetic/CAD-based, 3D reconstruction or mesh-based, floorplan-based and panorama-based datasets.

Synthetic 3D scene datasets. Embodied AI simulation engines often make use of synthetic scenes with rearrangeable objects . Often these scenes are limited to isolated rooms or individual rooms that are connected via a magic portal . There are also datasets of building-scale synthetic scenes . However, these authored scenes often do not reflect the variety of architectural layout as well as object arrangement and clutter in the real world. Typically, objects in datasets of synthetic scenes are limited in visual and geometric diversity, since the same set of objects are reused across scenes. In addition, there is a sim-to-real gap between the rendered appearance of synthetic objects and real-world objects. To limit this discrepancy between synthetic and real domains, there are a number of recent synthetic scene dataset efforts that are designed from real world counterpart environments . HM3D is a reconstruction dataset capturing the layout and appearance of a large number of real buildings.

3D reconstruction datasets. Existing reconstructions of indoor spaces are limited in scale. Common reconstruction datasets consist primarily of scans for regions of rooms and single rooms .ScanNet and Replica do contain a number of multi-room scenes. There exist datasets with building level reconstruction, but these are limited in the overall number of scenes and real-world spaces (BuildingParser , 2D-3D-S , Matterport3D ). The largest building-level reconstruction dataset is Gibson which consists of 571571 scenes. However, many scans from Gibson suffer from reconstruction artifacts and ‘holes’ due to partially reconstructed surfaces. Prior work performed manual inspection and found that only 106106 / 571571 Gibson scans are of acceptable quality (i.e., ≥4\geq 4 on a scale of 0-5) . Dehghan et al. have recently released reconstructions of 1@6611@661 scenes but most are single room-scale regions, from rented homes in three European cities. HM3D contains 1@0001@000 building-scale reconstructions spanning a diverse set of locations around the world.

Floorplan and panorama datasets. Floorplan datasets can be converted to 3D floorplans outlining the architectural layout of buildings and rooms using heuristics. However, the architectural layout tends to be oversimplified as there is typically no specification of wall height or ground level (i.e. all rooms have equal height and simple flat ceilings). Most importantly, these datasets do not provide the textured appearance of the environments or of furniture and other objects present in the rooms. Recently, the Zillow indoor dataset provides floorplans and panorama images captured from a variety of properties. However, almost all captured properties are unfurnished, and even with panorama images available it is not easy to produce 3D mesh reconstructions of the interiors. In contrast, HM3D provides a large number of interiors that are reconstructed to a higher surface completeness and with higher visual fidelity than prior reconstruction datasets.

Dataset

The Habitat-Matterport 3D Dataset (HM3D) is a collection of 1@0001@000 3D reconstructions and consists of multi-floor residences, stores, and other private indoor spaces. All spaces were scanned using a Matterport Pro2 tripod-based depth sensor.https://matterport.com/cameras/pro2-3D-camera Alignment of the RGB-D data, surface meshing, and texturing were carried out using the reconstruction pipeline provided by Matterport, Inc. Scans were taken from spaces in 3838 countries, and 181181 geographic regions (states, provinces etc.) across those countries. In the United States, spaces are located in 4343 states. Figure 2 shows some example scenes.

The set of 1@0001@000 scenes in HM3D were curated from a larger pool of candidate scenes through a two-stage annotation and verification process. First, a group of 15 volunteer annotators rated each scene on a 1-5 quality scale assessing scene quality. The annotators visually inspected each scene for reconstruction artifacts such as holes/cracks, the presence of realistic and dense furnishing, the number of closed doors (which prevent access to some rooms), and the ‘interactive potential’ of the scene (based on objects with which a person might interact). The annotators also had access to quantitative metrics characterizing the navigability of the scene so that they could detect reconstruction issues causing disconnectedness in the building floors. After a round of rating, a second set of three expert annotators collated scenes by first sorting highly ranked scenes and then selecting so as to preserve diversity in the number of floors per scene. In total, HM3D curation, annotation, and release represents an estimated 800800+ hours of human effort.

The final set of scenes spans a broad spectrum of total area, with the smallest scene having a floor area of 49m249\text{m}^{2} and the largest scene an area of 2@172m22@172\text{m}^{2}. The architectural layout of the scenes also spans a broad spectrum with buildings of between one and eight floors, and between one ‘room’ and 9393 rooms.Room statistics were obtained using the mesh chunk meta-data from the Matterport reconstruction pipeline. Each mesh chunk is created by the reconstruction pipeline from a set of tripod locations in the same room. More detailed statistics regarding the dataset composition are visualized in Figure 3.

We compare the scale of HM3D to other datasets using a number of metrics that measure the overall floor area, navigable area, and structural complexity of the scenes.

Navigation complexity measures the difficulty of navigating in a scene. This is computed as the maximum ratio of geodesic path to euclidean distances between any two navigable locations in the scene. This is the same metric as reported for the original Gibson dataset to again make the statistics comparable . Higher values indicate more complex layouts with navigation paths that deviate significantly from straight-line paths.

Table 1 reports the values of these metrics for HM3D as well as a number of other indoor datasets, primarily focusing on existing 3D reconstruction datasets. We also compute the metrics for the RoboTHOR dataset which is synthetic but based on real-world layouts. The chosen comparison points span a spectrum of total sizes and complexities. For Gibson, note a second set of metric values for the restricted subset of fewer “high quality” Gibson scenes that were rated as at least 4/5 by a set of human annotators. This subset of Gibson exhibits fewer reconstruction artifacts than the full Gibson dataset (see Savva et al. for a description of the original rating process).

We can make a number of observations. First, HM3D provides 1.7×1.7\times higher floor area and 1.4×1.4\times higher navigable area compared to the Gibson (previously the largest). In particular, HM3D provides 20×20\times higher floor area and 15.6×15.6\times higher navigable area if we only consider the high quality reconstructions in Gibson 4+. Second, the scene clutter of HM3D exceeds that of most other datasets by ∼ ⁣ ⁣1.2×\sim\!\!1.2\times with the exception of RoboTHOR which is a significantly smaller dataset. Finally, the navigation complexity metric shows that HM3D scenes are relatively complex to navigate, close to other building-scale datasets such as MP3D and Gibson (by a factor of 0.8 \mathchar45 1.1×0.8~{}\mathchar 45\relax~{}1.1\times), and higher than room-scale datasets such as Replica, RoboTHOR and ScanNet (by a factor of 2.2 \mathchar45 6.4×2.2~{}\mathchar 45\relax~{}6.4\times).

2 Reconstruction completeness comparison

We compare the distribution of this metric for HM3D and other datasets in Figure 4 (right). Overall, we see that HM3D scenes exhibit fewer artifacts (more scenes with lower ‘% defect’ values). While ScanNet offers more scans than HM3D, almost all scans from ScanNet exhibit severe reconstruction artifacts. Other large-scale datasets such as Gibson, and MP3D exhibit broader distributions with a significant number of scenes having fairly high reconstruction defect values. HM3D has more than three times as many scenes with less than 5% of views exhibiting artifacts compared to Gibson (560 scenes vs 175 scenes). As expected, Gibson 4+ provides a smaller but higher-quality subset of Gibson scenes with fewer reconstruction artifacts. While RoboTHOR scans have very few artifacts, they are not photorealistic and are small in quantity. The Replica scans are much smaller in number, and some scans have ∼ ⁣80%\sim\!80\% defects since the roof is missing. Overall, HM3D offers the largest number of scans with high completeness.

3 Visual fidelity comparison

We also compare the overall visual quality of rendered images from HM3D with prior datasets. For each dataset, we use the RGB images from Section 3.2 to ensure that we assess the visual fidelity of rendered images from all parts of a scene. We compare the image quality against a set of real RGB images generated from high-resolution panoramas (i.e., 360∘360^{\circ} field-of-view equirectangular images) in Gibson and MP3D using the FID and KID metrics. We refer to these sets of real RGB images as ‘Gibson real’ and ‘MP3D real’. Figure 5 summarizes the results of this comparison.

The quality of images rendered from HM3D is much closer to real images when compared to the other datasets. Out of all datasets, images rendered from HM3D exhibit the lowest FID / KID scores when compared with both MP3D real (20.53/15.7820.53/15.78) and Gibson real (20.49/12.7620.49/12.76). Note that we observe a domain shift between datasets that leads to non-zero FID and KID scores even for real images. Comparing images from Gibson real with MP3D real provides a ‘lower bound’ of 6.16−6.236.16-6.23 against which we can compare the metric values for rendered dataset images. As expected, images rendered from RoboTHOR have the highest FID and KID scores since they are not photorealistic. Images rendered from ScanNet also have high FID and KID scores since they exhibit significant mesh artifacts (see Sec. 3.2). Images rendered from Replica have relatively high distance scores (despite having high quality scans) due to the lack of textured ceilings in multiple scans (eg., 34.9434.94 FID vs. Gibson real, and 42.7642.76 FID / 19.3119.31 KID vs. MP3D real). Images rendered from the remaining datasets have significantly higher FID and KID values when compared to HM3D, showing that they have lower visual fidelity as measured against the real images from Gibson and MP3D.

Experiments

A popular downstream application for large-scale 3D reconstruction datasets has been to use them with 3D simulation platforms to study embodied AI tasks such as visual navigation . As described in the previous sections, HM3D improves over existing datasets both in terms of size and quality. In this section, we perform experiments to show that navigation agents trained on HM3D benefit from its scale and quality, and generalize better when transferred to other datasets.

We train and evaluate PointNav agents on Gibson 4+, Gibson, MP3D, and HM3D datasets. We divide the 1@0001@000 HM3D scenes into disjoint sets of 800800 train / 100100 val / 100100 test scenes. We use the standard train / val / test splits for Gibson 4+ and MP3D . We create new PointNav episode datasets for the full Gibson train scenes and HM3D using the generation script from Savva et al. . Specifically, we generate 4.11M train episodes for Gibson, and 8.0M train / 2500 val / 2500 test episodes for HM3D.10@00010@000 episodes per train scene. 25 episodes per val/test scene. These splits are publicly available to aid reproducibility: https://github.com/facebookresearch/habitat-lab. In general, MP3D has the hardest episodes and Gibson has the easiest episodes. See Appendix A7 for a comparison between the different PointNav episode datasets.

We use a standard agent architecture for training on different datasets . A ResNet-50 backbone extracts visual features , and an MLP extracts location features from GPS+compass readings. An LSTM state-encoder aggregates these features over time , and fully-connected layers are used to predict action logits (i.e., the policy) and state values (i.e., the value function). Actions are then stochastically sampled from the predicted action logits. We train the agent using DD-PPO for 1.5 billion frames which was shown to be sufficient to achieve near state-of-the-art performance .

We separately benchmark agents for two types of inputs. For ‘RGB inputs’, the agent navigates using RGB and GPS+compass sensors. For ‘depth inputs’, the agent navigates using depth and GPS+compass sensors. For brevity, ‘X agent’ refers to an agent trained on dataset X (e.g., Gibson agent), and ‘X agent (R, D)’ denotes the SPL performance of X agent with RGB (R) and depth (D) inputs. Table 2 and Table 2 present results for all agents and datasets. We analyze these results next to answer 3 key questions.

1) Is HM3D beneficial for training PointNav agents? Consider the validation curves in Table 2. Both HM3D agents (RGB and depth inputs) converge faster and perform better than corresponding agents trained on other datasets. Specifically, the HM3D agent closely follows the Gibson agent on Gibson (val) and outperforms it on MP3D (val) and HM3D (val). The validation performance of the HM3D agent rapidly outpaces the MP3D and Gibson 4+ agents on all cases. The test performance in Table 2 confirms the above trends. On Gibson (test) with depth inputs, the HM3D agent matches the Gibson agent achieving 0.93 SPL and 1.0 success. On all other cases, the HM3D agents outperform the other agents by a large margin. For example, on MP3D (test), the HM3D agent (0.71, 0.83) significantly outperforms the second-best Gibson agent (0.68, 0.80). Thus, HM3D is pareto-optimal since the HM3D agents achieve the best performance on all test sets.

2) Are HM3D scenes diverse in terms of visual appearance and 3D layouts? Diversity in visual appearance and 3D layouts in the training scenes is essential for generalization to novel scenes and datasets, and adaptability to difficult PointNav episodes. From Table 2, HM3D agents (both RGB and depth) outperform the next best method on MP3D (test) by 3 SPL points, and achieve perfect success on Gibson (test). This is impressive generalization since the HM3D agent had not observed any Gibson or MP3D scenes during training, and yet was able to overcome the domain gap in appearance and layouts of the scenes. This attests to the visual richness and layout diversity of HM3D which enables good generalization to previously unseen scenes and datasets.

Next, we compare the performance of different agents as a function of the episode difficulty in Figure 7. We quantify episode difficulty using the geodesic distance between the start and goal locations . We group the MP3D (test) episodes into different bins based on the above metric, and plot the mean and standard deviation of an agent’s performance on all episodes in each bin. We select MP3D (test) since it has highest diversity of difficulty levels (see Appendix A7). We observe that the HM3D agent adapts much better compared to other agents as the episode difficulty increases. This is yet another indicator that the layouts in HM3D are complex and diverse. 3) Does PointNav benefit from scaling up 3D datasets? Prior work has verified the data scaling hypothesis for passive perception, i.e., scaling up the dataset can significantly improve performance on various passive perception problems . However, this is not well-established in the embodied perception literature due to the lack of large-scale 3D datasets with high quality. As discussed in earlier sections, HM3D offers large-scale, high visual fidelity, and high-quality reconstructions. Thus, we use HM3D to test the data scaling hypothesis for embodied AI. Specifically, we test the relationship between the training dataset size and the PointNav performance. For this purpose, we additionally train agents on two random subsets of HM3D containing 10%10\% and 50%50\% of the scans in the HM3D train split (i.e., 80 and 400 scans respectively). We refer to these agents as HM3D (10%10\%) and HM3D (50%50\%). Figure 8 shows the PointNav SPL on the test splits as a function of the total navigable area in the training scenes. We observe that the navigation performance is strongly correlated with the total navigation area (Pearson coefficient ρ=0.88\rho=0.88), and that the performance scales near-linearly as the total navigable area increases. This result is also helpful to decide the data budget for training PointNav agents. Using more data leads to better performance (particularly on the harder episodes in MP3D), but requires more computational resources and time. Depending on the task difficulty and availability of computational resources, researchers can choose the appropriate dataset(s) for experimentation.

Conclusion

We presented the Habitat-Matterport 3D (HM3D) dataset consisting of 1@0001@000 building-scale reconstructions from the real world. To our knowledge, HM3D offers the largest dataset of high-quality 3D reconstructions of interiors for academic research. Through a series of quantitative analyses we showed that HM3D improves upon existing 3D reconstruction datasets in three ways: significantly larger spatial scale, improved reconstruction completeness, and higher visual fidelity. We also carried out experiments with PointGoal navigation for embodied AI agents to show that agents trained on HM3D match or outperform agents trained on other datasets even when evaluated on other datasets. This demonstrates the value of HM3D as a dataset for embodied AI. Extension of HM3D with object semantics and physical attributes in future work will enable even more embodied AI tasks such as ObjectGoal navigation and object rearrangement. We hope that HM3D will catalyze research in the area of embodied AI.

References

Acknowledgements

We thank all the volunteers who contributed to the dataset curation effort: Harsh Agrawal, Sashank Gondala, Rishabh Jain, Shawn Jiang, Yash Kant, Noah Maestre, Yongsen Mao, Abhinav Moudgil, Sonia Raychaudhuri, Ayush Shrivastava, Andrew Szot, Joanne Truong, Madhawa Vidanapathirana, Joel Ye. We thank our collaborators at Matterport for their contributions to the dataset: Conway Chen, Victor Schwartz, Nicole Rogers, Sachal Dhillon, Raghu Munaswamy, Mark Anderson.

Licenses for referenced datasets

Gibson: http://svl.stanford.edu/gibson2/assets/GDS_agreement.pdf Matterport3D: http://kaldir.vc.in.tum.de/matterport/MP_TOS.pdf ScanNet: http://kaldir.vc.in.tum.de/scannet/ScanNet_TOS.pdf Replica: https://github.com/facebookresearch/Replica-Dataset/blob/master/LICENSE

Appendix A1 Limitations of Habitat-Matterport 3D

Data acquisition: The dataset is currently limited to scans from 38 countries. The dataset is restricted to contain data from building-owners who can afford to purchase the Matterport Pro2 sensor (which costs \sim 3@000\$) and have internet access to upload data to the cloud. The dataset also excludes regions where the Matterport Pro2 is not available to purchase. Due to these factors, we are limited in the types of regions and neighborhoods which can be included in the dataset. This can introduce an unintended bias into the algorithms developed based on our dataset, where the algorithms work only in a subset of buildings that we encounter in the real world. Nevertheless, this dataset is a significant leap from past building-scale datasets that were restricted to labs, residences, and offices. We hope to expand our data set in the future to include scans from many more diverse backgrounds and countries.

Task support: The dataset only supports geometric tasks in static (i.e., unchanging) environments and does not include semantic annotations. We plan to investigate augmenting the dataset with semantic annotations to tackle high-level understanding tasks like object retrieval. We also plan to study dynamic and changing environments so that our simulations will be fluid rather than static. This would bring simulated training environments closer to the real world, where people and pets freely move around and where everyday objects such as mobile phones, wallets, and shoes are not always in the same spot throughout the day.

Appendix A2 Hyperparameters for PointNav experiments

We use the publicly available implementation of DD-PPO from Habitat Lab. We use the same hyperparameters as Wijmans et al. for our experiments. We use a ResNet-50 backbone and an LSTM with 512-D hidden states and 2 layers. Following Wijmans et al. , we replace BatchNorm with GroupNorm layers in the ResNet-50 backbone. We use a PPO clip parameter of 0.20.2, 2 PPO epochs, 2 mini-batches, a value loss coefficient of 0.50.5, entropy coefficient of 0.010.01 and a learning rate of 0.000250.00025. Please see the default configuration here for more details. We train each model for 1.5 billion steps (sufficient for convergence) with 256 parallel environments divided between 8 nodes, 4 workers (i.e., 4 GPUs) per node and 8 environments per worker.

Appendix A3 Computational requirements

The PointNav experiments were the most computationally expensive of all our experiments. Each experiment is run in our internal cluster in a distributed fashion over 8 nodes, with 4 GPUs per node. Each GPU (32 in total) is a Volta 16/32 GB. Training an agent takes 2-3 days with depth inputs, and 4-5 days with RGB inputs.

Appendix A4 Accessing Habitat-Matterport 3D dataset

HM3D is free and available for academic, non-commercial research here: https://matterport.com/habitat-matterport-3d-research-dataset The terms of use are available here: https://matterport.com/matterport-end-user-license-agreement-academic-use-model-data

Appendix A5 Habitat-Matterport 3D dataset collection process

The 1000 scans in HM3D were collected by Matterport Inc. in collaboration with the Habitat team at Facebook AI Research. Matterport directly contacted its users explicitly requesting them to contribute their scans for open-sourced Embodied AI research (see mailer snippet below).

Imagine if firefighters could ask a robot to detect where smoke is coming from within your house, then command it to find people who need help. Or, if you could ask an AI assistant to locate your car keys. To realize innovations like these, robots and AI assistants need to be trained in how to act in multiple environments. They must learn to recognize and navigate through 3D spaces. That’s where you come in. We have identified your Matterport 3D model as an ideal space for an open-source AI project focused on furthering such causes.

Each user agreed to the following terms while contributing their scans for the dataset.

I agree to allow Matterport to use the Space(s) (including all related imagery) that I have designated in this form for academic and/or non-commercial purposes as further provided in the Matterport Terms and Conditions for Academic and Non-Commercial Use of Spaces, without payment by Matterport for such use. I affirm that I have all necessary rights, consents and permissions relating to my Space(s) necessary to grant the foregoing permission. By checking this box, I specifically agree to all of the provisions of the Matterport Terms and Conditions for Academic and Non-Commercial Use of Spaces &\& Matterport Privacy Policy.

After obtaining scans from users, Matterport used commercially reasonable efforts to try and obscure personally identifiable information such as pictures of people or faces, names, documents with personal information, diplomas, driver’s licenses, email addresses, phone numbers, street addresses, personal notes / letters / envelopes, employer information, license plates, and street names. Human reviewers were asked to preview images from every scanned location to check for the above information and annotated each instance of personal information using a label. Any personal information identified in the previous step was blurred using a pixel-wise blur mask. The blurred data was used to recreate the 3D scans. A helpful FAQ regarding this process can be found here: https://go.matterport.com/ProjectHabitat.html. Overall, the dataset contains scans from users in 38 countries (see Figure A1).

Appendix A6 PointNav validation results

In Figure 6 from the main paper, we presented the validation performance as a function of the training steps. Now, we present the final validation performance of the best checkpoint (analogous to Table 2 in the main paper). See Table A1. We observe trends similar to the ones observed in the main paper. The HM3D agent matches the Gibson agent on Gibson (val) with RGB inputs. On all other cases, the HM3D agents outperform the other agents by a good margin, particularly in the RGB case.

Appendix A7 Comparing PointNav episode datasets

In Figure A2, we compare the difficulties of the Gibson (val), MP3D (val), and HM3D (val) episode datasets. MP3D has the hardest episodes and Gibson has the easiest episodes.

Appendix A8 Dataset characteristics that impact PointNav performance

In the main paper, we compared datasets along different characteristics such as navigable area, visual fidelity, 3D reconstruction quality, and the utility for training agents for the PointNav task. In Section 4 and Figure 6, we analyzed the impact of dataset size on the PointNav performance, and noted that total navigable area in the training scans are highly correlated with the PointNav results. Here, we perform a complete analysis of how the following factors affect the PointNav performance (on the val splits).

We compute the above metrics for all the train datasetsWe compute EMD(train, val) for all pairs of train and val sets.. For a given PointNav val set, we measure the Pearson’s correlation between each of the above metrics for a train dataset and the navigation SPL achieved by agents trained on the same dataset (see Table A2). As noted in the main paper, we observe that the navigable area is highly correlated with the PointNav performance (0.820.82 to 0.970.97) indicating that large-scale datasets are critical for achieving high-quality navigation. For both MP3D (val) and HM3D (val), there is a strong negative correlation of −0.4-0.4 to −0.7-0.7 between EMD (train, val) and SPL. This indicates that large distribution shifts between the train and val episodes leads to worse performance. Next, we observe that KID (mean) is weakly correlated with the SPL (−0.30-0.30 to 0.10.1), indicating that visual fidelity may not strongly impact PointNav performance in simulation. In most cases, % defects is generally uncorrelated with SPL. However, we find a slightly positive correlation (0.17\mathchar450.320.17\mathchar 45\relax 0.32) with SPL on MP3D (val). This may be due to the fact that MP3D val scenes have significantly more mesh reconstruction artifacts than Gibson val scenesOnly Gibson 4+ scenes are used for Gibson (val) and Gibson (test). (see Figure 4 in main paper). Agents trained on scenes with more mesh reconstruction artifacts adapt better to such testing conditions. Note that these results are not very indicative of the transfer performance to a real robot. It is possible that higher visual fidelity and lower % defects may be necessary for real-world transfer.

Appendix A9 Example scenes from Habitat-Matterport 3D

We provide more examples of scenes from Habitat-Matterport 3D in the same style as Figure 2 in the main paper. In Figure A3, we visualize 5 residences, and in Figure A4, we visualize 5 diverse scenes such as offices, gyms, restaurants, and nightclubs. All 900 scenes from the train and val splits of HM3D can be visualized on the dataset website: https://aihabitat.org/datasets/hm3d/

Appendix A10 PointNav qualitative results

We show sample PointNav episodes of the HM3D agents in Figure A5 and Figure A6. We present the qualitative results in a format similar to Wijmans et al. . The episodes are categorized based on the difficulty (i.e., the geodesic distance b/w start and goal), and the agent performance (in SPL).