Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at Scale

Ram Ramrakhya, Eric Undersander, Dhruv Batra, Abhishek Das

Introduction

General-purpose robots that can perform a diverse set of embodied tasks in a diverse set of environments have to be good at visual exploration. Consider the canonical example of asking a household robot, ‘Where are my keys?’. To answer this (assuming the robot does not remember the answer from memory), the robot would have to search the house, often guided by intelligent priors – e.g. peeking into the washroom or kitchen might be sufficient to be reasonably sure the keys are not there, while exhaustively searching the living room might be much more important since keys are more likely to be there. While doing so, the robot has to internally keep track of where all it has been to avoid redundant search, and it might also have to interact with objects, e.g. check drawers and cabinets in the living room (but not those in the washroom or kitchen!).

This example illustrates fairly sophisticated exploration, involving a careful interplay of various implicit objectives (semantic priors, exhaustive search, efficient navigation, interaction, etc.). Many recent tasks of interest in the embodied AI community – e.g. ObjectGoal Navigation , rearrangement , language-guided navigation and interaction , question answering – involve some flavor of this visual exploration. With careful reward engineering, reinforcement learning (RL) approaches to these tasks have achieved commendable success . However, engineering the ‘right’ reward function so that the learned policy exhibits desired behavior is unintuitive and frustrating (even for domain experts), expensive (requiring multiple rounds of retraining under different rewards), and not scalable to new tasks or behaviors. For complex tasks (e.g. object rearrangement or tasks specified in open-ended natural language), RL from scratch may not even get off the ground.

In this work, we advance the alternative research agenda of imitation learning – i.e. collecting a large dataset of human demonstrations (that implicitly capture intelligent behavior we wish to impart to our agents) and learning policies directly from these human demonstrations.

First, we develop a safe scalable virtual teleoperation data-collection infrastructure – connecting the Habitat simulator running in a browser to Amazon Mechanical Turk (AMT). We develop this in way that enables collecting human demonstrations for a variety of tasks being studied within the Habitat ecosystem (e.g. PointNav , ObjectNav , ImageNav , VLN-CE , MultiON , etc.).

We use this infrastructure to collect human demonstration datasets for 22 tasks requiring visual search – 1) ObjectGoal Navigation (e.g. ‘find & go to a chair’) and 2) Pick&Place (e.g. ‘find mug, pick mug, find counter, place on counter’). In total we collect 92k92k human demonstrations, 80k80k demonstrations for ObjectNav and 12k12k demonstrations for Pick&Place. In contrast, the largest existing datasets have 33-10k10k human demonstrations in simulation or on real robots , an order of magnitude smaller. This virtual teleoperation data contains 29.3M29.3M actions, which is equivalent to 22,60022,600 hours of real-world teleoperation time assuming a LoCoBot motion model from (details in appendix (Sec. A.3)). The first thing this data provides is a ‘human baseline’ with sufficiently tight error-bars to be taken seriously. On the ObjectNav validation split, humans achieve 93.7<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>±</mo><mn>0.1</mn></mrow><annotationencoding="application/x−tex">±0.1</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.7278em;vertical−align:−0.0833em;"></span><spanclass="mord">±</span><spanclass="mord">0.1</span></span></span></span></span>%93.7<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>±</mo><mn>0.1</mn></mrow><annotation encoding="application/x-tex">\pm 0.1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">±</span><span class="mord">0.1</span></span></span></span></span>\% success and 42.5<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>±</mo><mn>0.5</mn></mrow><annotationencoding="application/x−tex">±0.5</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.7278em;vertical−align:−0.0833em;"></span><spanclass="mord">±</span><spanclass="mord">0.5</span></span></span></span></span>%42.5<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>±</mo><mn>0.5</mn></mrow><annotation encoding="application/x-tex">\pm 0.5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">±</span><span class="mord">0.5</span></span></span></span></span>\% Success Weighted by Path Length (SPL) (vs. 34.6%34.6\% success and 7.9%7.9\% SPL for the 2021 Habitat ObjectNav Challenge winner ). The success rate (93.7%93.7\%) suggests that this task is largely doable for humans (but not 100%100\%). The SPL (42.5%42.5\%) suggests that even humans need to explore significantly.

Beyond scale, the data is also rich and diverse in the strategies that humans use to solve the tasks. Fig. 1 shows an example trajectory of an AMT user controlling a LoCoBot looking for a ‘plant’ in a new house – notice the peeking into rooms, looping around the dining table – all of which is (understandably) absent from the shortest path to the goal.

We use this data to answer the question – how does large-scale imitation learning (IL) (which has not been hitherto possible) compare to large-scale reinforcement learning (RL) (which is the status quo)? On ObjectNav, we find that IL (with no bells or whistles) using only 70k70k human demonstrations outperforms RL using 240k240k agent-gathered trajectories. This effectively establishes an ‘exchange rate’ – a single human demonstration appears to be worth ∼4{\sim}4 agent-gathered ones. More importantly, we find the IL-trained agent learns efficient object-search behavior – as shown in Fig. 1 and Sec. 7. The IL agent learns to mimic human behavior of peeking into rooms, checking corners for small objects, turning in place to get a panoramic view – none of these are exhibited as prominently by the RL agent. Finally, the accuracy vs. training-data-size plot (Fig. 1b) shows promising scaling behavior, suggesting that simply collecting more demonstrations is likely to advance the state of art further. On Pick&Place, the comparison is even starker – IL-agents achieve ∼18%{\sim}18\% success on episodes with new object-receptacle locations when trained with 9.5k9.5k human demonstrations, while RL agents fail to get beyond 0%0\%.

On both tasks, we find that demonstrations from humans are essential; imitating shortest paths from an oracle produces neither accuracy nor the strategic search behavior. In hindsight, this is perfectly understandable – shortest paths (e.g. Fig. 1(a3)) do not contain any exploration but the task requires the agent to explore. Essentially, a shortest path is inimitable, but imitation learning is invaluable. Overall, our work provides compelling evidence for investing in large-scale imitation learning of human demonstrations.

Related work

Embodied Demonstrations from Humans. Prior expert demonstration datasets for embodied tasks combining vision and action (and optionally language) can be broadly categorized into either consisting of shortest-path trajectories from a planner with privileged information , or consisting of human-provided trajectories . While some works in the former collect natural language data from humans , we contend that collecting navigation data from humans is equally crucial. Datasets with human-provided navigation trajectories are typically small. TEACh , CVDN and WAY have <10k{<}10k episodes, while the EmbodiedQA dataset has ∼700{\sim}700 human-provided episodes – all prohibitively small for training proficient agents. A key contribution of our work is a scalable web-based infrastructure for collecting human navigation and interaction demonstrations, that is easily extensible to any task situated in the Habitat simulator, including language-based tasks. We have collected ∼13{\sim}13x more demonstrations (in total 92k92k) compared to prior publicly available datasets. In a similar vein, Abramson et al. study large-scale imitation learning on ∼600k{\sim}600k human demonstrations, but their dataset is not publicly available and the environments used lack visual realism compared to Matterport3D .

Exploration. Learning how to explore an environment to gather sufficient information for use in downstream tasks has a rich history . Curiosity-based approaches typically use reinforcement learning to maximize intrinsic rewards that capture the surprise or state prediction error of the agent . State visitation count rewards are also popular for learning exploration . We refer the reader to Ramakrishnan et al. for a review of exploration objectives for embodied agents. For improving exploration in ObjectNav specifically, SemExp made use of a modular policy for semantic mapping and path planning, Ye et al. used time-decaying state visitation count reward, and Maksymets et al. used area coverage reward.

Most relatedly, Chen et al. used ∼700{\sim}700 human navigation trajectories from the EmbodiedQA dataset (ignoring the questions) to learn task-independent exploration using imitation learning. We likewise train agents via imitation learning on human demonstrations, but rather than encouraging task-agnostic exploration, we consider human demonstrations to be a rich task-specific mix of exploration and efficient navigation, that simple architectures without explicit mapping and planning modules can be trained on.

Habitat-WebGL Infrastructure

To be able to train agents via imitation learning on human demonstrations, we first need a reliable pipeline to collect human demonstrations at scale. To this end, we develop a web-based setup to connect the Habitat simulator to AMT users, building on the work of Newman et al. .

Interface. Fig. 2 shows a screenshot of the interface an AMT user interacts with to complete a data collection task. This web application renders assets from Habitat-Sim running on the user’s browser via WebGL. All data collection in this work was done in Matterport3D scans, but any Habitat-compatible asset may be used in future. Users can see the agent’s first-person RGB view, and can move around and grab / release objects using keyboard controls. On the task page, users are provided an instruction and details about keyboard controls to complete the task. For ObjectNav, we provide an instruction of the form ‘Find and go to the ’. For tasks requiring interaction with objects (e.g. Pick&Place), we highlight the object under the user’s gaze by drawing a 33D bounding box around it (pointed to by a crosshair as in video games). In our initial pilots, we found this to improve user experience when grabbing objects instead of users having to guess when objects are available to be picked up. When an object is successfully grabbed, it disappears from the first-person view and immediately appears in the ‘inventory’ area on the task interface. When a grabbed object is released, it is dropped at the center of the user’s screen where the crosshair would be pointing to. If the crosshair points to a distance, the object is dropped on the floor from a height at a distance of 1m1m from the agent’s location. Upon completion, users submit the task by clicking ‘Submit’. At this point, the sequence of keyboard actions, agent, and object states are recorded in our backend server.

Habitat simulator and PsiTurk. Our Habitat-WebGL application is developed in Javascript, and allows us to access all C++ simulator APIs through Javascript bindings. This lets us use the full set of simulation features available in Habitat. To simulate physics, we use the physics APIs from Habitat 2.02.0 , including rigid body dynamics support (C++ APIs exposed as Javascript bindings). Our interface executes actions entered by users every 5050ms (rendering 2020 frames per second) and then steps physics for 5050ms in the simulator. All of our tasks on AMT are served using PsiTurk and an NGINX reverse proxy, and all data stored in a MySQL database. We use PsiTurk to manage the tasks as it provides us with useful helper functions to log task-related metadata, as well as launch and approve tasks.

See Section A.6 for details on how we validate human-submitted AMT tasks and ensure data quality.

Tasks and Datasets

Using our web infrastructure, we collect demonstration datasets for two embodied tasks – ObjectNav and Pick&Place, an instantiation of object rearrangement .

In the ObjectGoal Navigation (ObjectNav) task, an agent is tasked with navigating to an instance of a specified object category (e.g. ‘chair’) in an unseen environment. The agent does not have access to a map of the environment and must navigate using an RGBD camera and a GPS+Compass sensor which provides location and orientation information relative to the start of the episode. The agent also receives the goal object category ID as input. The full action space is discrete and consists of move_forward (0.25m0.25m), turn_left (30∘30^{\circ}), turn_right (30∘30^{\circ}), look_up (30∘30^{\circ}), look_down (30∘30^{\circ}), and stop actions. For the episode to be considered successful, the agent must stop within 1m1m Euclidean distance of the goal object within a maximum of 500500 steps and be able to turn to view the object from that end position .

Human Demonstrations (ObjectNav-HD). We collect 70k70k demonstrations on the 5656 training scenes from Matterport3D following the standard splits defined in . For each scene, we collect ∼59{\sim}59 demonstration episodes for each unique goal object category with a randomly set start location of the human demonstrator for each episode. This amounts to an average of ∼1250{\sim}1250 demonstrations per scene. Additionally, we collect 10k10k demonstrations on 2525 training scenes from Gibson . For each Gibson scene, we collect ∼66{\sim}66 demonstration episodes for each unique goal object category. This amounts to ∼396{\sim}396 demonstrations per scene. Similar to when training artificial agents, humans can view first-person RGB on the task interface, but unlike artificial agents, humans do not get access to Depth and GPS+Compass. We assume humans are sufficiently proficient at inferring depth and odometry from vision, to the extent required to accomplish the goal. In total, we collect 80k80k ObjectNav demonstrations amounting to ∼19.5M{\sim}19.5M steps of experience, each episode averaging 243243 steps.

Shortest Path Demonstrations. To compare against prior embodied datasets of shortest paths and to demonstrate the unique advantage of human demonstrations, we also generate a dataset of shortest paths. The analysis in this section was performed on a subset of 3535k demonstrations of ObjectNav-HD (collected in first phase). These demonstrations are generated by greedily fitting actions to follow the geodesic shortest path to the nearest navigable goal object viewpoint. Since shortest paths are (by design) shorter than human demonstrations (average 6767 vs. 243243 steps per demonstration), we compensate by generating a larger number of shortest paths to roughly match the steps with 35k35k human demonstrations (7.6M7.6M steps from 114k114k shortest paths vs. 8.4M8.4M steps from 35k35k human demonstrations).

Analysis. Table 3a reports statistics of our human and shortest path demonstration datasets. Recall that an episode is considered a failure if the target object is not found within 500500 navigation steps. Under this definition, humans fail on 11.1%11.1\% training set episodes; they fail on 0%0\% episodes if we relax the step-limit. Surprisingly, SPL for humans is 39.9%39.9\% for training split episodes, significantly lower than 94.9%94.9\% for shortest paths underscoring the difficulty in searching for objects in in unseen environments.

We additionally report two metrics to demonstrate that the ObjectNav task requires significant exploration. Occupancy Coverage (OC) measures percentage of total area covered by the agent when navigating. To compute OC, we first divide the map into voxel grids of 2.5m×2.5m×2.5m2.5m\times 2.5m\times 2.5m and increment a counter for each visited voxel. Sight Coverage (SC) measures the percentage of total navigable area visible to the agent in its field of view (FOV) during an episode. To compute SC, we project a mask on the top-down map of the environment using the agent’s FOV, that is iteratively updated at every step to update the area seen by the agent. OC and SC metrics for human demonstrations show that humans traverse 3-43\text{-}4x and observe 22x the area of the environment when performing this task compared to shortest paths.

Fig. 3b,c show episode length and action histograms for human and shortest path demonstrations. Human demonstrations are longer (average ∼243{\sim}243 vs. ∼67{\sim}67 steps per demonstration) and have a slightly more uniform action distribution.

2 Object Rearrangement – Pick&Place

In the pick-and-place task (Pick&Place), an agent must follow an instruction of the form ‘Place the on the ’, without being told the location of the or in a new environment. The agent must explore and navigate to the object, pick it up, explore and navigate to the receptacle, and place the previously picked-up object on it. Similar to ObjectNav, agents are not equipped with a map of the environment, and only have access to an RGBD camera and a GPS+compass sensor. At a high level, Pick&Place can be thought of as a natural extension of ObjectNav, performing it twice in the same episode – once to find the specified object and again to find the specified receptacle – delimited by grab and release actions. For object interaction, we use the ‘magic pointer’ abstraction defined in . If the agent is not holding any object, the grab/release action will pick the object pointed to by its crosshair (at the center of its viewpoint) if within 1.5m1.5m of the object. If the agent is already holding an object, the grab/release action will drop the object at the crosshair location. If there is no drop-off point within 1.5m1.5m in the direction of the crosshair, the object will be dropped on the floor 1m1m in front of the agent. The full action space is discrete and consists of move_forward (0.15m0.15m), move_backward (0.15m0.15m), turn_left (5∘5^{\circ}), turn_right (5∘5^{\circ}), look_up (5∘5^{\circ}), look_down (5∘5^{\circ}), grab_release, no_op (step physics 50ms50ms), and stop. For the episode to be considered successful, the agent must place the object on top of the receptacle – i.e. the object center should be at a height greater than the receptacle center, and within 0.7m0.7m of the receptacle object center – within 15001500 steps. We picked this 0.7m0.7m threshold distance between the object and receptacle based on pilots on AMT. 0.7m0.7m was sufficiently strict for avoiding false positives in the collected demonstrations where users are able to submit the task without necessarily placing the object on top of the receptacle.

Human Demonstrations (Pick&Place-HD). We collect human demonstrations for Pick&Place on 99 scenes from Matterport3D . In each episode, objects and receptacles are instantiated by randomly sampling from 457457 possible object-receptacle pairs. We initialize the object and receptacle at randomly sampled locations in the environment, and collect 33 demonstrations for each object-receptacle pair. The agent, object, and receptacle locations are randomized across all episodes (including the 33 we collect for each object-receptacle pair). In total, we have 457×3457\times 3 unique object-receptacle-agent position initializations per scene, amounting to 457×3×9=∼12k457\times 3\times 9={\sim}12k demonstrations, which is ∼11.5{\sim}11.5M steps in experience, each episode averaging 932932 steps.

Shortest Path Demonstrations. Similar to ObjectNav, we generate shortest path demonstrations for Pick&Place. These demonstrations are generated by first using the geodesic shortest-path follower to the object, then using a heuristic action planner to face and pick up the object, then following the geodesic shortest-path to the receptacle, and again using a heuristic action planner to drop the object on the receptacle. We generated 25.725.7k shortest path demonstrations for Pick&Place, each averaging 342342 steps, amounting to a total of ∼8.8{\sim}8.8 million steps of experience.

Analysis. Table 3a reports statistics for human and shortest path demonstrations. Similar to ObjectNav, humans have significantly lower SPL, and 22x higher occupancy and sight coverage compared to shortest paths, suggesting the need for exploration. Comparing episode lengths and action histograms (see appendix (Sec. A.1.1) for figure), human demonstrations are longer and make use of all 99 actions. Interestingly, humans often use the move_backward action to backtrack, which the shortest path agents do not use (by design), instead of turning 180∘180^{\circ} and moving forward. This behavior does not appear in ObjectNav shortest path demonstrations because there is just one target object, and so the geodesic shortest path would never involve backtracking or making 180∘180^{\circ} turns.

Imitation Learning from Demonstrations

We use behavior cloning to learn a policy from demonstrations. Let πθ(at ∣ ot)\pi_{\theta}(a_{t}\,|\,o_{t}) denote a policy parametrized by θ\theta that maps observations oto_{t} to a distribution over actions ata_{t}. Let τ\tau denote a trajectory consisting of state, observation, human action tuples: \tau=\big{(}s_{0},o_{0},a_{0},\ldots,s_{T},o_{T},a_{T}\big{)} and \mathcal{T}=\big{\{}\tau^{(i)}\big{\}}_{i=1}^{N} denote a dataset of human demonstrations. The learning problem can be summarized as:

Inflection weighting introduced in Wijmans et al. , adjusts the loss function to upweight timesteps where actions change (i.e. at−1≠ata_{t-1}\neq a_{t}). Specifically, the inflection weighting loss coefficient is computed as total no. of actions in the dataset divided by the total no. of inflection points, and this coefficient is multiplied with the loss at each inflection timestep where at−1≠ata_{t-1}\neq a_{t}. This approach was found to be useful for tasks like navigation with long sequences of the same actions, e.g. several ‘forward’ actions when navigating corridors . We use inflection weighting in all our experiments and found it to help over vanilla behavior cloning.

Our base policy is a simple CNN+RNN architecture. We first embed all sensory inputs using feed-forward modules. For RGB, we use a randomly initialized ResNet18 . For depth, we use a ResNet50 that was pretrained on PointGoal navigation using DD-PPO . Then these RGB and depth features (and optionally other task-specific features) are concatenated and fed into a GRU to predict a distribution over actions at+1a_{t+1}. Task-specific architectural choices over this base policy are described in the next sections.

Fig. 4(a) shows our ObjectNav architecture. Similar to Anand et al. , we feed in RGBD inputs of size 640×480640\times 480 passed through a 2x2-AvgPool layer to reduce the resolution (performing low-pass filtering + downsampling). The agent also has a GPS+Compass sensor, which provides location and orientation relative to start of the episode. GPS+Compass inputs are pass through fully-connected layers to embed them to 3232-d vectors. In addition to RGBD and GPS+Compass, following Ye et al. , we use two additional semantic features – semantic segmentation (SemSeg) of the input RGB and a ‘Semantic Goal Exists’ (SGE) scalar which is the fraction of the visual input occupied by the goal category. These semantic features are computed using a pretrained and frozen RedNet that was pretrained on SUN RGB-D and finetuned on 100k100k randomly sampled front-facing views rendered in the Habitat simulator. Finally, we also feed in the object goal category embedded into a 3232-d vector. All of these input features are concatenated to form an observation embedding, and fed into a 22-layer, 512512-d GRU at every timestep. We train this policy for ∼400{\sim}400M steps (=∼21={\sim}21 epochs on ∼70k{\sim}70k demonstration episodes). We evaluate checkpoints at every ∼15{\sim}15M steps for the last 5050M steps of training, and report metrics for checkpoints with the highest success on the validation split.

2 Pick&Place

Fig. 4(b) shows our Pick&Place architecture. We feed in RGBD inputs of size 256×256256\times 256. In addition to RGBD observations, the policy gets as input language instructions of the form ‘Place the on the ’ encoded using a single-layer LSTM . RGBD and instruction features are concatenated to form an observation embedding, which is fed into a 22-layer, 512512-d GRU at every timestep. We train this policy for ∼90{\sim}90M steps (=∼10={\sim}10 epochs on ∼9.5k{\sim}9.5k demonstration episodes). We evaluate checkpoints at every ∼10{\sim}10M steps during training, and report metrics for checkpoints with the highest success on the validation split.

Experiments & Results

Table 4c reports results on the MP3D val split for several baselines. First, we compare our approach with two state-of-the-art RL approaches from prior work. Maksymets et al. (row 11) train their policy using a reward structure that breaks ObjectNav into two subtasks – exploration and direct navigation to goal object once it is spotted. This agent gets a positive reward for maximizing area coverage until it sees the goal object. It then receives a navigation reward to minimize distance-to-object. This policy achieves 20.0%20.0\% success and 6.5%6.5\% SPL (row 11). then combine this reward structure with Treasure Hunt Data Augmentation (THDA) – inserting arbitrary 33D target objects in the scene to augment the set of training episodes. With THDA, this achieves 28.4%28.4\% success and 11.0%11.0\% SPL (row 22), 7.0%7.0\% worse and 0.8%0.8\% better respectively than behavior cloning on 70k70k human demonstrations (row 99). Ye et al. (row 33) train their policy with a combination of exploration and distance-based navigation rewards, and their representations with several auxiliary tasks (e.g. inverse dynamics and predicting map coverage). This achieves 34.6%34.6\% success and 7.9%7.9\% SPL (row 33), which is 0.8%0.8\% worse on success and 2.3%2.3\% worse on SPL than our approach (row 99). Khandelwal et al. (row 44) train a policy using CLIP as a visual backbone with simple distance-based navigation rewards. This achieves 21.6%21.6\% success and 8.7%8.7\% SPL (row 44), which is 13.8%13.8\% worse on success and 1.5%1.5\% worse on SPL than our approach (row 99). IL on a dataset of shortest paths achieves 4.4%4.4\% success and 2.2%2.2\% SPL (row 55), significantly worse than training on 35k35k human demonstrations (31.6%31.6\% success, 8.5%8.5\% SPL). Recall that comparison of shortest path demonstrations was done with a subset of 35k35k ObjectNav-HD demonstrations that were collected in the first phase of the project. Next, we also collected 10k10k human demonstrations on ObjectNav episodes generated in THDA fashion – i.e. asking humans to find randomly inserted objects. Notice that this involves pure exhaustive search, since there are no semantic priors that humans can leverage in this setting. An IL agent trained on 10k10k THDA demonstrations combined with the original 40k40k demonstrations achieves 33.2%33.2\% success and 9.5%9.5\% SPL (row 88) which is 0.8%0.8\% better on success and 0.4%0.4\% better on SPL than 50k50k non-THDA demonstrations (row 77), i.e. adding these THDA demonstrations with exhaustive search behavior helps. We also collected 10k10k demonstrations on Gibson ObjectNav episodes to compare effect of different scene datasets. An agent trained on 10k10k Gibson demonstrations combined with 60k60k MP3D demonstrations achieves 33.9%33.9\% success and 9.7%9.7\% SPL (row 1111), which is 1.5%1.5\% worse on success and 0.5%0.5\% worse on SPL compared to when we use MP3D-only demonstrations (row 99).

Finally, we also benchmark human performance on the MP3D val split – 93.7%93.7\% success, 42.5%42.5\% SPL (row 1212).

ObjectNav Sensor Ablations. Table 1 reports results on the MP3D val split for various ablations of our approach trained on 35k35k human demonstrations. First, without any visual input (row 11), i.e. no RGBD and semantic inputs, the agent fails to learn anything (0%0\% success, 0%0\% SPL). Second, without SemSeg and SGE features (and keeping only RGB and Depth features) to the policy, performance drops by 8.9%8.9\% success and 2.4%2.4\% SPL (row 22 vs. 33).

Habitat ObjectNav Challenge Results. Table 2 compares our results with prior approaches from the 20202020 and 20212021 Habitat Challenge leaderboards. Our approach (IL w/ 70k70k demonstrations) achieves 27.8%27.8\% success and 9.9%9.9\% SPL (row 88), outperforming prior RL-trained counterparts – 3.3%3.3\% better success, 3.5%3.5\% better SPL than Red Rabbit (6-Act Base) (row 5), and 6.7%6.7\% better success, 1.1%1.1\% better SPL than ExploreTillSeen + THDA (row 7).

Performance vs. Dataset size. To investigate scaling behavior, we plot val success against the size of the human demonstrations dataset in Fig. 1b. We created splits of the human demonstrations’ dataset of increasing sizes, from 4k4k to 70k70k, and trained models with the same set of hyperparameters on each split. All hyperparameters were picked early in the course of the data collection (on the 4k4k and 12k12k subsplits) and fixed for later experiments. So val performance in the small-data regime may be an optimistic estimate and in the large data regime a pessimistic estimate. True scaling behavior may be even stronger. Increasing dataset size consistently improves performance and has not yet saturated, suggesting that simply collecting more demonstrations is likely to lead to further gains.

Sample Efficiency. Fig. 5 plots val success against no. of training steps of experience (in millions) in Fig. 5(a) and against unique steps of experience in Fig. 5(b). Recall that IL involves ∼21{\sim}21 epochs on a static dataset of ∼70k{\sim}70k demos, while RL (from ) gathers unique agent-driven trajectories on-the-fly. Fig. 5(a) shows that IL behaves like supervised learning (as expected) with improvements coming from long training schedules; unfortunately, this means that wall-clock training times are not lower than RL. Fig. 5(b) shows that IL requires 77x fewer unique steps of experience to outperform success and is thus much more sample-efficient.

Zero-shot results on Gibson are in Section A.2.

2 Pick&Place

Results. We report results in Table 3 across three evaluation splits. 1) New Initializations: new locations of objects and receptacles. This tests generalization to unseen locations in seen environments. 2) New Instructions: compositionally novel object-receptacle combinations of objects and receptacles individually seen during training. 3) New Environments: generalization to 22 scenes held out from training. Similar to ObjectNav and as described in Section 4, we also report results with shortest paths. Again, these paths are significantly shorter (average 342342 vs. 932932 steps per demonstration) and hence, we generate a larger dataset of 25.7k25.7k episodes roughly matching the cumulative steps of experience with human demonstrations (8.8M8.8M shortest path steps vs. 11.5M11.5M human steps). Training on 9.5k9.5k human demonstrations achieves 17.5%17.5\% success, 9.8%9.8\% SPL on new object-receptacle initializations (row 22). Across splits, training on shortest paths hurts success by 88-16%16\%. Going to new object-receptacle pairs, success drops by 2.4%2.4\% (row 55 vs. 22), and then going to new environments further hurts success by 6.8%6.8\% (row 88 vs. 55). We also trained an RL policy with the exploration and distance-based rewards from , but it failed to get beyond 0%0\% success on new object-receptacle intializations. See the appendix (Sec. A.1.2) for training details.

Performance vs. Dataset size. Similar to ObjectNav, we trained policies on 2.5k2.5k to 9.5k9.5k subsets of our Pick&Place data, and found that performance continues to improve with more data. Figure in appendix (Sec. A.1.3).

Characterizing Learned Behaviors

To characterize the behaviors learnt by our best IL agents, we first sample 300300 validation ObjectNav episodes for each method and manually categorize the behavior observed. A subset of observed behaviors are visualized in Fig. 6. Our agents demonstrate sophisticated object-search behaviors e.g. peeking into rooms to maximize sight coverage (SC), instead of occupancy coverage (OC), checking corners of rooms for small objects, beelining to goal object once seen, exhaustive search (ES), turning in place to get a panoramic view (PT), and looping back to recheck some areas. Amusingly, unlike shortest path / RL agents, these IL agents also stand idle and ‘look around’ i.e. turn in place, like humans. Table 4 quantifies these behaviors. See appendix (Sec. A.4) for details on how these were computed. Agents trained with IL on human demonstrations have higher coverage (both occupancy and sight), peeking behavior, panoramic turns, beelines, and exhaustive search than RL. RL-trained agents achieve higher average Goal Room Time Spent (GRTS) – i.e. time spent in the room containing the target object – but also have significantly higher variance in GRTS across scenes compared to IL agents. See appendix (Sec. A.4) for a per-scene breakdown of GRTS as well as histograms of time spent in each room (instead of just target room) when searching for a target object. We also discuss limitations of our approach in the appendix Sec. A.7.

Conclusion

We developed the infrastructure to collect human demonstrations at scale and using this, trained imitation learning (IL) agents on 92k+92k+ human demonstrations for ObjectNav and Pick&Place. On ObjectNav, we found that IL using 70k70k human demonstrations outperforms RL using 240k240k agent-gathered trajectories, and on Pick&Place, IL agents get to ∼18%{\sim}18\% success while RL fails to get beyond 0%0\%. Qualitatively, we found that IL agents pick up on sophisticated object-search behavior implicitly captured in human demonstrations, much more prominently than RL agents. Overall, we believe our work makes a compelling case for investing in large-scale imitation learning of human demonstrations.

Acknowledgements. We thank Devi Parikh for help with brainstorming and direction, Joel Ye for answering questions about his Red Rabbit , Oleksandr Maksymets for the THDA episode generation pipeline, and Erik Wijmans for help with debugging DDP. The Georgia Tech effort was supported in part by NSF, ONR YIP, and ARO PECASE. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government, or any sponsor.

References

Appendix A Appendix

Recall that in the pick-and-place task (Pick&Place), an agent must follow an instruction of the form ‘Place the on the ’, without being told the location of the or in a new environment. The agent must explore and navigate to the object, pick it up, explore and navigate to the receptacle, and place the previously picked-up object on it. In this section, we go over statistics of the human demonstrations dataset, how our Pick&Place imitation learning (IL) agents scale as a function of training dataset size, and details of our reinforcement learning baseline for Pick&Place.

Fig. 8 compares the episode length and action histograms for human and shortest path demonstrations for Pick&Place. Human demonstrations are longer (average 932932 vs 342342 steps per demonstration) and have a more uniform action distribution compared to shortest paths. Human demonstrations also make use of all 99 actions whereas shortest path demonstrations use only 66 actions. Notice, humans also tend to stand idle and do nothing (5050ms of idle time is translated to a NO_OP action). They likely use this time to strategize their next set of actions to explore the environment, which is not the case in shortest path demonstrations (by design).

A.1.2 RL Baseline

Similar to the imitation learning baseline, our base policy is a simple CNN+RNN architecture. We first embed all sensory inputs using feed-forward modules. For RGB, we use a randomly initialized ResNet18 . For depth, we use a ResNet50 that was pretrained on PointGoal navigation using DDPPO . In addition to RGBD observations, the policy gets as input language instructions of the form ‘Place the on the ’ encoded using a single-layer LSTM . RGBD and instruction features are concatenated to form an observation embedding, which is fed into a 22-layer, 512512-d GRU at every timestep. We train this policy for ∼100{\sim}100M steps on ∼9.5k{\sim}9.5k episodes.

Rewards. The agent receives a sparse success reward rsuccessr_{success}, a slack reward rslackr_{slack} to motivate faster goal-seeking, an exploration reward rexplorer_{explore}, an object seen reward rseenr_{seen}, a grab/release success reward rgrab_releaser_{grab\_release}, and a drop penalty reward rdrop_penaltyr_{drop\_penalty} to penalize dropping the object far from the receptacle. For incentivizing exploration, we use a visitation-based coverage reward from Ye et al. . We first divide the map into a voxel grid of 2.5m×2.5m×2.5m2.5m\times 2.5m\times 2.5m voxels and reward the agent for visiting each voxel. Similar to , we smooth rexplorer_{explore} by decaying it by number of steps the agent has spent in the voxel (visit count vv). To ensure that the agent prioritizes Pick&Place (and not just exploration), we decay rexplorer_{explore} based on episode timestep tt with a decay constant of d=0.995d=0.995. The agent is provided a reward for exploration until it sees the object. Once it sees the object, it receives a significant positive reward rseenr_{seen}, and then the reward switches to a path-efficiency based navigation reward. In addition, the agent also receives a significant positive reward when it successfully grabs a object or releases the object close to the receptacle.

Results. A policy trained with this reward for 100100M steps fails to get beyond 0%0\% success on the Pick&Place task. The agent learns to pick up the object at the start of training if it sees the object while navigating but it fails to search for the receptacle and place the object on top of receptacle. Overall, throughout training, the agent doesn’t solve the task successfully even once demonstrating the difficulty of the task and inadequacy of the above reward structure.

A.1.3 Performance vs. Dataset Size

Fig. 7 plots Val success of our IL agent vs. the size of the Pick&Place human demonstrations dataset. We trained policies on 2.5k2.5k to 9.5k9.5k subsets of the data. Performance continues to improve with more data and has not saturated.

A.2 Zero-shot ObjectNav results on Gibson

To test generalization of the IL agents trained on human demonstrations, we report zero-shot results by transferring our policy trained on 40k40k human demonstrations to the Gibson dataset val split in Table 5. To enable zero-shot transfer of semantic features, we remap the common goal categories (chair, couch, potted plant, bed, toilet, TV, dining-table) from Matterport3D to Gibson goal category IDs. Our IL agent achieves 49.4%49.4\% success and 16.4%16.4\% SPL (row 7) with no finetuning on Gibson dataset. Comparing our zero-shot results to approaches trained on Gibson, our IL agent is 33.6%33.6\% better on success and 11.5%11.5\% better on SPL than an RL baseline that takes RGBD + Semantics as input (row 3 vs. row 7). Next, we compare our approach with SemExp which builds explicit semantic maps and learns a goal-oriented semantic exploration policy which learns semantic priors for efficient navigation. Our approach is 4.9%4.9\% worse on success and 3.5%3.5\% worse on SPL compared to SemExp (row 6 vs. row 7). uses a modular framework for ObjectNav by using a potential function conditioned on a semantic map which is used to decide where to look for unseen object in an environment. Our approach is 24.1%24.1\% worse on success and 24.6%24.6\% worse on SPL.

A.3 Estimating time using a LoCoBot motion model

To estimate the time a robot would take to execute the collected human trajectories in the real world, we use the LoCoBot motion model from Krantz et al. . This model consists of a rotation function that maps turn angle to time and a translation function that maps straight-line distance to time. For estimating time required for grab/release actions, we replace them with 0.15m0.15m forward steps and use the straight-line distance translation function. We use the MoveBase controller from for all our time estimates, with the following rotation and translation equations:

A.4 Characterizing Learnt Behaviors

In this section, we describe the metrics used to characterize the exploration behavior exhibited by these agents in Sec. 7 in the main paper. These include 1) Occupancy Coverage (OC) and 2) Sight Coverage (SC) introduced in Sec. 4 in the main paper, as well as 3) Goal Room Time Spent (GRTS) – the number of steps as a fraction of total episode length an agent takes within the room bounding box containing the target object, 4) Peeks – check if the agent steps back into the last visited room after taking just ∼10{\sim}10 steps in another room, 5) Panoramic Turn (PT) – whether the agent stands at one place and turns left and right to get sweeping views, 6) Beeline – if the agent takes 1010 continuous forward actions before reaching the goal in the last 1515 steps, 7) Exhaustive Search (ES) – ≥75%\geq 75\% sight coverage. To compute these metrics, we use the semantic annotations in Matterport3D. These annotations provide 33D bounding box coordinates for each room category in an environment. We use these bounding box coordinates to track the rooms an agent visits during an episode. GRTS gives us a measure of how often the agent ends up reaching goal object room but doesn’t successfully locate the object. A higher GTRS suggests that the agent is at least good at reaching semantically meaningful locations in search of the goal object. We find that RL agents have higher average GRTS but also significantly higher variance in GRTS across scenes while our IL agents have lower average GRTS but more consistently spend time in the target room (see Fig. 9). To evaluate not just the final room the agent ends up at, but all the rooms it visits through the course of an episode, we also plot distributions of the time spent per room category for each goal object (see Fig. 15) for human demonstrations vs. IL agents trained on human demonstrations vs. RL agents.

A.5 Inter-human Variance in ObjectNav

To get a sense for the variance in ObjectNav human demonstrations, we collected 2020 unique human-provided trajectories for the same initial location and target object (‘cabinet’). This is visualized in Fig. 10. We see that there is quite a bit of diversity in navigation trajectories across humans. They often navigate to different instances of the goal object category ‘cabinet’, and even when multiple humans go to the same object instance, the routes taken are different (red vs. blue trajectory).

We also plot the average SPL per AMT user in our dataset in Fig. 11. We find that human performance has a lot of variability, ranging from 25.2%25.2\% to 68.2%68.2\% (Fig. 11(a)). The SPL range that has the most humans is ∼50%{\sim}50\%. The best-performing human annotator achieves an SPL of 68.2%68.2\% averaged over 66 episodes (Fig. 11(b)), which is particularly close to shortest paths and arguably super-human.

A.6 AMT Interface

Fig. 12 shows a screenshot of our AMT interface for collecting Pick&Place demonstrations. For the Pick&Place task, we provide humans with an instruction of the form ‘Place the on the ’, without being told the location of the or in a new environment, and they can see agent’s first-person view of the environment. They can make the agent move, look around, and interact with the environment using keyboard controls. Once the AMT user completes the task they can submit the task by clicking the ‘Submit’ button. We then run task-specific validation checks to ensure only successful tasks get submitted.

Validation. To ensure data quality, every submitted AMT task goes through a set of validation checks. For ObjectNav, we use the same set of validation checks as the Habitat challenge evaluation setup, i.e. a task is considered successful only when the user has moved the agent to within 1m1m of the goal object. We do not limit the maximum number of steps to allow users on AMT to explore the environment. This captures key human exploration behavior necessary to succeed at these tasks. Similarly, for Pick&Place, a task is considered successful when the target object is placed on a receptacle object. Specifically, we check if the Euclidean distance between the centers of the target and receptable objects is less than 0.7m0.7m, and that the target object is at a height greater than the receptacle center.

A.7 Limitations

Our approach is fundamentally limited by the limitations of imitation learning as our approach uses vanilla behavior cloning with inflection weighting. Additionally, these agents trained on human demonstrations exhibit some common failure cases. Some examples of common failure cases are – reaching close to the goal object but not within goal radius and ending episode early, trying to move straight when agent is colliding and getting stuck, looping around multiple instances of the goal object and as a result, exceeding maximum episode steps, and exploring the environment and not finding the goal object. Our approach is also limited by the amount of human demonstrations we can gather and the agent architecture being trained on this dataset. Currently, we use a vanilla CNN+RNN architecture to learn imitation learning policies but we can build better architecture which make full use of the rich semantic information these human demonstrations have.