PIRLNav: Pretraining with Imitation and RL Finetuning for ObjectNav
Ram Ramrakhya, Dhruv Batra, Erik Wijmans, Abhishek Das
Introduction
Since the seminal work of Winograd winograd_cogpsy72, designing embodied agents that have a rich understanding of the environment they are situated in, can interact with humans (and other agents) via language, and the environment via actions has been a long-term goal in AI smith_al05; hermann_arxiv17; hill_arxiv17; chaplot_aaai18; anderson_cvpr18; jain2019two; das_phd_thesis_2020; abramson2020imitating; weihs2021learning; lynch2022interactive. We focus on ObjectGoal Navigation anderson_arxiv18; objectnav_tech_report, wherein an agent situated in a new environment is asked to navigate to any instance of an object category (‘find a plant’, ‘find a bed’, etc.); see Fig. 2. ObjectNav is simple to explain but difficult for today’s techniques to accomplish. First, the agent needs to be able to ground the tokens in the language instruction to physical objects in the environment (e.g. what does a ‘plant’ look like?). Second, the agent needs to have rich semantic priors to guide its navigation to avoid wasteful exploration (e.g. the microwave is likely to be found in the kitchen, not the washroom). Finally, it has to keep track of where it has been in its internal memory to avoid redundant search.
Humans are adept at ObjectNav. Prior work ramrakhya2022 collected a large-scale dataset of human demonstrations for ObjectNav, where human subjects on Mechanical Turk teleoperated virtual robots and searched for objects in novel houses. This first provided a human baseline on ObjectNav of success rate on the Matterport3D (MP3D) dataset mp3d On val split, for 21 object categories, and a maximum of 500 steps. compared to success rate of the best performing method ramrakhya2022. This dataset was then used to train agents via imitation learning (specifically, behavior cloning).
While this approach achieved state-of-art results ( success rate on MP3D val dataset), it has two clear limitations. First, behavior cloning (BC) is known to suffer from poor generalization to out-of-distribution states not seen during training, since the training emphasizes imitating actions not accomplishing their goals. Second and more importantly, it is expensive and thus not scalable. Specifically, Ramrakhya et al. ramrakhya2022 collected demonstrations on scenes in Matterport3D Dataset, which took hours of human teleoperation and $ dollars. A few months after ramrakhya2022 was released, a new higher-quality dataset called HM3D-Semantics v0.1 yadav2022habitat became available with annotated 3D scenes, and a few months after that HM3D-Semantics v0.2 added additional scenes. Scaling Ramrakhya et al.’s approach to continuously incorporate new scenes involves replicating that entire effort again and again.
On the other hand, training with reinforcement learning (RL) is trivially scalable once annotated 3D scans are available. However, as demonstrated in Maksymets et al. thda_iccv21, RL requires careful reward engineering, the reward function typically used for ObjectNav actually penalizes exploration (even though the task requires it), and the existing RL policies overfit to the small number of available environments.
Our primary technical contribution is PIRLNav, an approach for pretraining with BC and finetuning with RL for ObjectNav. BC pretrained policies provide a reasonable starting point for ‘bootstrapping’ RL and make the optimization easier than learning from scratch. In fact, we show that BC pretraining even unlocks RL with sparse rewards. Sparse rewards are simple (do not involve any reward engineering) and do not suffer from the unintended consequences described above. However, learning from scratch with sparse rewards is typically out of reach since most random action trajectories result in no positive rewards.
While combining IL and RL has been studied in prior work schaal_neurips96; eqa_modular; rajeswaran_RSS_18; vpt22; lynch2019play, the main technical challenge in the context of modern neural networks is that imitation pretraining results in weights for the policy (or actor), but not a value function (or critic). Thus, naively initializing a new RL policy with these BC-pretrained policy weights often leads to catastrophic failures due to destructive policy updates early on during RL training, especially for actor-critic RL methods jsrl. To overcome this challenge, we present a two-stage learning scheme involving a critic-only learning phase first that gradually transitions over to training both the actor and critic. We also identify a set of practical recommendations for this recipe to be applied to ObjectNav. This leads to a PIRLNav policy that advances the state-the-art on ObjectNav from success rate (in chaplot_neurips20) to (, relative improvement).
Next, using this BCRL training recipe, we conduct an empirical analysis of design choices. Specifically, an ingredient we investigate is whether human demonstrations can be replaced with ‘free’ (automatically generated) sources of demonstrations for ObjectNav, e.g. (1) shortest paths (SP) between the agent’s start location and the closest object instance, or (2) task-agnostic frontier exploration frontier (FE) of the environment followed by shortest path to goal-object upon observing it. We ask and answer the following:
‘Do human demonstrations capture any unique ObjectNav-specific behaviors that shortest paths and frontier exploration trajectories do not?’ Yes. We find that BC / BCRL on human demonstrations outperforms BC / BCRL on shortest paths and frontier exploration trajectories respectively. When we control the number of demonstrations from each source such that BC success on train is the same, RL-finetuning when initialized from BC on human demonstrations still outperforms the other two.
‘How does performance after RL scale with BC dataset size?’ We observe diminishing returns from RL-finetuning as we scale BC dataset size. This suggests, by effectively leveraging the trade-off curve between size of pretraining dataset size vs. performance after RL-Finetuning, we can achieve closer to state-of-the-art results without investing into a large dataset of BC demonstrations.
‘Does BC on frontier exploration demonstrations present similar scaling behavior as BC on human demonstrations?’ No. We find that as we scale frontier exploration demonstrations past trajectories, the performance plateaus.
Finally, we present an analysis of the failure modes of our ObjectNav policies and present a set of guidelines for further improving them. Our policy’s primary failure modes are: a) Dataset issues: comprising of missing goal annotations, and navigation meshes blocking the path, b) Navigation errors: primarily failure to navigate between floors, c) Recognition failures: where the agent does not identify the goal object during an episode, or confuses the specified goal with a semantically-similar object.
Related Work
ObjectGoal Navigation. Prior works on ObjectNav have used end-to-end RL mousavian2018sem_midlevel; ye_iccv21; thda_iccv21, modular learning chaplot_neurips20; liang2020sscnav; ramakrishnan2022poni, and imitation learning ramrakhya2022; ovrl. Works that use end-to-end RL have proposed improved visual representations mousavian2018sem_midlevel; wei2019, auxiliary tasks ye_iccv21, and data augmentation techniques thda_iccv21 to improve generalization to unseen environments. Improved visual representations include object relation graphs wei2019 and semantic segmentations mousavian2018sem_midlevel. Ye et al. ye_iccv21 use auxiliary tasks like predicting environment dynamics, action distributions, and map coverage in addition to ObjectNav and achieve promising results. Maksymets et al. thda_iccv21 improve generalization of RL agents by training with artificially inserted objects and proposing a reward to incentivize exploration.
Modular learning methods for ObjectNav have also emerged as a strong competitor chaplot_neurips20; liang2020sscnav; ramakrishnan_arxiv20. These methods rely on separate modules for semantic mapping that build explicit structured map representations, a high-level semantic exploration module that is learned through RL to solve the ‘where to look?’ subproblem, and a low-level navigation policy that solves ‘how to navigate to ?’.
The current state-of-the-art methods on ObjectNav ramrakhya2022; ovrl make use of BC on a large dataset of human demonstrations. with a simple CNN+RNN policy architecture. In this work, we improve on them by developing an effective approach to finetune these imitation-pretrained policies with RL.
Imitation Learning and RL Finetuning. Prior works have considered a special case of learning from demonstration data. These approaches initialize policies trained using behavior cloning, and then fine-tune using on-policy reinforcement learning schaal_neurips96; rajeswaran_RSS_18; vpt22; lynch2019play; jan2008; jens2008, On classical tasks like cart-pole swing-up schaal_neurips96, balance, hitting a baseball jan2008, and underactuated swing-up jens2008, demonstrations have been used to speed up learning by initializing policies pretrained on demonstrations for RL. Similar to these methods, we also use a on-policy RL algorithm for finetuning the policy trained with behavior cloning. Rajeswaran et al. rajeswaran_RSS_18 (DAPG) pretrain a policy using behavior cloning and use an augmented RL finetuning objective to stay close to the demonstrations which helps reduce sample complexity. Unfortunately DAPG is not feasible in our setting as it requires solving a systems research problem to efficiently incorporate replaying demonstrations and collecting experience online at our scale. rajeswaran_RSS_18 show results of the approach on a dexterous hand manipulation task with a small number of demonstrations that can be loaded in system memory and therefore did not need to solve this system challenge. This is not possible in our setting, just the 256256 RGB observations for the demos we collect would occupy over 2 TB memory, which is out of reach for all but the most exotic of today’s systems. There are many methods for incorporating demonstrations/imitation learning with off-policy RL awac; awopt2021corl; kalashnikov2018scalable; AWRPeng19; marwil. Unfortunately these methods were not designed to work with recurrent policies and adapting off-policy methods to work with recurrent policies is challenging r2d2_iclr19. See the Appendix A for more details. The RL finetuning approach that demonstrates results with an actor-critic and high-dimensional visual observations, and is thus most closely related to our setup is proposed in VPT vpt22. Their approach uses Phasic Policy Gradients (PPG) cobbe2020ppg with a KL-divergence loss between the current policy and the frozen pretrained policy, and decays the KL loss weight over time to enable exploration during RL finetuning. Our approach uses Proximal Policy Gradients (PPO) schulman_arxiv17 instead of PPG, and therefore does not require a KL constraint, which is compute-expensive, and performs better on ObjectNav.
ObjectNav and Imitation Learning
In ObjectNav an agent is tasked with searching for an instance of the specified object category (e.g., ‘bed’) in an unseen environment. The agent must perform this task using only egocentric perceptions. Specifically, a RGB camera, Depth sensor We don’t use this sensor as we don’t find it helpful., and a GPS+Compass sensor that provides location and orientation relative to the start position of the episode. The action space is discrete and consists of move_forward (), turn_left (), turn_right (), look_up (), look_down (), and stop actions. An episode is considered successful if the agent stops within Euclidean distance of the goal object within steps and is able to view the object by taking turn actions objectnav_tech_report.
We use scenes from the HM3D-Semantics v0.1 dataset yadav2022habitat. The dataset consists of scenes and unique goal object categories. We evaluate our agent using the train/val/test splits from the 2022 Habitat Challenge https://aihabitat.org/challenge/2022/.
2 ObjectNav Demonstrations
Ramrakhya et al. ramrakhya2022 collected ObjectNav demonstrations for the Matterport3D dataset mp3d. We begin our study by replicating this effort and collect demonstrations for the HM3D-Semantics v0.1 dataset yadav2022habitat. We use Ramrakhya et al.’s Habitat-WebGL infrastructure to collect demonstrations, amounting to human annotation hours.
3 Imitation Learning from Demonstrations
We use behavior cloning to pretrain our ObjectNav policy on the human demonstrations we collect. Let denote a policy parametrized by that maps observations to a distribution over actions . Let denote a trajectory consisting of state, observation, action tuples: and denote a dataset of human demonstrations. The optimal parameters are
We use inflection weighting eqa_matterport to adjust the loss function to upweight timesteps where actions change (i.e. ).
Our ObjectNav policy architecture is a simple CNN+RNN model from ovrl. To encode RGB input CNN, we use a ResNet50 he_cvpr16. Following ovrl, the CNN is first pre-trained on the Omnidata starter dataset eftekhar2021omnidata using the self-supervised pretraining method DINO caron2021emerging and then finetuned during ObjectNav training. The GPS+Compass inputs, , and , are passed through fully-connected layers FC FC to embed them to 32-d vectors. Finally, we convert the object goal category to one-hot and pass it through a fully-connected layer FC, resulting in a 32-d vector. All of these input features are concatenated to form an observation embedding, and fed into a 2-layer, 2048-d GRU at every timestep to predict a distribution over actions - formally, given current observations , GRU. To reduce overfitting, we apply color-jitter and random shifts yarats2021mastering to the RGB inputs.
RL Finetuning
Our motivation for RL-finetuning is two-fold. First, finetuning may allow for higher performance as behavior cloning is known to suffer from a train/test mismatch – when training, the policy sees the result of taking ground-truth actions, while at test-time, it must contend with the consequences of its own actions. Second, collecting more human demonstrations on new scenes or simply to improve performance is time-consuming and expensive. On the other hand, RL-finetuning is trivially scalable (once annotated 3D scans are available) and has the potential to reduce the amount of human demonstrations needed.
The RL objective is to find a policy that maximizes expected sum of discounted future rewards. Let be a sequence of object, action, reward tuples (, , ) where is the action sampled from the agent’s policy, and is the reward. For a discount factor , the optimal policy is
To solve this maximization problem, actor-critic RL methods learn a state-value function (also called a critic) in addition to the policy (also called an actor). The critic represents the expected value of returns when starting from state and acting under the policy , where returns are defined as . We use DD-PPO wijmans_iclr20, a distributed implementation of PPO schulman_arxiv17, an on-policy RL algorithm. Given a -parameterized policy and a set of rollouts, PPO updates the policy as follows. Let , be the advantage estimate and be the ratio of the probability of action under current policy and under the policy used to collect rollouts. The parameters are updated by maximizing:
We use a sparse success reward. Sparse success is simple (does not require hyperparameter optimization) and has fewer unintended consequences (e.g. Maksymets et al. thda_iccv21 showed that typical dense rewards used in ObjectNav actually penalize exploration, even though exploration is necessary for ObjectNav in new environments). Sparse rewards are desirable but typically difficult to use with RL (when initializing training from scratch) because they result in nearly all trajectories achieving reward, making it difficult to learn. However, since we pretrain with BC, we do not observe any such pathologies.
2 Finetuning Methodology
We use the behavior cloned policy weights to initialize the actor parameters. However, notice that during behavior cloning we do not learn a critic nor is it easy to do so -- a critic learned on human demonstrations (during behavior cloning) would be overly optimistic since all it sees are successes. Thus, we must learn the critic from scratch during RL. Naively finetuning the actor with a randomly-initialized critic leads to a rapid drop in performance After the initial drop, the performance increases but the improvements on success are small. (see Fig. 8) since the critic provides poor value estimates which influence the actor’s gradient updates (see Eq.(3)). We address this issue by using a two-phase training regime:
Phase 1: Critic Learning. In the first phase, we rollout trajectories using the frozen policy, pre-trained using BC, and use them to learn a critic. To ensure consistency of rollouts collected for critic learning with RL training, we sample actions (as opposed to using argmax actions) from the pre-trained BC policy: . We train the critic until its loss plateaus. In our experiments, we found steps to be sufficient. In addition, we also initialize the weights of the critic’s final linear layer close to zero to stabilize training.
Phase 2: Interactive Learning. In the second phase, we unfreeze the actor RNN The CNN and non-visual observation embedding layers remain frozen. We find this to be more stable. and finetune both actor and critic weights. We find that naively switching from phase 1 to phase 2 leads to small improvements in policy performance at convergence. We gradually decay the critic learning rate from to while warming-up the policy learning rate from to between to steps, and then keeping both at through the course of training. See Fig. 3. We find that using this learning rate schedule helps improve policy performance. For parameters that are shared between the actor and critic (i.e. the RNN), we use the lower of the two learning rates (i.e. always the actor’s in our schedule). To summarize our finetuning methodology:
First, we initialize the weights of the policy network with the IL-pretrained policy and initialize critic weights close to zero. We freeze the actor and shared weights. The only learnable parameters are in the critic.
Next, we learn the critic weights on rollouts collected from the pretrained, frozen policy.
After training the critic, we warmup the policy learning rate and decay the critic learning rate.
Once both critic and policy learning rate reach a fixed learning rate, we train the policy to convergence.
3 Results
Comparing with the RL-finetuning approach in VPT vpt22. We start by comparing our proposed RL-finetuning approach with the approach used in VPT vpt22. Specifically, vpt22 proposed initializing the critic weights to zero, replacing entropy term with a KL-divergence loss between the frozen IL policy and the RL policy, and decay the KL divergence loss coefficient, , by a fixed factor after every iteration. Notice that this prevents the actor from drifting too far too quickly from the IL policy, but does not solve uninitialized critic problem. To ensure fair comparison, we implement this method within our DD-PPO framework to ensure that any performance difference is due to the fine-tuning algorithm and not tangential implementation differences. Complete training details are in the Section C.3. We keep hyperparameters constant for our approach for all experiments. Table 1 reports results on HM3D val for the two approaches using human demonstrations. We find that PIRLNav achieves Success compared to VPT and comparable SPL.
Ablations. Next, we conduct ablation experiments to quantify the importance of each phase in our RL-finetuning approach. Table 2 reports results on the HM3D val split for a policy BC-pretrained on human demonstrations and RL-finetuned for steps, complete training details are in Section C.4. First, without a gradual learning transition (row ), i.e. without a critic learning and LR decay phase, the policy improves by on success and on SPL. Next, with only a critic learning phase (row ), the policy improves by on success and on SPL. Using an LR decay schedule only for the critic after the critic learning phase improves success by and SPL by , and using an LR warmup schedule for the actor (but no critic LR decay) after the critic learning phase improves success by and SPL by . Finally, combining everything (critic-only learning, critic LR decay, actor LR warmup), our policy improves by on success and on SPL.
ObjectNav Challenge 2022 Results. Using our overall two-stage training approach of BC-pretraining followed by RL-finetuning, we achieve state-of-the-art results on ObjectNav– success and SPL on both the test-standard and test-challenge splits and success and SPL on val. Table 3 compares our results with the top-4 entries to the Habitat ObjectNav Challenge 2022 habitat_challenge2022. Our approach outperforms Stretch chaplot_neurips20 on success rate on both test-standard and test-challenge and is comparable on SPL ( worse on test-standard, better on test-challenge). ProcTHOR procthor, which uses procedurally-generated environments for training, achieves success and SPL on test-standard split, which is worse at success and worse at SPL than ours. For sake of completeness, we also report results of two unpublished entries uploaded to the leaderboard – Populus A. and ByteBOT. Unfortunately, there is no associated report yet with these entries, so we are unable to comment on the details of these approaches, or even whether the comparison is meaningful.
Role of demonstrations in BC→\rightarrowRL transfer
Our decision to use human demonstrations for BC-pretraining before RL-finetuning was motivated by results in prior work ramrakhya2022. Next, we examine if other cheaper sources of demonstrations lead to equally good BCRL generalization. Specifically, we consider sources of demonstrations:
Shortest paths (SP). These demonstrations are generated by greedily sampling actions to fit the geodesic shortest path to the nearest navigable goal object, computed using the ground-truth map of the environment. These demonstrations do not capture any exploration, they only capture success at the ObjectNav task via the most efficient path. Task-Agnostic Frontier Exploration (FE) chaplot_neurips20. These are generated by using a 2-stage approach: 1) Exploration: where a task-agnostic strategy is used to maximize exploration coverage and build a top-down semantic map of the environment, and 2) Goal navigation: once the goal object is detected by the semantic predictor, the developed map is used to reach it by following the shortest path. These demonstrations capture ObjectNav-agnostic exploration.
Human Demonstrations (HD) ramrakhya2022. These are collected by asking humans on Mechanical Turk to control an agent and navigate to the goal object. Humans are provided access to the first-person RGB view of the agent and tasked to reach within m of the goal object category. These demonstrations capture human-like ObjectNav-specific exploration.
Using the BC setup described in Sec. 3.3, we train on SP, FE, and HD demonstrations. Since these demonstrations vary in trajectory length (e.g. SP are significantly shorter than FE), we collect steps of experience with each method. That amounts to SP, FE, and HD demonstrations respectively. As shown in Table 4, BC on SP demonstrations leads to success and SPL. We believe this poor performance is due to an imitation gap advisor, i.e. the shortest path demonstrations are generated with access to privileged information (ground-truth map of the environment) which is not available to the policy during training. Without a map, following the shortest path in a new environment to find a goal object is not possible. BC on FE demonstrations achieves success and SPL, which is significantly better than BC on shortest paths ( success, SPL). Finally, BC on HD obtains the best results – success, SPL. These trends suggest that task-specific exploration (captured in human demonstrations) leads to much better generalization than task-agnostic exploration (FE) or shortest paths (SP).
2 Results with RL Finetuning
Using the BC-pretrained policies on SP, FE, and HD demonstrations as initialization, we RL-finetune each using our approach described in Sec. 4. These results are summarized in Fig. 4. Perhaps intuitively, the trends after RL-finetuning follow the same ordering as BC-pretraining, i.e. RL-finetuning from BC on HD FE SP. But there are two factors that could be leading to this ordering after RL-finetuning – 1) inconsistency in performance at initialization (i.e. BC on HD is already better than BC on FE), and 2) amenability of each of these initializations to RL-finetuning (i.e. is RL-finetuning from HD init better than FE init?).
We are interested in answering (2), and so we control for (1) by selecting BC-pretrained policy weights across SP, FE, and HD that have equal performance on a subset of train success. This essentially amounts to selecting BC-pretraining checkpoints for FE and HD from earlier in training as success is the maximum for SP.
Fig. 5 shows the results after BC and RL-finetuning on a subset of the HM3D train and on HM3D val. First, note that at BC-pretraining train success rates are equal (), while on val FE is slightly better than HD followed by SP. We find that after RL-finetuning, the policy trained on HD still leads to higher val success () compared to FE () and SP (). Notice that RL-finetuning from SP leads to high train success, but low val success, indicating significant overfitting. FE has smaller train-val gap after RL-finetuning but both are worse than HD, indicating underfitting. These results show that learning to imitate human demonstrations equips the agent with navigation strategies that enable better RL-finetuning generalization compared to imitating other kinds of demonstrations, even when controlled for the same BC-pretraining accuracy.
Results on SP-favoring and FE-favoring episodes. To further emphasize that imitating human demonstrations is key to good generalization, we created two subsplits from the HM3D val split that are adversarial to HD performance – SP-favoring and FE-favoring. The SP-favoring val split consists of episodes where BC on SP achieved a higher performance compared to BC on HD, i.e. we select episodes where BC on SP succeeded but BC on HD did not or both BC on SP and BC on HD failed. Similarly, we also create an FE-favoring val split using the same sampling strategy biased towards BC on FE. Next, we report the performance of RL-finetuned from BC on SP, FE, and HD on these two evaluation splits in Table 5. On both SP-favoring and FE-favoring, BC on HD is at success (by design), but after RL-finetuning, is able to significantly outperform RL-finetuning from the respective BC on SP and FE policies.
3 Scaling laws of BC and RL
In this section, we investigate how BC-pretraining RL-finetuning success scales with no. of BC demonstrations.
Human demonstrations. We create HD subsplits ranging in size from to episodes, and BC-pretrain policies with the same set of hyperparameters on each split. Then, for each, we RL-finetune from the best-performing checkpoint. The resulting BC and RL success on HM3D val vs. no. of HD episodes is plotted in Fig. 1. Similar to ramrakhya2022, we see promising scaling behavior with more BC demonstrations.
Interestingly, as we increase the size of of the BC pretraining dataset and get to high BC accuracies, the improvements from RL-finetuning decrease. E.g. at BC demonstrations, the BCRL improvement is success, while at BC demonstrations, the improvement is . Furthermore, with BC-pretraining demonstrations, the RL-finetuned success is only worse than RL-finetuning from BC demonstrations ( vs. ). Both suggest that by effectively leveraging the trade-off between the size of the BC-pretraining dataset vs. performance gains after RL-finetuning, it may be possible to achieve close to state-of-the-art results without large investments in demonstrations.
How well does FE Scale? In Section 5.1, we showed that BC on human demonstrations outperforms BC on both shortest paths and frontier exploration demonstrations, when controlled for the same amount of training experience. In contrast to human demonstrations however, collecting shortest paths and frontier exploration demonstrations is cheaper, which makes scaling these demonstration datasets easier. Since BC performance on shortest paths is significantly worse even with x more demonstrations compared to FE and HD ( SP vs. FE and HD demos, Sec. 5.1), we focus on scaling FE demonstrations. Fig. 6 plots performance on HM3D val against FE dataset size and a curve fitted using demonstrations to predict performance on FE dataset-sizes . We created splits ranging in size from to . Increasing the dataset size doesn’t consistently improve performance and saturates after demonstrations, suggesting that generating more FE demonstrations is unlikely to help. We hypothesize that the saturation is because these demonstrations don’t capture task-specific exploration.
Failure Modes
To better understand the failure modes of our BCRL ObjectNav policies, we manually annotate failed HM3D val episodes from our best ObjectNav agent. See Fig. 7. The most common failure modes are: Missing Annotations (): Episodes where the agent navigates to the correct goal object category but the episode is counted as a failure due to missing annotations in the data. Inter-Floor Navigation (): The object is on a different floor and the agent fails to climb up/down the stairs. Recognition Failure (): The agent sees the object in its field of view but fails to navigate to it. Last Mile Navigation wasserman2022lastmile (). Repeated collisions against objects or mesh geometry close to the goal object preventing the agent from reaching close to it. Navmesh Failure (). Hard-to-navigate meshes blocking the path of the agent. E.g. in one instance, the agent fails to climb stairs because of a narrow nav mesh on the stairs. Looping (). Repeatedly visiting the same location and not exploring the rest of the environment. Semantic Confusion (). Confusing the goal object with a semantically-similar object. E.g. ‘armchair’ for ‘sofa’. Exploration Failure (). Catch-all for failures in a complex navigation environment, early termination, semantic failures (e.g. looking for a chair in a bathroom), etc.
As can be seen in Fig. 7, most failures () are due to issues in the ObjectNav dataset – due to missing object annotations due to holes / issues in the navmesh. failures are due to the agent being unable to climb up/down stairs. We believe this happens because climbing up / down stairs to explore another floor is a difficult behavior to learn and there are few episodes that require this. Oversampling inter-floor navigation episodes during training can help with this. Another failure mode is failing to recognize the goal object – where the object is in the agent’s field of view but it does not navigate to it, and where the agent navigates to another semantically-similar object. Advances in the visual backbone and object recognition can help address these. Prior works ramrakhya2022; chaplot_neurips20 have used explicit semantic segmentation modules to recognize objects at each step of navigation. Incorporating this within the BCRL training pipeline could help. failures are due to last mile navigation, suggesting that equipping the agent with better goal-distance estimators could help. Finally, only failures are due to looping and lack of exploration, which is promising!
Conclusion
To conclude, we propose PIRLNav, an approach to combine imitation using behavior cloning (BC) and reinforcement learning (RL) for ObjectNav, wherein we pretrain a policy with BC on human demonstrations and then finetune it with RL, leading to state-of-the-art results on ObjectNav ( success, improvement over previous best). Next, using this BCRL training recipe, we present a thorough empirical study of the impact of different demonstration datasets used for BC-pretraining on downstream RL-finetuning performance. We show that BC / BCRL on human demonstrations outperforms BC / BCRL on shortest paths and frontier exploration trajectories, even when we control for same BC success on train. We also show that as we scale the pretraining dataset size for BC and get to higher BC success rates, the improvements from RL-finetuning start to diminish. Finally, we characterize our agent’s failure modes, and find that the largest sources of error are 1) dataset annotation noise, and inability of the agent to 2) navigate across floors, and 3) recognize the correct goal object.
Acknowledgements. We thank Karmesh Yadav for OVRL model weights ovrl, and Theophile Gervet for answering questions related to the frontier exploration code chaplot_neurips20 used to generate demonstrations. The Georgia Tech effort was supported in part by NSF, ONR YIP, and ARO PECASE. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government, or any sponsor.
References
Appendix A Prior work in RL Finetuning
Preliminaries. Rajeswaran et al. rajeswaran_RSS_18 proposed DAPG, a method which incorporates demonstrations in RL, and thus quite relevant to our methodology. DAPG first pretrains a policy using behavior cloning then finetunes the policy using an augmented RL objective (shown in Eq. (4)). DAPG proposes to use different parts of demonstrations dataset during different stages of learning for tasks involving sequence of behaviors. To do so, they add an additional term to the policy gradient objective:
Here is a trajectory obtained by executing the current policy, denotes a trajectory obtained by replaying a demonstration, and is a weighting function to alternate between imitation and reinforcement learning. DAPG uses a heuristic weighting scheme to set to decay the auxiliary objective:
where and are hyperparameters and is the update iteration counter. The decaying weighting term is used to avoid biasing the gradient towards the demonstrations data towards the end of training.
Implementation Details. rajeswaran_RSS_18 showed results of using DAPG on dexterous hand manipulation tasks for object relocation, in-hand manipulation, tool use, etc. To train the policy with behavior cloning, they use demonstrations for each task gathered using the Mujoco HAPTIX system mujocohaptix. The small size of the demonstrations dataset and the observation input allows DAPG to load the demonstrations dataset in system memory which makes it feasible to compute the augmented RL objective shown above.
Challenges in adopting rajeswaran_RSS_18’s setup. Compared to rajeswaran_RSS_18, our setup uses high-dimensional visual input (256256 RGB observations) and ObjectNav demonstrations for training. Following DAPG’s training implementation, storing the visual inputs for demonstrations in system memory would require TB, which is significantly higher than what is possible on today’s systems. An alternative is to leverage on-the-fly demonstration replay during RL training. However, efficiently incorporating demonstration replay with experience collection online requires solving a systems research problem. Naively switching between online experience collection using the current policy and replay demonstrations would require x the current experience collection time, overall hurting the training throughput.
A.2 Feasibility of Off-Policy RL finetuning
There are several methods for incorporating demonstrations with off-policy RL awac; awopt2021corl; kalashnikov2018scalable; AWRPeng19; marwil. Algorithm 1 shows the general framework of off-policy RL (finetuning) methods.
Unfortunately, most of these methods use feedforward state encoders, which is ill-posed for partially observable settings. In partially observable settings, the agent requires a state representation that combines information about the state-action trajectory so far with information about the current observation, which is typically achieved using a recurrent network.
To train a recurrent policy in an off-policy setting, the full state-action trajectories need to be stored in a replay buffer to use for training, including the hidden state of the RNN. The policy update requires a sequence input for multiple time steps where is sampled sequence length. Additionally, it is not obvious how the hidden state should be initialized for RNN updates when using a sampled sequence in the off-policy setting. Prior work DRQNhausknecht_aaai15 compared two training strategies to train a recurrent network from replayed experience:
Bootstrapped Random Updates. The episodes are sampled randomly from the replay buffer and the policy updates begin at random steps in an episode and proceed only for the unrolled timesteps. The RNN initial state is initialized to zero at the start of the update. Using randomly sampled experience better adheres to DQN’s dqn random sampling strategy, but, as a result, the RNN’s hidden state must be initialized to zero at the start of each policy update. Using zero start state allows for independent decorrelated sampling of short sequences which is important for robust optimization of neural networks. Although this can help RNN to learn to recover predictions from an initial state that mismatches with the hidden state from the collected experience but it might limit the ability of the network to rely on it’s recurrent state and exploit long term temporal correlations.
Bootstrapped Sequential Updates. The full episode replays are sampled randomly from the replay buffer and the policy updates begin at the start of the episode. The RNN hidden state is carried forward throughout the episode. Eventhough this approach avoids the problem of finding the correct initial state it still has computational issues due to varying sequence length for each episode, and algorithmic issues due to high variance of network updates due to highly correlated nature of the states in the trajectory.
Even though using bootstrapped random updates with zero start states performed well in Atari which is mostly fully observable, R2D2r2d2_iclr19 found using this strategy prevents a RNN from learning long-term dependencies in more memory critical environments like DMLab. r2d2_iclr19 proposed two strategies to train recurrent policies with randomly samples sequences:
Stored State. In this strategy, the hidden state is stored at each step in the replay and use it to initialize the network at the time of policy updates. Using stored state partially remedies the issues with initial recurrent state mismatch in zero start state strategy but it suffers from ‘representational drfit’ leading to ‘recurrent state staleness’, as the stored state generated by a sufficiently old network could differ significantly from a state from the current policy.
Burn-in. In this strategy the initial part of the replay sequence is used to unroll the network and produce a start state (‘burn-in period’) and update the network on the remaining part of the sequence.
While R2D2 r2d2_iclr19 found a combination of these strategies to be effective at mitigating the representational drift and recurrent state staleness, this increases computation and requires careful tuning of the replay sequence length and burn-in period .
Both r2d2_iclr19; hausknecht_aaai15 demonstrate the issues associated with using a recurrent policy in an off-policy setting and present approaches that mitigate issues to some extent. Applying these techniques for Embodied AI tasks and off-policy RL finetuning is an open research problem and requires empirical evaluation of these strategies.
Appendix B Prior work in Imitation Learning
In Imitation Learning (IL), we use demonstrations of successful behavior to learn a policy that imitates the expert (demonstrator) providing these trajectories. The simplest approach to IL is behavior cloning (BC), which uses supervised learning to learn a policy to imitate the demonstrator. However, BC suffers from poor generalization to unseen states, since the training mimics the actions and not their consequences. DAgger ross_aistats11 mitigates this issue by iteratively aggregating the dataset using the expert and trained policy to learn the policy . Specifically, at each step , the new dataset is generated by:
where, is a queryable expert, and is the trained policy at iteration . Then, we aggregate the dataset and train a new policy on the dataset . Using experience collected by the current policy to update the policy for next iteration enables DAgger ross_aistats11 to mitigate the poor generalization to unseen states caused by BC. However, using DAgger ross_aistats11 in our setting is not feasible as we don’t have a queryable human expert for policies being trained with human demonstrations.
Alternative approaches ho_nips16; bahdanau_iclr19; abbeel_icml04; max_ent_irl; fu2018learning for imitation learning are variants of inverse reinforcement learning (IRL), which learn reward function from expert demonstrations in order to train a policy. IRL methods learn a parameterized reward function, which models the behavior of the expert and assigns a scalar reward to a demonstration. Given the reward , a policy is learned to map states to distribution over actions at each time step. The goal of IRL methods is to learn a reward function such that a policy trained to maximize the discounted sum of the learned reward matches the behavior of the demonstrator. Compared to prior works ho_nips16; bahdanau_iclr19; abbeel_icml04; max_ent_irl; fu2018learning, our setup uses a partially-observable setting and high-dimensional visual input for training. Following training implementation from prior works, storing visual inputs of demonstrations for reward model training would require system memory, which is significantly higher than what is possible on today’s systems. Alternatively, efficiently replaying demonstrations during RL training with reward model learning in the loop requires solving an open systems research problem. In addition, applying these methods for tasks in a partially observable setting is an open research problem and requires empirical evaluation of these approaches.
Appendix C Training Details
We use a distributed implementation of behavior cloning by ramrakhya2022 for our imitation pretraining. Each worker collects frames of experience from environments parallely by replaying actions from the demonstrations dataset. We then perform a policy update using supervised learning on mini batches. For all of our BC experiments, we train the policy for steps on GPUs using Adam optimizer with a learning rate which is linearly decayed after each policy update. Table 6 details the default hyperparameters used in all of our training runs.
C.2 Reinforcement Learning
To train our policy using RL we use PPO with Generalized Advantage Estimation (GAE) schulman_iclr16. We use a discount factor of and set GAE parameter to 0.95. We do not use normalized advantages. To parallelize training, we use DD-PPO with workers on GPUs. Each worker collects frames of experience from environments parallely and then performs epochs of PPO update with mini batches in each epoch. For all of our experiments, we RL finetune the policy for steps. Table 7 details the default hyperparameters used in all of our training runs.
C.3 RL Finetuning using VPT
To compare with RL finetuning approach proposed in VPT vpt22 we implement the method in DD-PPO framework. Specifically, we initialize the critic weights to zero, replace the entropy term in PPO schulman_arxiv17 with a KL-divergence loss between the frozen IL policy and RL policy, and decay the KL divergence loss coefficient, , by a fixed factor after every iteration. This loss term is defined as:
where is the frozen behavior cloned policy, is the current policy, and is the loss weighting term. Following, VPT vpt22 we set to at the start of training and decay it by after each policy update. We use learning rate of without a learning rate decay for our VPT vpt22 finetuning experiments.
C.4 RL Finetuning Ablations
For ablations presented in Sec. 4.3 of the main paper (also shown in Table 8) we use a policy pretrained on human demonstrations using BC and finetuned for steps using hyperparameters from Table 7. We try learning rates (, , and ) for both BC RL (row 2) and BC RL (+ Critic Learning) (row 3) and we report the results with the one that works the best. For PIRLNav we use a starting learning rate of and decay it to , consistent with learning rate schedule of our best performing agent. For ablations we do not tune learning rate parameters of PIRLNav, we hypothesize tuning the parameters would help improve performance.
We find BC RL (row 2) works best with a smaller learning rate but the training performance drops significantly early on, due to the critic providing poor value estimates, and recovers later as the critic improves. See Fig. 8. In contrast when using proposed two phase learning setup with the learning rate schedule we do not observe a significant drop in training performance.