Offline Visual Representation Learning for Embodied Navigation
Karmesh Yadav, Ram Ramrakhya, Arjun Majumdar, Vincent-Pierre Berges, Sachit Kuhar, Dhruv Batra, Alexei Baevski, Oleksandr Maksymets
Introduction
We are interested in teaching embodied AI agents to see (i.e. understand the structure and semantics of their environments) and move (i.e. strategically explore and navigate to accomplish goals). This endeavour is of fundamental importance from a practical perspective (e.g. for building home assistant robots) and a scientific perspective (e.g. what are the right visuomotor inductive biases?).
Broadly speaking, three different goal specifications have emerged in the embodied visual navigation literature: 1) point-goal navigation (PointNav ), where an agent must navigate to a relative goal coordinate (‘go to , ’), 2) image-goal navigation (ImageNav ), where an agent must navigate to the location of a goal image (‘go to where this image was taken’) and 3) object-goal navigation (ObjectNav ), where an agent must navigate to an instance of a goal category (‘find bed’) – all in a new environment without a prebuilt map.
The status quo for solving these visual navigation tasks is to train agents tabula rasa, i.e. visual representations and navigation policies are trained from scratch for each specific task, optionally augmented with auxiliary tasks (e.g., predicting the action taken between two successive observations). This line of work has yielded some very exciting results. For instance, Wijmans et al. showed that the PointNav task situated in Gibson environments can reach 99.6% success rate with reinforcement learning (RL) by training an agent for 2.5 billion steps in the Habitat simulator ; Ramakrishnan et al. further sharpened this result and achieved 100% success on this dataset. Similarly, Jaderberg et al. achieved near-human-performance in certain maze exploration tasks in DM Control by augmenting deep RL with auxillary tasks.
However, for tasks requiring semantic scene understanding, results have been much weaker. The winning entries of the 2020 and 2021 ObjectNav challenges attained 18% and 24% success rate respectively. In ImageNav, the best known result (in the single-RGB-camera setting) is 29.2% success rate .
Clearly, what’s true in 2D image understanding is also true in embodied scene understanding – representations matter. So how should we learn useful visual representations for embodied tasks?
In this paper, we propose Offline Visual Representation Learning (OVRL, pronounced Overall). As the name suggests, OVRL (Fig 1) divides visuomotor learning into two stages : 1) pretraining visual representation offline and 2) downstream finetuning. The offline representation learning stage involves using a self-supervised learning (SSL) technique (DINO ) to train a vision model on a large-scale pre-rendered dataset of images from indoor environments called Omnidata . In downstream finetuning, these representations are finetuned on individual task like ImageNav and ObjectNav in the Habitat simulator .
We find that OVRL significantly outperforms the tabula rasa status quo. Specifically, on ImageNav we advance the state of art from 29.2% to 54.2% (+25% absolute, 86% relative) in the single RGB camera setting on the Gibson dataset . On ObjectNav, we advance state of art from 18.1% to 23.2% (+5.1% absolute, 28% relative) in the RGBD camera + known-pose setting. Importantly, we observe that the same pretrained vision backbone generalizes across different scene datasets. Specifically, in ObjectNav, OVRL outperforms our imitation learning (IL) from-scratch baseline by 5.3%, which is impressive because the OVRL encoder never saw MP3D scenes during pretraining. Performance gains of pretrained models are known to diminish (or entirely disappear) when finetuned with long schedules . Surprisingly, we find that the benefits of OVRL pretraining are not only sustained, but continue to increase (rather than decrease) over 2 billion frames of training on ImageNav using the HM3D dataset; this suggests a significant rethinking might be needed in what we consider to be ‘standard’ training schedules for these tasks. Finally, we conduct extensive empirical ablations of OVRL’s different components and find that finetuning the encoder with image augmentations is quite important for achieving good performance.
Related Work
Self Supervised Learning (SSL) in RL: Prior work has proposed contrastive learning (with image augmentations) as an auxiliary loss with RL, although it was later shown that the performance boost was due to image augmentation . CPC , CPCAction and ST-Dim propose different variants of temporal contrastive losses, however these methods add complexity and require a sequence of images for training. ATC demonstrated the first decoupling of the representation learning and RL objective, using only pairs of images to train a temporal contrast objective. In comparison, our method does not require any form of temporal objective and is capable of learning representations from an IID collection of images. PBL , SPR and SGI employ non-contrastive temporal losses similar to BYOL , however they require extra loss terms to prevent their representation from collapsing, which our method doesn’t need.
Visual Navigation: Both classical SLAM-based as well as learning-based approaches have been proposed for embodied visual navigation. End-to-end learning methods typically use fewer hand-crafted modules and have shown more promise. Memory-Augmented RL uses an attention-based model that leverages episodic memory to learn navigation and obtains SOTA results in ImageNav with 4 RGB cameras. In comparison, we use a simpler model architecture while achieving higher performance. On the single camera setup, improves performance using a combination of a goal-view reward and goal-view sampling. We find that using this reward and view sampling leads to further improvements for OVRL models as well.
Similarly, end-to-end RL methods exist for ObjectNav and use data augmentation and auxiliary rewards to improve generalization. In contrast, modular methods , disentangle navigation and semantic mapping. Recently, a competitive imitation learning approach powered by large scale dataset was proposed, which we build upon. improves visual representations by including semantic segmentations, while we focus on RGB representations.
SSL in Embodied AI: EmbCLIP showed that using the CLIP encoder can provide useful representations for EAI tasks. CLIP is pretrained on an unreleased dataset containing 400M image-caption pairs (WebImageText dataset (WIT)). In contrast, we pretrain on a much smaller (14.5M images) and public dataset called Omnidata Starter Dataset . CRL proposes learning visual representations online using samples collected with a curiosity-based navigation policy that is incentivized to find images with a high SSL loss; we compare against a scaled-up version of CRL in our experiments and find that OVRL significantly outperforms it. EPC used self-supervised learning to learn environment level representations by training to predict the missing zones in a zone segmented video sequence; however this requires pose information, OVRL pre-training does not (making it potentialy applicable to videos from the web in the future). Lastly, works from Ye et al. have showed that using auxiliary objectives during training can help on tasks like PointNav and ObjectNav. OVRL outperforms these results on ObjectNav without using any auxiliary losses, leaving open possible future improvements by combining the two ideas.
Approach
In this section we describe OVRL, our two-stage learning approach, which includes an encoder pretraining step using DINO, followed by downstream policy learning in Habitat Simulator on the ImageNav and ObjectNav tasks.
We pretrain our visual encoder using DINO , a recently proposed self-supervised learning (SSL) algorithm. DINO uses knowledge distillation, where a student network is trained to match the outputs of a teacher network. As illustrated in Fig 2 (left), an input image is first transformed using data augmentations. Specifically, we use the multi-crop data augmentation strategy introduced in to produce two global views ( and ) at a resolution and eight local views () at a lower resolution (). All of the views are processed by the student network, but the teacher network only processes the global views. The student and teacher networks both output dimensional feature vectors for each view, which are converted into probability distributions ( and ) using a temperature scaled softmax function (using the parameters and , respectively). The student network is trained to match the outputs of the teacher via stochastic gradient descent (SGD) with the cross-entropy loss: The teacher network parameters are updated as an exponential moving average of the student network parameters.
Representations learned using a self-distillation loss are prone to collapse, meaning that the network converges to a trivial solution like predicting the same representation for every image. To avoid collapse, DINO centers and sharpens the teacher’s output before the softmax operation. Specifically, the centering operation adds the term to the teacher’s output, which is updated as follows: where is a momentum parameter and is the batch size. Sharpening is achieved by setting for the teacher softmax normalization function.
We use the convolutional layers in a modified ResNet50 architecture combined with a projection head for the student and teacher networks. Specifically, we modify the ResNet50 by 1) reducing the number of output channels at every layer by half (i.e. using 32 ResNet baseplanes instead of 64) and 2) using GroupNorm instead of BatchNorm in our backbone, similar to DDPPO . The projection head is a 3-layer MLP (Multi-Layer Perceptron) with 2048 hidden units followed by norm and a weight normalized fully connected layer of K dimension. We keep the BatchNorm in the MLP, however the whole projection head is discarded after training.
2 Downstream Learning
The following paragraphs discuss the data augmentations used for policy learning and network architecture used for ImageNav and ObjectNav experiments.
Data Augmentation: Prior works have shown that using image augmentations during policy learning can help improve overall performance and leads to better generalization on the test set. We experimented with 4 different augmentations including color jitter, random rotation, random resized crop and translate . Similar to RAD , we use augmentations that are consistent over time, thus enabling the augmentation to retain temporal information. Also, in line with , we use different augmentations when action sampling from the policy (for RL), and during the forward-backward pass (in both RL and IL).
ImageNav Policy Learning: Fig 2 shows our ImageNav architecture. At each timestep , the policy receives the current observation and goal observation . These observations consist of RGB images, which are first passed through the data augmentation module and then featured using the observation and goal visual encoders. We initialize both encoders using ResNet50 weights from the pretraining stage. The feature vectors are then passed through a set of fully-connected layers, concatenated together and finally fed into a 2-layer, 512-dimensional LSTM network, along with the representation of the action from the previous timestep. The model is trained using DD-PPO .
ObjectNav Policy Learning: Fig 3 shows our ObjectNav architecture. We build upon Habitat-Web , which uses imitation learning (IL) for training the agent. The policy receives the current observation at every timestep along with the category ID of the goal object . The current observation consists of a RGB + depth image and the agent’s location and orientation obtained from a GPS+Compass sensor. The RGB image is first passed through the data augmentation module and then featurized using the RGB encoder , while the depth image is directly passed to the depth encoder . We initialize the RGB encoder with our pretrained ResNet50 and finetune it during the task, while the depth encoder weights are initialized with a ResNet50 pretrained on the PointNav task and kept frozen. These RGB and depth features are then concatenated together with the GPS+Compass sensor and goal embeddings and fed into a 2-layer 2048-dimensional GRU to predict a distribution over actions . The model is trained using a distributed version of behavior cloning .
Experiments
In this section, we provide implementation details for our pretraining approach using DINO. Then, we provide results of finetuing OVRL representation for the ImageNav and ObjectNav tasks alongside comparisons with several baselines.
Dataset. We pretrain the ResNet encoder using Omnidata Starter Dataset (OSD) , consisting of approximately 14.5 million rendered images from 3D scenes: Replica , Replica+GSO , Hypersim , Taskonomy , BlendedMVG and Habitat-Matterport3D (HM3D) . We exclude images from Gibson test scenes (which are part of Taskonomy) to not contaminate our downstream experiments on ImageNav. We do not use the images from the CLEVR dataset due to their visual and domain dissimilarity from other scenes.
Implementation Details. We use the LARS optimizer with weight decay and batch size 8192, distributed over 64 GPUs. The learning rate starts at 0, is linearly increased during the first 10 epochs to its base value 0.15, followed by a decay to the minimum value of using a cosine schedule . The student and teacher temperatures and are set to 0.1 and 0.04. The network is trained for a total of 100 epochs on OSD.
2 Downstream Learning: ImageNav
Datasets. We conduct experiments using the Habitat simulator and the Gibson dataset , on the standard set of 72 training and 14 testing scenes. Unlike PointNav and ObjectNav, ImageNav does not yet have an authoritative benchmark and prior works often report results in settings with subtle differences, making results incommensurable. We report results under multiple settings and datasets to allow direct comparisons to a variety of prior work. Specifically, we report results on two sets of evaluation episodes – (1) split “A” contains 4,200 episodes (300 / testing scene) and was generated by , (2) split “B” contains 3,000 episodes (approximately 200 / testing scene) and was generated by .
Task Details. Our ImageNav agents are equipped with only RGB cameras (-resolution), with no access to depth or GPS+Compass sensors (though we do compare to and outperform prior work that uses depth cameras). We consider two variants – (1) one front-facing camera (1 RGB) and (2) four cameras (4 RGB) facing ‘front’, ‘back’, ‘left’, and ‘right’ directions (as in ). The agent’s action space consists of four actions: \moveforward(), \turnleft(), \turnright() and \stopac. An episode is terminated if the agent calls the \stopacaction or after the agent has taken 1000 steps in the environment. If the \stopacaction is called within of the goal location, the episode is a success. We report the success rate (SR) and Success weighted by Path Length (SPL) , which is a measure of path efficiency. Agents are trained for 500M steps (25k updates) using 32 GPUs with 10 environments each. Every worker collects (up to) 64 frames of experience, followed by 2 PPO epochs with 2 mini-batches using a learning rate of .
Data Augmentation. For ImageNav, we first found the best set of augmentations by conducting sweeps over individual augmentation parameters and then trying the best sequence of two data augmentations. We found applying color jitter with a value of 0.3 for brightness, contrast, saturation and hue levels followed by translation with a pad of 4 pixels on all the images gave us the best results. We also found data augmentations useful during testing and thus, we report all our results with test-time augmentations.
Baselines. We compare OVRL against the following baselines:
Scratch: A baseline identical in structure and training to OVRL except trained from scratch without any pretraining. Comparing OVRL to Scratch establishes the value of pretraining.
Curious Representation Learning (CRL): CRL is a two-stage approach that pretrains a visual encoder under the SimCLR objective using images collected online while training with an exploration policy that is rewarded to maximize the SSL loss. The encoder is then frozen and used for downstream ImageNav task. Unfortunately, the CRL manuscript reported results with 10M frames of training, which we find to be deficient for 2 reasons – in our experience, training does not converge till 500M frames and results at 10M are highly sensitive to initialization. Thus, we reimplemented and engineered CRL to work on multiple GPUs and scaled up CRL pretraining and training phases from 10M steps to 500M in Gibson scenes. Consequently, the results reported here are significantly stronger than those presented in . Comparing OVRL to CRL establishes the value of pretraining on an IID-sampled (vs agent gathered) dataset of images, while disentangling the effects of scaling.
Zero-Experience Required (ZER): ZER extends the Scratch baseline by introducing a new reward structure that encourages matching the orientation of the goal image. Additionally, ZER randomly samples goal image orientations during training to increase the number of novel episodes.
Mem-Aug RL : Memory-augmented RL proposes an attention based model that leverages episodic memory for navigation with panoramic images.
NRNS : NRNS uses passive videos to learn a geodesic distance estimator and a target prediction model to create a topological map and then navigates an agent using a combination of global and local policy. It requires both an RGBD image and egocentric pose information to navigate.
Results. In Table 1, we report ImageNav results averaged over 3 random seeds. In row 6, we show that OVRL attains 41.3% SR and 26.9% SPL on split “A”, while the scratch baseline (row 1) only achieves 17.9% SR and 9.3% SPL. We emphasize that the only difference between these two methods is that OVRL is initialized with a pretrained encoder. This clearly shows that using pretrained encoders can make a big difference to the performance of an embodied agent.
We also find that OVRL outperforms all prior work. Specifically, OVRL outperforms ZER (row 2) by 12.1% in terms of success rate and 5.3% in SPL. In row 3, we find that reproducing ZER with a deeper CNN (ResNet50 vs. ResNet9) to match the visual encoder used in our approach leads to a small drop in ZER performance. However, in row 7, we find that using the ZER reward structure and goal view sampling with our pretraining approach (OVRL + ZER-Reward) further improves our agent’s success rate by 12.9% and SPL by 0.1%. These results highlight that our pretraining approach can be used alongside other improvements to embodied agents to achieve additional performance gains. In row 4 and 5, we present results on CRL , as reported in the paper and from our re-implementation. According to , CRL achieves 5.8% success rate and 3.2% SPL when pretrained on the MP3D dataset followed by finetuning and testing on the Gibson dataset using the PointNav episode splits. In comparison, our reimplementation achieves 20.4% success rate and 10.2% SPL on the test split “A”. Comparing CRL with OVRL (row 6), we observe that CRL’s alternative pretraining technique underperforms us by 20.9% in success rate and 16.7% in SPL.
In row 9, we extend OVRL to the multi-view setting (4 RGB) for direct comparison with Mem-Aug RL (details in Section 0.A.1). We find that OVRL outperforms Mem-Aug RL by 10.8% in success rate and 6.5% in SPL despite not using an additional memory mechanism within the agent’s architecture. Finally, in row 11, we evaluate on episodes used in (split “B”) and find that our approach outperforms NRNS by 21.5% in success and 16.0% in SPL even without having access to depth information or an egocentric pose estimate.
3 Downstream Learning: ObjectNav
Dataset. We conduct ObjectNav experiments in the Habitat simulator using the Matterport3D scene dataset , with the standard set of 61 training, 11 validation and 18 testing scenes. The results are reported on the MP3D val and test-std split from Habitat Challenge ObjectNav benchmark.
Task Details. In the ObjectNav task, an agent is tasked with navigating to an instance of a specified object category (e.g. ‘sofa’) in a unseen environment. The agent does not have access to a map of the environment and must navigate with RGBD camera and GPS+Compass sensor which provides location and orientation information relative to the start of the episode. Visual observations from the simulator are first downsampled by 0.5, from to , and then resized and center cropped to while preserving the aspect ratio. Additionally, the agent also receives the goal object category ID as input. The agent’s action space consists of \moveforward(), \turnleft(), \turnright(), \lookup(), \lookdown(), and \stopacactions. For an episode to be considered success the agent has to stop within euclidean distance of the goal object within steps and be able to turn to view the object from the end position. We train all agents for M steps. We evaluate checkpoints at every M steps for last M steps of training and report metrics for the checkpoints with the highest success on the validation split.
Imitation Learning. We use the dataset of ObjectNav human demonstrations collected by Habitat-Web . For each of the Matterport3D train scenes and each goal category, it provides approximately human demonstrations from randomly-chosen start locations. We perform behavior cloning on human demonstrations, which amounts to million frames of experience.
Data Augmentation. Similar to ImageNav, we use a combination of color jitter followed by translation. To find out the best hyperparameters, we did a grid search and found that color jitter values of 0.4 for brightness, contrast, saturation and hue levels followed by a pad of 16 pixels gave the best results.
Baselines. We compare our method against the following baselines:
DDPPO: A standard RL baseline that uses a single ResNet18 encoder with 32 baseplanes for RGB+Depth observations trained with DD-PPO .
EmbCLIP: EmbCLIP investigates the effectiveness of CLIP’s visual representation for EAI tasks. The policy encodes the RGB observations using CLIP’s frozen ResNet50, which is passed through a two-layer CNN to obtain the visual embedding. This embedding is concatenated with the goal embedding and then passed through a recurrent model and the policy is trained using DD-PPO .
Habitat-Web : We reimplement the agents from Habitat-Web to use identical architecture and image augmentations as OVRL, except the RGB encoder is trained from scratch. This comparison establishes the value of pretraining.
Results. We compare our approach with existing ObjectNav methods in the RGB and RGBD observation settings and present our results for 1 seed in Table 2. On the MP3D test-std split, in the RGB setting, we find that OVRL improves over Habitat-Web by % in success rate and % in SPL (row 2 vs. 3). OVRL also improves over EmbCLIP by in success rate while performing worse in SPL in the RGB setting (row 1 vs. 3).
In the RGBD setting, OVRL’s success rate and SPL on the test set improve to 23.2% and 7.6% respectively. Our implementation of Habitat-Web performs slightly better than the reported results (row 5 vs. 6) since it uses a bigger network (ResNet18 vs ResNet50) and image augmentations. We observe that the gap between OVRL and Habitat-Web reduces to 5.3% and 1.5% for success rate and SPL respectively (rows 6 vs. 7); likely since the policy now relies lesser on the RGB image representations for navigation. OVRL also performs much better than DDPPO, achieving higher success rate and higher SPL (rows 4 vs. 7) in the RGBD seting.
Analysis
In this section, we deconstruct our approach and systematically evaluate different design choices. All analyses in this section are conducted on the ImageNav task in the single RGB camera setting described in Section 4.2.
First, we evaluate the choice of the pretraining dataset and then study scaling laws for the best-performing dataset.
We compare OSD with two alternative datasets: a) ImageNet-1k , and b) Gibson-ShortestPath dataset, which consists of 12 million images collected (by us) from 307 Gibson scenes using an oracle agent navigating the the PointNav training episodes via the shortest paths. Table 3 (rows 1-3) shows that representations learnt on the ImageNet-1k dataset achieve 20% success rate improving the Scratch baseline (from Table 1, row 1) by only 2.1%. Using the Gibson-ShortestPath dataset gives a significant boost to the performance, taking the success rate to 36.9%. This suggests that if the pretraining dataset is visually similar (though notice not identical) to the downstream task, the SSL objective is able to capture useful features. Also interestingly, we observe that the success rate of OVRL trained on OSD-100% (row 3) is higher than that Gibson-ShortestPath dataset by 4.4%. We attribute this improved performance to the diverse sampling scheme of the Omnidata pipeline.
Next, we study scaling laws for pretraining on OSD. We randomly subsample images from OSD and create smaller datasets containing only 1% (145K), 10% (1.45M) and 25% (3.625M) of the original, while ensuring that each successively smaller dataset is a subset of the larger ones. The results are summarized in Table 3 (rows 4-6). We notice that peak performance is achieved rather quickly, with essentially no difference between OSD-10%, 25%, and 100%. Notice that OSD contains images uniformly sampled from 2400 scenes; thus, we need somewhere between 60 images (1%) and 600 images (10%) per scene to learn useful representation for embodied tasks.
Finally, since OSD contains images from multiple 3D scene datasets (Taskonomy/Gibson, HM3D, Replica+GSO, BlenderMVG), we study the contribution of its constituents in Table 3 (rows 7-9). We find that using only HM3D or Taskonomy subset for pretraining gives similar performance as using the full dataset, while using the combination (Taskonomy + HM3D) leads to a slight increase in the performance. We hypothesize this gain in performance between OSD-100% (row 3) and OSD-Taskonomy+HM3D (row 9) is due to the removal of the images from the Replica+GSO and BlendedMVG datasets, which contain a large number of images focused on household objects and sculptures, therefore deteriorating the quality of the learnt representation for ImageNav.
2 How do augmentations and finetuning impact performance?
In this section, we disentangle the contributions of finetuning the vision encoder from image augmentations (during finetuning stage). We do this by freezing the visual encoder during ImageNav training and refer to this as OVRL-Frozen. We then analyse performance of models re-trained and tested without augmentations. Table 4 shows the results (notice that Scratch and OVRL results in the presence of augmentations are identical to those in Table 1). First, we notice that finetuning is an unconditionally good idea (with and without augmentation), with significant gains in the presence of augmentation – +11.2% success rate and +9.9% SPL (row 2 vs. 3, col. “Augmentations”). This result is not surprising, since finetuning allows the visual representations to adapt to a novel image distribution and capture task-specific details.
On the other hand, augmentations seem to be conditionally good. Adding augmentations during finetuning while keeping the vision encoder frozen hurts – the success rate and SPL drop by 2.9% and 2.0% respectively (row 2). We speculate that augmentations are harmful for OVRL-Frozen since the network is incapable of learning the useful invariances from the image augmentations. However, if the vision encoder is finetuned, using augmentations improve success rate and SPL by 6.0% and 3.8 (row 3)%. Finally, we see that the pretraining step is leading to 12-16% success rate improvement over Scratch (row 1 vs. 2). This shows that both finetuning and image augmentations are important components for achieving the best performance of OVRL, but augmentations should not be used alone.
3 How does the choice of SSL algorithm affect performance?
While our approach is agnostic to the choice of SSL algorithm, we conduct an ablation to show that our choice of SSL algorithm, DINO, is superior to other existing approaches. We compare against MOCO-v2 and SimCLR, two well-known contrastive learning techniques. On the ImageNet classification task, DINO is known to outperform both these approaches in terms of the linear and k-NN classification accuracy. In Table 5.3, we find that DINO surpasses both algorithms in terms of success and SPL on the ImageNav tasks as well. In the future, OVRL can be coupled with improved SSL algorithms.
4 Does model size matter?
We study the choice of our vision model (ResNet50-32B) by comparing it against a ResNet18 and a ResNet101 network with 32 baseplanes and a ResNet50 with 64 baseplanes. We see in Table 5.3 that using a ResNet18 leads to more than 7% decline in success rate and 5.0% decline in SPL (row 1 vs. 2) as the model fails to consolidate all the knowledge in a smaller network. On the other hand, increasing the number of layers or baseplanes leads to increased compute requirements, with only a slight improvement in performance for ResNet50-64B (row 3).
5 Is OVRL effective when the agent is trained for much longer?
Performance gains of pretrained models are known to diminish (or entirely disappear) when finetuned with long schedules . So, are OVRL representations also just useful for short finetuning schedules? To answer this question, we train an OVRL agent on the ImageNav task using the 800 training scenes from the HM3D dataset and compare against an agent trained from scratch. We use the 8 million PointNav training episodes generated by to train for 2 billion simulation steps and then present the test performance on the Gibson test split “A” in Table 5 for 1 seed. We find a high correlation between the training and testing success rates of the two methods in Fig 5, indicating positive transfer from the HM3D to Gibson dataset. We further observe that OVRL-HM3D even surpasses the performance of OVRL trained on the Gibson scenes (Table 1 row 6). Finally, in Fig 5, we observe that the benefits of OVRL pretraining over scratch are not only sustained, but continue to increase over the course of this long training schedule, which suggests that a significant rethinking might be needed in what is consider to be the ‘standard’ training schedule for these tasks.
Conclusions
We presented OVRL, a simple two-stage technique for using self-supervised learning techniques in EAI tasks. We show that OVRL learns generalizable representations and present results on the ImageNav and ObjectNav tasks in Habitat. OVRL achieves a success rate of 54% on ImageNav in Gibson (+25% absolute improvement in SOTA) in the single-RGB setting and 23% on ObjectNav in MP3D (+5.1% absolute improvement in SOTA) in the RGB/RGBD-only setting. Finally, we conduct extensive ablation studies to understand the usefulness of the different components in our approach.
In the future, we plan to investigate the importance of different components of DINO to performance on embodied AI tasks. It would also be interesting to look at better data collection strategies, inspired by , that could further improve the learnt representations. Finally, we are also interested in extending our work to interactive EAI tasks like Rearrangement .
References
Appendix 0.A Approach
To compare against Mem-Aug Nav , we take the architecture proposed in Section 3 and extend it to work in the multi-view (4 RGB Cameras) setting. At each timestep , the policy receives the current observation and goal observation . Both and tuples consist of 4 RGB images from 4 cameras facing facing ‘front’, ‘back’, ‘left’, and ‘right’ directions . These observations are first passed through the data augmentation module and then featurized using the observation and goal visual encoders. We initialize both encoders using ResNet50 weights from the pretraining stage. The current observation feature vectors are then passed through a fully-connected (FC) layer, concatenated together and passed through another FC layer. For the goal representation, the feature vectors are passed through an FC layer and the outputs are added together similar to to keep the representation rotation-invariant. Finally, the current and goal observation representations are concatenated together and fed into a 2-layer, 512-dimensional LSTM network, along with the representation of the action from the previous timestep.
Appendix 0.B Additional Experiments
In this section, we present results on indoor scene classification experiments along with additional experiments we conducted on the ImageNav task.
In this section, we study the representations learned with OVRL pretraining and after downstream finetuning to gain insights into our two-stage learning framework. Specifically, we evaluate our pretrained models on an indoor scene classification task.
Implementation Details. For this experiment, we use a subset of the Places365-Standard dataset containing images from categories corresponding to indoor room scenes (details in Section 0.E in the Appendix). To evaluate the representation, we use the pretrained visual encoders (ResNet50 with 32 baseplanes) from different methods, freeze their weights and learn a linear classifier on the final average pooled features obtained from the encoder. We train all baselines for epochs using Adam optimizer with a learning rate of .
Baselines. We compare OVRL against visual encoders obtained from the ImageNav from Scratch baseline and CRL presented in Section 4.2. We also compare with two encoders pretrained on the ImageNet dataset . The first is pretrained with DINO, the same SSL algorithm used in OVRL pretraining. The second is pretrained with supervised learning using the ImageNet labels.
Results. We report the top-1 and top-5 accuracy on the Places dataset in Table 7. We find that OVRL pretraining (OVRL-Frozen – row 5) outperforms alternative pretraining methods or pretraining on other datasets. Specifically, OVRL-Frozen achieves top-5 accuracy and top-1 accuracy, which is better in terms of top-1 accuracy compared to the representations learnt by the Scratch baseline on ImageNav task (row 5 vs. 1). Our approach also performs better in top-1 accuracy when compared to our implementation of the pretraining approach proposed in CRL . Next, we show that using DINO (the same SSL method as OVRL) to pretrain on ImageNet dataset doesn’t lead to representations that transfer as well as us to this task. The ImageNet-DINO baseline performs worse compared to OVRL-Frozen (row 4 vs. 5). Interestingly, we find that finetuning on the ImageNav task reduces scene classification performance by 8% (rows 5 vs. 6). This suggests that some of the visual features learned through SSL that are useful for high-level scene classification, might not be helpful for ImageNav. Finally, we compare OVRL with representations learnt using supervised learning on ImageNet, and find that ImageNet-SL is better in top-1 accuracy (row 5 vs. 7).
B.2 Downstream Learning: ImageNav
ImageNet Pretrained Models In Section 5.1, we compared the performance of models pretrained with DINO using ImageNet-1k dataset, Gibson-ShortestPath dataset and OSD. We find it surprising that ImageNet pretraining leads to only a slight improvement in performance (+2.1% success and +4.6% SPL) over Scratch baseline and therefore investigate this issue further. We start by freezing the pretrained visual encoder and see a 3.7% improvement in the success rate (row 1 vs. 2, Table 8) and 0.6% deterioration in the SPL. We follow this by removing the image augmentations from the frozen baseline and find an overall improvement of 6.1% in success rate and 1% in the SPL over ImageNet-Finetuned (row 2 vs. 4). Overall, these results demonstrate that ImageNet representations are not easily adapted to embodied tasks such as ImageNav using simple finetuning and data augmentation techniques. In contrast, the performance of our approach (row 6, Table. 1), which is pretrained on OSD, improves with downstream finetuning and can take advantage of data augmentation.
Appendix 0.C Additional Analysis
In Section 5.5, we studied OVRL’s properties while training on the HM3D dataset for a long time horizon (2 billion (B) simulation steps). Here, we conduct a similar study for the Gibson training scenes in the ImageNav task, by training OVRL and Scratch agent for 1.5B steps. Gibson being a much smaller dataset compared to HM3D (72 vs 800 training scenes), we expect to see a faster convergence of the two methods on Gibson. In Fig. 6, we observe that the OVRL starts to converge after 500 million (M) steps on the training set, finally reaching a success rate of 95% and SPL of 89% at the end of training. Meanwhile, Scratch baseline continues to improve after 500M steps, with both the success rate and SPL nearly doubling in the next 1B steps of training to final values of 70% and 60% respectively. However, on the test set we see much weaker gains, with Scratch only improving by 5% in success rate and 4.2% in SPL, reaching final values of 22.9% and 13.5% respectively. In comparison, OVRL’s success rate of 41.3% and SPL of 26.9% at 500M steps (row 6, Table 1), is still approximately double the current performance of Scratch. Overall, this experiment demonstrates that Scratch is unable to match OVRL’s performance even with significantly longer training horizon and supports the hypothesis about importance of offline visual representation.
C.2 How do SSL representations compare against the representations learnt with RL on a given task?
To answer this question, we take the visual encoder learnt by the Scratch agent on ImageNav and use it as the pretrained visual encoder for a downstream task. We choose the same ImageNav task for this downstream experiment which eliminates the possibility of encountering any kind of distribution shift in the dataset. We try both finetuning and freezing the RL pretrained visual encoder and call these methods RL-Pretrained-Finetuned and RL-Pretrained-Frozen respectively. Looking at the results in Fig 9, we observe a significant improvement in the success rate and SPL of both the methods over the Scratch baseline on the training episodes. We also see that the gap between the success rate of OVRL and RL-Pretrained baselines reduces to only 8%. However, as we look at the test performance, we observe a significant generalization gap in the performance of RL-Pretrained methods. Finally, comparing RL-Pretrained methods and OVRL (Table 8) on the test set, we see a difference of approximately 14% in success rate and 11% in SPL. Overall, this result demonstrates that pretraining visual encoders using self-supervised learning is beneficial for generalization in embodied AI tasks.
Appendix 0.D Compute Resources
We use the 32GB Tesla V100 GPU for all our experiments. Pre-training our vision encoder on the 12 million images from the Omnidata Starter Dataset takes a total of 65 hours on 64 GPUs – i.e. 173 GPU days of training. A single run of OVRL using the default settings on the ImageNav task takes 40 hours to train parallelly on 32 GPUs – equivalent to 53 GPU days. The ImageNav experiments in Section 4.2 take approximately 960 GPU days to train in total. In ObjectNav, we train for 2 days in our RGB only experiments and 3 days in the RGBD experiments, taking a total of 640 GPU days. Finally, the experiments across the analysis section total over 3500 GPU days of training time.
Appendix 0.E Places Dataset
For the scene classification experiments in Section 0.B.1, we use 55 indoor scene classes from the Places365 dataset adapted from . Specifically, we use the following classes: airport_terminal, apartment_building-outdoor, art_gallery, art_studio, attic, auditorium, ballroom, banquet_hall, bar, basement, beauty_salon, bedroom, bookstore, cafeteria, classroom, closet, cockpit, coffee_shop, conference_center, conference_room, corridor, dining_room, dorm_room, engine_room, fire_escape, home_office, hospital_room, hotel-outdoor, hotel_room, inn-outdoor, kitchen, living_room, lobby, locker_room, mansion, martial_arts_gym, motel, museum-indoor, music_studio, nursery, office, office_building, patio, railroad_track, residential_neighborhood, restaurant, restaurant_kitchen, restaurant_patio, shed, shopfront, shower, stageindoor, staircase, swimming_pool-outdoor, waiting_room.