OVRL-V2: A simple state-of-art baseline for ImageNav and ObjectNav

Karmesh Yadav, Arjun Majumdar, Ram Ramrakhya, Naoki Yokoyama, Alexei Baevski, Zsolt Kira, Oleksandr Maksymets, Dhruv Batra

Introduction

Imagine a home assistant robot that can find things in the house. For instance, we might ask it in natural or templated language to ‘find a sweatshirt’. Or we may show the agent a picture of our favorite sweatshirt and ask it to find it. Designing systems for autonomous navigation to semantic goals is a challenge of broad scientific and societal interest.

In the embodied AI research community, two concrete goal specifications have emerged for semantic navigation. In ImageNav , an agent is spawned in an unseen environment and asked to find a location ‘described’ by a goal image. In ObjectNav , it is asked to find any instance of an object category given its name ‘find a <<name>>’.

Both problems test the agent’s semantic understanding and episodic memory – what objects or parts of a scene are visible in the goal image (is it part of a kitchen or bathroom)? Where are these objects (seen in the goal image in ImageNav or mentioned by name in ObjectNav) typically found in a house? Where does the agent find itself at initialization? And how should it strategically search the environment to find the object or scene that it is looking for without looping over the same area multiple times?

The classical sense-plan-act pipeline from the robotics literature approaches this problem via a sequence of modules – detecting objects (or extracting semantic features) in 2D images or in 3D point clouds , accumulating detections (or features) into a 2D (top-down) or 3D map and planning a path or waypoints on this map , executed by a low-level controller. These approaches have been quite prevalent in prior work and have proved to be a strong baselines in both the tasks.

In this work, we advance an alternative research program – training generalist agents constructed from task-agnostic neural components without any task-specific modules. Such general-purpose methods offer advantages of simplicity in design, positive scaling with available compute (incorporating the ‘bitter lesson’ ), and versatile applicability to multiple tasks.

A flurry of recent work on image and video understanding has found that visual transformers (ViTs) powered by self-supervised representation learning can provide general-purpose visual representations for recognition and generation tasks. However, while the training recipes for convolutional networks are mature and robust, the recipes for ViTs are contingent and brittle, and in the case of ViTs for visual navigation, yet to be fully discovered – and that discovery is the focus of our work.

Our key technical contributions and findings are as follows:

Compression layers are needed for ViTs in Visual Navigation. We find that ViT-based agents trained from scratch perform poorly compared to ResNets (e.g. achieving only 36.1% success rate (SR) on ImageNav vs. 59.9% for ResNets). This is despite a substantially higher model capacity (ViT-Small has ∼\sim4 times more parameters than a half-width ResNet50). We find that a key issue with using ViTs for navigation problems is that both the [CLS] token embedding and global-average pooling remove a spatial structure that is important for the task. We propose using a compression layer (consisting of a 2D convolution plus flattening) operating over ViT patch representations to preserve spatial information, and find that it leads to ViTs outperforming ResNets (67.4% vs. 59.9% SR on ImageNav).

Visual pretraining unlocks positive scaling laws for the first time. We demonstrate, for the first time, positive scaling laws with ViT-based agents on ImageNav. Specifically, we find that visual representation learning (using masked autoencoding (MAE) ) not only improves performance, but also enables model scaling with ViTs. With this pretraining, we are able to increase the model size from ViT-Small to ViT-Base and observe gains in success rate from 80.5% to 82.0% (+1.5%) and SPL (success weighted by path efficiency) from 55.2% to 58.7% (+3.5%).

Single architecture achieves SoTA on ImageNav and ObjectNav. Putting it all together (ViTs, compression layers, pretraining, policy training improvements and scaling), we present OVRL-v2 (Offline Visual Representation Learning v2), a simple ViT+compression-layer+LSTM architecture as a successor to the state-of-the-art method, OVRL . OVRL-v2 pushes the state-of-art success rate on ImageNav from 54.2% (in ) to 82.0% (+27.8% absolute and 51.3% relative improvement) and on ObjectNav achieves 64.0% success rate comparable to state-of-the-art (65.0%, obtained by concurrent but orthogonal work ). OVRL-v2 agents use only RGB and GPS+Compass sensors; no egocentric depth (as used by ), no semantic segmentation (as used by ), no object detection (as used by ), no semantic or geometric mapping (as used by ).

Overall, this work does not present a fundamentally new approach, but rather recommendations for training a general-purpose architecture that achieves state-of-art performance today and could serve as a strong baseline for future methods.

Related Work

Visual Navigation. Visual Navigation approaches can be divided into three categories: a) SensePlanAct pipelines from classical robotics literature (typically without any learned components) , b) Modular pipelines with learned modules , and c) Monolithic neural network approaches . While many modular learning methods build explicit semantic maps , others simply use object detectors or segmentors without mapping . In comparison, we show that it is possible to achieve state-of-art performance on semantic navigation tasks without using semantic mapping, object detection, or segmentation of any kind. Such an embodied agent, composed of task-agnostic components, not only forms a strong baseline for any task but also provides the foundation for a generalist agent, which is the goal of embodied AI.

ViTs in Embodied AI. Very few works in Embodied AI have used ViTs as their vision backbone. In the Vision and Language Navigation literature, uses CLIP ViT models but keep them frozen during training. History Aware Multimodal Transformer takes a ViT model pretrained on ImageNet and further pretrains the ViT along with the rest of the model using a 2 step process with proxy tasks. In comparison, OVRL-v2 demonstrates how to finetune a SSL pretrained model using gradients from either imitation learning or reinforcement learning losses. In robotics, and showed the benefits of using pretrained ViT but do not see any benefits of finetuning.

Self Supervised Learning (SSL) for Embodied AI. Recent works have explored self-supervised visual representations for visuomotor control and visual navigation . demonstrated the efficacy of frozen MAE representations for motor control tasks, while proposed incorporating a reconstruction-based self-supervised objective alongside online model-based RL training. EmbCLIP uses off-the-shelf CLIP encoders, which are frozen during policy learning. In our work, we show the effectiveness of finetuning, not just for our method but also for ViT-based CLIP models. Both OVRL and EmbCLIP use ResNet-based visual encoders. By contrast, in this work we focus on adapting more recent ViT-based backbones for visual navigation.

Background: Tasks and Visual Pretraining

We study two visual navigation tasks: image-goal navigation (ImageNav) and object-goal navigation (ObjectNav) . To address these tasks, we design an embodied agent leveraging a vision transformer (ViT) . This section provides an overview of each task, then describes an approach that we use for pretraining ViTs.

Fig. 2 illustrates the ImageNav and ObjectNav tasks. In both, an agent starts at a random position and orientation in an unknown 3D scene. The agent must explore the environment to find a goal location. In ImageNav, the goal is an image (e.g., a picture of a sofa) that is taken from the goal position. In ObjectNav, the agent is given the name of an object (e.g., ‘sofa’) that it has to find.

In these tasks, agents perceive the environment using an egocentric RGB camera. Agents navigate using a discrete action space. In ImageNav, the standard set of actions includes: move_forward (0.25m0.25m), turn_left (30∘30^{\circ}), turn_right (30∘30^{\circ}) and stop to indicate that the agent thinks it has reached the goal. In ObjectNav, agents can also look_up (30∘30^{\circ}) and look_down (30∘30^{\circ}).

Agents are evaluated in previously unseen environments, which allows measuring how well navigation behaviors generalize. Two standard metrics are used to assess the agent’s navigation performance: success rate (SR) and success weighted by (inverse) path length (SPL) . SPL rewards agents that take shorter paths to the goal, thus measuring how efficiently the agent explores new environments.

2 Masked Autoencoders (MAEs)

Visual navigation tasks require understanding visual cues to navigate in new environments. Thus, agents require strong visual representations. We use masked autoencoding (MAE) – an efficient self-supervised visual representation learning algorithm designed for pretraining vision transformers (ViTs) – to improve the performance of our ViT-based agent. MAE derives its efficiency from an asymmetric encoder-decoder design. Specifically, an input image is first divided into non-overlapping patches, a high fraction (75%) of which are randomly masked during pretraining. The encoder only processes the remaining unmasked patches, which reduces the computational burden during pretraining. A small decoder is tasked with reconstructing the full input image. Both the encoder and decoder are ViTs, which naturally handle processing the variable number of patches. The high masking percentage is achievable due to the natural redundancy across patches in real-world images, which makes the full image predictable from only a small subset of the constituent parts. After pretraining, the decoder is discarded, and only the encoder is used for downstream tasks.

Approach

We use a general-purpose agent architecture for both visual navigation tasks (ImageNav and ObjectNav). As shown in Fig. 4, both agents primarily consist of a visual encoder (a ViT initialized randomly or pretrained with MAE), a goal encoder, and a recurrent policy network. This section describes several key components of our approach.

Compression layers for ViTs. As shown in Fig. 4, our visual navigation agents process the RGB observation OtO_{t} with a ViT-based visual encoder fθobsf_{\theta_{obs}}. Specifically, the input images are converted into non-overlapping 16×\times16 patches after data augmentation, concatenated with a [CLS] token, and then processed with a ViT, which outputs a representation for each patch and the [CLS] token. In tasks such as image classification, it is common to represent the image using either (a) the [CLS] token output or (b) the average pooling of the patch representations (i.e., global average pooling).

Notice that both solutions remove the spatial structure present in the patch layout; we contend that this removal is detrimental for navigation. Thus, as illustrated in Fig. 3, we use a layer that first reshapes the patch representations (generated by a ViT) back into a grid. Next, we process them with a convolutional layer that reduces the dimensionality of (i.e. compresses) each patch.The convolution layer is a 3×\times3 Conv + GroupNorm + ReLU. Finally, we flatten the output to maintain spatially distinct features. This architecture has also been used in prior work (e.g., ) to process grid features produced by a ResNet, which we adapt here for ViTs and refer to as a compression layer. The PyTorch implementation for this is presented in Appendix C.

Visual Navigation with ViTs As illustrated in Fig. 4, the output from the visual encoder fobsf_{obs} is concatenated with a goal representation and an embedding of a GPS+Compass sensor (only used for ObjectNav) that provides pose information. The concatenated output is processed by a recurrent LSTM-based policy network, which predicts actions.

The difference between agents for each task is the method used to encode the goal. In ImageNav, the image-goal OgO_{g} is encoded with a visual encoder fθgoalf_{\theta_{goal}} with an identical architecture to fθobsf_{\theta_{obs}}. For ObjectNav, the goal object category (e.g. ‘sofa’) is encoded via a learned embedding layer.

We train our ImageNav agent with reinforcement learning (RL) using DD-PPO with the reward function described in a future subsection. For ObjectNav, we train our agent using human demonstrations with a distributed version of behavior cloning . Further training and evaluation details are provided in Appendix A.

Visual Encoder Pretraining. Our proposed approach of using a ViT-based visual encoder within a model-free navigation agent (Fig. 4) can be trained end-to-end from scratch (e.g., using the RL rewards described in the next section). In addition, we investigate pretraining the ViT-based visual encoder using the masked autoencoding (MAE) algorithm described in Sec. 3.2. For pretraining, we collect an in-domain image dataset from HM3D and Gibson scenes. This follows the observation in prior work (e.g., ), which demonstrates that pretraining on in-domain data (as opposed to datasets like ImageNet) improves downstream performance. Further details about the pretraining dataset and hyperprameters are provided in Appendix A.

ImageNav Rewards. The reward used for visual navigation is typically composed of three components: (a) a sparse reward csc_{s} for successfully completing the task, (b) a per time-step penalty γ\gamma to incentivize efficiency, and (c) one or more reward-shaping terms to simplify the optimization problem. A common reward-shaping term is the change in (geodesic) distance to goal. Formally, let dtd_{t} indicate the agent’s geodesic distance to goal at time tt; now, the reward-shaping term can be written as: dt−1−dtd_{t-1}-d_{t}. Putting all three reward terms together, this reward is defined as:

where rgr_{g} is the goal radius and ata_{t} is the agent’s action.

One limitation of the reward function in Eq. 1 is that it is indifferent to the agent’s ‘heading’ at termination – the agent is neither rewarded for looking at the goal object (which is a desirable behavior since navigation is typically a precursor to manipulation), nor is the agent penalized for ending the episode looking away from the object. To resolve this issue, proposed two additional angle reward terms that incentivizes, 1) turning towards the goal (using an angle-to-goal (θt\theta_{t}) reward shaping term) and 2) stopping while looking at the goal (using a terminal reward). Both these rewards are only awarded after the agent has entered within a goal radius rgr_{g}. While demonstrated that their reward can improve ImageNav performance, we found that our OVRL-v2 agent is able to hack the reward function, by never ending the episode, moving into the goal radius, turning to look at the goal, moving outside the goal radius, turning back and repeating. We provide more details about this reward and visualize the behaviour of the agent in Appendix F. We hypothesize that prior work did not notice this exploitability because it only becomes apparent when the experiments are appropriately scaled.

We propose a principled fix to the reward function in . Our key insight is that we can transform the angle-to-goal reward shaping term into a difference of potential functions, which are provably optimal for reward-shaping . Specifically, we define an angle-to-goal function θ^t\hat{\theta}_{t} that is equal to π\pi outside of the goal radius and equal to the angle-to-goal otherwise:

With this definition, agents can be appropriately rewarded (or penalized) for entering (or exiting) the goal radius using the difference in the per timestep measure Δθ^t=θ^t−1−θ^t\Delta_{\hat{\theta}_{t}}=\hat{\theta}_{t-1}-\hat{\theta}_{t}. This term will be outside the goal radius, ⩾0\geqslant 0 when entering (thus encouraging the agent), and ⩽0\leqslant 0 when exiting (thus discouraging the agent). Furthermore, Δθ^t\Delta_{\hat{\theta}_{t}} has the desirable property that any positive reward accumulated inside the goal radius will entirely be lost if the agent exits, resulting in a zero-sum path. We take advantage of these properties, and propose the full reward structure as follows:

where θg\theta_{g} is an angle success threshold (set to 25∘ in our case). In Sec. 5, we find that the reward defined in Eq. 3 substantially improves ImageNav performance (particularly in terms of path efficiency as measured by SPL).

Experimental Findings

In this section, we first establish an ImageNav baseline that is competitive with existing SoTA methods. We then use this strong baseline to systematically address the following research questions:

Do ViTs work out-of-the-box for ImageNav? No. We discover that despite higher model capacity, ViT-based agents trained from-scratch underperform smaller ResNet agents by considerable margins.

How does adding a compression layer affect performance? We find that using a compression layer to maintain spatial structure in image representations significantly improves navigation performance on ImageNav.

Does performance scale with larger ViTs? When trained from scratch we observe mixed results. However, self-supervised visual pretraining results in consistent across-the-board improvements, along with scaling

Can strong visual navigation agents ‘hack’ the new reward function in Eq. 3? No. Agents that can ‘hack’ the ZER reward are no longer able to ‘hack’ the reward function with our proposed corrections.

How does OVRL-v2 performance compare with the ImageNav SoTA? OVRL-v2 significantly improves over prior work, including approaches that use additional cameras that provide panoramic views of the environment.

Do the architectural improvements transfer to ObjectNav? Yes. OVRL-v2 outperforms ObjectNav SoTA in terms of SR without even using a depth sensor or segmentation module, as commonly used for ObjectNav.

As a starting point, we use a baseline agent from with a similar architecture to the model-free navigator described in Sec. 4. Instead of the ViT-based visual encoder used in our approach, this baseline uses a half-width ResNet50 with GroupNorm that is trained from scratch. We train this agent with ZER rewards (Eq. 4) and report results in Tab. 1 row 1.

Next, we make two improvements to strengthen this baseline. First, we discover that removing an AvgPool operation that is used in numerous prior works (e.g., ) to downsample images before agent processing, significantly improves performance. Specifically, removing this AvgPool operation improves ImageNav SR by +32.9% absolute and SPL by +9.1% (Tab. 1 row 1 vs. 2)! We do not use this AvgPool in all the following experiments.

Finally, to avoid the potential for reward hacking (discussed in Fig. 4), we switch to the reward function from Eq. 3. In Tab. 1 row 3, we observe that this leads to a small (-0.7%) drop in SR, but a large improvement in SPL of +5.6% absolute (a +20.1% relative improvement). Unless otherwise specified, we use the corrected reward function from Eq. 3 in the remaining experiments. We use this strengthened baseline from Tab. 1 row 3 to study the effects of switching to a ViT-based visual encoder, next.

2 Using ViTs in a Visual Navigation Agent

Negative result. In Tab. 2 (rows 2 and 3), we switch the visual encoder in the baseline agent from a ResNet50 to a ViT-Small. In row 2 we use the [CLS] token representation produced by the ViT to represent images. In row 3, we use global average pooling over the patch representations generated by ViT. As compared with the scratch ResNet50 baseline (row 1), we find that both solutions lead to substantially reduced performance despite the increase in model capacity as measured by number of parameters (50.9M for ViT-Small and 21.5M for ResNet50 agent). Specifically, the better of the two options ([CLS] in row 2), results in a drop of -23.8% in SR and -9.8% in SPL. These results indicate that, unfortunately, ViTs are not drop-in replacements for the ResNets commonly used for visual navigation.

ViTs require Compression layer. In Tab. 2 row 4, we discover that using the compression layer described in Sec. 4 reverses the negative results in rows 2 and 3, improving ImageNav SR by +7.5% and SPL by +3.6% over the ResNet50 baseline in row 1. This suggests that preserving spatial structure (using a compression layer) is critical for visual navigation. Thus, we use compression layer in all remaining experiments.

3 Scaling with and without Visual Pretraining

The encouraging performance of our ViT-Small agent (from Tab. 2 row 4, replicated in Tab. 3 row 1) suggests that further model scaling may lead to additional improvements. Unfortunately, we find that this is not the case. In Tab. 3, we observe that simply switching from the 50.9M parameter ViT-Small (row 1) to the 179.2M parameter ViT-Base (row 2) produces another negative result: SR drops -2.4% while SPL minimally increases by +0.7%.

In contrast, we find that visual pretraining (rows 3 and 4) resolves this issue. Specifically, we pretrain ViTs using MAE (as described in Sec. 3) on the HGSP dataset (details in Appendix A). First, with pretraining we observe a large boost in navigation performance for the ViT-Small agent. Specifically, SR improves by +13.1% and SPL by +18.1% (row 1 vs. 3). Next, we find the negative scaling in rows 1 vs. 2 is reversed. With pretraining, switching from ViT-Small to ViT-Base results in a +1.5% improvement in SR and +3.5% gain in SPL (rows 3 vs. 4) – i.e., we finally see positive results from model scaling. Such positive scaling was not observed in prior works such as .

When compared to our intial baseline, our accumulated improvements are +54.3% in SR and +39.9% in SPL (Tab. 1 row 1 vs. Tab. 3 row 4).

4 Comparing Reward Functions

In Tab. 4, we revisit the reward function used to train our visual navigation agents to confirm that using our corrected reward function (Eq. 3) is indeed necessary. In row 1, we find that switching back to the ZER reward (Eq. 4) results in a substantial performance drop of -5.5% in SR and a -14.6% in SPL. In Fig. 5, we observe that this drop is due in-part to reward hacking. Specifically, after 300M steps of training the agent using ZER rewards (orange) learns to hack the reward. This corresponds to a dramatic increase in the training reward (Fig. 5 left), yet precipitous drops in SPL (middle) and SR (right). Using our fix (blue), OVRL-v2 does not hack the reward and we observe steady improvements in performance over the course of training.

5 Comparisons with the ImageNav SoTA

In Tab. 5, we compare OVRL-v2 (row 8) with prior work on ImageNav (rows 1 - 5)Details about each method is provided in Appendix B. and two baselines: a version of our full approach that uses the CLIP visual encoder (similar to EmbCLIP ) instead of MAE (row 6) and the ViT-Small baseline first presented in Tab. 2 row 4, which uses our full method with the ViT-Small architecture and without MAE pretraining (row 7). We observe that OVRL-v2 outperforms all of these methods, exceeding the next best single camera (1 Cam) method, OVRL (row 4) by +27.8% in SR and +31.7% in SPL. OVRL-v2 also outperforms Mem-Aug Nav , a method that uses 4 cameras providing panoramic views of the environment and goal, by +13.0% in SR and 2.7% in SPL.

Additionally, we observe that the ViT-Small baseline (row 7) is competitive with state-of-the-art methods, achieving the second highest SR and third highest SPL. Finally, the CLIP ViT-Base baseline (row 6) significantly underperforms OVRL-v2 with a drop of -30.3% in SR and -21.3% in SPL. This indicates that pretraining with SSL on in-domain images – and not the large-scale vision-and-language pre-training on out of domain dataset – is a key ingredient in the success of OVRL-v2.

6 Transferring Improvements to ObjectNav

This section presents a comparison of OVRL-v2 with prior works on ObjectNav. In Tab. 6, we find that OVRL-v2 achieves 64.0%64.0\% SR and 29.0%29.0\% SPL on ObjectNav Test-Standard split (row 9). This represents a 4.0% improvement in SR and a 2.0% improvement in SPL over our implementation of OVRL. OVRL-v2 is only 1% behind PIRLNav on SR, the current state-of-the-art approach. PIRLNav is a concurrent work to OVRL-v2 which uses the pretrained ResNet encoder from OVRL and learns a policy first by imitation learning (IL), followed with RL finetuning. We believe the improvements proposed by PIRLNav are orthogonal and OVRL-v2 will further benefit from their RL finetuning strategy.

OVRL-v2 performs better than Stretch in terms of SR by +4.0% (64.0% vs. 60.0% in row 9 vs. row 5). Stretch uses a depth camera and an explicit semantic prediction and mapping module, and outperforms OVRL-v2 in terms of SPL by +5% (34.0% vs. 29.0% in rows 5 vs. 9). Similarly, OVRL-v2 outperforms Habitat-Web (row 2), which uses an explicit semantic predictor, by +9.0% on SR and +7.0% on SPL (rows 2 vs. 9). We also compare to ProcTHOR (row 4), which uses procedural scene generation to pretrain a policy in 10k10k environments, then finetunes on HM3D ObjectNav. ProcTHOR achieves 54.0% SR and 32.0% SPL which is 10% worse on SR and +3.0% better on SPL compared to OVRL-v2. Our approach also improves over OVRL by +4.0% in SR and +2.0% in SPL. We observe that the SPL is lower for all agents that learn from only human demonstrations and hypothesize that OVRL-v2 can benefit from further RL finetuning.

Additionally, we compare OVRL-v2 with other end-to-end methods that do not use pretrained visual encoders. First, we compare with the Habitat-Web RGB only baseline which uses a ResNet backbone and is trained using the same human demonstrations as us. We find that OVRL-v2 is +27.7% better on SR and +14.1% better on SPL, demonstrating the value in using a pretrained visual encoder for IL. We also compare with a DD-PPO RGBD RL baseline (row 1), and find OVRL-v2 is +36.8% better on SR and +13.9% better on SPL.

Analysis and Ablations

This section presents ablations and a failure analysis for our ImageNav agent.

SSL Algorithm. In Tab. 7(a), we use different pretraining algorithms to initialize the visual encoder. We find that Data2Vec (row 3) attains similar performance to the MAE (row 4) initialization we use in Sec. 5. Both methods substantially outperform a variation of OVRL-v2 initialized with CLIP weights and then finetune (row 2) or frozen (row 1) as done in EmbCLIP .

Using the Visual Representations. In Tab. 7(b), we study the impact of using a compression layer with a pretrained visual encoder (augmenting the analysis in Sec. 5.2 without pretraining). First, we find that using global average pooling (row 2) outperforms [CLS] token representation (row 1). Next, we observe that while global average pooling and compression layer perform similarly in terms of SPL (58.7% vs. 59.3%), compression layers lead to a substantially higher SR (82.0% vs. 76.9%). We hypothesize that preserving spatial structure with compression layers is particularly useful for recognizing the goal location (and stopping in a correct place), thus leading to a higher SR.

Visual Encoder Learning Rate. In initial experiments, we found finetuning representations pretrained with MAE or Data2Vec led to overfitting and poor generalization. Thus, we experimented with tuning the learning rate (LR) used specifically for weights of the visual encoder. In Tab. 7(c), we observe that tuning the LR leads to massive improvements in SR of +31.0% and SPL of +20.9% (row 1 vs. 2). In fact, we find finetuning with a bad LR (row 1) is worse than simply freezing the representations (row 3).

Pretraining Dataset. In Tab. 7(e) we compare the effect of pretraining with a dataset (OSD ) that was used in prior work . We observe similar performance with both datasets, indicating that both choices are similarly ‘in-domain’ with respect to the downstream ImageNav task.

Image Augmentations. In LABEL:tab:augmentations, we ablate the use of image augmentations during policy learning. We discover that augmentations play a vital role in preventing overfitting of the pretrained ViT agent. Without augmentations, OVRL-v2’s SR drops by -24.2% and SPL drops by -15.5%.

ImageNav Failures We present qualitative analysis of the failure cases of OVRL-v2 in Fig. 6. We perform this analysis by manually labeling each case with a cause (e.g., ‘exploration failure’). We find that in approximately one-third of the failures the goal-image is not semantically meaningful (‘bad goals’) such as blank walls. These failures account for nearly 6% of the dataset, suggesting the upper bound on ImageNav success is 94%. Other key issues include ‘unexplained stop’ where the reason for stopping is unclear, ‘nearly reached’ when the agent stops short, and ‘looked similar’ finding a location very similar to the goal-image (but incorrect). Additional details in Appendix D.

Several of these failures may be reduced with simple techniques. For example, the ‘nearly reached’ cases may be reduced by using a stronger success criteria during training. Similarly, increasing image resolution may reduce ‘looked similar’ failures as the agent may recognize additional visual cues to distinguish the correct goal. For other cases, such as ‘unexplained stop’ further investigation is required.

Conclusion

In this paper, we demonstrate that a model-free navigation agent (OVRL-v2) composed of task-agnostic components (ViTs, convolutions, and LSTMs) can achieve state-of-the-art results on both ImageNav and ObjectNav. To achieve this, we show that a compression layer operating over ViT patch representations is required, which preserves the spatial information. Finally, we discover that visual pretraining with MAE enables positive scaling trends with larger ViT architectures.

Acknowledgements

The Georgia Tech effort was supported in part by NSF, ONR YIP, and ARO PECASE. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of the U.S. Government, or any sponsor.

References

Appendix A Experimental Details

Pretraining Dataset. For pretraining, we create a dataset using the 800 HM3D and 72 Gibson training scenes. We collect a total of 1.45M RGB images using a camera with a resolution of 512×\times512 and a 90∘ FoV, attached to an oracle agent that navigates in a scene from a random start position to a random goal position using the shortest path finding algorithm. We refer to this dataset as the HM3D-Gibson Shortest Path (HGSP) dataset. We determined the HGSP dataset size (of 1.45M) based on the observation in that pretraining on 10% (1.45M) of the Omnidata Starter Dataset (OSD) leads to the same downstream performance as training on 100% of the dataset totalling 14.5M images.

Pretraining Details. We pretrain visual encoders using MAE using the same hyperparameters as . We pretrain ViT-Small and ViT-Base encoders, initialized from scratch, for 800 epochs on the HGSP dataset. We use these pretrained encoders to initialization the visual encoder for downstream task. We use the same pretrained encoder for both the ImageNav and ObjectNav experiments.

Data Augmentation For downstream tasks (ImageNav and ObjectNav), we use image augmentations during training and evaluation (i.e. at test-time). Specifically, we apply random color-jitter followed by random shifts (i.e. random translations). We apply the same image augmentation across time and to all of the samples on each GPU in our distributed training setup. For ImageNav, we use a color-jitter of 0.3 for brightness, contrast, saturation and hue levels followed by random shifts with a padding of 4 pixels. For ObjectNav, we use a value of 0.4 for color jitter and a padding of 16 for random shifts. These data augmentation settings follow the ImageNav experiments in .

ImageNav Benchmark. We perform ImageNav experiments using a standard dataset released by . This benchmark uses the Habitat simulator and situates the task in the Gibson environments, which include 72 training and 14 validation scenes. The validation set includes 300 episode for each scene (4,200 episodes total). In this benchmark, agents are simulated as a cylinder with a height of 1.5m, radius of 0.1m, and sensors placed 1.25m above the center of the base. The RGB camera has a resolution of 128×\times128 and a 90∘ field-of-view (FoV). Agents can take up to 1000 steps in the environment, and an agent is successful if it calls stop within 1m of the goal position.

ImageNav Training Details. We train agents in the Gibson environments for 500M timesteps (25k updates) using a total of 320 environment running in parallel. Every environment collects (up to) 64 frames of experience which is followed by 2 PPO epochs with 2 mini-batches. Unless specified explicitly, we use a learning rate of 2.5×10−42.5\times 10^{-4} for training the agent and update the parameters using the AdamW optimizer with a weight decay of 10−610^{-6}. We train agents with the reward functions in Eq. 4 and Eq. 3 from the main paper, using the following settings: success weighting cs=5.0c_{s}=5.0, angle success weighting ca=5.0c_{a}=5.0, goal radius rg=1.0r_{g}=1.0, angle threshold θg=25∘\theta_{g}=25^{\circ}, and slack penalty γ=0.01\gamma=0.01. We evaluate performance every 25M steps of training. We report metrics based on the highest success rate (SR) achieved on the validation set.

ObjectNav Benchmark. We conduct ObjectNav experiments using the HM3DSem dataset , which uses the Habitat simulator and HM3D environments. The dataset consists of 80 training, 20 validation, and 20 testing scenes. We report results on the v0.1 HM3DSem val and test-std splits that were used in the 2022 Habitat Challenge ObjectNav benchmark. In this benchmark, the agent models a LocoBot with a height of 0.88m, radius of 0.18m, and sensors placed at the top of the agent’s head. The RGB camera has a 640×\times480 resolution and a 79∘ horizontal FoV. In each episode, agents must find an object drawn from one of 6 categories: ‘chair’, ‘bed’, ‘plant’, ‘toilet’, ‘tv/monitor’, and ‘sofa’. Agents are allowed 500 steps in the environment, and episodes are considered successful if the agent stops within 0.1m of a viewpoint that is (a) within 1m of any instance of an object from the goal category and (b) from which that instance is visible, following the evaluation protocol laid out in .

ObjectNav Human Demonstrations. For training our imitation learning agent, we use the dataset of ObjectNav human demonstrations collected by Habitat-Web for HM3DSem dataset using Amazon Mechanical Turk. The dataset consists of 77k77k human demonstrations for 80 HM3DSem scenes that were released by . For each scene, we have ∼158{\sim}158 episodes for each unique goal object category with a randomly set start location amounting to ∼950{\sim}950 demonstrations per scene. The dataset amounts to a total of ∼12.1{\sim}12.1M steps in experience, with each episode averaging ∼159{\sim}159 steps.

ObjectNav Training Details. We train all ObjectNav agents in the HM3D environment for ∼400{\sim}400M steps (25K updates) using 512 parallel environments. Similar to ImageNav, we use a weight decay of 10-6, a learning rate of 10-4 for the visual encoder and 10-3 LR for everything else with the AdamW optimizer. We evaluate checkpoints after every 10001000 policy updates and report metrics for the checkpoints with the highest validation SPL.

Appendix B ImageNav Baselines

This section describes the state-of-the-art ImageNav methods from prior work listed in Tab. 5 (rows 1−51-5).

Zero Experience Required (ZER) proposes a novel reward function (detailed in Eq. 4) for ImageNav, which is used to train a general purpose agent composed of an ResNet-9 visual encoder (for observations and goals) and a GRU-based policy network that are trained from scratch. Our initial from-scratch ResNet-50 ImageNav baseline presented in Tab. 1 row 1 uses a similar architecture and the same reward function as .

Zero-Shot ObjectNav (ZSON) uses the reward function from to train an ImageNav agent consisting of a pretrained ResNet-50 encoder from for visual observations and a CLIP visual encoder for processing goal-images. The agents in are trained in HM3D and then evaluated in Gibson .

Curious Representation Learning (CRL) uses curiosity-based exploration to collect data for visual representation learning. The visual encoder is then used within a general purpose agent that is finetuned for ImageNav. We report results reproduced by .

Offline Visual Representation Learning (OVRL) pretrains a ResNet-50 encoder using the DINO self-supervised representation learning algorithm on images from the Omnidata Starter Dataset (OSD) . The pretrained ResNet-50 is used within a general purpose agent to process visual observations and goal images while finetuing for ImageNav with the reward function from .

Memory-Augmented RL for ImageNav (Mem-Aug Nav) enhances a general purpose agent with an attention-based memory module. In , agents use four RGB sensors for observations and goals, which provide a panoramic view of the environment. By contrast, our agents and the agents in Tab. 5 (rows 1−41-4) use a single RGB camera. Thus, we highlight results from in gray.

Appendix C Compression Layer Implementation

PyTorch-style pseudocode for creating the compression layer described in Sec. 4 is presented in Algorithm 1. The layer operates on patch representations generated by a ViT-based visual encoder. It reshapes the patches into a grid and then uses a convolutional layer to compress them to a lower dimension. The dimension is determined such that when the compressed patches are subsequently concatenated (i.e., flattened) into a vector the size is approximately 2,048.

Appendix D ImageNav Failure Analysis Details

Here we describe each of the categories used for the ImageNav failure analysis within the main paper (which shows the distribution of these failures) in greater detail:

Bad Goals. We found that a large number goals in the validation dataset face the wall or capture a noisy parts of the scene. These goals are not semantically meaningful – i.e., they either do not ‘describe’ a unique location in the environment or do not provide any indication of where the goal might be found.

Looked Similar. The goal image and agent observation at the stopping time have some visual similarities that may have caused the agent to believe it has reached the goal.

Unexplainable Stop. The reason why the agent called stop is unclear to the annotator.

Nearly Reached. The agent’s distance to the goal is between 1.0m to 1.5m and the agent is looking in the direction of the goal image.

Slightly Far. Distance to goal is farther than 1.5m but the agent is looking towards the goal image.

Exploration Failure. The exploration strategy of the agent fails. For example, when the agent keeps looping in one small region of the environment or stop after reaching an area that is far from the goal.

Didn’t Stop. The agent saw the goal but did not stop.

Navmesh Issues. The agent needs to pass through a narrow passage that is nearly the same size as the agent.

Scene Issues. There is a hole in the scene (i.e., a missing part of the underlying 3D geometry of the environment), which confuses the agent.

Appendix E Additional Ablations

In this section, we present the results of additional ablations conducting on the ImageNav task with randomly initialized and pretrained model-free navigators.

Image Augmentations. In Tab. 8, we ablate the use of image augmentations for a randomly initialized and pretrained policy. We find that in both cases agents benefit from using augmentation during policy learning. Without pretraining, augmentations improve SR by +17.4% and SPL by +5.5% (rows 1 vs. 2). With pretraining, augmentations lead to gains in SR of +24.2% and gains in SPL of +15.5% (rows 3 vs. 4). Additionally, we find that using augmentations without pretraining (row 2) leads to a higher SR of 67.4% than using pretraining without augmentations (row 3), which results in a SR of 57.8%.

Increasing Model Size. The default version of our randomly initialized navigation agent uses a ViT-Small as the vision backbone. In Table 9, we demonstrate that when trained from scratch, ViT-Small agents are more successful than ResNet counterparts, including a full-width ResNet-50 model that uses 8.2M more parameters than the ViT-Small variant. Specifically, we find that using a full-width ResNet-50 leads to +4.5% improvement in SPL compared to the half-width ResNet-50. However, even with fewer parameters, ViT-Small attains 7.4% higher SR, while only being 0.9% worse in terms of SPL.

Appendix F ZER Hacking

Let θt\theta_{t} denote the angle between the agent’s center-of-mass and goal location and cac_{a} is angle success weighting; the ZER reward is written as:

where θg\theta_{g} is an angle success threshold (set to 25∘ in our experiments).

While demonstrated that the reward in Eq. 4 can improve ImageNav performance, it has a subtle but significant flaw: the reward is hackable. The culprit is that the angle-to-goal reward shaping term is not a difference of potential functions as recommended by the theory . Specifically, agents can enter the goal radius (looking away from the goal), accumulate reward by turning towards the goal, exit the goal radius, no longer face any penalty for turning around, and repeat the process.

In Fig. 7, we visualize the reward hacking behaviour of an ImageNav agent trained with the ZER reward. Towards the end of the trajectory near the goal (left), the agent repeatedly enters and exits the goal radius to accumulate the angle-to-goal reward term in Eq. 4 rather than successfully complete the episode.