Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation

Ziad Al-Halah, Santhosh K. Ramakrishnan, Kristen Grauman

Introduction

In visual navigation, an agent must intelligently move around in an unfamiliar environment to reach a goal, using its egocentric camera to avoid obstacles and decide where to go next. As a fundamental research problem in embodied AI, visual navigation has many potential applications—such as service robots in the home or workplace, mobile search and rescue robots, assistive technology for the visually impaired, and augmented reality systems to help people navigate or find objects.

Recent work in computer vision explores visual navigation from many different fronts. In PointNav, an agent is asked to go to a specific position in an unmapped environment (e.g., go to (x,y)(x,y)) . In ObjectNav, the agent must find an object by name (e.g., go to the nearest telephone) . In RoomNav, the agent must find a room (e.g., go to the kitchen) . In AudioNav, the agent must find a sounding target (e.g., find the ringing phone) . In ImageNav, the agent must go to where a given photo was taken . Each case presents a distinct goal to the agent. Accordingly, researchers have pursued task-specific models to treat each one, typically training policies with deep reinforcement learning (RL).

Despite exciting advances, learning task-specific navigation policies has inherent limitations. Training embodied agents from scratch for each new task and relying on special-purpose architectures and priors (e.g., room layout maps for RoomNav, object co-occurrence priors for ObjectNav, directional cues for AudioNav, etc.) requires repeated access to training environments for gathering new agent experience in the context of each task, greatly hindering sample efficiency. Even with today’s fast simulators and photorealistic scanned environments , this typically amounts to days and weeks of computation on a small army of GPU servers to train a single policy. Moreover, by tackling each variant in isolation, agents fail to capture what is common across the tasks. Finally, some tasks require manual annotations such as object labels in 3D space, which naturally limits how extensively they can be trained.

In this work, we challenge the assumption that distinct navigation tasks require distinct policies. Intuitively, finding a good policy for one navigation task should help with the rest. For example, if we know how to find a microwave, then finding a kitchen should be easy too; if we know how to find an object by name, then finding it based on a hand-drawn sketch—or the sounds it emits—should be possible too. In short, it should be beneficial to learn one navigation task and then apply the accumulated experience to many.

To that end, we propose a modular transfer learning approach for semantic visual navigation that enables zero-shot experience learning. See Figure 1. First, we develop a general-purpose semantic search policy. Specifically, using a novel reward and task augmentation strategy, we train a source policy for the image-goal task, where the agent receives a picture taken at some unknown camera pose somewhere in the environment, and must travel to find it. Next, we develop a joint goal embedding that is trained offline (i.e., no interactive agent experience) to relate various target goal types to image-goals. Finally, we address target downstream tasks either by zero-shot transfer with no new agent experience, or by fine-tuning with a limited amount of agent experience on the target task.

Zero-shot learning traditionally focuses on supervised tasks such as image recognition , where models forgo using labeled samples for the new class. Instead, the proposed zero-shot experience learning (ZSEL) focuses on reinforcement learning tasks, where models forgo using interactions in the physical environment for the new navigation task. ZSEL is important for lifelong learning, where an agent will face novel tasks once it is deployed and must solve them while using no or few training episodes.

Using hundreds of multi-room environments from Matterport3D , Gibson , and HM3D , we demonstrate our approach for four challenging tasks and goals expressed with five different modalities—images, category names, audio, hand-drawn sketches, and edgemaps. Our ImageNav results advance the state of the art, and our modular transfer approach outperforms the best existing methods of transfer based on self-supervision, supervision, and RL. Finally, our ZSEL performance on 55 semantic navigation tasks is equivalent to 507507 million interactions required by task-specific policies learned from scratch.

Related Work

Traditional methods in visual navigation often rely on mapping the 3D space and then planning their movements . However, fueled by fast simulators and large-scale photorealistic datasets there have been great advances in learning-based navigation approaches leading to a near-perfect agent for tasks like point-goal navigation . In this work, we consider semantic visual navigation, where the agent is given a semantic description of the goal (e.g., object-goal , image-goal , room-goal , audio-goal ) but, unlike point-goal, the goal location is unknown. Hence, the agent needs to leverage learned scene priors to explore the environment efficiently to find and navigate to the target. Current approaches tackle each navigation task separately: a new model is trained for each task and each target modality , which has the disadvantages discussed above. In contrast, we propose a unified approach to semantic visual navigation, in which a single trained policy can handle diverse tasks and goal modalities.

Transfer Learning in Navigation

Pre-learning a representation from large-scale image datasets and transferring it to a downstream task proved to be very successful for visual recognition . We observe a similar trend in embodied navigation, where pre-learning good representations of the 3D environment or primitive skills help the agent to learn a downstream task better while using fewer training samples. Recent methods focus on pre-training the observation encoder of the agent, either in a supervised or self-supervised fashion. While this leads to improved performance on the target tasks, a new policy is still learned from scratch for each task, resulting in low sample efficiency. In contrast, our approach enables a full transfer paradigm where all the agent’s components can be reused efficiently on the downstream tasks. Prior work shows that transferring a strong point-goal policy to non-goal driven tasks (e.g., flee and exploration) can lead to better performance . Differently, we propose to learn and transfer a general-purpose semantic search policy. Our policy can find semantic goals presented in different modalities for a diverse set of goal-driven navigation tasks.

Sharing knowledge between multiple tasks can be achieved in a multi-task learning setup where all tasks are learned jointly in a supervised manner, or via meta-RL where a meta policy learned from a distribution of tasks is finetuned on the target. Unlike these methods, our policy is learned from one task that does not require manual annotations, and it can be transferred in a zero-shot setup where the policy does not receive any interactive training on the target.

Zero-Shot Learning

Zero-shot learning (ZSL) can be seen as an extreme case of transfer learning where the target task has zero training samples. Prior ZSL work focuses on supervised learning, e.g., image classification . In contrast, the proposed zero-shot experience learning (ZSEL) setup learns behaviors rather than classifiers; the policy learned on a source task needs to perform a set of target tasks, without receiving any new interactive experiences on the target. Further, unlike where a world model for synthetic environments is constructed, and control policies are trained on ‘imagined’ episodes, we consider a model-free approach and a ZSEL setup in realistic environments where the policy receives zero target interactions (i.e., neither imagined nor real). To our knowledge, we are first to propose a ZSEL model for embodied navigation.

Plug & Play Modular Transfer Learning

We introduce a novel transfer learning approach for visual navigation. Our model has three main components: 1) we start by learning a semantic search policy for image goals using a novel reward and task augmentation (Fig. 2a); 2) we leverage the image goal encoder to learn a joint goal embedding space for the different goal modalities (Fig. 2b); and finally, 3) we transfer the learned agent modules to downstream tasks in a plug and play fashion (Fig. 2c).

In the following, we consider an agent with 33 main modules: 1) an observation encoder (fOf_{O}) that encodes the received observations oto_{t} from the environment; 2) a goal encoder (fGf_{G}) that encodes the task’s goal; 3) a policy (π\pi) that uses the output of fOf_{O} and fGf_{G} to navigate and find the goal.

The policy is a key component in the modern end-to-end visual navigation agent. It guides the agent towards solving a task given a set of sequential observations and a goal. Such policies are often learned with reinforcement learning (RL) where the agent interacts with its environment (by moving about) and attempts to solve the task in a trial and error fashion. If the agent succeeds in its attempt, then it receives a reward to encourage such behavior from the policy in the future.

A main challenge for this learning paradigm is that the policy requires a large number of interactions with the environment in order to find a proper way to solve the task. This usually amounts to tens and hundreds of millions and up to billions of interactions, and correspondingly days or weeks of GPU cluster time. Furthermore, for each new task a policy is typically learned from scratch, which further increases the learning cost substantially.

We propose to learn a general-purpose semantic search policy that can be transferred and perform well on a variety of navigation tasks. Our idea is to learn such a policy with the image-goal task, where the agent receives a picture taken at some unknown camera pose somewhere in the environment, and must travel to find it. Our choice of image-goal for the source policy is significant. It requires no manual annotations, and image-goals can be sampled freely anywhere in a training environment. As a result, the policy can be trained on large-scale experience (e.g., collected from a fleet of robots deployed in various environments) which can improve its generalization to new tasks and domains. Furthermore, an image-goal encourages the learned policy to capture semantic priors for finding things in a 3D space. For example, by seeking images of couches and chairs, the agent learns implicitly to leverage these objects’ context and the room layout in order to find the image views effectively.

In an episode of image goal navigation, the agent starts from a random position p0p_{0} in an unexplored scene, and it is tasked to find a certain location pGp_{G} given an image IGI_{G} sampled with the camera at pGp_{G}. The agent receives an RGB observation oto_{t} at each step tt and needs to perform the best sequence of actions at∈a_{t}\in {move_forward,turn_left,turn_right,stop}\{\texttt{move\_forward},\texttt{turn\_left},\texttt{turn\_right},\texttt{stop}\} that would bring it to the goal within a maximum number of steps SS. Unlike the common point-goal task where the goal location is known , here pGp_{G} is unknown, and the agent needs to leverage the learned semantic priors to search and find where IGI_{G} could have been sampled from.

Our setup differs from recent methods in ImageNav where panoramic FoV sensors are required . Here, we consider a standard FoV for the agent’s view . While having a complete FoV sensor simplifies localization, this strong requirement is often not available in common robotic platforms and leads to high computational cost. This reduces the scalability and adoption of such methods by diverse agent configurations. In addition, our task setup allows our model to transfer to a diverse set of semantic navigation tasks in a plug and play fashion without the need for modifications to the target tasks (for which the literature does not use panoramic images).

View Reward

It is common to use the reduced distance to the goal to reward the agent for getting closer to pGp_{G} in addition to the success reward of finding and stopping within a small distance dsd_{s} of pGp_{G}. However, while this reward proves to be quite successful for navigation tasks like point-goal, we argue it is less suited for semantic goals like images. Since the reward does not carry a signal about the semantic goal itself, the agent may fail or require much more experience in order to capture the implicit relation between the goal and the distance to goal reward (DTG). For example, if the goal shows an image of an oven, the agent may get close and stop nearby while looking at a book on the counter, and nonetheless receive a full success reward. This may lead to capturing trivial or incoherent associations between the goal and the agent’s observations.

In order to encourage the agent to leverage the information provided in the goal description IGI_{G} and effectively capture useful semantic priors that may help it in finding pGp_{G}, we propose a new reward function that rewards the agent for looking at IGI_{G} when getting closer to pGp_{G}, so it can better draw the association between its oto_{t} and IGI_{G}. Specifically, we define the reward function at step tt as:

where rdr_{d} is the reduced distance to the goal from the current position relative to the previous one, rαr_{\alpha} is the reduced angle in radians to the goal view from the current view relative to the previous one, [⋅][\cdot] is the indicator function, and γ=0.01\gamma=0.01 is a slack reward to encourage efficiency. Note, this reward will encourage the agent to look at IGI_{G} when it gets near the goal, since it is rewarded to reduce the angle between its current view vtv_{t} and the view of the goal vGv_{G} (see Fig. 2a). Finally, the agent receives a maximum success reward of 1010 if it reaches the goal and stops within a distance of dsd_{s} from pGp_{G} and an angle αs\alpha_{s} from vGv_{G}:

View Augmentation

In addition to the view reward introduced above, we also provide a simple task augmentation method to promote generalization by increasing the diversity of goals presented to the agent. For each training episode, rather than having a fixed IGI_{G}, we sample a view from a random angle at location pGp_{G} and provide the agent with the associated IGI_{G} from the sampled view as the goal descriptor. This has a regularization effect on the model learning; the agent will be less likely to overfit due to the changing goal description each time the agent experiences a given goal. Furthermore, with the start p0p_{0} and the goal location pGp_{G} fixed in a training episode but not IGI_{G}, this encourages the agent to capture implicit spatial semantic priors of things that usually appear near each other as viewed from pGp_{G}. For example, the agent would learn that an image of chairs as seen when peeking from the door is likely to be at the same location of the current image goal that is showing a dining table, since the agent experienced the same episode before but with IGI_{G} showing the chairs, hence prompting the agent to explore the dining room.

Policy Training

We train our policy using reinforcement learning (RL), Fig. 2a. For each training episode, we sample an image-goal IGI_{G} from pGp_{G}. The agent encodes its current observation oto_{t} (an RGB image) with fOf_{O} and the image-goal with fGIf_{G}^{I} and passes these encodings to the policy π\pi. The policy further encodes these information along with the history of observations so far to produce a state embedding sts_{t}. An actor-critic network leverages sts_{t} to predict state value ctc_{t} and the agent’s next action ata_{t}. Based on the agent’s state in the environment, it receives a reward (Eq. 1 and Eq. 2). The model is trained end-to-end using PPO .

2 Joint Goal Embedding Learning

Having learned the semantic search policy, we can now transfer our model to downstream tasks. Specifically, we consider downstream navigation tasks where the goals are object categories (ObjectNav ), room types (RoomNav ), or view encodings (ViewNav), and they may be expressed by the modalities of a label name, a sketch, an audio clip, or an edgemap; see Sec. 4.2. A key advantage of learning the semantic search policy using RGB image-goals is that these goals contain rich information about the target visual appearance and context. Furthermore, in order to solve the image-goal navigation task, our model learns to encode these visual cues via a compact dense representation produced by the image-goal encoder fGIf_{G}^{I}.

Our idea is to leverage fGIf_{G}^{I} to learn a joint embedding space of different goal modalities for the various tasks. In other words, we upgrade the image-goal embedding space to be a joint goal embedding space to draw associations between the images and the different goal modalities like sketches, category names, and audio (Fig. 2b). This step can be carried out quite efficiently and using an offline dataset. For example, to learn about object-goals that are represented with a label (e.g., a chair) we only need to annotate a set of images with chairs. Then we train an object-goal encoder to produce an embedding similar to the image-goal encoder for compatible image-label pairs. In our experiments, we use offline datasets of size 2020K images or less in which the “annotations” are actually automatic object detections. This is several orders of magnitude smaller than the amount of interactions usually needed to train a target-specific policy (tens to hundreds of millions) .

Formally, let D={(xi,gi)}D=\{(x_{i},g_{i})\} be a set of images xix_{i} and their associated goals gig_{i}, where gig_{i} can be of any goal modality (e.g., audio, sketch, image, category name, edgemap) depending on the downstream task specifications. We learn a joint goal embedding space by minimizing the loss:

During goal embedding learning, we freeze fGIf_{G}^{I} and learn fGMf_{G}^{M} using Eq. 3 such that fGMf_{G}^{M} learns to encode its goal similar to the corresponding image embedding from fGIf^{I}_{G}.

3 Transfer and Zero-Shot Experience Learning

Having learned the semantic search policy and the joint goal embedding as described above, now we can transfer our model to downstream navigation tasks (Fig. 2c). For that, we only need to replace fGIf_{G}^{I} with the suitable goal encoder for the task, such that:

Our plug and play modular transfer approach has multiple advantages. Since all modules are compatible with each other, this means the model can perform the target task out of the box, i.e., it does not require any further task-specific interactions to solve the target task. We refer to this setup as zero-shot experience learning (ZSEL). Training policies with modern RL frameworks is the most expensive part of the model learning, and with ZSEL we manage to circumvent this requirement. Furthermore, due to the modular nature of our approach, it is easy to generalize to a wide variety of tasks and goal modalities. For a new task, only the respective goal encoder is trained with an offline dateset then integrated in the full model in a plug&play fashion.

Finally, our model can be easily finetuned for the downstream task to capture any additional cues specific to the task to reach a better performance. Unlike the common approach in the literature where only fOf_{O} is pretrained and transferred , here the full model is transferred to the target task. This leads to higher initial performance, faster convergence, and better overall performance, as we will show in Sec. 4.

Evaluation

In the following experiments, we first evaluate our semantic search policy performance in the source task (image-goal navigation) compared to state-of-the-art methods (Sec. 4.1); then we show how our model transfers to a diverse set of downstream navigation tasks (Sec. 4.2).

For fair comparisons, we adopt the same architecture and training pipeline for our model and all RL baselines, and we note any deviations from this shared setup in the respective sections. We use a ResNet9 for fOf_{O} and a GRU of 22 layers and embedding size of 128128 for π\pi. For goal encoders fGf_{G}, we use a ResNet9 to encode image-, sketch-, edgemap- and audio-goal modalities. We transform an audio clip to a spectrogram before encoding it by fGf_{G}. If the goal is a category name, we use a 22 layer MLP for fGf_{G}. We train the policy using DD-PPO and allocate the same computation resources for all models. We use input augmentation (random cropping and color-jitter) during training to improve the stability and performance of the RL methods . See Supp for more details. We adopt end-to-end RNN-based RL since it is a common , generic architecture, does not require hand-crafted modules, shows good performance on real-data , and learned in sim has the potential to generalize well to real . However, our contributions are orthogonal to the RL architecture used.

Agent Configuration

1 Image-Goal Navigation

Datasets

We use the Habitat simulator and the Gibson environments to train our model. We use the dataset from . The training split contains 99K episodes sampled from each of the 7272 training scenes. Following the setup from , all RL models are trained for 5050K updates (500500 million frames) on the training split. The test split has 4.24.2K episodes sampled uniformly from 1414 disjoint (unseen) scenes. For direct comparison with , we also test our model on a second split (“split B”) provided by that has 33K episodes and the same structure as the test split from (“split A”).

Baselines

We compare our image-goal model to the following baselines and SoTA methods: 1) Imitation Learning: This model’s policy is trained using supervised learning to predict the ground truth best action on the shortest path to the goal given its current observation. 2) Zhu et al. : The model uses a ResNet50 shared between fOf_{O} and fGf_{G}, pretrained on ImageNet and frozen. 3) Mezghani et al. : This is the SoTA panoramic image-goal navigation model. It uses a ResNet18 for fOf_{O} and fGf_{G}, a 22 layer LSTM for π\pi, and a specialized episodic memory. We adapt this model to our FoV for oto_{t} and IGI_{G} and train it using the author’s code. 4) DTG-RL: This model uses the shared architecture along with the common distance to goal dense reward for training. 5) Hahn et al. : This model learns from a passive dataset of videos collected from the Gibson training scenes and uses a customized architecture based on topological maps (see for details).

Results and Analysis

Table 4 reports the overall performance in terms of average success rate (Succ) and Success weighted by inverse Path Length (SPL) over 33 random seeds. Our model outperforms strong baselines and the SoTA in image-goal navigation by a significant margin. In split A, our model gains +6.6%+6.6\% in Succ and +2.6%+2.6\% in SPL over the best baseline. The method designed for panoramic sensors tends to underperform in this challenging setting. We see a drop in Succ from 69% to 9% when using and FoV, respectively, since such methods rely heavily on the FoV for accurate localization. In split B, our model gains +9%+9\% in Succ and +11.2%+11.2\% in SPL over . It is important to note that the model from uses a much more complete sensor configuration than our method (pose sensor, RGB and Depth sensors of 480×640480\times 640 resolution, and a FoV) and it is trained offline from passive videos sampled from the simulator. Nonetheless, our model outperforms by a large margin, showing that interactive learning of end-to-end RL models still has an advantage over heuristic and passive approaches.

Ablations

To validate our contributions from Sec. 3.1, we test our model performance when removing the view reward or the view augmentation. As shown in Table 4, we see a degradation in performance whenever one of these components are removed, and the largest gain is realized when they work in tandem. Additionally, we test our model under noisy actuation. While methods in split A do not provide results under noisy conditions, does. Following the setup from , we use the noise model from that simulates actions learned from a Locobot . Our model shows robustness to noise and maintains its advantage over the baselines (Table 4 bottom).

Cross-Domain Generalization

Next we test the models trained on Gibson on datasets from Matterport3D (MP3D) and HM3D . In addition to the visual domain gap between these datasets, MP3D has more complex and larger scenes than Gibson, and HM3D has high diversity in terms of scene types. This poses a very challenging cross-domain evaluation setting. The test split from each dataset has in total 33K episodes sampled uniformly from 100100 and 1818 scenes for HM3D and MP3D, respectively. Table 2 shows the results. Overall, we see a drop in performance for all models in this challenging setting, especially for HM3D since there is high diversity in the 100100 val scenes. Nonetheless, our model outperforms all baselines on both datasets, showing that our contributions lead to better generalization by encouraging the agent to pay closer attention to the semantic information provided by the goal.

2 Transfer to Downstream Tasks

We consider 33 target tasks and 44 goal modalities:

2) RoomNav: The agent is tasked with finding the nearest room of 66 types: living-room, kitchen, bedroom, office, bathroom, and dining-room. The goal is a label (find an office) and the episode is successful if the agent steps inside the room with the maximum episode length S=500S=500 .

Datasets

For all target tasks, we use 2424 train / 55 test scenes from the Gibson tiny set that has semantic annotations . These scenes are disjoint from those used in Sec. 4.1. We train all methods for up to 2020 million steps on the target tasks and report evaluation performance averaged over 33 random seeds. See Supp for details.

To train the goal embedding for the modalities in ObjectNav and RoomNav, we sample 1414K images of objects and 2020K of rooms from the training scenes. We use the object labels generated by a model from to draw the associations between the images and each modality. For ViewNav, we sample 170170K views from the training scenes and generate their edgemaps using an edge texture model . The number of samples for the offline datasest is driven by the available instances of each goal type in the training scenes. While there is a finite number of rooms and objects, we can sample views freely from any location in a scene.

Baselines

We compare our model to a set of baselines and SoTA models in transfer learning: 1) Task Expert which learns from scratch on the downstream task. 2) MoCo v2 initializes fOf_{O} using MoCo training on ImageNet (IMN) or on a set of images randomly sampled from the Gibson (Gib.) training scenes. 3) CRL pretrains fOf_{O} (ResNet50) using a combination of curiosity-based exploration and self-supervised learning. We initialize fOf_{O} from a pretrained model provided by the authors. 4) Visual Priors uses a set of 44 ResNet50s pretrained encoders as fOf_{O}. The encoders are trained in a supervised manner to predict 44 features (e.g. semantic segmentation, surface normal) that provide maximum coverage for downstream navigation tasks . 5) Zhou et al. transfers 22 pretrained ResNet50s for depth prediction and semantic segmentation; however, unlike , these are used with an RGB encoder (ResNet9) trained from scratch. 6) SplitNet pretrains fOf_{O} (customized CNN) using a mix of 66 auxiliary tasks (motion and visual tasks) and point-goal navigation. We initialize fOf_{O} from a pretrained model provided by the authors. 7) DD-PPO (PN) pretrains the model for point-goal navigation (PN) and both fOf_{O} and π\pi are transferred.

Transfer Learning

Table 3 shows the results. Our approach outperforms all baselines by a significant margin. Interestingly, the self-supervised methods reach a competitive performance to those that rely on the availability of dense annotations (like semantic segmentation and ground truth depth) for supervised representation learning . Furthermore, methods that learn a curiosity-based representation (CRL ) or via auxiliary tasks and RL (SplitNet ) do not transfer as well as the SSL and SL methods. Additionally, compared to the strong DD-PPO approach which was trained on the same data as our policy but for the PointNav task, our model achieves substantial gains in success rate (from +5%+5\% and up to +14%+14\%) across all tasks. This indicates that our semantic search policy is much better suited to transfer to diverse downstream tasks compared to the PointNav policy. Moreover, when looking at the test performance over the course of training for the best transfer methods compared to ours (Fig. 3), we notice that our approach has a much higher start and improves faster to better performance. Our model reaches the top performance of the best competitor up to 12.5×12.5\times faster.

Zero-Shot Experience Learning

A unique feature of our approach is its ability to perform the downstream task without receiving any new experiences from it. Our model shows excellent performance under the challenging ZSEL setting. Our ZSEL model outperforms the Task Expert in 44 out 55 of the tasks despite receiving zero new experience on the target, and even after training the Task Expert for up to 2020 million steps (Ours-ZSEL in Table 3).

In addition, we see in Fig. 3 that the majority of transfer learning models struggle to reach our ZSEL performance. Note that our model does not have any advantages in terms of architecture, which is shared with the rest of the models. Thus, the high ZSEL performance is attributable to our modular transfer approach. In ObjectNav and RoomNav, the best competitor requires between 22 to 1616 million steps to reach our ZSEL performance, and with the exception of ObjectNav-Audio the competitors show little improvement over that level. In ViewNav, we notice that none of the baselines are capable of reaching our ZSEL level. This can be attributed to the challenging goal modality where estimating distances for successful stopping is difficult, and to the close proximity of this task to the source ImageNav task that our semantic policy is most familiar with.

Modular Transfer Ablation

Fig. 4 (left) shows a modular ablation of our approach on the ObjectNav-Label task. Transferring individual modules separately has mixed impact on performance. While transferring fGf_{G} and π\pi only does not improve over the ‘No Transfer’ case, fOf_{O} leads to positive transfer effect. This is expected since in this model fOf_{O} is a deep CNN with the largest portion of parameters. Having a good initialization of this component is beneficial. Nonetheless, when combining the modules together with our plug and play modular approach we see substantial gains. Our full model demonstrates the best performance and enables ZSEL, thus validating our contributions.

Scalability

We evaluate our model’s ability to scale in terms of experience gathered on the source task. We find a strong correlation between the experience gathered on the source task and the ZSEL performance on downstream ones. As our semantic search policy receives more experience on the source task (ImageNav), its ZSEL performance on the target task gets better (Fig. 4 right). This is important since our source task requires no annotations and can be easily scaled to more scenes and large datasets. For an analysis of our model in terms of the used sensors, see Supp.

Long-Term Task Expert Training

We saw above that our model scales well and its transfer improves when more experience is gathered on the source task. However, does a task expert become competitive if it simply gets longer training on the target task? How long does that model take to catch up with our approach? To find out, we train the Task Expert on each of the target tasks for up to 500500M steps. Fig. 5 shows the results. The Task Expert requires on average more than 2222M steps on ObjectNav and RoomNav, and up to 416416M on ViewNav (in total 507507M steps over the 55 tasks) to reach our ZSEL performance. It never reaches our model’s top performance when our model is finetuned on the target task. Moreover, our model reaches the best performance of the Task Expert 34.7×34.7\times faster. The Task Expert needs task-specific experience with task-specific annotations, which can be expensive and limits the available training data. In contrast, our model learns in the source task using more diverse goals that can be sampled randomly from the (unannotated) scene, thus scaling more effectively.

Additional Results and Discussions

Please see Supp for qualitative results, an analysis of failure cases, and a discussion of limitations and the societal impact of our approach.

Conclusion

We introduce a plug&play modular transfer learning approach that provides a unified model for a diverse set of semantic visual navigation tasks with different goal modalities. Our semantic search policy outperforms the SoTA in the source task of image-goal navigation, as well as the SoTA in transfer learning for visual navigation by a significant margin. Furthermore, our model is able to perform new tasks effectively with zero-shot experience—to our knowledge, a completely new functionality for visual navigation. This is a stepping stone for future work, especially for tasks with high-cost training data. Being able to do ZSEL and learn from few experiences is a crucial skill for an agent in open-world and lifelong learning settings.

Acknowledgements: UT Austin is supported in part by DARPA L2M, the UT Austin IFML NSF AI Institute, and the FRL Cog Sci Consortium. K.G is paid as a Research Scientist by Meta AI. Thanks to Lina Mezghani for providing access to data and code.

References

Supplementary Materials

Additional information presented in this supplementary:

Details on the shared implementation (Sec. 7).

Details of the image-goal navigation dataset (Sec. 8).

Detailed results on image-goal navigation across 33 levels of episode difficulties (Table 4).

Examples of the visual goal modalities used in target tasks (Fig. 6).

Dataset details for the target tasks and goal embedding space (Sec. 9).

Detailed results with standard deviations on target tasks (Table 5).

Qualitative results for our model in all tasks and goal modalities (Fig. 10).

Examples of failure cases for our model (Fig. 11).

Performance curves for all tasks and goal modalities in transfer learning setup (Fig. 7) and long-term training of Task Expert (Fig. 8).

Ablation on the sensor configuration used by the agent (Fig. 9)

Discussion of potential societal impact (Sec. 10) and limitations (Sec. 11).

Shared Setup

All RL methods are trained with the following setup. We use input augmentation of random cropping and color jitter for both observations and goals. The models are trained with DD-PPO . We set the number of PPO epochs 22, the forward steps 128128, the entropy coefficient 0.010.01, clipping of 0.20.2, and train the model end-to-end using the Adam optimizer . We allocate the same number of processes and resources to all methods.

We use the Habitat simulator along with the Gibson , Matterport3D , and HM3D datasets. These datasets are photorealistic and scans of real-world environments with varying complexities, sizes, room layouts, types. In all our experiments, the test scenes are disjoint from those used for training to assess the agent ability to generalize to previously unseen environments.

Image-Goal Navigation

Detailed Results

We show in Table 4 the detailed results of all models across the three levels of episode difficulties (easy, medium and hard). Our model shows better performance across the different levels and in both split A and B. For a qualitative result, see Fig. 10 A.

Transfer Learning to Downstream Tasks

We use 2929 scenes from Gibson and split them into 2424 scenes for training and 55 for testing. In the following, we present the details of the datasets used for each of the downstream tasks.

ObjectNav: We sample 2424K episodes for training and 11K episodes for testing. For the sketch-goals (Fig. 6 middle), we sample 8080 sketches from for each object category and split them to 7070 used during training and 1010 for testing. For the audio-goals, we sample 1212 audio clips from of lengths ranging from 1313 to 5353 seconds and split them 50/5050/50 for training and testing. At the start of each episode of the ObjectNav (Audio) task a random 44 seconds duration is sampled from the respective audio clip and split, and presented to the agent as the goal descriptor.

RoomNav: We sample 2525K episodes for training and 290290 for testing.

ViewNav: We sample 2424K episodes for training and 1.51.5K for testing. We generate edgemaps for 300300 random views per scene, and we randomly assign one of those per episode as the goal (Fig. 6 bottom).

Goal Embedding Space

In the joint goal embedding space, we aim to learn goal encoders that are compatible to the image-goal encoder. For example, an image view from a living room with a TV detected in it will be used as the positive anchor for a sketch of a TV, a sound clip from a TV, the TV label, the living room label, and the edgemap of the view. The annotations for the sampled image view are based on model predictions from . During training, the parameters of fGIf_{G}^{I} are kept frozen, and we train the various goal encoders defined in Main/Sec.4 using the loss from Main/Sec.3.2.

Detailed Results for All Tasks

In Table 5 we show the average success rate and standard deviation for all methods over 33 random seeds. Fig. 7 shows the performance of the best transfer learning methods and our approach across all tasks and goal modalities. Fig. 8 shows the Task Expert performance when trained for up to 500500M steps on each of the respective tasks and in comparison to our model performance under the ZSEL setting or when it is finetuned. Furthermore, we show example navigation episodes from all tasks and goal modalities for our approach in Fig. 10. Our plug and play modular transfer learning approach enable our model to perform a diverse set of tasks effectively.

Scalability (Sensors)

We evaluate our model’s ability to scale across the sensor suite. Fig. 9 shows our model performance when varying the sensors’ configuration in the source task (ImageNav) and evaluating on ObjectNav (Label) under the ZSEL setup. As expected, when enriching the agent sensors to include depth and pose sensors in addition to vision, we see an additional improvement in performance. More importantly, when increasing the vision sensor resolution from 128128 to 256256 we see a significant bump in ZSEL success rate that exceeds the one from diversifying the sensory suite. Our model seems to benefit from an enhanced vision channel as it carries the important semantic cues needed for our semantic search policy and goal embedding space.

Failure Cases

We show in Fig. 11 few examples of failure cases encountered by our model. We notice that some of these failure cases are related to the type of the goal modality. For example, in ImageNav the agent sometimes finds the object described in the image however misestimate the view point the image is taken from, hence stops a bit far from the goal location (Fig. 11 A). In ViewNav (Edgemap), the goal modality lacks distinctive texture and color information which leads the agent to sometime stops at a location with similar edge structure, but it is actually not the goal (Fig. 11 B). A type of failure cases spotted in multiple tasks are the early stopping cases. In these cases, the agent fails to estimate the distance to the goal correctly and stops early resulting in an unsuccessful episode (Fig. 11 C).

Potential Societal Impact

Our approach’s application domain is semantic visual navigation. Here, autonomous agents are trained to find semantic objects in a 3D environment. Such a technology can have positive societal impact by improving people’s life, especially in domains like elder care, with robots that can aid in daily life tasks (e.g. find my keys, go to the bedroom and bring me my medicine). On the other side, the datasets used in this study are 3D scans of building and houses from certain geographic and cultural areas (western style houses from well-off areas). This creates certain biases in the type of building architectures, room, and object types the agent is familiar with. Consequently, this may limit the availability of this technology to a small section of the population. More diverse datasets and methods with robust adaptation to strong shifts in building layouts and object types are needed to mitigate these effects.

Discussion and Limitations

We propose a novel approach for modular transfer learning that enables the agent to handle multiple tasks with diverse goal modalities effectively. Our model can solve the downstream tasks out-of-the-box in zero-shot experience learning setup alleviating the need for expensive interactive training of the policy. Alternatively, our model can be finetuned on the downstream task to learn task-specific cues where it showed to learn faster, generalize better and reach higher performance than the baselines. While we focused in this work on semantic navigation tasks, this can be seen as a first step in this exciting direction. Additional research is needed to generalize this method to tasks that require a series of goals and a compatible policy that can plan effectively in a multi-goal setup (e.g. VLN ). Further, our results and evaluation demonstrate strong transfer learning performance for our method. However, as usual in transfer learning, there is not a theoretical guarantee that a transfer effect will always be beneficial. Target tasks with significant differences to the source task may not benefit from transferring the accumulated experience in the source.