Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual Navigation
Ziad Al-Halah, Santhosh K. Ramakrishnan, Kristen Grauman
Introduction
In visual navigation, an agent must intelligently move around in an unfamiliar environment to reach a goal, using its egocentric camera to avoid obstacles and decide where to go next. As a fundamental research problem in embodied AI, visual navigation has many potential applications—such as service robots in the home or workplace, mobile search and rescue robots, assistive technology for the visually impaired, and augmented reality systems to help people navigate or find objects.
Recent work in computer vision explores visual navigation from many different fronts. In PointNav, an agent is asked to go to a specific position in an unmapped environment (e.g., go to ) . In ObjectNav, the agent must find an object by name (e.g., go to the nearest telephone) . In RoomNav, the agent must find a room (e.g., go to the kitchen) . In AudioNav, the agent must find a sounding target (e.g., find the ringing phone) . In ImageNav, the agent must go to where a given photo was taken . Each case presents a distinct goal to the agent. Accordingly, researchers have pursued task-specific models to treat each one, typically training policies with deep reinforcement learning (RL).
Despite exciting advances, learning task-specific navigation policies has inherent limitations. Training embodied agents from scratch for each new task and relying on special-purpose architectures and priors (e.g., room layout maps for RoomNav, object co-occurrence priors for ObjectNav, directional cues for AudioNav, etc.) requires repeated access to training environments for gathering new agent experience in the context of each task, greatly hindering sample efficiency. Even with today’s fast simulators and photorealistic scanned environments , this typically amounts to days and weeks of computation on a small army of GPU servers to train a single policy. Moreover, by tackling each variant in isolation, agents fail to capture what is common across the tasks. Finally, some tasks require manual annotations such as object labels in 3D space, which naturally limits how extensively they can be trained.
In this work, we challenge the assumption that distinct navigation tasks require distinct policies. Intuitively, finding a good policy for one navigation task should help with the rest. For example, if we know how to find a microwave, then finding a kitchen should be easy too; if we know how to find an object by name, then finding it based on a hand-drawn sketch—or the sounds it emits—should be possible too. In short, it should be beneficial to learn one navigation task and then apply the accumulated experience to many.
To that end, we propose a modular transfer learning approach for semantic visual navigation that enables zero-shot experience learning. See Figure 1. First, we develop a general-purpose semantic search policy. Specifically, using a novel reward and task augmentation strategy, we train a source policy for the image-goal task, where the agent receives a picture taken at some unknown camera pose somewhere in the environment, and must travel to find it. Next, we develop a joint goal embedding that is trained offline (i.e., no interactive agent experience) to relate various target goal types to image-goals. Finally, we address target downstream tasks either by zero-shot transfer with no new agent experience, or by fine-tuning with a limited amount of agent experience on the target task.
Zero-shot learning traditionally focuses on supervised tasks such as image recognition , where models forgo using labeled samples for the new class. Instead, the proposed zero-shot experience learning (ZSEL) focuses on reinforcement learning tasks, where models forgo using interactions in the physical environment for the new navigation task. ZSEL is important for lifelong learning, where an agent will face novel tasks once it is deployed and must solve them while using no or few training episodes.
Using hundreds of multi-room environments from Matterport3D , Gibson , and HM3D , we demonstrate our approach for four challenging tasks and goals expressed with five different modalities—images, category names, audio, hand-drawn sketches, and edgemaps. Our ImageNav results advance the state of the art, and our modular transfer approach outperforms the best existing methods of transfer based on self-supervision, supervision, and RL. Finally, our ZSEL performance on semantic navigation tasks is equivalent to million interactions required by task-specific policies learned from scratch.
Related Work
Traditional methods in visual navigation often rely on mapping the 3D space and then planning their movements . However, fueled by fast simulators and large-scale photorealistic datasets there have been great advances in learning-based navigation approaches leading to a near-perfect agent for tasks like point-goal navigation . In this work, we consider semantic visual navigation, where the agent is given a semantic description of the goal (e.g., object-goal , image-goal , room-goal , audio-goal ) but, unlike point-goal, the goal location is unknown. Hence, the agent needs to leverage learned scene priors to explore the environment efficiently to find and navigate to the target. Current approaches tackle each navigation task separately: a new model is trained for each task and each target modality , which has the disadvantages discussed above. In contrast, we propose a unified approach to semantic visual navigation, in which a single trained policy can handle diverse tasks and goal modalities.
Transfer Learning in Navigation
Pre-learning a representation from large-scale image datasets and transferring it to a downstream task proved to be very successful for visual recognition . We observe a similar trend in embodied navigation, where pre-learning good representations of the 3D environment or primitive skills help the agent to learn a downstream task better while using fewer training samples. Recent methods focus on pre-training the observation encoder of the agent, either in a supervised or self-supervised fashion. While this leads to improved performance on the target tasks, a new policy is still learned from scratch for each task, resulting in low sample efficiency. In contrast, our approach enables a full transfer paradigm where all the agent’s components can be reused efficiently on the downstream tasks. Prior work shows that transferring a strong point-goal policy to non-goal driven tasks (e.g., flee and exploration) can lead to better performance . Differently, we propose to learn and transfer a general-purpose semantic search policy. Our policy can find semantic goals presented in different modalities for a diverse set of goal-driven navigation tasks.
Sharing knowledge between multiple tasks can be achieved in a multi-task learning setup where all tasks are learned jointly in a supervised manner, or via meta-RL where a meta policy learned from a distribution of tasks is finetuned on the target. Unlike these methods, our policy is learned from one task that does not require manual annotations, and it can be transferred in a zero-shot setup where the policy does not receive any interactive training on the target.
Zero-Shot Learning
Zero-shot learning (ZSL) can be seen as an extreme case of transfer learning where the target task has zero training samples. Prior ZSL work focuses on supervised learning, e.g., image classification . In contrast, the proposed zero-shot experience learning (ZSEL) setup learns behaviors rather than classifiers; the policy learned on a source task needs to perform a set of target tasks, without receiving any new interactive experiences on the target. Further, unlike where a world model for synthetic environments is constructed, and control policies are trained on ‘imagined’ episodes, we consider a model-free approach and a ZSEL setup in realistic environments where the policy receives zero target interactions (i.e., neither imagined nor real). To our knowledge, we are first to propose a ZSEL model for embodied navigation.
Plug & Play Modular Transfer Learning
We introduce a novel transfer learning approach for visual navigation. Our model has three main components: 1) we start by learning a semantic search policy for image goals using a novel reward and task augmentation (Fig. 2a); 2) we leverage the image goal encoder to learn a joint goal embedding space for the different goal modalities (Fig. 2b); and finally, 3) we transfer the learned agent modules to downstream tasks in a plug and play fashion (Fig. 2c).
In the following, we consider an agent with main modules: 1) an observation encoder () that encodes the received observations from the environment; 2) a goal encoder () that encodes the task’s goal; 3) a policy () that uses the output of and to navigate and find the goal.
The policy is a key component in the modern end-to-end visual navigation agent. It guides the agent towards solving a task given a set of sequential observations and a goal. Such policies are often learned with reinforcement learning (RL) where the agent interacts with its environment (by moving about) and attempts to solve the task in a trial and error fashion. If the agent succeeds in its attempt, then it receives a reward to encourage such behavior from the policy in the future.
A main challenge for this learning paradigm is that the policy requires a large number of interactions with the environment in order to find a proper way to solve the task. This usually amounts to tens and hundreds of millions and up to billions of interactions, and correspondingly days or weeks of GPU cluster time. Furthermore, for each new task a policy is typically learned from scratch, which further increases the learning cost substantially.
We propose to learn a general-purpose semantic search policy that can be transferred and perform well on a variety of navigation tasks. Our idea is to learn such a policy with the image-goal task, where the agent receives a picture taken at some unknown camera pose somewhere in the environment, and must travel to find it. Our choice of image-goal for the source policy is significant. It requires no manual annotations, and image-goals can be sampled freely anywhere in a training environment. As a result, the policy can be trained on large-scale experience (e.g., collected from a fleet of robots deployed in various environments) which can improve its generalization to new tasks and domains. Furthermore, an image-goal encourages the learned policy to capture semantic priors for finding things in a 3D space. For example, by seeking images of couches and chairs, the agent learns implicitly to leverage these objects’ context and the room layout in order to find the image views effectively.
In an episode of image goal navigation, the agent starts from a random position in an unexplored scene, and it is tasked to find a certain location given an image sampled with the camera at . The agent receives an RGB observation at each step and needs to perform the best sequence of actions that would bring it to the goal within a maximum number of steps . Unlike the common point-goal task where the goal location is known , here is unknown, and the agent needs to leverage the learned semantic priors to search and find where could have been sampled from.
Our setup differs from recent methods in ImageNav where panoramic FoV sensors are required . Here, we consider a standard FoV for the agent’s view . While having a complete FoV sensor simplifies localization, this strong requirement is often not available in common robotic platforms and leads to high computational cost. This reduces the scalability and adoption of such methods by diverse agent configurations. In addition, our task setup allows our model to transfer to a diverse set of semantic navigation tasks in a plug and play fashion without the need for modifications to the target tasks (for which the literature does not use panoramic images).
View Reward
It is common to use the reduced distance to the goal to reward the agent for getting closer to in addition to the success reward of finding and stopping within a small distance of . However, while this reward proves to be quite successful for navigation tasks like point-goal, we argue it is less suited for semantic goals like images. Since the reward does not carry a signal about the semantic goal itself, the agent may fail or require much more experience in order to capture the implicit relation between the goal and the distance to goal reward (DTG). For example, if the goal shows an image of an oven, the agent may get close and stop nearby while looking at a book on the counter, and nonetheless receive a full success reward. This may lead to capturing trivial or incoherent associations between the goal and the agent’s observations.
In order to encourage the agent to leverage the information provided in the goal description and effectively capture useful semantic priors that may help it in finding , we propose a new reward function that rewards the agent for looking at when getting closer to , so it can better draw the association between its and . Specifically, we define the reward function at step as:
where is the reduced distance to the goal from the current position relative to the previous one, is the reduced angle in radians to the goal view from the current view relative to the previous one, is the indicator function, and is a slack reward to encourage efficiency. Note, this reward will encourage the agent to look at when it gets near the goal, since it is rewarded to reduce the angle between its current view and the view of the goal (see Fig. 2a). Finally, the agent receives a maximum success reward of if it reaches the goal and stops within a distance of from and an angle from :
View Augmentation
In addition to the view reward introduced above, we also provide a simple task augmentation method to promote generalization by increasing the diversity of goals presented to the agent. For each training episode, rather than having a fixed , we sample a view from a random angle at location and provide the agent with the associated from the sampled view as the goal descriptor. This has a regularization effect on the model learning; the agent will be less likely to overfit due to the changing goal description each time the agent experiences a given goal. Furthermore, with the start and the goal location fixed in a training episode but not , this encourages the agent to capture implicit spatial semantic priors of things that usually appear near each other as viewed from . For example, the agent would learn that an image of chairs as seen when peeking from the door is likely to be at the same location of the current image goal that is showing a dining table, since the agent experienced the same episode before but with showing the chairs, hence prompting the agent to explore the dining room.
Policy Training
We train our policy using reinforcement learning (RL), Fig. 2a. For each training episode, we sample an image-goal from . The agent encodes its current observation (an RGB image) with and the image-goal with and passes these encodings to the policy . The policy further encodes these information along with the history of observations so far to produce a state embedding . An actor-critic network leverages to predict state value and the agent’s next action . Based on the agent’s state in the environment, it receives a reward (Eq. 1 and Eq. 2). The model is trained end-to-end using PPO .
2 Joint Goal Embedding Learning
Having learned the semantic search policy, we can now transfer our model to downstream tasks. Specifically, we consider downstream navigation tasks where the goals are object categories (ObjectNav ), room types (RoomNav ), or view encodings (ViewNav), and they may be expressed by the modalities of a label name, a sketch, an audio clip, or an edgemap; see Sec. 4.2. A key advantage of learning the semantic search policy using RGB image-goals is that these goals contain rich information about the target visual appearance and context. Furthermore, in order to solve the image-goal navigation task, our model learns to encode these visual cues via a compact dense representation produced by the image-goal encoder .
Our idea is to leverage to learn a joint embedding space of different goal modalities for the various tasks. In other words, we upgrade the image-goal embedding space to be a joint goal embedding space to draw associations between the images and the different goal modalities like sketches, category names, and audio (Fig. 2b). This step can be carried out quite efficiently and using an offline dataset. For example, to learn about object-goals that are represented with a label (e.g., a chair) we only need to annotate a set of images with chairs. Then we train an object-goal encoder to produce an embedding similar to the image-goal encoder for compatible image-label pairs. In our experiments, we use offline datasets of size K images or less in which the “annotations” are actually automatic object detections. This is several orders of magnitude smaller than the amount of interactions usually needed to train a target-specific policy (tens to hundreds of millions) .
Formally, let be a set of images and their associated goals , where can be of any goal modality (e.g., audio, sketch, image, category name, edgemap) depending on the downstream task specifications. We learn a joint goal embedding space by minimizing the loss:
During goal embedding learning, we freeze and learn using Eq. 3 such that learns to encode its goal similar to the corresponding image embedding from .
3 Transfer and Zero-Shot Experience Learning
Having learned the semantic search policy and the joint goal embedding as described above, now we can transfer our model to downstream navigation tasks (Fig. 2c). For that, we only need to replace with the suitable goal encoder for the task, such that:
Our plug and play modular transfer approach has multiple advantages. Since all modules are compatible with each other, this means the model can perform the target task out of the box, i.e., it does not require any further task-specific interactions to solve the target task. We refer to this setup as zero-shot experience learning (ZSEL). Training policies with modern RL frameworks is the most expensive part of the model learning, and with ZSEL we manage to circumvent this requirement. Furthermore, due to the modular nature of our approach, it is easy to generalize to a wide variety of tasks and goal modalities. For a new task, only the respective goal encoder is trained with an offline dateset then integrated in the full model in a plug&play fashion.
Finally, our model can be easily finetuned for the downstream task to capture any additional cues specific to the task to reach a better performance. Unlike the common approach in the literature where only is pretrained and transferred , here the full model is transferred to the target task. This leads to higher initial performance, faster convergence, and better overall performance, as we will show in Sec. 4.
Evaluation
In the following experiments, we first evaluate our semantic search policy performance in the source task (image-goal navigation) compared to state-of-the-art methods (Sec. 4.1); then we show how our model transfers to a diverse set of downstream navigation tasks (Sec. 4.2).
For fair comparisons, we adopt the same architecture and training pipeline for our model and all RL baselines, and we note any deviations from this shared setup in the respective sections. We use a ResNet9 for and a GRU of layers and embedding size of for . For goal encoders , we use a ResNet9 to encode image-, sketch-, edgemap- and audio-goal modalities. We transform an audio clip to a spectrogram before encoding it by . If the goal is a category name, we use a layer MLP for . We train the policy using DD-PPO and allocate the same computation resources for all models. We use input augmentation (random cropping and color-jitter) during training to improve the stability and performance of the RL methods . See Supp for more details. We adopt end-to-end RNN-based RL since it is a common , generic architecture, does not require hand-crafted modules, shows good performance on real-data , and learned in sim has the potential to generalize well to real . However, our contributions are orthogonal to the RL architecture used.
Agent Configuration
1 Image-Goal Navigation
Datasets
We use the Habitat simulator and the Gibson environments to train our model. We use the dataset from . The training split contains K episodes sampled from each of the training scenes. Following the setup from , all RL models are trained for K updates ( million frames) on the training split. The test split has K episodes sampled uniformly from disjoint (unseen) scenes. For direct comparison with , we also test our model on a second split (“split B”) provided by that has K episodes and the same structure as the test split from (“split A”).
Baselines
We compare our image-goal model to the following baselines and SoTA methods: 1) Imitation Learning: This model’s policy is trained using supervised learning to predict the ground truth best action on the shortest path to the goal given its current observation. 2) Zhu et al. : The model uses a ResNet50 shared between and , pretrained on ImageNet and frozen. 3) Mezghani et al. : This is the SoTA panoramic image-goal navigation model. It uses a ResNet18 for and , a layer LSTM for , and a specialized episodic memory. We adapt this model to our FoV for and and train it using the author’s code. 4) DTG-RL: This model uses the shared architecture along with the common distance to goal dense reward for training. 5) Hahn et al. : This model learns from a passive dataset of videos collected from the Gibson training scenes and uses a customized architecture based on topological maps (see for details).
Results and Analysis
Table 4 reports the overall performance in terms of average success rate (Succ) and Success weighted by inverse Path Length (SPL) over random seeds. Our model outperforms strong baselines and the SoTA in image-goal navigation by a significant margin. In split A, our model gains in Succ and in SPL over the best baseline. The method designed for panoramic sensors tends to underperform in this challenging setting. We see a drop in Succ from 69% to 9% when using and FoV, respectively, since such methods rely heavily on the FoV for accurate localization. In split B, our model gains in Succ and in SPL over . It is important to note that the model from uses a much more complete sensor configuration than our method (pose sensor, RGB and Depth sensors of resolution, and a FoV) and it is trained offline from passive videos sampled from the simulator. Nonetheless, our model outperforms by a large margin, showing that interactive learning of end-to-end RL models still has an advantage over heuristic and passive approaches.
Ablations
To validate our contributions from Sec. 3.1, we test our model performance when removing the view reward or the view augmentation. As shown in Table 4, we see a degradation in performance whenever one of these components are removed, and the largest gain is realized when they work in tandem. Additionally, we test our model under noisy actuation. While methods in split A do not provide results under noisy conditions, does. Following the setup from , we use the noise model from that simulates actions learned from a Locobot . Our model shows robustness to noise and maintains its advantage over the baselines (Table 4 bottom).
Cross-Domain Generalization
Next we test the models trained on Gibson on datasets from Matterport3D (MP3D) and HM3D . In addition to the visual domain gap between these datasets, MP3D has more complex and larger scenes than Gibson, and HM3D has high diversity in terms of scene types. This poses a very challenging cross-domain evaluation setting. The test split from each dataset has in total K episodes sampled uniformly from and scenes for HM3D and MP3D, respectively. Table 2 shows the results. Overall, we see a drop in performance for all models in this challenging setting, especially for HM3D since there is high diversity in the val scenes. Nonetheless, our model outperforms all baselines on both datasets, showing that our contributions lead to better generalization by encouraging the agent to pay closer attention to the semantic information provided by the goal.
2 Transfer to Downstream Tasks
We consider target tasks and goal modalities:
2) RoomNav: The agent is tasked with finding the nearest room of types: living-room, kitchen, bedroom, office, bathroom, and dining-room. The goal is a label (find an office) and the episode is successful if the agent steps inside the room with the maximum episode length .
Datasets
For all target tasks, we use train / test scenes from the Gibson tiny set that has semantic annotations . These scenes are disjoint from those used in Sec. 4.1. We train all methods for up to million steps on the target tasks and report evaluation performance averaged over random seeds. See Supp for details.
To train the goal embedding for the modalities in ObjectNav and RoomNav, we sample K images of objects and K of rooms from the training scenes. We use the object labels generated by a model from to draw the associations between the images and each modality. For ViewNav, we sample K views from the training scenes and generate their edgemaps using an edge texture model . The number of samples for the offline datasest is driven by the available instances of each goal type in the training scenes. While there is a finite number of rooms and objects, we can sample views freely from any location in a scene.
Baselines
We compare our model to a set of baselines and SoTA models in transfer learning: 1) Task Expert which learns from scratch on the downstream task. 2) MoCo v2 initializes using MoCo training on ImageNet (IMN) or on a set of images randomly sampled from the Gibson (Gib.) training scenes. 3) CRL pretrains (ResNet50) using a combination of curiosity-based exploration and self-supervised learning. We initialize from a pretrained model provided by the authors. 4) Visual Priors uses a set of ResNet50s pretrained encoders as . The encoders are trained in a supervised manner to predict features (e.g. semantic segmentation, surface normal) that provide maximum coverage for downstream navigation tasks . 5) Zhou et al. transfers pretrained ResNet50s for depth prediction and semantic segmentation; however, unlike , these are used with an RGB encoder (ResNet9) trained from scratch. 6) SplitNet pretrains (customized CNN) using a mix of auxiliary tasks (motion and visual tasks) and point-goal navigation. We initialize from a pretrained model provided by the authors. 7) DD-PPO (PN) pretrains the model for point-goal navigation (PN) and both and are transferred.
Transfer Learning
Table 3 shows the results. Our approach outperforms all baselines by a significant margin. Interestingly, the self-supervised methods reach a competitive performance to those that rely on the availability of dense annotations (like semantic segmentation and ground truth depth) for supervised representation learning . Furthermore, methods that learn a curiosity-based representation (CRL ) or via auxiliary tasks and RL (SplitNet ) do not transfer as well as the SSL and SL methods. Additionally, compared to the strong DD-PPO approach which was trained on the same data as our policy but for the PointNav task, our model achieves substantial gains in success rate (from and up to ) across all tasks. This indicates that our semantic search policy is much better suited to transfer to diverse downstream tasks compared to the PointNav policy. Moreover, when looking at the test performance over the course of training for the best transfer methods compared to ours (Fig. 3), we notice that our approach has a much higher start and improves faster to better performance. Our model reaches the top performance of the best competitor up to faster.
Zero-Shot Experience Learning
A unique feature of our approach is its ability to perform the downstream task without receiving any new experiences from it. Our model shows excellent performance under the challenging ZSEL setting. Our ZSEL model outperforms the Task Expert in out of the tasks despite receiving zero new experience on the target, and even after training the Task Expert for up to million steps (Ours-ZSEL in Table 3).
In addition, we see in Fig. 3 that the majority of transfer learning models struggle to reach our ZSEL performance. Note that our model does not have any advantages in terms of architecture, which is shared with the rest of the models. Thus, the high ZSEL performance is attributable to our modular transfer approach. In ObjectNav and RoomNav, the best competitor requires between to million steps to reach our ZSEL performance, and with the exception of ObjectNav-Audio the competitors show little improvement over that level. In ViewNav, we notice that none of the baselines are capable of reaching our ZSEL level. This can be attributed to the challenging goal modality where estimating distances for successful stopping is difficult, and to the close proximity of this task to the source ImageNav task that our semantic policy is most familiar with.
Modular Transfer Ablation
Fig. 4 (left) shows a modular ablation of our approach on the ObjectNav-Label task. Transferring individual modules separately has mixed impact on performance. While transferring and only does not improve over the ‘No Transfer’ case, leads to positive transfer effect. This is expected since in this model is a deep CNN with the largest portion of parameters. Having a good initialization of this component is beneficial. Nonetheless, when combining the modules together with our plug and play modular approach we see substantial gains. Our full model demonstrates the best performance and enables ZSEL, thus validating our contributions.
Scalability
We evaluate our model’s ability to scale in terms of experience gathered on the source task. We find a strong correlation between the experience gathered on the source task and the ZSEL performance on downstream ones. As our semantic search policy receives more experience on the source task (ImageNav), its ZSEL performance on the target task gets better (Fig. 4 right). This is important since our source task requires no annotations and can be easily scaled to more scenes and large datasets. For an analysis of our model in terms of the used sensors, see Supp.
Long-Term Task Expert Training
We saw above that our model scales well and its transfer improves when more experience is gathered on the source task. However, does a task expert become competitive if it simply gets longer training on the target task? How long does that model take to catch up with our approach? To find out, we train the Task Expert on each of the target tasks for up to M steps. Fig. 5 shows the results. The Task Expert requires on average more than M steps on ObjectNav and RoomNav, and up to M on ViewNav (in total M steps over the tasks) to reach our ZSEL performance. It never reaches our model’s top performance when our model is finetuned on the target task. Moreover, our model reaches the best performance of the Task Expert faster. The Task Expert needs task-specific experience with task-specific annotations, which can be expensive and limits the available training data. In contrast, our model learns in the source task using more diverse goals that can be sampled randomly from the (unannotated) scene, thus scaling more effectively.
Additional Results and Discussions
Please see Supp for qualitative results, an analysis of failure cases, and a discussion of limitations and the societal impact of our approach.
Conclusion
We introduce a plug&play modular transfer learning approach that provides a unified model for a diverse set of semantic visual navigation tasks with different goal modalities. Our semantic search policy outperforms the SoTA in the source task of image-goal navigation, as well as the SoTA in transfer learning for visual navigation by a significant margin. Furthermore, our model is able to perform new tasks effectively with zero-shot experience—to our knowledge, a completely new functionality for visual navigation. This is a stepping stone for future work, especially for tasks with high-cost training data. Being able to do ZSEL and learn from few experiences is a crucial skill for an agent in open-world and lifelong learning settings.
Acknowledgements: UT Austin is supported in part by DARPA L2M, the UT Austin IFML NSF AI Institute, and the FRL Cog Sci Consortium. K.G is paid as a Research Scientist by Meta AI. Thanks to Lina Mezghani for providing access to data and code.
References
Supplementary Materials
Additional information presented in this supplementary:
Details on the shared implementation (Sec. 7).
Details of the image-goal navigation dataset (Sec. 8).
Detailed results on image-goal navigation across levels of episode difficulties (Table 4).
Examples of the visual goal modalities used in target tasks (Fig. 6).
Dataset details for the target tasks and goal embedding space (Sec. 9).
Detailed results with standard deviations on target tasks (Table 5).
Qualitative results for our model in all tasks and goal modalities (Fig. 10).
Examples of failure cases for our model (Fig. 11).
Performance curves for all tasks and goal modalities in transfer learning setup (Fig. 7) and long-term training of Task Expert (Fig. 8).
Ablation on the sensor configuration used by the agent (Fig. 9)
Discussion of potential societal impact (Sec. 10) and limitations (Sec. 11).
Shared Setup
All RL methods are trained with the following setup. We use input augmentation of random cropping and color jitter for both observations and goals. The models are trained with DD-PPO . We set the number of PPO epochs , the forward steps , the entropy coefficient , clipping of , and train the model end-to-end using the Adam optimizer . We allocate the same number of processes and resources to all methods.
We use the Habitat simulator along with the Gibson , Matterport3D , and HM3D datasets. These datasets are photorealistic and scans of real-world environments with varying complexities, sizes, room layouts, types. In all our experiments, the test scenes are disjoint from those used for training to assess the agent ability to generalize to previously unseen environments.
Image-Goal Navigation
Detailed Results
We show in Table 4 the detailed results of all models across the three levels of episode difficulties (easy, medium and hard). Our model shows better performance across the different levels and in both split A and B. For a qualitative result, see Fig. 10 A.
Transfer Learning to Downstream Tasks
We use scenes from Gibson and split them into scenes for training and for testing. In the following, we present the details of the datasets used for each of the downstream tasks.
ObjectNav: We sample K episodes for training and K episodes for testing. For the sketch-goals (Fig. 6 middle), we sample sketches from for each object category and split them to used during training and for testing. For the audio-goals, we sample audio clips from of lengths ranging from to seconds and split them for training and testing. At the start of each episode of the ObjectNav (Audio) task a random seconds duration is sampled from the respective audio clip and split, and presented to the agent as the goal descriptor.
RoomNav: We sample K episodes for training and for testing.
ViewNav: We sample K episodes for training and K for testing. We generate edgemaps for random views per scene, and we randomly assign one of those per episode as the goal (Fig. 6 bottom).
Goal Embedding Space
In the joint goal embedding space, we aim to learn goal encoders that are compatible to the image-goal encoder. For example, an image view from a living room with a TV detected in it will be used as the positive anchor for a sketch of a TV, a sound clip from a TV, the TV label, the living room label, and the edgemap of the view. The annotations for the sampled image view are based on model predictions from . During training, the parameters of are kept frozen, and we train the various goal encoders defined in Main/Sec.4 using the loss from Main/Sec.3.2.
Detailed Results for All Tasks
In Table 5 we show the average success rate and standard deviation for all methods over random seeds. Fig. 7 shows the performance of the best transfer learning methods and our approach across all tasks and goal modalities. Fig. 8 shows the Task Expert performance when trained for up to M steps on each of the respective tasks and in comparison to our model performance under the ZSEL setting or when it is finetuned. Furthermore, we show example navigation episodes from all tasks and goal modalities for our approach in Fig. 10. Our plug and play modular transfer learning approach enable our model to perform a diverse set of tasks effectively.
Scalability (Sensors)
We evaluate our model’s ability to scale across the sensor suite. Fig. 9 shows our model performance when varying the sensors’ configuration in the source task (ImageNav) and evaluating on ObjectNav (Label) under the ZSEL setup. As expected, when enriching the agent sensors to include depth and pose sensors in addition to vision, we see an additional improvement in performance. More importantly, when increasing the vision sensor resolution from to we see a significant bump in ZSEL success rate that exceeds the one from diversifying the sensory suite. Our model seems to benefit from an enhanced vision channel as it carries the important semantic cues needed for our semantic search policy and goal embedding space.
Failure Cases
We show in Fig. 11 few examples of failure cases encountered by our model. We notice that some of these failure cases are related to the type of the goal modality. For example, in ImageNav the agent sometimes finds the object described in the image however misestimate the view point the image is taken from, hence stops a bit far from the goal location (Fig. 11 A). In ViewNav (Edgemap), the goal modality lacks distinctive texture and color information which leads the agent to sometime stops at a location with similar edge structure, but it is actually not the goal (Fig. 11 B). A type of failure cases spotted in multiple tasks are the early stopping cases. In these cases, the agent fails to estimate the distance to the goal correctly and stops early resulting in an unsuccessful episode (Fig. 11 C).
Potential Societal Impact
Our approach’s application domain is semantic visual navigation. Here, autonomous agents are trained to find semantic objects in a 3D environment. Such a technology can have positive societal impact by improving people’s life, especially in domains like elder care, with robots that can aid in daily life tasks (e.g. find my keys, go to the bedroom and bring me my medicine). On the other side, the datasets used in this study are 3D scans of building and houses from certain geographic and cultural areas (western style houses from well-off areas). This creates certain biases in the type of building architectures, room, and object types the agent is familiar with. Consequently, this may limit the availability of this technology to a small section of the population. More diverse datasets and methods with robust adaptation to strong shifts in building layouts and object types are needed to mitigate these effects.
Discussion and Limitations
We propose a novel approach for modular transfer learning that enables the agent to handle multiple tasks with diverse goal modalities effectively. Our model can solve the downstream tasks out-of-the-box in zero-shot experience learning setup alleviating the need for expensive interactive training of the policy. Alternatively, our model can be finetuned on the downstream task to learn task-specific cues where it showed to learn faster, generalize better and reach higher performance than the baselines. While we focused in this work on semantic navigation tasks, this can be seen as a first step in this exciting direction. Additional research is needed to generalize this method to tasks that require a series of goals and a compatible policy that can plan effectively in a multi-goal setup (e.g. VLN ). Further, our results and evaluation demonstrate strong transfer learning performance for our method. However, as usual in transfer learning, there is not a theoretical guarantee that a transfer effect will always be beneficial. Target tasks with significant differences to the source task may not benefit from transferring the accumulated experience in the source.