GNM: A General Navigation Model to Drive Any Robot
Dhruv Shah, Ajay Sridhar, Arjun Bhorkar, Noriaki Hirose, Sergey Levine
I Introduction
Machine learning methods have enabled broad generalization with real-world applicability in natural language processing , visual perception , and other domains by leveraging Internet-scale data. Such generalization typically requires learning general patterns from diverse datasets, which are usually collected once and then reused for various purposes. Such large-scale models also support the ability to be adapted for new tasks by reusing the representations learned from broader, larger, and more general datasets, for example by or zero-shot transfer , or fine-tuning on target-domain data. Although this paradigm has been very successful, it is difficult to apply in robotics due to the sheer diversity of environments and platforms across researchers. Control policies learned end-to-end usually require separate data collection for each robotic platform, leading to “fragmentation” in progress, where every researcher works with their own robot-specific dataset and policies, making it infeasible to accumulate large enough datasets. Can we overcome this challenge by training models on more general and reusable cross-robot datasets?
We study this question in the context of visual navigation, where heterogeneity between robots might include different camera hardware, viewpoints, dynamics, and more broadly, embodiments, but where the over-arching navigation objective looks similar irrespective of these differences. A wheeled robot, quadruped, or a drone all have the same abstract objectives: to explore the environment, plan a path to the goal, and avoid collisions. Leveraging this shared abstraction across robots and training a general navigational omnipolicy from large-scale data could enable broad generalization to novel environments, unseen sensor parameters (e.g., camera intrinsics and extrinsics), and new robot configurations.
In this paper, we propose to take a step towards this kind of data sharing by training an embodiment-agnostic general navigation model (GNM) from an aggregated multi-robot dataset. The primary contribution of our work is a framework for training a general omnipolicy from multi-robot datasets, with empirical evidence that such an omnipolicy can effectively learn from heterogeneous datasets and generalize to novel robot platforms. To facilitate this, we aggregate a large heterogeneous dataset of navigation trajectories collected across 6 robots, spanning 60 hours of interactions in challenging indoor and outdoor environments. We train the GNM on this dataset and deploy it on 4 distinct robot platforms, including 2 new robots. We show that a single learned policy can be used across multiple robots to perform goal-reaching in challenging indoor and outdoor environments, outperforming policies trained with any single dataset. We also report robustness to degradation in camera parameters, tire damage, and other gradual changes that the robot may experience over its lifetime.
We have publicly released the trained GNM policy, code used to train and deploy our models on various popular robot platforms, as well as the dataset used to train these models at our project page. We hope that this represents a step towards both general-purpose multi-robot datasets and general-purpose visual navigational models that can be deployed on a wide range of robots — similar to how practitioners currently use pre-trained models in vision and language, such models could constitute pre-trained backbones for visual navigation.
II Related Work
Learning from large, diverse robotic datasets has been studied for various robotic applications where data sharing across similar robots helps scale learning to challenging environments . However, for applications such as ground or aerial navigation, with different sensors and robot dynamics, current approaches tend to rely on learning from small datasets which are only representative of a single robotic platform. Our paper proposes learning navigation behavior from heterogeneous robot datasets, collected across multiple embodiments.
Our work is closely related to transfer learning, where the objective is to train policies that transfer across domains, such as across dynamics , environments , morphologies , viewpoints , and embodiments . Our focus is not on designing specific domain adaptation algorithms or hand-engineered augmentations for transfer, but rather studying how direct generalization of simple, high-capacity models trained on real-world data can provide a path to broadly applicable navigational policies. Towards this, our work is also closely related to DroNet , which imitates expert on-road driving data to control a quadrotor. We take this paradigm one step further, showing that we can train goal-conditioned policies on data from multiple robots and control new ones, including a quadrotor.
Prior work has also explored learning of visual representations or end-to-end policies from passive data, such as YouTube videos, which can be scaled up massively without real-world data collection . We explore a complementary direction, studying how readily available on-robot data (also passive) can lead to generalizable policies. This is particularly relevant for navigation, where data is plentiful, and trajectories from multiple robots can directly train a policy, as opposed to two-stage methods that use Internet data for representation learning followed by in-domain adaptation.
Following a large body of research in visual navigation , we use a combination of topological graphs for high-level planning and image-goal policies for low-level control, which gives us an efficient way to scale reactive policies for long-range navigation . Prior work has also extended this framework for complex tasks beyond goal-reaching, such as exploration , instruction following , and reinforcement learning . We show that that our GNM can be coupled with such topological graphs to scale image-goal navigation to new robots.
III Multi-Robot Training Dataset
Our aim is to train a general visual navigation model that can learn broadly applicable navigational affordances across a variety of distinct robotic systems. To facilitate such large-scale policy learning, we aggregated a heterogeneous dataset of navigation trajectories sourced from 8 datasets collected on robotic platforms with varying dynamics, sensors, and behaviors. The datasets contain a variety of challenging indoor and off-road environments (Table I and Fig. ). We have publicly released this dataset on the project page.
The GNM dataset contains over 60 hours of real-world navigation trajectories: a combination of tele-operated and autonomous navigation behaviors collected across 6 distinct robotic platforms, including 4 commercially available platforms (TurtleBot, Clearpath Jackal, Warthog and Spot) and 2 custom platforms (Yamaha Viking ATV, RC Car). The trajectories contain widely varying robot dynamics and top speeds ranging between 0.2 and 10m/s, operating in a diverse set of environments (e.g., office buildings, hallways, suburban, off-road trails, university campus etc.).
To train navigation policies that can operate solely from egocentric visual observations, the dataset contains forward-facing RGB images paired with the robot’s commanded actions and local odometry measurements. Each robot has different camera parameters, necessitating any successful policy to generalize across variations in camera pose and intrinsic parameters, though all platforms use the same type of sensor (monocular RGB camera). It is straightforward to further expand GNM by adding other datasets of relevant navigation behaviors , or mix-and-match subsets of the dataset based on the desired application,
IV Training a General Navigation Model
To study a common navigation task across robots and environments, we consider the problem of image-goal navigation , where a robot is tasked with navigating to a goal location specified as an image observation taken at . Unlike PointGoal , GPS navigation, or semantic objectives , image-goal navigation is a general framework that does not rely on ground truth localization or semantic labels, and allows us to formulate a very general navigation task that can be trained with any visual navigation dataset. Our goal is to train a goal-reaching policy that can navigate solely from egocentric visual observations. To provide a general task representation for this policy, we condition it on the desired goal and integrate it into a navigational system based on topological graphs .
Such systems have shown great navigation results in a variety of indoor and outdoor environments — what would it take to train such a policy across robots, with varying controllers, dynamics and sensor placements? We highlight two key ingredients in training multi-robot policies: (i) carefully choosing the right action representation that facilitates transfer across robots, and (ii) conditioning the policies on a “summary” vector that allows it to deduce the properties of the robot it is controlling, so different robots can exhibit different, valid capabilities. Although we found the particular design decisions described in this section to be important for good performance, as we discuss in our experiments (Sec. V-C), we emphasize that the primary contribution of our work is not a novel learning algorithm, but an empirical demonstration that policies learned from heterogeneous datasets can generalize broadly to new environments and new robots.
While the general task of navigation from egocentric images is common across robots, the specific inputs (camera observations) and outputs (actions, dynamics) can vary substantially: a TurtleBot is differential-drive, expects low-level velocity commands, and has a top speed of 0.5m/s, whereas an ATV uses Ackermann steering, expects throttle and steering commands, and drives up to 20 faster. Learning a common control policy that operates directly on these raw, unstructured outputs can be challenging due to these inconsistencies and high-variance outputs (e.g., speed m/s). This is further exacerbated when generalizing to new robots, where the policy might need to “guess” how fast it should move.
To this end, we propose using a shared abstraction to allow the goal-reaching policies to operate in a transformed action space that is consistent across robots, making the data points look “similar” and easier to learn common patterns from. In our experiments, we found this to be important to be able to learn from multiple datasets (see Sec. V-C1 for analysis). We use a combination of relative waypoints and yaw change as a mid-level action space. Labels for these actions can be obtained by using local odometry, which are easily available across datasets. Additionally, the policy also predicts the temporal distance to the goal , as a measure of traversability, which is used by the navigation system to estimate the connectivity of the topological graph.
IV-B Embodiment Context
When deployed on an arbitrary robot, the policy must infer the capabilities of that particular robot. For instance, a TurtleBot can spin in-place but not go over bumps on the road, whereas an RC Car can easily traverse small bumps but has a limited turning radius. A simple way to provide such awareness to the policy is to condition it on hand-designed parameters that provide a concise “summary” of capabilities, such as its size, turning radius etc. Defining these parameters by hand presents a barrier to fast and easy deployment of the policy to new robots, and requires human intuition to identify and define a relevant set of parameters. Instead, we propose a simple and automatic approach: rather than manually defining parameters that fully identify the robot, we use a sequence of consecutive past observations from the robot’s viewpoint to infer a learned embodiment context , and condition the learned policy on this context in addition to the observations. This context contains information about the robot’s configuration and dynamics, which can be used to condition the behavior of the policy.
While this context may not contain all information to fully identify the robot, we hypothesize that it is sufficient to effectively control the robot. Our experiments show that the embodiment context allows the same policy to be deployed on novel robot configurations without designing any hand-engineered robot representation. We empirically evaluate different ways of providing context in Sec. V-C2 and find that the most effective representation is achieved by using a temporally consistent context that conditions the policy on consecutive past observations .
IV-C Implementation Details
A combination of conditioning the policies on embodiment context and transforming the action space can allow a simple goal-reaching policy to be trained from heterogeneous datasets. It is important to note that the proposed modifications are orthogonal to the choice of downstream policy architecture and learning algorithm, and we could use different encoders or train with reinforcement learning.
V Deploying the GNM Across Robots
We deploy our learned GNM omnipolicy in a variety of challenging indoor and outdoor environments on four different robot platforms. We designed out experiments to answer the following questions:
Can multi-robot training enable generalization to novel robots and environments?
Do GNM policies outperform policies trained solely on single-domain data?
How important are the design choices made in Sec. IV for attaining good performance with the GNM?
Are policies trained with multiple datasets more robust to degradation than single-domain policies?
We deploy the GNM on four distinct robotic platforms, including a quadrotor and two other novel robots with no corresponding training data, as shown in Fig 3.
Vizbot: A custom-built robot platform inspired by the design of Niwa et. al. , based on a Roomba. It is equipped with an off-the-shelf PCB-mounted fisheye camera. There is no training data from a Vizbot or any other Roomba-like robot.
DJI Tello: A commercially available quadrotor equipped with a forward-facing camera. There is no training data from any quadrotor for GNM. We restrict the drone to a horizontal plane 1m off the ground, to mimic ground navigation.
Clearpath Jackal UGV: A commercially available off-road platform equipped with an off-the-shelf PCB-mounted fisheye camera. This system resembles the data collection platform used for the RECON, Berkeley, and SCAND-J datasets, but has a different camera and mounting height.
LoCoBot: A popular open-source platform based on a Kobuki, equipped with an off-the-shelf PCB-mounted fisheye camera. There is no training data from a LoCoBot, although GS was collected on a similar TurtleBot2, albeit with a different spherical camera at a lower height.
V-B Zero-Shot Deployment
Towards answering Q1, we deploy the same trained GNM on four distinct robotic platforms without any fine-tuning per robot. Fig. 3 and Table II summarize our evaluation in a variety of indoor and outdoor environments on 4 different robots, all using the same model. Most notably, the GNM can control a Tello, despite never having seen any trajectories from aerial robots in GNM. A GNM policy consistently outperforms single robot policies across all tested robots, performing up to 5x better in some cases. We also observe generalization to massively out-of-distribution (OOD) settings, like a LoCoBot navigating outdoors on a sidewalk, or a Jackal navigating inside an office building, which were not present in the training data. This suggests that training on heterogeneous datasets can enable generalization to novel environment-robot pairs, as well as entirely new robots.
To better understand how data sharing benefits performance (Q2), we quantitatively evaluate the navigation performance of policies trained with heterogeneous datasets in an assortment of 20 indoor and outdoor environments on the Jackal and LoCoBot platforms (Tables III, IV). To project the performance trends with varying amounts of data, we train policies from increasingly diverse subsets the training data — “Small”, “Mid”, and “Large”, corresponding to data from the first 2, 4, and 6 datasets listed in Table I. We quantify performance using success rates, measured as the mean progress made towards the goal. For videos of our experiments and more information on the testing environments, please check out the supplementary video and project page.
Deploying on a LoCoBot, which is an unseen robot with no corresponding data present in the dataset, we find that policies trained on a single dataset (e.g., GoStanford (GS) or CoryHall ) fail to generalize to a new embodiment with different sensors. Fine-tuning visual representations trained for task-agnostic datasets like ImageNet, which is a popular strategy for pre-training in many vision-based applications , improves a bit but still struggles in a majority of the environments. However, policies trained by sharing task-relevant datasets across robots significantly outperform these single-domain policies, as shown in Table III. We also observe that adding more and diverse datasets (GNM-Large) contributes towards improvements in performance, despite the additional data coming from seemingly unrelated tasks (e.g., off-road driving). Fig. 4 shows an example office environment where increasing the diversity of training data improves performance.
We observe similar trends on a Jackal, which is deployed on a variety of previously unseen outdoor and indoor environments (Table IV). Unsurprisingly, a single-domain policy trained on off-road RECON data performs well for many outdoor environments, but struggles with navigating indoors, which is OOD for the RECON dataset. Similarly, a GS policy struggles in outdoor environments but succeeds in some easy indoor environments. GNM omnipolicies are able to generalize better to a variety of indoor and “Hard” outdoor environments, which can be over 100m long, significantly outperforming the single-domain policies (Fig. 4).
V-C A Systematic Analysis of the Design Space
Towards answering Q3, we perform a systematic analysis of the design choices presented in Sec. IV. We evaluate each design choice on a LoCoBot, which is an unseen robot with no corresponding training data, in indoor environments with varying levels of complexity, where “Easy” environments have wide passages and smooth turns, “Moderate” environments have tight passages or sharp turns, and “Hard” environments are larger (up to 50m) with a combination of tight passages and multiple turns.
We compare the three action spaces discussed in Sec. IV-A by training three different policies on GNM-Mid and evaluating them in 10 environments (Table V). While using velocities as an action space works well for most easy environments, often outperforming the policy using waypoints, both these policies struggle in environments requiring dynamic maneuvers like sharp turns. A policy based on normalized waypoints, on the other hand, significantly outperforms the others, including in the challenging environments. This suggests that normalizing the action space indeed allows the policies to learn more effectively and generalize to new robots.
V-C2 Embodiment Context
We consider two ways to represent the embodiment context: (i) temporally consistent context containing consecutive past observations , and (ii) static context, containing a fixed set of past observations from the robot in the target environment. Comparing these choices in environments of varying complexities (Table V), we find that adding either form of context significantly boosts the navigation performance in the harder environments, which require the robot to navigate tight passages with multiple obstacles and sharp turns. This suggests that the context helps the polices generalize better due to the additional information about the embodiment (e.g., viewpoint, speed etc.). Between the two, we found the temporal variant superior, suggesting that the temporal information (e.g., speed, turning radius etc.) is important to enable this generalization. In our main experiments discussed in Sec. V-B and Fig. 3, we use a temporally consistent context with .
V-C3 Policy Architecture
We also compared different policy architectures to encode the goal information: (i) single-encoder stacking, where the observation and goal images are stacked along the channel dimension , (ii) a Siamese architecture, where the images are processed with independent encoders and the resulting embeddings are combined , and (iii) the conditional architecture in Fig. 2, with an additional pathway from the observation to the policy outputs . We found that the choice of architecture significantly affects the navigation performance, with the conditional model being the most performant. We hypothesize that this is due to the additional pathway that allows the learned embeddings to be conditioned on the current observations, leading to more generalizable representations, as studied in prior work .
V-D Robustness to Degradation
A key strength of training on heterogeneous datasets is that learning across varied parameters encourages the policy to learn shared affordances across robots, thus being robust to small variation in robot parameters, such as sensor placement and mechanical properties. We show that the shared GNM can indeed offer such robustness by testing it under some example degradation scenarios shown in Fig. 5.
When testing the trained policy with a steering degradation (Fig. 5a), where the robot’s maximum angular velocity is clipped, we find that the GNM can compensate for the degradation by taking a longer, smoother path towards the goal without any localization failures. We also tested the GNM while perturbing the position of the camera and physically affecting the dynamics by damaging the robot during navigation, and find that it can successfully reach the goals despite the degradation (Fig. 5d). Please see the supplemental video for these experiments.
VI Discussion
In this paper, we demonstrated that a general goal-conditioned navigation policy trained from navigation datasets collected by multiple distinct robots, ranging from RC cars to ATVs, can control new robots in challenging environments. The design of our learning framework is simple, and largely follows prior work: the novel observation is that a set of relatively simple decisions, such as including a temporal context and standardizing the action space, is sufficient to enable broad generalization from heterogeneous data. Empirically, we show that our approach can enable real-world navigation for a range of robots, including some not seen in training, and even an underactuated quadrotor.
Our specific instantiation of this principle does have some limitations. Most prominently, our system does not explicitly account for differences in capabilities: we assume all robots are ground robots (though we study generalization to a quadrotor) with a forward-facing RGB camera. Handling diverse sensing, actuation (beyond variability in speed and steering), and traversability, would be an exciting direction for future work. Secondly, our dataset could be much larger: while we observe exciting generalization from 60 hours of data, a much larger and broader dataset could enable even better generalization in the future.
The promise of such a general navigation model trained on diverse data is that it may provide a pre-trained base model for a variety of downstream navigation applications. In the same way that computer vision researchers and practitioners typically start off by downloading a pre-trained backbone to use for their task, we hope that future navigation projects might use a pre-trained navigational omnipolicy that generalizes broadly enough to offer a “universal” starting point.
Acknowledgments
This research was supported by the DARPA RACER program, ARO W911NF-21-1-0097, ARL DCIST CRA W911NF-17-2-0181, AFOSR FA9550-22-1-0273, Toyota Motor North America, and Toyota Research Institute. The authors would like to thank Haresh Karnan, Xuesu Xiao, Gregory Kahn, Xiangyun Meng, and Byron Boots, for their help in aggregating the heterogeneous dataset used for training the GNM. The authors would also like to thank Brian Ichter, Antonio Loquercio, Jie Tan, Devendra Singh Chaplot, Tingnan Zhang, Laura Smith, Nick Rhinehart, Frederik Ebert, and Kelvin Xu, for useful discussions and feedback on an earlier draft of the paper.