Collaborative Visual Navigation

Haiyang Wang, Wenguan Wang, Xizhou Zhu, Jifeng Dai, Liwei Wang

Introduction

Human intelligence would collectively solve the problem that otherwise is unthinkable by a single person. For instance, Condorect’s Jury Theorem indicates that, under certain assumption, adding more voters in a group would increase the probability that the majority chooses the right answer. Thus it is believed that a next grand challenge of AI is to answer how multiple intelligent agents could learn human-level collaborations . With the recent rapid advances of machine learning and robotic technologies, the collaboration of AI agents gathers significantly increasing attention from researchers, and finds great potential in a wide range of real-world applications, such as search-and rescue , area coverage , harbour protection , and so on. In short, such multi-agent system (MAS) is of central importance to AI and still in its infancy.

Multi-agent reinforcement learning (MARL)​ and embodied systems​ are the main driving power behind MAS. MARL field is rapidly expanding, and some of approaches even show super-human level performance in various games , such as Magent​ , StarCraft​ , and Dota2 . However, most of existing algorithms operate on low-dimensional observations, \eg, grid-world like or game environments. As for embodied systems, the availability of large-scale, visually-rich 3D datasets and community interest in active perception led to a recent surge of simulation platforms . Based on these high-performance 3D simulators, various active vision tasks are investigated, such as embodied navigation , instruction following , embodied question-answering and perception , \etc. However, most of them either do not support multi-agent setting , or little consider planning .

Thus there is still a huge gap between MAS and realistic visual environments. To close this gap, we introduce a collaborative visual navigation (MAVN) task for image-goal based, multi-agent navigation, under visually realistic environments, and develop a corresponding dataset, called CollaVN. Autonomous navigation forms a core building block in embodied systems, and long received research interest across many fields, such as computer vision , robotics , linguistics​ , and cognition​ . In MAVN, agents cooperate to reach a target position (determined by a corresponding image), with their individual perception. Three task settings are explored to cover more challenges in MAS/MARL: i) CommonGoal (Fig. 1(a)), where agents are assigned a same goal picture; ii) SpecificGoal (Fig. 1(b)), where agents are assigned different goal pictures; and iii) Ad-hoCoop (Fig. 1(c)), where the agent numbers in training and testing are different. CommonGoal serves as a basic setting where agents learn how to cooperate to reach a goal faster. SpecificGoal captures more complex scenarios where agents need to collaborate with individual targets. Ad-hoCoop is built upon CommonGoal but agents are required to adapt to different team sizes during deployment, simulating a more open MAS situation .

Our CollaVN dataset is built upon iGibsonV1 simulator . Besides inheriting the advantages of iGibsonV1 in high-quality rendering, rich diversity and good accessibility, CollaVN has the following unique characteristics:

Multi-agent operation: CollaVN supports MAVN with arbitrary group sizes (2-4 agents used in current setting). Agents can be safely initialized with collision detection.

Multi-task setup: All the three task variants are involved and companied with standard evaluation tools.

Flexible target generation: CollaVN allows to generate different target pictures flexibly, well supporting the three MAVN settings and future dataset extension.

Panoramic ⁣{}_{\!} observation: ⁣{}_{\!} A ⁣{}_{\!} panoramic ⁣{}_{\!} camera ⁣{}_{\!} is ⁣{}_{\!} equipped ⁣{}_{\!} for rendering agent egocentric panorama in real-time.

We further propose a novel memory-augmented communication approach for MAVN. Communication is viewed as one of the most desired and efficient ways of multi-agent interaction/coordination . Communication is especially crucial when a set of cooperative agents are placed in a partially observable environment, such as our MAVN setting – the agents need to exchange information (\eg, past experience, current perception, future plan, \etc), coordinate their actions, behave as a group and achieve the target goal better . Thus learning communication has become a center topic in MARL field. From earlier natural-language-based to recent hidden-state-based communication forms and pre-defined communication protocols to learnable strategies , communication based MARL made great advance. Some most recent ones are concerned with practical communication-bandwidth-limited situations and conduct studies mainly around the theme of learning what, when and who to communicate. However, in previous methods, agents simply discard all the information after each communication round. This is problematic, as some communication information, even not useful in present, may be valuable in the future. This raises the risk of losing essential information and easily leads to suboptimal behaviors driven by short-termism. To address these limitations, in our memory-augmented communication approach, each agent is equipped with a private, compartmentalized memory, for safely storing and accurately recalling its past communication information. During each communication round, each agent can send a more reasonable request to other agents, by considering its current states and stored information. Each of other agents will check its current state and past communication information to give a better response. The agents can make full use of the rich information stored in the memory during navigation. Overall, our approach helps agents conduct more efficient communication and enhances their long-term planning ability, eventually leading to better collaboration and navigation performance.

In a nutshell, our contributions are three-fold:

To narrow the gap between MAS and visual perception, we develop a large-scale and publicly-accessible dataset, CollaVN, for collaborative visual navigation (MAVN), in photo-realistic, multi-agent environments.

Diverse MAVN task settings are explored to cover many core challenges in MAS/MARL, with packaged evaluation tools and baselines.

A novel memory-augmented communication approach is proposed to address more efficient multi-agent collaborations and robust long-term planning.

Several baselines are evaluation tools are developed for comprehensive benchmarking. Experimental results show that our memory-augmented communication framework achieves promising performance over the three MAVN settings on CollaVN dataset. This work is expected to provide insights into the critical role of MAS in embodied perception tasks, and foster research on the open issues raised.

Related Work

Vision-based Navigation. As a crucial step towards building intelligent robots, autonomous navigation has been long studied in robotics community. Recently vision-based navigation attained growing attention in computer vision community and was explored in many other forms, such as point-goal based or object-oriented (\ie, reaching a specific position or finding an object) , natural-language-guided , audio-assisted navigation , and from indoor environments to street scenes . Prominent early methods rely on a pre-given/-computed global map for path planing, while later ones refer to simultaneous localization and mapping (SLAM) techniques that reconstruct map on the fly. Some more recent methods learn planning and mapping jointly .

Although great advance has been achieved, visual navigation in multi-agent setting has remained largely unexamined. Most existing studies focused on single-agent navigation, excepting considering two-agent collaboration in finding and lifting bulky items. In this work, we develop a large-scale dataset for indoor, collaborative, multi-agent visual navigation. The agents are required to work together to perform same or different tasks, or even without any assumption on the number of involved agents. Thus our task setting is more novel and general.

Embodied AI Environments. The recent surge of research in embodied perception is greatly driven by new 3D environments and simulation platforms, such as iGibsonV1 , AI2-THOR , and Habitat . Compared with grid-like or game environments , these open accessible, photo-realistic platforms bring perception and planning in a close loop and make the training and testing of embodied agents reproducible . Based on these platforms, numerous embodied vision datasets were proposed , while most of them are designed for single-agent tasks. We build our dataset upon iGibsonV1 which is featured by large scale, rich visual diversity and strong extensibility. Although iGibsonV2 also involves multiple agents, it addresses obstacle avoidance in social scenes. It does not investigate collaboration among agents, while which is the core nature and research focus in MAS/MARL. Moreover, our dataset supports diverse essential MARL task settings, making it unique in the filed of visual navigation.

Multi-Agent Reinforcement Learning (MARL). MARL tackles ⁣{}_{\!} the ⁣{}_{\!} sequential ⁣{}_{\!} decision-making ⁣{}_{\!} problem ⁣{}_{\!} of ⁣{}_{\!} multiple ⁣{}_{\!} agents in a common environment, each of which aims to achieve its own long-term goal by interacting with the environment and other agents . Basically, MARL algorithms can be placed into four categories . i) Analysis of emergent behaviors. Some early studies analyze single-agent RL algorithms in multi-agent environments. ii) Learning communication. Methods in this category address collaboration through explicit communication, which attracts increasing attention recently. iii) Learning cooperation. Many other efforts indirectly arrive at cooperation via, for example, policy parameter sharing or experience replay buffer . iv) Agents modeling agents. These methods build models to reason the behaviors of other agents for better decision-making.

These algorithms mostly focus on grid-world like scenes, despite the role of perception to some degree. Our CollaVN dataset provides a visually-rich platform for fostering studies about MARL in computer vision. We consider MAVN is of a cooperative nature and agents are situated in a partially observable environment, and concern with learning efficient communication for long-term and collaborative planning.

Learn to Communicate. Communication is a fundamental aspect of intelligence, enabling agents to behave as a group, rather than a collection of individuals. Along with the direction of communication based collaboration, MARL researchers first used predefined communication protocols , and then learnable strategies for information exchanging. However, sharing information among all agents is problematic, as communication bandwidth is limited the real world. To reduce bandwidth wasting, some recent methods adopt pre-defined communication groups or set “what, when and/or who to communicate” as a part of policy learning .

As previous methods only have an “immediate memory” during communication, they suffer from a bias towards short-sighted behaviors. We instead equip each agent with a private, external memory that persistently stores the information in all past communication rounds and enables the reuse of these information during future navigation. This enhances the long-term planning ability of our agents. This idea is also distinctively different from , where all the agents need to transmit messages to a central memory, incurring a major communication bottleneck.

Ad-Hoc Cooperation. The problem of ad-hoc team play in multi-agent cooperative games was raised in the early 2000s and is mostly studied in the robotic soccer domain . In many close MARL environments agents learn a fixed and team-dependent policy . While in ad-hoc setting agents are desired to assess and adapt to the capabilities of others to behave optimally in an open environment. See for a survey. Existing work in ad-hoc MARL mainly focuses on game settings, such as hanabi (card game) . Our work makes a further step towards exploring ad-hoc setting in visually-rich MARL; the navigation agents are required to adapt to different team sizes during deployment.

Multi-Agent Visual Navigation Task

CollaVN is a visual MARL dataset, created for the development of intelligent, cooperative navigation agents.

iGibsionV1 . CollaVN is built upon iGibsonV1, a recently released open-source 3D simulator. iGibsonV1 supports fast visual rendering in 400 fps and allows researchers to easily train and evaluate embodied navigation agents.

New Modules. To support multi-agent navigation task, three modules are developed: i) An agent initialization module is responsible for “safely” placing multiple agents at the beginning of each navigation episode, \ie, randomly placing agents but being aware of constraints of scene and physics, such as avoiding collision. ii) As the panoramic visual observation space is widely adopted in current embodied vision tasks , a panoramic camera is equipped to render agents 360∘ egocentric views in real-time. iii) A target generation module is built for generating a target picture as the goal of navigation. This module adopts a front-view camera, taking pictures around the agent height (\ie, 0.9 m) under the constraints of scene and physics.

Scenes. All the 572 full buildings in iGibsonV1 are involved in CollaVN, covering a total area of 211K m2\text{m}^{2}. The area of a physical navigation space considered in CollaVN is from 10 m ×\times 10 m to 30 m ×\times 30 m.

Agent. In CollaVN, we use Locobot, a widely used simulation agent . Its body width is 0.36 m, controlled by two wheels. The maximum translational velocity is 70 cm/s and rotational velocity is 180 deg/s.

Data Generation. Following Gibson-tiny split , we use 25/5/5 scenes for creating train/val/test data. We build three sub-dasets for different MAVN tasks, \ie, CommonGoal, SpecificGoal, Ad-hoCoop. A total of 1M/60K/120K episodes/samples are generated for train/ val/test. According to initial agent-target distance at the beginning of each episode, fine-grained annotations for val/ test data are given: {\{easy (1.5-3 m), medium (3-5 m), hard (5-10 m)}\}. We provide more details in §3.2.

Comparison ⁣{}_{\!} to ⁣{}_{\!} Previous ⁣{}_{\!} Datasets. ⁣{}_{\!} As ⁣{}_{\!} shown ⁣{}_{\!} in ⁣{}_{\!} Table 1, ⁣{}_{\!} Habitat only supports the single-agent setting. AI2THOR just supported multi-agent setting recently. However, its main focus is task planning of sequential actions to change object states and its environment space is small (only 9 m ×\times 4 m), making it not suitable for studying challenging, long-range navigation. Note that recently extended AI2THOR to involve a collaborative task, \ie, finding and lifting furniture. They are limited to highly-correlated actions between two agents and cannot explore complex multi-agent collaborative behaviors under more general situations (\eg, larger team sizes, different goals). iGibsonV2 , also built upon iGibsonV1, mainly focuses on obstacle avoidance in social scenes, \ie, only one agent is required to execute point-goal navigation while other agents are moving around (acting as pedestrians). Thus iGibsonV2 does not address collaboration among multiple navigators.

2 Task Setup

Basic Setting. MAVN assumes there are NN autonomous agents situated in a CollaVN environment. Each agent is required to navigate the environment to reach a target location (indicated by a goal image), under a cooperative setting. Formally, at the beginning of each navigation episode, each agent nn will be assigned a target goal image gng_{n}, \ie, a 3 ⁣× ⁣128 ⁣× ⁣1283\!\times\!128\!\times\!128 RGB image. Each agent does not know other agents’ positions. At each time step tt, each agent nn receives its visual observation OntO^{t}_{n}, \ie, a first-person panoramic-view 3 ⁣× ⁣512 ⁣× ⁣1283\!\times\!512\!\times\!128 RGB image. As in many communication based MARL settings , agents operate in a partially observable environment and perceive the environment from their own view. They can exchange information via a low bandwidth communication channel. Sensing and control errors and communication delay are not considered.

Action Space. The agents can control its wheels, and the maximum allowed translational and rotational velocities are 70 cm/s and 180 deg/s, respectively. At each time step tt, each agent nn takes a continuous action ant ⁣∈ ⁣2a^{t}_{n}\!\in\!^{2}, corresponding to the normalized velocities. For example, ant ⁣= ⁣(1,0.5)a^{t}_{n}\!=\!(1,0.5) means the agent will move at the maximum translational velocity (\ie, 70 cm/s) and half of the maximum rotational velocity (\ie, 90 deg/s) during [tt, t ⁣+ ⁣1t\!+\!1]. The interval between tt and t ⁣+ ⁣1t\!+\!1 is 1 s.

MAVN Task Settings. To better cover challenges in MARL and MAS, three different MAVN task settings are designed:

CommonGoal: NN robot agents are assigned a common goal, \ie, g1 ⁣= ⁣⋯ ⁣= ⁣gn ⁣= ⁣⋯ ⁣= ⁣gNg_{1}\!=\!\cdots\!=\!g_{n}\!=\!\cdots\!=\!g_{N}. Agents need to collaboratively navigate to the target goal within a maximum time step length TT. This is the most basic MAVN task setting with the purpose of investing how several agents can learn, from visual environment, to communicate so as to effectively and collaboratively solve a same given task.

SpecificGoal: The difference from CommonGoal is that, SpecificGoal assigns NN agents with NN different goal images, \ie, g1 ⁣≠ ⁣⋯ ⁣≠ ⁣gn ⁣≠ ⁣⋯ ⁣≠ ⁣gNg_{1}\!\neq\!\cdots\!\neq\!g_{n}\!\neq\!\cdots\!\neq\!g_{N}. In this setting, we are particularly interested in how agents learn communication and collaboration with different long-term targets.

Ad-hoCoop: In CommonGoal and SpecificGoal, the team size remains fixed during both training and deployment phases. The learned agents may not generalize to new configurations of teams. Ad-hoCoop is a more open world setting, in which agents must adapt to different team sizes. We test the ad-hoc team play by adding or removing agents at test time. Note that in Ad-hoCoop agents are assigned a same goal in each episode.

Task-Specific Sub-Dataset. Our CollaVN has three sub-datasets, correspond to the three MAVN tasks:

CommonGoal-CollaVN: Three subsets are created for different agent numbers (\ie, N ⁣= ⁣{2,3,4} ⁣N\!=\!\{2,3,4\}\!) and agents in each episode are assigned a same target image. For each subset, we generate 5K episodes per training scene. For each val/test scene, we generate 0.5K/1K episodes per configuration, \ie, {\{easy (1.5-3 m), medium (3-5 m), hard (5-10 m)}\}. Finally, each subset has 125K/7.5K/15K samples for train/val/test.

SpecificGoal-CollaVN: Three subsets are created for different agent numbers (\ie, N ⁣= ⁣{2,3,4} ⁣N\!=\!\{2,3,4\}\!) and agents in each episode are assigned different target images. Similarly, each subset has 125K/7.5K/15K samples for train/val/ test.

Ad-Hoc-CollaVN: Two subsets are created for different ad-hoc team size changing situations (\ie, N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3 and N ⁣ ⁣: ⁣3 ⁣→ ⁣2N\!\!:\!3\!\rightarrow\!2), and each has 125K/7.5K/15K samples for train/val/test. Agents in each episode are assigned a same target image. For N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3, its train set is the one in CommonGoal-CollaVN with N ⁣= ⁣2N\!=\!2, but its val and test sets are new generated. Each training sample in N ⁣= ⁣2N\!=\!2 contains 2 agents while each val/test sample has 3 agents. The subset of N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3 is built in a similar way.

3 Evaluation Measures

For performance evaluation, we use four metrics, where the former three are multi-agent augmented versions of SR (Success Rate), DTS (Distance to Success) and SPL (Success weighted by Path Length ), and the last one, named SSR (Success weighted by Step Ratio), is new proposed.

SR. Following , we consider an episode is successful for an agent if it reaches the target location within 1 m radius. Let τk,n ⁣∈ ⁣{0,1}\tau_{k,n}\!\in\!\{0,1\} be a binary indicator of success in episode kk for agent nn. SR is given as: 1KN ⁣∑k ⁣∑n ⁣τk,n\frac{1}{KN}\!\sum\nolimits_{k}\!\sum\nolimits_{n\!}\tau_{k,n}.

DTS. Let dk,ngd^{g}_{k,n} denote the geodesic distance between agent nn and its goal location at the end of episode kk, and dsd^{s} the success threshold (\ie, 1 m). DTS ⁣= ⁣1KN ⁣∑k ⁣ ⁣∑n ⁣dk,n ⁣g ⁣− ⁣ds\text{DTS}\!=\!\frac{1}{KN}\!\sum\nolimits_{k\!}\!\sum\nolimits_{n\!}d^{g}_{k,n\!}\!-\!d^{s}.

SPL. Let lk,ngl^{g}_{k,n} be the geodesic distance from starting position to the goal of agent nn in episode kk, and lk,nl_{k,n} the length of the path actually taken by agent nn in episode kk. We have: SPL ⁣= ⁣1KN ⁣∑k ⁣∑nτk,nlk,ng/max(lk,ng, lk,n)\text{SPL}\!=\!\frac{1}{KN}\!\sum\nolimits_{k}\!\sum\nolimits_{n}\tau_{k,n}{l^{g}_{k,n}}/{\text{max}(l^{g}_{k,n},~{}l_{k,n})}.

SSR. Let TT denote the allowed maximum time step and Tk,nT_{k,n} the number of navigation steps used by agent nn in episode kk. SSR ⁣= ⁣1KN ⁣∑k ⁣∑n ⁣τk,nT/min(T,Tk,n)\text{SSR}\!=\!\frac{1}{KN}\!\sum\nolimits_{k}\!\sum\nolimits_{n\!}\tau_{k,n}{T}/{\text{min}(T,T_{k,n})}. As MAVN is a very difficult task, in our experiments (§5.3), we found agents usually need to take many steps to reach the targets, \ie, lk,ng ⁣≪ ⁣lk,nl^{g}_{k,n}\!\ll\!l_{k,n}, making SPL very small and less discriminative. So we design SSR as a complementary.

Methodology

Task Statement. With our three MAVN tasks, we explore how agents learn from visually-rich environments to achieve collaborative navigation. In CommonGoal (SpecificGoal), agents learn to cooperatively navigate to same (different) locations, determined by target images {gn ⁣ ⁣∈ ⁣Gn}n=1N\{g_{n\!}\!\in\!\mathcal{G}_{n}\}^{N}_{n=1}. In Ad-hoCoop, agents are even expected to learn more robust policies π\bm{\pi} that are adaptive to the team size NN.

2 absent{}_{\!}Memory-Augmentedabsent{}_{\!} Communicationabsent{}_{\!} forabsent{}_{\!} MAVNabsent{}_{\!\!\!\!}

Overview. Our agents are connected by a low-bandwidth communication network without any central controller, through which they can coordinate their actions and behave as a group. As shown in Fig. 2(a), each agent mainly has two components: i) Map Building Module (§4.2.1) and ii) Memory-Augmented Communication Module (§4.2.2). For i), map building is a crucial step in navigation; an environment representation S^n\hat{\mathcal{S}}_{n} is estimated for path planning. For ii), each agent nn learns to generate useful information for communication and gather messages wn ⁣ ⁣∈ ⁣Wnw_{n\!}\!\in\!\mathcal{W}_{n} from other agents simultaneously. It is equipped with a private memory for accurately storing and recalling past communication information. For each agent nn, based on its private observation, estimated local map, communication information and target goal, its policy becomes πθn ⁣ ⁣: ⁣On ⁣ ⁣× ⁣S^n ⁣ ⁣× ⁣Wn ⁣ ⁣× ⁣Gn ⁣↦ ⁣An\bm{\pi}_{\theta_{n}\!}\!:\!\mathcal{O}_{n\!}\!\times\!\hat{\mathcal{S}}_{n\!}\!\times\!\mathcal{W}_{n\!}\!\times\!\mathcal{G}_{n\!}\mapsto\!\mathcal{A}_{n}. See §3.2 for the detailed definitions of G\mathcal{G}, A\mathcal{A}, and O\mathcal{O}. Our whole system is built as trainable end-to-end; neural networks are used as approximations for policies and modules. Corresponding descriptions are presented below.

Learning environment layouts is crucial for autonomous navigation . Inspired by , we equip each agent with a map building module Fmap\mathcal{F}^{\text{map}}, which online estimates environment structures from agent’s past percepts. For each agent nn, we denote its pose as pn ⁣= ⁣(xn,yn,on)p_{n}\!=\!(x_{n},y_{n},o_{n}), where (xn,ynx_{n},y_{n}) indicates its xy co-ordinate and ono_{n} represents its orientation in radians . We set pn0 ⁣= ⁣(0,0,0)p^{0}_{n}\!=\!(0,0,0) and let each agent build map on its ego-centric coordinate system.

For each agent nn at time step tt, Fmap\mathcal{F}^{\text{map}} takes its local observation OntO^{t}_{n}, current and last pose pnt−1:tp_{n}^{t-1:t}, prior predicted map Snt−1S^{t-1}_{n} as input, and outputs an updated map:

SnS_{n} is a (256 ⁣+ ⁣4) ⁣× ⁣L ⁣× ⁣L(256\!+\!4)\!\times\!L\!\times\!L matrix where L ⁣× ⁣LL\!\times\!L denotes the map size. The first 256 channels store visual features over all the explored locations. To save the map size, the grid size is set as 9 m2(3 m ⁣× ⁣3 m)9~{}\text{m}^{2}(3~{}\text{m}\!\times\!3~{}\text{m}) in the physical world. The last four channels include a probability map of obstacle and three binary maps storing explored area, agent past trajectory and current location, with 100 cm2(10 cm ⁣× ⁣10 cm)100~{}\text{cm}^{2}(10~{}\text{cm}\!\times\!10~{}\text{cm}) gride size. At the beginning of an episode, Sn0S^{0}_{n} is initialized with all zeros and the agent nn is put in the center of Sn0S^{0}_{n}.

2.2 Memory-Augmented Communication Module

In current communication based MARL algorithms , all the information generated in ttht^{th}-round communication are directly abandoned at the beginning of time step t ⁣+ ⁣1t\!+\!1 and new messages are generated for (t ⁣+ ⁣1)th(t\!+\!1)^{th}-round communication. However, some discarded information may be useful in the future. In addition, the communication is mainly about agents’ current observations, failing to involving their past experience explicitly. It is difficult if an agent wants to know another agent’s status in a previous time step. We instead propose a memory-augmented communication strategy, that allows agent to accurately store and recall their past communication information. It is instantiated as a handshake communication based MARL

Handshake Strategy. Three stages are involved in vanilla handshake communication, \ie, request, match, and select, and each agent acts both requester and supporter:

request. ⁣{}_{\!} At ⁣{}_{\!} the ⁣{}_{\!} beginning ⁣{}_{\!} of ⁣{}_{\!} each ⁣{}_{\!} time/communication ⁣{}_{\!} step tt, each agent nn first compresses its local observation OntO^{t}_{n} to a query μnt\bm{\mu}^{t}_{n}, key κnt\bm{\kappa}^{t}_{n}, and value υnt\bm{\upsilon}^{t}_{n}:

where μn ⁣t\bm{\mu}^{t}_{n\!} and κn ⁣t\bm{\kappa}^{t}_{n\!} are compact vectors while υn ⁣t\bm{\upsilon}^{t}_{n\!} is a high-dimension vector to fully preserve the information of OntO^{t}_{n}. Then the agent nn broadcasts μn ⁣t\bm{\mu}^{t}_{n\!} to other agents. As μn ⁣t\bm{\mu}^{t}_{n\!} is small-size, this only causes little bandwidth transmission.

match. Each of other agents derives a matching score sn,ms_{n,m} between the received query μnt\bm{\mu}^{t}_{n} and its own key κmt\bm{\kappa}^{t}_{m}:

where ψ\bm{\psi} is a learnable similarity function. sn,ms_{n,m} can be intuitively viewed as the importance of the information provided by the supporter mm for the requester nn.

select. Some connections can be removed if their matching scores are small. Then the requester nn collects information from its temporally connected supporters mtm^{t} at time step tt:

where υmtt\bm{\upsilon}^{t}_{m^{t}} is the value vector returned from the supporter mtm^{t}, and the integrated message ωnt{\bm{\omega}}^{t}_{n} is used to help the requester nn take an action decision anta_{n}^{t}.

Memory-Augmented Communication. As shown in Fig. 2(b), each agent nn maintains a private, external memory Mnt\mathcal{M}^{t}_{n}, which stores all its past generated communication information: {υn1:t−1,κn1:t−1}\{\bm{\upsilon}^{1:t-1}_{n},\bm{\kappa}^{1:t-1}_{n}\}. Then our memory-augmented handshake communication has the following four stages:

request. At the beginning of each time step tt, each agent nn generates query, key, and value vectors:

where P(⋅)(\cdot) is avg-pooling and both private observation OntO^{t}_{n} and map SntS^{t}_{n} are considered. In addition, past generated communication information, stored in Mnt\mathcal{M}^{t}_{n}, are also involved for generating a more reasonable μnt\bm{\mu}^{t}_{n}.

match. Then the agent nn broadcasts μnt\bm{\mu}^{t}_{n} to other agents. Each of other agents looks up its current key κmt\bm{\kappa}^{t}_{m}, and all the past self-generated keys κm1:t−1\bm{\kappa}^{1:t-1}_{m} stored in Mmt\mathcal{M}^{t}_{m}:

Then the final matching score and returned message are:

The agent nn also computes the correlations between the query μnt\bm{\mu}^{t}_{n} and all its current and past keys κn1:t\bm{\kappa}^{1:t}_{n} stored in Mnt\mathcal{M}^{t}_{n}:

select. To filter out some less connections, we apply an activation function Γ(sn,mt,η)\Gamma(s^{t}_{n,m},\eta) for each score sn,mts^{t}_{n,m}:

As in , only the supporters with non-zero Γ\Gamma scores are allowed to send messages to the requestor nn (\ie, learning who to communicate). Moreover, when Γ(sn,nt,η) ⁣= ⁣0\Gamma(s^{t}_{n,n},\eta)\!=\!0, that means the requestor nn thinks there is no need to establish any communication at this time step tt, (\ie, learning when to communicate). We set η ⁣= ⁣1/N\eta\!=\!1/N following . If Γ(sn,nt,η) ⁣> ⁣0\Gamma(s^{t}_{n,n},\eta)\!>\!0, the requester nn will collect information from itself and other connected supporters mtm_{t} (with Γ(sn,mtt,η) ⁣> ⁣0\Gamma(s^{t}_{n,m^{t}},\eta)\!>\!0):

store. Next, the agent nn will store its value and key information {υnt,κnt}\{\bm{\upsilon}^{t}_{n},\bm{\kappa}^{t}_{n}\} in the memory, getting Mnt\mathcal{M}^{t}_{n} for the next round communication. In our current implementation, we simply assume the private memory for each agent is large enough to store all its generated communication information over the whole navigation episode. There are several ways to improve the efficiency. One can perform store operation in a certain interval of time, and/or build the memory as a stack with fixed capacity, and/or even learn a specific write operation to see if it is necessary to store current communication information by comparing its uniqueness with existing memory slots. We put this as a part of our future work.

After ttht^{th}-round communication, each agent nn takes an navigation action through:

3 Fully Decentralized Learning

where the advantage AntA^{t}_{n} is estimated from Vϕn ⁣V_{\phi_{n}\!} and RnR_{n} , and V^nt\hat{V}_{n}^{t} indicates the expected discounted return .

Experiments

We conduct experiments on the three MAVN tasks, \ie, CommonGoal, SpecificGoal, Ad-hoCoop on the corresponding subdatasets of CollaVN, \ie, CommonGoal-CollaVN, SpecificGoal-CollaVN, and Ad-Hoc-CollaVN.

Network Details. The embedding of local visual observation OnO_{n} is from a pretrained ResNet18 (compressed in to 1024-d). The implementation of map building module Fmap\mathcal{F}^{\text{map}} (§4.2.1) mainly follows and the map size is 38.4m ⁣× ⁣38.4m38.4m\!\times\!38.4m in the physical world. Query μ\bm{\mu}, key κ\bm{\kappa}, and value υ\bm{\upsilon} in §4.2.2 are 256-d, 256-d, 2048-d vectors, respectively. All the functions in §4.2.2 are one-layer MLPs. The policy π\bm{\pi} and value function VV are GRUs .

Evaluation. For different MALN tasks, we adopt test sets of corresponding sub-datasets in CollaVN for benchmarking, measured by (§3.3) SR, DTS, SPL and SSR. The maximum episode length is set as T ⁣= ⁣80T\!=\!80 steps.

Training. Our model was trained using dense reward of geodesic distance reduced to the goal location and constant -0.05 step penalty to collision.

Reproducibility. Our model is implemented in PyTorch and trained on four NVIDIA Tesla V100 GPUs with a 32GB memory per-card. To reveal full details of our algorithm, our implementations are released.

2 Baseline

Following baselines are used in our experiments (their sub-modules are the same with ours unless specified):

Random walk is the simplest heuristic for navigation: agent randomly selects an action at each step.

IL learns an off policy by imitation learning (IL). One agent is trained and several copies are used for testing.

MARL w/o com learns target-driven MARL policy without explicit communication, using IPPO .

MARL w/o mem can be viewed as a variant of our model without using memory.

3 Experimental Results

Performance on CommonGoal Task. For the settings with different agent numbers, \ie, N ⁣= ⁣2,3,4N\!=\!2,3,4 on the CommonGoal-CollaVN sub-dataset, all the benchmarked models are trained on the corresponding train sets and tested on the corresponding test sets. The results are summarized in Table 2. As seen, our model outperforms over all other baselines across all the metrics. Interestingly, we can also observe IL based baseline, \ie, IL-M, even shows comparable performance with a MARL baseline, \ie, MARL w/o com, revealing the difficulty of this task.

Performance on SpecificGoal Task. Also, for different agent numbers on the SpecificGoal-CollaVN sub-dataset, we train and evaluate the models on the corresponding train and test sets. From Table 3 we can find that our method sill gets better performance. However, compared with the results in Table 2, all the methods suffer from performance drop, which indicates that it is harder to learn communication when agents perform different tasks.

Performance on Ad-hoCoop Task. For N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3 setting in the Ad-Hoc-CollaVN dataset, each of train scenes contains two agents while each of test scenes have three agents. Similarly, for N ⁣: ⁣3 ⁣→ ⁣2N\!:\!3\!\rightarrow\!2 setting, each of train scenes contains three agents while each of test scenes have two agents. For the two settings, all the methods are trained and test on the corresponding train and test sets. The results in Table 4 show again our model performs better on the two ad-hoc team play settings, mainly due to our fully decentralized learning strategy.

Effect on Memory-Based Communication. Compared with MARL w/o mem, our model shows significantly better performance over all the three MAVN settings. MARL w/o mem only shows limited improvements over MARL w/o com. These observations suggest the importance of long-term memory in MAVN.

Conclusion

A new dataset was introduced for multi-agent collaborative navigation in complex environments. A memory-based communication model was presented to address the reuse of past communication information and facilitate cooperation. Our experiments showed that there is still huge room for further improvement. This work is expected to inspire more future efforts towards this promising direction.

References

Appendix A Details of CollaVN

Our CollaVN dataset has three subdatasets, corresponding to the three MAVN tasks: CommonGoal, SpecificGoal and Ad-HoCoop.

CommonGoal-CollaVN: It contains three subsets, which are created for different agent numbers (\ie, N ⁣= ⁣{2,3,4} ⁣N\!=\!\{2,3,4\}\!), and agents in each episode are assigned a same target image. For each subset of different agent number, we randomly sample 5K episodes per training scene. For each scene in val/test split, we generate 0.5K/1K episodes per configuration, \ie, {\{easy (1.5-3 m), medium (3-5 m), hard (5-10 m)}\}. The distance range means that the initial start and target position of all agents are within this limit. Finally, each subset has 125K/7.5K/15K samples for train/val/test, \ie, 25 ⁣× ⁣5K/3 ⁣× ⁣5 ⁣× ⁣0.5K/3 ⁣× ⁣5 ⁣× ⁣1K ⁣= ⁣125K/7.5K/15K25\!\times\!5\text{K}/3\!\times\!5\!\times\!0.5\text{K}/3\!\times\!5\!\times\!1\text{K}\!=\!125\text{K}/7.5\text{K}/15\text{K}. Detailed statistics of CommonGoal-CollaVN subdataset are shown in Table 5.

SpecificGoal-CollaVN: Three subsets are created for different agent numbers (\ie, N ⁣= ⁣{2,3,4} ⁣N\!=\!\{2,3,4\}\!), but agents in each episode are assigned different target images. Similarly, each subset has 125K/7.5K/15K samples for train/val/ test. Detailed statistics of SpecificGoal-CollaVN subdataset are shown in Table 6.

Ad-Hoc-CollaVN: Two subsets are created for different ad-hoc team size changing situations (\ie, N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3 and N ⁣ ⁣: ⁣3 ⁣→ ⁣2N\!\!:\!3\!\rightarrow\!2), and each has 125K/7.5K/15K samples for train/val/test. Agents in each episode are assigned a same target image. For N ⁣ ⁣: ⁣2 ⁣→ ⁣3N\!\!:\!2\!\rightarrow\!3, its train set is the one in CommonGoal-CollaVN with N ⁣= ⁣2N\!=\!2, but its val and test sets are new generated. Each training sample in N ⁣= ⁣2N\!=\!2 contains 2 agents while each val/test sample has 3 agents. The subset of N ⁣ ⁣: ⁣3 ⁣→ ⁣2N\!\!:\!3\!\rightarrow\!2 is built in a similar way. Detailed statistics of Ad-Hoc-CollaVN subdataset are shown in Table 7.

Appendix B Implementation Details

During training, rewards are provided to all agents individually at every step. The reward includes two terms: i) +1.00 ⁣× ⁣△d+1.00\!\times\!\triangle d, where dd is the geodesic distance between agent’s current location and target location, and △d\triangle d refers to the distance change after adapting a navigation action at this step. +1.00+1.00 is the weight of goal-driven reward. ii) A penalty of −0.05-0.05 whenever the agent collides with environment or other agents. The maximum total reward achievable for a single agent is +D+D, where DD is the initial distance between agent and goal, \ie, reaching the goal through the shortest path without any collision. Our models are trained to maximize the expected discounted cumulative gain with a discounting factor γ=0.99\gamma=0.99.

B.2 Hyperparameter Details

We use PyTorch for implementing and training our model. As for map building, we follow , maintain a FIFO memory of size 5000 for training. After one step in each thread, we perform 10 updates to the map building module with a batch size of 64.

During training, we train our agents with 25 parallel threads, with each thread using one scene in the training set. We use Proximal Policy Optimization (PPO) , with 5 mini-batches, and 8 epochs in each PPO update. Our PPO implementation is based on . We use Adam optimizer with a learning rate of 0.00001, a discount factor of γ\gamma = 0.99, an entropy coefficient of 0.001, value loss coefficient of 0.5 for training.

Appendix C Evaluation Details

Why SSR instead of SPL? SPL was introduced in :

where NN is the agent number, KK equals the number of test episodes, τk,n ⁣∈ ⁣{0,1}\tau_{k,n}\!\in\!\{0,1\} is a binary indicator of success in episode kk for agent nn, lk,ngl^{g}_{k,n} is the geodesic distance from starting position to the goal, and lk,nl_{k,n} the length of the path actually taken by agent.

The value of lk,ng/max(lk,ng, lk,n){l^{g}_{k,n}}/{\text{max}(l^{g}_{k,n},~{}l_{k,n})} is always less than one (<<1). If we perform a long-term or difficult navigation task, the success rate is usually low and the navigation path is very long. In this case, SPL will be very small and lose discriminability. This phenomenon can be observed in Table 8, which shows the performance of our model on CommonGoal task with medium difficulty and different team sizes. SPL basically stays still with the success rate changing.

To resolve this problem, we introduce SSR (Success weighted by Step Ratio):

TT denotes the allowed maximum time step and Tk,nT_{k,n} the number of navigation steps used by agent nn in episode kk. As shown in Table 8, SSR is more discriminative than SPL.

Appendix D Qualitative Results

In this section, we provide more intuitive results of our model.

The visualization of navigation trajectories of three task settings with different team size (N ⁣= ⁣2,3N\!=\!2,3) are shown in Fig. 3, Fig. 5, and Fig. 6. The start and target locations are represented by blue circles and orange circles, respectively. From these figures, we can observe that the agents trained with our method is able to learn a robust policy to collaboratively explore the environment and successfully reaches the target location. From Fig. 6, we find that our method can adapt to different team sizes at test times well.

D.2 Effect of Communication

To show the effectiveness of communication in MVAN task, we demonstrate the trajectory of a single agent system in Fig. 4. We compare the effect by deploying this agents for a particular initialization of an episode, \ie, the scene, agents’ start location and target location are same as CommonGoal task with team size of two (Top figure in Fig. 3). As shown in Fig. 4, the single agent failed to reach the target position within maximum time step (it spends too much time on exploring a wrong room). However, in CommonGoal task, we find that implicit communication can allow agents to avoid this wrong exploration as much as possible based on the experience of other agents and achieve the task faster.