TarMAC: Targeted Multi-Agent Communication

Abhishek Das, Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Michael Rabbat, Joelle Pineau

Introduction

Effective communication is a key ability for collaboration. Indeed, intelligent agents (humans or artificial) in real-world scenarios can significantly benefit from exchanging information that enables them to coordinate, strategize, and utilize their combined sensory experiences to act in the physical world. The ability to communicate has wide-ranging applications for artificial agents – from multi-player gameplay in simulated (e.g. DoTA, StarCraft) or physical worlds (e.g. robot soccer), to self-driving car networks communicating with each other to achieve safe and swift transport, to teams of robots on search-and-rescue missions deployed in hostile, fast-evolving environments.

A salient property of human communication is the ability to hold targeted interactions. Rather than the ‘one-size-fits-all’ approach of broadcasting messages to all participating agents, as has been previously explored (Sukhbaatar et al. 2016; Foerster et al. 2016; Singh et al. 2019), it can be useful to direct certain messages to specific recipients. This enables a more flexible collaboration strategy in complex environments. For example, within a team of search-and-rescue robots with a diverse set of roles and goals, a message for a fire-fighter (e.g. “smoke is coming from the kitchen”) is largely meaningless for a bomb-defuser.

We develop TarMAC, a Targeted Multi-Agent Communication architecture for collaborative multi-agent deep reinforcement learning. Our key insight in TarMAC is to allow each individual agent to actively select which other agents to address messages to. This targeted communication behavior is operationalized via a simple signature-based soft attention mechanism: along with the message, the sender broadcasts a key which encodes properties of agents the message is intended for, and is used by receivers to gauge the relevance of the message. This communication mechanism is learned implicitly, without any attention supervision, as a result of end-to-end training using task reward.

The inductive bias provided by soft attention in the communication architecture is sufficient to enable agents to 1) communicate agent-goal-specific messages (e.g. guide fire-fighter towards fire, bomb-defuser towards bomb, etc.), 2) be adaptive to variable team sizes (e.g. the size of the local neighborhood a self-driving car can communicate with changes as it moves), and 3) be interpretable through predicted attention probabilities that allow for inspection of which agent is communicating what message and to whom.

Our results however show that just using targeted communication is not enough. Complex real-world tasks might require large populations of agents to go through multiple rounds of collaborative communication and reasoning, involving large amounts of information to be persistent in memory and exchanged via high-bandwidth communication channels. To this end, our actor-critic framework combines centralized training with decentralized execution (Lowe et al. 2017), thus enabling scaling to large team sizes. In this context, our inter-agent communication architecture also supports multiple rounds of targeted interactions at every time-step, wherein the agents’ recurrent policies persist relevant information in internal states.

While natural language, i.e. a finite set of discrete tokens with pre-specified human-conventionalized meanings, may seem like an intuitive protocol for inter-agent communication – one that enables human-interpretability of interactions – forcing machines to communicate among themselves in discrete tokens presents additional training challenges. Since our work focuses on machine-only multi-agent teams, we allow agents to communicate via continuous vectors (rather than discrete symbols), as has been explored in (Sukhbaatar et al. 2016; Singh et al. 2019), and agents have the flexibility to discover and optimize their communication protocol as per task requirements.

We provide extensive empirical evaluation of our approach across a range of tasks, environments, and team sizes.

We begin by benchmarking TarMAC and its ablation without attention on a cooperative navigation task derived from the SHAPES environment (Andreas et al. 2016) in Section 5.1. We show that agents learn intuitive attention behavior across task difficulties.

Next, we evaluate TarMAC on the traffic junction environment (Sukhbaatar et al. 2016) in Section 5.2, and show that agents are able to adaptively focus on ‘active’ agents in the case of varying team sizes.

We then demonstrate its efficacy in 33D environments with a cooperative first-person point-goal navigation task in House33D (Wu et al. 2018) (Section 5.3).

Finally, in Section 5.4, we show that TarMAC can be easily combined with IC3Net (Singh et al. 2019), thus extending its applicability to mixed and competitive environments, and leading to significant improvements in performance and sample complexity.

Related Work

Multi-agent systems fall at the intersection of game theory, distributed systems, and Artificial Intelligence in general (Shoham & Leyton-Brown 2008), and thus have a rich and diverse literature. Our work builds on and is related to prior work in deep multi-agent reinforcement learning, the centralized training and decentralized execution paradigm, and emergent communication protocols.

Multi-Agent Reinforcement Learning (MARL). Within MARL (see Busoniu et al. 2008 for a survey), our work is related to efforts on using recurrent neural networks to approximate agent policies (Hausknecht & Stone 2015), stabilizing algorithms for multi-agent training (Lowe et al. 2017; Foerster et al. 2018), and tasks in novel domains e.g. coordination and navigation in 3D environments (Peng et al. 2017; OpenAI 2018; Jaderberg et al. 2018).

Centralized Training & Decentralized Execution. Both Sukhbaatar et al. 2016 and Hoshen 2017 adopt a centralized framework at both training and test time – a central controller processes local observations from all agents and outputs a probability distribution over joint actions. In this setting, the controller (e.g. a fully-connected network) can be viewed as implicitly encoding communication. Sukhbaatar et al. 2016 propose an efficient controller architecture that is invariant to agent permutations by virtue of weight-sharing and averaging (as in Zaheer et al. 2017), and can, in principle, also be used in a decentralized manner at test time since each agent just needs its local state vector and the average of incoming messages to take an action. Meanwhile, Hoshen 2017 proposes to replace averaging by an attentional mechanism to allow targeted interactions between agents. While closely related to our communication architecture, this work only considers fully-supervised one-next-step prediction tasks, while we study the full reinforcement learning problem with tasks requiring planning over long time horizons.

Moreover, a centralized controller quickly becomes intractable in real-world tasks with many agents and high-dimensional observation spaces e.g. navigation in House3D (Wu et al. 2018). To address these weaknesses, we adopt the framework of centralized learning but decentralized execution (following Foerster et al. 2016; Lowe et al. 2017) and further relax it by allowing agents to communicate. While agents can use extra information during training, at test time, they pick actions solely based on local observations and communication messages.

Emergent Communication Protocols. Our work is also related to recent work on learning communication protocols in a completely end-to-end manner with reinforcement learning – from perceptual input (e.g. pixels) to communication symbols (discrete or continuous) to actions (e.g. navigating in an environment). While (Foerster et al. 2016; Jorge et al. 2016; Das et al. 2017; Kottur et al. 2017; Mordatch & Abbeel 2017; Lazaridou et al. 2017) constrain agents to communicate with discrete symbols with the explicit goal to study emergence of language, our work operates in the paradigm of learning a continuous communication protocol in order to solve a downstream task (Sukhbaatar et al. 2016; Hoshen 2017; Jiang & Lu 2018; Singh et al. 2019). Jiang & Lu 2018; Singh et al. 2019 also operate in a decentralized execution setting and use an attentional communication mechanism, but in contrast to our work, they use attention to decide when to communicate, not who to communicate with. In Section 5.4, we discuss how to potentially combine the two approaches.

Table 1 summarizes the main axes of comparison between our work and previous efforts in this exciting space.

Technical Background

Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs). A Dec-POMDP is a multi-agent extension of a partially observable Markov decision process (Oliehoek 2012). For NN agents, it is defined by a set of states SS describing possible configurations of all agents, a global reward function RR, a transition probability function TT, and for each agent i∈1,...,Ni\in{1,...,N} a set of allowed actions AiA_{i}, a set of possible observations Ωi\Omega_{i} and an observation function OiO_{i}. At each time step every agent picks an action aia_{i} based on its local observation ωi\omega_{i} following its own stochastic policy πθi(ai∣ωi)\pi_{\theta_{i}}(a_{i}|\omega_{i}). The system randomly transitions to the next state s′s^{\prime} given the current state and joint action T(s′∣s,a1,...,aN)T(s^{\prime}|s,a_{1},...,a_{N}). The agent team receives a global reward r=R(s,a1,...,aN)r=R(s,a_{1},...,a_{N}) while each agent receives a local observation of the new state Oi(ωi∣s′)O_{i}(\omega_{i}|s^{\prime}). Agents aim to maximize the total expected return J=∑t=0TγtrtJ=\sum_{t=0}^{T}\gamma^{t}r_{t} where γ\gamma is a discount factor and TT is the episode time horizon.

where Qπ(s,a)Q_{\pi}(s,a) is the action-value. It is the expected remaining discounted reward if we take action aa in state ss and follow policy π\pi thereafter. Actor-Critic algorithms learn an approximation Q^(s,a)\hat{Q}(s,a) of the unknown true action-value function by e.g. temporal-difference learning (Sutton & Barto 1998). This Q^(s,a)\hat{Q}(s,a) is the Critic and πθ\pi_{\theta} is the Actor.

Multi-Agent Actor-Critic. Lowe et al. 2017 propose a multi-agent Actor-Critic algorithm adapted to centralized learning and decentralized execution wherein each agent learns its own policy πθi(ai∣ωi)\pi_{\theta_{i}}(a_{i}|\omega_{i}) conditioned on local observation ωi\omega_{i} using a central Critic that estimates the joint action-value Q^(s,a1,...,aN)\hat{Q}(s,a_{1},...,a_{N}) conditioned on all actions.

TarMAC: Targeted Multi-Agent Communication

We now describe our multi-agent communication architecture in detail. Recall that we have NN agents with policies {π1,...,πN}\{\pi_{1},...,\pi_{N}\}, respectively parameterized by {θ1,...,θN}\{\theta_{1},...,\theta_{N}\}, jointly performing a cooperative task. At every timestep tt, the iith agent for all i∈{1,...,N}i\in\{1,...,N\} sees a local observation ωit\omega_{i}^{t}, and must select a discrete environment action ait∼πθia_{i}^{t}\sim\pi_{\theta_{i}} and send a continuous communication message mitm_{i}^{t}, received by other agents at the next timestep, in order to maximize global reward rt∼Rr_{t}\sim R. Since no agent has access to the underlying complete state of the environment sts_{t}, there is incentive in communicating with each other and being mutually helpful to do better as a team.

Policies and Decentralized Execution. Each agent is essentially modeled as a Dec-POMDP augmented with communication. Each agent’s policy πθi\pi_{\theta_{i}} is implemented as a 11-layer Gated Recurrent Unit (Cho et al. 2014). At every timestep, the local observation ωit\omega_{i}^{t} and a vector citc_{i}^{t} aggregating messages sent by all agents at the previous timestep (described in more detail below) are used to update the hidden state hith_{i}^{t} of the GRU, which encodes the entire message-action-observation history up to time tt. From this internal state representation, the agent’s policy πθi(ait ∣ hit)\pi_{\theta_{i}}\left(a_{i}^{t}\,|\,h_{i}^{t}\right) predicts a categorical distribution over the space of actions, and another output head produces an outgoing message vector mitm_{i}^{t}. Note that for our experiments, agents are symmetric and policy parameters are shared across agents, i.e. θ1=...=θN\theta_{1}=...=\theta_{N}. This considerably speeds up learning.

Note that compared to an individual Critic Q^i(hit,ait)\hat{Q}_{i}(h_{i}^{t},a_{i}^{t}) per agent, having a centralized Critic leads to considerably lower variance in policy gradient estimates since it takes into account actions from all agents. At test time, the Critic is not needed and policy execution is fully decentralized.

used to compute cjt+1c_{j}^{t+1}, the input message for agent jj at t+1t+1:

Intuitively, attention weights are high when both sender and receiver predict similar signature and query vectors respectively. Note that Equation 2 also includes αii\alpha_{ii} corresponding to the ability to self-attend (Vaswani et al. 2017), which we empirically found to improve performance, especially in situations when an agent has found the goal in a coordinated navigation task and all it is required to do is stay at the goal, so others benefit from attending to this agent’s message but return communication is not necessary. Note that the targeting mechanism in our formulation is implicit i.e. agents implicitly encode properties without addressing specific recipients. For example, in a self-driving car network, a particular message may be for “cars travelling on the west to east road" (implicitly encoding properties) as opposed to specifically for “car 2” (explicit addressing).

For multi-round communication, aggregated message vector cjt+1c_{j}^{t+1} and internal state hjth_{j}^{t} are first used to predict the next internal state h′jt{h^{\prime}}_{j}^{t} taking into account the first round:

Next, the updated hidden state h′jt{h^{\prime}}_{j}^{t} is used to predict signature, query, value followed by repeating Equations 1-4 for multiple rounds until we get a final aggregated message vector cjt+1c_{j}^{t+1} to be used as input at the next timestep. Number of rounds of communication is treated as a hyperparameter.

Our entire communication architecture is differentiable, and message vectors are learnt through backpropagation.

Experiments

We evaluate TarMAC on a variety of tasks and environments. All our models were trained with a batched synchronous version of the multi-agent Actor-Critic described above, using RMSProp with a learning rate of 7×10−47\times 10^{-4} and α=0.99\alpha=0.99, batch size 1616, discount factor γ=0.99\gamma=0.99 and entropy regularization coefficient 0.010.01 for agent policies. All our agent policies are instantiated from the same set of shared parameters; i.e. θ1=...=θN\theta_{1}=...=\theta_{N}. Each agent’s GRU hidden state is 128128-d, message signature/query is 1616-d, and message value is 3232-d (unless specified otherwise). All results are averaged over 55 independent seeds (unless noted otherwise), and error bars show standard error of means.

The SHAPES dataset was introduced by Andreas et al. 2016 github.com/jacobandreas/nmn2/tree/shapes, and originally created for testing compositional visual reasoning for the task of visual question answering. It consists of synthetic images of 22D colored shapes arranged in a grid (3×33\times 3 cells in the original dataset) along with corresponding question-answer pairs. There are 33 shapes (circle, square, triangle), 33 colors (red, green, blue), and 22 sizes (small, big) in total (see Figure 2).

We convert each image from SHAPES into an active environment where agents can now be spawned at different regions of the image, observe a 5×55\times 5 local patch around them and their coordinates, and take actions to move around – {\{up, down, left, right, stay}\}. Each agent is tasked with navigating to a specified goal state in the environment within a max no. of steps – {\{‘red’, ‘blue square’, ‘small green circle’, etc. }\} – and the reward for each agent at every timestep is based on team performance i.e. rt=# agents on goal# agentsr_{t}=\frac{\textnormal{\# agents on goal}}{\textnormal{\# agents}}. Having a symmetric, team-based reward incentivizes agents to cooperate in finding each agent’s goal.

How does targeting work? Recall that each agent predicts a signature and value vector as part of the message it sends, and a query vector to attend to incoming messages. The communication is targeted because the attention probabilities are a function of both the sender’s signature and receiver’s query vectors. So it is not just the receiver deciding how much of each message to listen to. The sender also sends out signatures that affects how much of each message is sent to each receiver. The sender’s signature could encode parts of its observation most relevant to other agents’ goals (e.g. it would be futile to convey coordinates in the signature), and the message value could contain the agent’s own location. For example, in Figure 2, at t=6t=6, we see that when agent 2 passes by blue, agent 4 starts attending to agent 2. Here, agent 2’s signature likely encodes the color it observes (which is blue), and agent 4’s query encodes its goal (which is also blue) leading to high attention probability. Agent 2’s message value encodes coordinates agent 4 has to navigate to, which it ends up reaching by t=21t=21.

SHAPES serves as a flexible testbed for carefully controlling and analyzing the effect of changing the size of the environment, no. of agents, goal configurations, etc. Figure 2 visualizes learned protocols, and Table 2 reports quantitative evaluation for three different configurations – 1) 4 agents, all tasked with finding red in 30×3030\times 30 images, 2) 4 agents, all tasked with finding red in 50×5050\times 50 images, 3) 4 agents, tasked with finding [[red,red,green,blue]] respectively in 50×5050\times 50 images. We compare TarMAC against two baselines – 1) without communication, and 2) with communication but where broadcasted messages are averaged instead of attention-weighted, so all agents receive the same message vector, similar to Sukhbaatar et al. 2016. Benefits of communication and attention increase with task complexity (30×30→50×5030\times 30\rightarrow 50\times 50 & find[[red]] →\rightarrow find[[red,red,green,blue]]).

2 Traffic Junction

Environment and Task. The simulated traffic junction environments from (Sukhbaatar et al. 2016) consist of cars moving along pre-assigned, potentially intersecting routes on one or more road junctions. The total number of cars is fixed at NmaxN_{\textnormal{max}}, and at every timestep new cars get added to the environment with probability parrivep_{\textnormal{arrive}}. Once a car completes its route, it becomes available to be sampled and added back to the environment with a different route assignment. Each car has a limited visibility of a 3×33\times 3 region around it, but is free to communicate with all other cars. The action space for each car at every timestep is gas and brake, and the reward consists of a linear time penalty −0.01τ-0.01\tau, where τ\tau is the number of timesteps since car has been active, and a collision penalty rcollision=−10r_{\textnormal{collision}}=-10.

Quantitative Results. We compare our approach with CommNet (Sukhbaatar et al. 2016) on the easy and hard difficulties of the traffic junction environment. The easy task has one junction of two one-way roads on a 7×77\times 7 grid with Nmax=5N_{\textnormal{max}}=5 and parrive=0.30p_{\textnormal{arrive}}=0.30, while the hard task has four connected junctions of two-way roads on a 18×1818\times 18 grid with Nmax=20N_{\textnormal{max}}=20 and parrive=0.05p_{\textnormal{arrive}}=0.05. See Figure 4(a), 4(b) for an example of the four two-way junctions in the hard task. As shown in Table 3, a no communication baseline has success rates of 84.984.9±4.3%\pm 4.3\% and 74.174.1±3.9%\pm 3.9\% on easy and hard respectively. On easy, both CommNet and TarMAC get close to 100%100\%. On hard, TarMAC with 11-round significantly outperforms CommNet with a success rate of 84.684.6±3.2%\pm 3.2\%, while 22-round further improves on this at 97.197.1±1.6%\pm 1.6\%, which is an ∼\sim18%18\% absolute improvement over CommNet. We did not see gains going beyond 22 rounds in this environment.

Message size vs. multi-round communication. We study performance of TarMAC with varying message value size and number of rounds of communication on the hard variant of the traffic junction task. As can be seen in Figure 3, multiple rounds of communication leads to significantly higher performance than simply increasing message size, demonstrating the advantage of multi-round communication. In fact, decreasing message size to a single scalar performs almost as well as 6464-d, perhaps because even a single real number can be sufficiently partitioned to cover the space of meanings/messages that need to be conveyed.

Model Interpretation. Interpreting the learned policies of TarMAC, Figure 4(a) shows braking probabilities at different locations – cars tend to brake close to or right before entering traffic junctions, which is reasonable since junctions have the highest chances for collisions. Turning our attention to attention probabilities (Figure 4(b)), we can see that cars are most-attended to when in the ‘internal grid’ – right after crossing the 11st junction and before hitting the 22nd junction. These attention probabilities are intuitive – cars learn to attend to specific sensitive locations with the most relevant local observations to avoid collisions. Finally, Figure 4(c) compares total number of cars in the environment vs. number of cars being attended to with probability >0.1>0.1 at any time. Interestingly, these are (loosely) positively correlated, with Spearman’s σ=0.49\sigma=0.49, which shows that TarMAC is able to adapt to variable number of agents. Crucially, agents learn this dynamic targeting behavior purely from task rewards with no hand-coding! Note that the right shift between the two curves is expected, as it takes a few timesteps of communication for team size changes to propagate. At a relative time shift of 33, the Spearman’s rank correlation between the two curves goes up to 0.530.53.

3 House3D

Next, we benchmark TarMAC on a cooperative point-goal navigation task in House3D (Wu et al. 2018). House3D provides a rich and diverse set of publicly-available github.com/facebookresearch/house3d 33D indoor environments, wherein agents do not have access to the top-down map and must navigate purely from first-person vision. Similar to SHAPES, the agents are tasked with finding a specified goal (such as ‘fireplace’) within a max no. of steps, spawned at random locations in the environment and allowed to communicate and move around. Each agent gets a shaped reward based on progress towards the specified target. An episode is successful if all agents end within 0.50.5m of the target object in 500500 navigation steps.

Table 4 shows success rates on a find[fireplace] task in House3D. A no-communication navigation policy trained with the same reward structure gets a success rate of 62.162.1±5.3%\pm 5.3\%. Mean-pooled communication (no attention) performs slightly better with a success rate of 64.364.3±2.3%\pm 2.3\%, and TarMAC achieves the best success rate at 68.968.9±1.1%\pm 1.1\%. TarMAC agents take 82.582.5 steps to reach the target on average vs. 101.3101.3 for no attention vs. 186.5186.5 for no communication. Figure 5 visualizes a predicted navigation trajectory of 44 agents. Note that the communication vectors are significantly more compact (3232-d) than the high-dimensional observation space (224×224224\times 224 image), making our approach particularly attractive for scaling to large agent teams.

Note that House3D is a challenging testbed for multi-agent reinforcement learning. To get to ∼\sim100%100\% accuracy, agents have to deal with high-dimensional visual observations, be able to navigate long action sequences (up to ∼\sim500500 steps), and avoid getting stuck against objects, doors, and walls.

4 Mixed and Competitive Environments

Finally, we look at how to extend TarMAC to mixed and competitive scenarios. Communication via sender-receiver soft attention in TarMAC is poorly suited for competitive scenarios, since there is always “leakage” of the agent’s state as a message to other agents via a low but non-zero attention probability, thus compromising its strategy and chances of success. Instead, an agent should first be able to independently decide if it wants to communicate at all, and then direct its message to specific recipients if it does.

The recently proposed IC3Net architecture by Singh et al. 2019 addresses the former – learning when to communicate. At every timestep, each agent in IC3Net predicts a hard gating action to decide if it wants to communicate. At the receiving end, messages from agents who decide to communicate are averaged to be the next input message. Replacing this message averaging with our sender-receiver soft attention, while keeping the rest of the architecture and training details the same as IC3Net, should provide an inductive bias for more flexible communication strategies, since this model (IC3Net + TarMAC) can learn both when to communicate and whom to address messages to.

We evaluate IC3Net + TarMAC on the Predator-Prey environment from Singh et al. 2019, consisting of nn predators, with limited vision, moving around (with a penalty of rexplore=−0.05r_{\text{explore}}=-0.05 per timestep) in search of a stationary prey. Once a predator reaches a prey, it keeps getting positive reward rprey=0.05r_{\text{prey}}=0.05 till end of episode i.e. till other agents reach prey or maximum no. of steps. The prey gets 0.050.05 per timestep only till the first predator reaches it, so it has incentive to not communicate its location. We compare average no. of steps for agents to reach the prey during training (Figure 6) and at convergence (Table 5). Figure 6 shows that using TarMAC with IC3Net leads to significantly faster convergence than IC3Net alone, and Table 5 shows that TarMAC agents reach the prey faster. Results are averaged over 33 independent runs with different seeds.

Conclusions and Future Work

We introduced TarMAC, an architecture for multi-agent reinforcement learning that allows targeted continuous communication between agents via a sender-receiver soft attention mechanism and multiple rounds of collaborative reasoning. Evaluation on four diverse environments shows that our model is able to learn intuitive communication attention behavior and improves performance, even in non-cooperative settings, with task reward as sole supervision. While TarMAC uses continuous vectors as messages, it is possible to force these to be discrete, either during training itself (as in Foerster et al. 2016) or by adding a decoder after learning to ground these messages into symbols.

In future, we aim to exhaustively benchmark TarMAC on more challenging 33D navigation tasks because we believe this is where decentralized targeted communication is most crucial, as it allows scaling to a large number of agents with high-dimensional observation spaces. In particular, we are interested in investigating combinations of TarMAC with recent advances in spatial memory, planning networks, etc.

References