PECAN: Leveraging Policy Ensemble for Context-Aware Zero-Shot Human-AI Coordination

Xingzhou Lou, Jiaxian Guo, Junge Zhang, Jun Wang, Kaiqi Huang, Yali Du

Introduction

Reinforcement learning (RL) has shown remarkable success in various domains, such as gaming AI Silver et al. (2017); Vinyals et al. (2019a); Du et al. (2019); Han et al. (2019); Shi et al. (2023), robotic manipulation Liu et al. (2022); Fang et al. (2019), traffic control Du et al. (2021); Du et al. (2022), etc. However, a significant challenge remains in constructing agents that can collaborate effectively with unseen partners, which is especially important under human-AI coordination. Many real-world applications of human-AI coordination, such as cooperative games Games (2016), self-driving vehicles Resnick et al. (2018); Mariani et al. (2021) and AI assistants Kakish et al. (2019); Andrychowicz et al. (2020), can be modeled as zero-shot human-AI coordination tasks. By avoiding the expensive human data collection and human involvement during training, zero-shot human-AI coordination holds the promise of more accessible AI systems that can enhance human capabilities. In this approach, an ego agent is trained with partner agents and later interacts with human proxy models or real humans.

Existing methods mainly vary in how the partner agents are acquired and how the ego agent is trained. Self-play methods tried to train ego agents through self play Silver et al. (2017, 2018); Brown and Sandholm (2018, 2019), in which the ego agent is trained to collaborate with a copy of itself. However, this approach has been shown to result in over-fitting to a single cooperative pattern Hu et al. (2020); Lupu et al. (2021) and poor generalization to real human collaborations. To address this issue, population-based training (PBT) Strouse et al. (2021); Zhao et al. (2021); Lupu et al. (2021) has been employed, where a population of diverse partner agents is used to train the ego agent. The variety of behaviors exhibited by these diverse partners can prevent over-fitting to a single cooperative pattern and improve generalization ability of the ego agent when cooperating with real humans.

However, there are still two limitations of current PBT methods: 1) The finite number of partners in the population restricts the behavioral diversity, making it difficult for the ego agent to coordinate with new partners, particularly those with unique behaviors. Although increasing the population size can address this issue, it requires significant computational resources and decreases the learning efficiency of the ego agent. 2) The ego agent learns a common best response (BR) for every partner, regardless of their behavior patterns. This partner-specific common BR can lead to unsatisfactory human-AI coordination performance as the ego agent lacks the ability to adapt its policy based on the partner’s type and behavior pattern.

To address these issues, we propose Policy Ensemble Context-Aware zero-shot human-AI coordinatioN (PECAN), where the policy ensemble method is proposed to increase the diversity of partners without increasing the population size, and the context-aware module is proposed to identify whether the partner is good or poor at the given task, i.e. the level of coordination skills. Thus, the ego agent is able to learn level-based common BR rather than common BR for specific partners in the population, which allows the ego agent to acquire more universal coordination behaviors and better coordinate with novel partners.

Specifically, the proposed policy ensemble can generate a new partner whose policy is the weighted average of policy primitives in the population. Since the weights are randomly generated, the policy-ensemble partners are distinct in each iteration. This increases partner diversity and improves the ego agent’s ability to collaborate with unseen partners. Additionally, we find that partners created by mixing policies from the same level display better behavioral diversity than those created from the entire population (see Fig. 4 and section 5.2). The context-aware module in PECAN is designed to identify the partner’s level of coordination skills based on past trajectories. This is achieved through supervised learning. The training data is collected by rolling out various policy ensembles and assigning the corresponding levels of the ensembles’ policy primitives as labels. During evaluation, the ego agent updates its recognized context at the start of each episode based on the past trajectory and uses this context to condition its actions.

Our proposed approach, PECAN, presents three main contributions to zero-shot human-AI coordination. 1) PECAN trains generalizable agents without relying on human data. 2) The use of level-based policy-ensemble partners and a context-aware module enhances population diversity without increasing the population size, allowing the ego agent to learn level-based common best responses. 3) PECAN achieves superior performance compared to state-of-the-art baselines in the Overcooked environment Games (2016), as demonstrated by our experimental results and additional studies.

Related Work

Zero-shot Coordination Zero-shot coordination (ZSC) has been studied in multiple previous studies Cui et al. (2021); Treutlein et al. (2021); Ribeiro et al. (2022). In ZSC framework introduced by Hu et al. (2020), two independently trained agents are paired together to fulfill a common purpose in a cooperative game. The paired agents will never encounter each other during training. Thus, the agents must employ compatible policies and should not over-fit to any arbitrary partners or cooperative patterns. Training the agents with diverse partners is effective to alleviate over-fitting to specific partners and improve ZSC performance. Population-based training (PBT) methods Strouse et al. (2021); Lupu et al. (2021); Zhao et al. (2021) have achieved state-of-the-art performance in ZSC. In Fig. 1(b), by maintaining a diverse population of training partners, the ego agent in PBT is able to collaborate with a diverse set of partners. FCP Strouse et al. (2021) trains a diverse population by setting different random seeds and including partners of level of cooperation skills and architectures. TrajeDi Lupu et al. (2021) and MEP Zhao et al. (2021) adopt explicit diversity objective to generate diverse policies as partners and achieve state-of-the-art ZSC performance. Our PECAN also maintains a population of policies. But the population is used to provide policy primitives for the policy ensemble rather than partners for the ego agent.

Human-AI Coordination Much previous work in human-AI coordination focuses on planning and learning with human models Carroll et al. (2019); Sadigh et al. (2016); Nikolaidis and Shah (2013); Kazantzidis et al. (2022). However, human-AI coordination can be naturally modelled as ZSC tasks, because humans are usually not involved in the training process. And humans usually prefer adaptive AI partners Strouse et al. (2021). Thus, an adaptive agent with strong ZSC performance is more likely to succeed in human-AI coordination tasks Strouse et al. (2021); Zhao et al. (2021).

Diversity in RL Diversity is a widely discussed issue in RL. SAC Haarnoja et al. (2018) encourages maximum entropy over the action distributions to improve a single policy’s diversity. Hong et al. (2018) maximizes the KL divergence between the current policy and some recent policy to encourage diverse action choices. Instead of the diversity of a single policy, many works also focus on diversity of a policy group. Marcolino et al. (2013) shows that a team of weak yet diverse agents can even defeat teams of strong but uniform agents under certain conditions. Derek and Isola (2021) proposed to use generative model to generate diverse policies. MEP Zhao et al. (2021) maximizes the population entropy to encourage diversity among a group of policies. EPPO Yang et al. (2022) uses mean inner product as diversity enhancement regularization and obtains a mutually distinct policy group. Besides diversity over action distributions, multiple previous studies also focus on the diversity over the induced trajectories. DIPG Masood and Doshi-Velez (2019) adopts maximum mean discrepancy (MMD) of the induced trajectories as the metric for diversity among policies and encourage qualitatively distinct behaviors. Other measures of diversity such as Jensen-Shannon Divergence Lupu et al. (2021) and mutual information Chenghao et al. (2021) are also adopted to improve diversity among policies and agents in (multi-agent) reinforcement learning. The partner diversity in PECAN is provided by 1) diversity enhancement regularization when generating policy primitives; 2) random selection as well as the random weights of policy primitives when generating policy-ensemble partners.

Policy Ensemble Policy ensemble (mixture of experts) Jacobs et al. (1991) is the mixture of a group of policy primitives Sutton et al. (1999), which is able to work individually in the target task. PMOE Ren et al. (2021) models policy ensemble as a Gaussian Mixture Model (GMM) with learnable weights and practically shows the improved diversity over both action distribution and induced trajectories. EPPO Yang et al. (2022) uses the arithmetic mean of the policy primitives as the policy ensemble. The learned policy ensemble achieves both high sample-efficiency and strong performance. Instead of learnable weights or arithmetic mean, the weights of policy primitives in PECAN are randomly generated to further improve diversity of the policy ensemble.

Problem Setting

Methodology

In this section, we will give the details of PECAN. First, we will introduce the population of policy primitives in PECAN. Then, we will introduce the proposed policy-ensemble module and the context-aware module respectively.

Like previous PBT methods Zhao et al. (2021); Strouse et al. (2021); Lupu et al. (2021), PECAN first maintains a population of training partners. To obtain a population, multiple diversity enhancement regularizations have been proposed such as mean inner-product among policies Yang et al. (2022), KL divergence among policies Hong et al. (2018) and maximum mean discrepancy Masood and Doshi-Velez (2019). In our paper, we directly use the MEP method Zhao et al. (2021) to constitute the population considering its computational efficiency and effectiveness.

Specifically, the agent in the population is trained by self-play to learn cooperation ability and maximize the population entropy (PEPE) in Zhao et al. (2021), which quantifies the diversity of the population. The formulation of PE is

where πˉ(⋅∣st)=1n∑i=0nπi(⋅∣st)\bar{\pi}(\cdot|s_{t})=\frac{1}{n}\sum\limits_{i=0}^{n}\pi^{i}(\cdot|s_{t}) is the mean policy of the population, and H(πˉ(⋅∣st))\mathcal{H}\left(\bar{\pi}(\cdot|s_{t})\right) is the entropy of the mean policy.

For agent ii, the objective of self-play training J(πi)J(\pi^{i}) is

where π=[πi,πi]\bm{\pi}=[\pi^{i},\pi^{i}] is the joint policy, at\bm{a}_{t} is the joint-action sampled from π\bm{\pi}, RR is the reward function and α\alpha is the temperature parameter controlling the relative importance of entropy maximization. By maximizing Eq. 2, agent ii will master high-level cooperation skills, and its policy πi\pi^{i} will be encouraged to diversify the current population.

In previous PBT methods Strouse et al. (2021); Zhao et al. (2021); Carroll et al. (2019), after the population of training partners is obtained, the ego agent will select a partner from the population for training in each iteration. The partner is selected by either uniform sampling Carroll et al. (2019); Strouse et al. (2021) or prioritized sampling Zhao et al. (2021). However, although the diversity enhancement regularization is adopted, the population diversity is still limited because the population is finite. Moreover, in their training procedure, the ego agent only learns a common BR to every partner in the population, which limits its capacity to coordinate with a novel partner in zero-shot evaluation, since the novel partner is never exposed to the ego agent during training.

2. Diverse Partners with Policy Ensemble

In this subsection, we will introduce how we adopt policy ensemble to improve the diversity of partners without increasing the population size. Figure. 1(c) gives the idea of our policy ensemble.

Instead of directly using policy πi, i∈[1,n]\pi^{i},\ i\in[1,n] in a policy set G:{π1,π2,...,πn}G:\{\pi^{1},\pi^{2},...,\pi^{n}\}, we use an ensembled policy πp\pi_{p} over this set GG as the training partner for the ego agents. Specifically, πp\pi_{p} is the weighted average of policies in GG, i.e.,

where ωi\omega_{i} is the weight of πi\pi^{i}. Intuitively, ωi\omega_{i} decides the contribution of πi\pi^{i} to the policy ensemble πp\pi_{p}. In particular, if the weight of one policy is 1 and the weights of the other policies are 0, the policy ensemble will degenerate into a single policy in the population (as previous PBT method did Zhao et al. (2021); Strouse et al. (2021); Lupu et al. (2021)). In contrast to this special case, we randomly assign different weights to each training iteration, and we observe that the generated partners typically exhibit different policy behaviour than the policy in the population (refer to Fig. 4). In this way, the partners generated by our method are more diverse than those generated by the vanilla PBT method, even if the population size is the same.

We split the population into three groups G1G_{1},G2G_{2} and G3G_{3} for low, medium and high level of agents according to their self-play performance. Specifically, the checkpoints from different training stages (initial, middle and final) to the population of the corresponding level following Zhao et al. (2021); Strouse et al. (2021). For example, the low-level group consists of the initial models from the population, the medium-level group consists of the middle models and the high-level group consists of the final models. And we also provide a self-play group G4G_{4} as in Zhao et al. (2021), which only contains a copy of the ego agent itself, to improve the ego agent’s self-play performance.

The mixture of policy primitives from the same group tends to remain within the same group, giving partners greater control over their levels. If the ego agent wishes to learn the level-based common BR introduced in a later section, the controllable level of partners is essential. Please refer to section 5.2 for empirical experiments validating that level-based grouping can provide partners with controllable levels. It is interesting that we also discovered that level-based grouping can increase more partner diversity . It is consistent with the previous results in Strouse et al. (2021) that level of skills is an important factor of partner diversity.

In each training iteration, one group is chosen to generate a policy ensemble for training the collaboration ability of the ego agent. Instead of randomly sampling the group, we employ a group-level prioritised sampling strategy that selects groups based on the average performance of the ego agent with each group J(Gi)J(Gi). This can stabilise the training process and guarantee that the trained ego agent can collaborate effectively with all groups. Specifically, we use rank-based prioritized sampling to assign higher priority to the group with which the ego agent is hard to cooperate as in Vinyals et al. (2019b); Zhao et al. (2021). The probability of group GiG_{i} being sampled is

where β\beta is the hyperparameter controlling the strength of prioritization. When β=0\beta=0, the prioritized sampling degenerates to uniform sampling as all groups have the same priority. And when β→∞\beta\to\infty, the group with the worst average performance will be selected with probability 1. As a smooth approximation of the ”maximize minimal” paradigm, group-level prioritized sampling helps the ego agent learns to coordinate with partners at the level with the worst coordination performance, and thus can avoid the problem of over-exploiting easy-to-cooperate partners Zhao et al. (2021).

With the randomly assigned weight ω\omega for each policy primitives and the rank-based prioritized group sampling, the policy-ensemble partners generated in each training iteration are different by design, and thus more diverse than previous methods. To stabilize the learning process, at the beginning of training, the partners have a high probability to be agents directly from the population rather than policy ensembles, during which the ego agent can learn basic cooperation skills. As training goes on, the partners are more likely to be policy ensembles, and the ego agent will collaborate with partners that have more diverse behaviours.

3. Context-Aware Ego Agent

In current methods, the ego agent only learns a common BR to every partner in the population, which limits its capacity to coordinate with an unfamiliar novel partner. Regardless of diversity of partners, this problem occurs whenever the novel partner in zero-shot evaluation is different from those during training. To alleviate this problem, we propose to learn level-based common BR for the ego agent instead of common BR to specific partners.

In order to learn a level-based BR, we devise a context encoder ff to help the ego agent analyze and predict the level of its partner as policy context. Context identification by trajectories has been previously studied Rakelly et al. (2019). But different from their context whose distribution is Gaussian, our contexts are class labels indicating the corresponding group G{1,2,3}G_{\{1,2,3\}} for low, medium and high level partners, which is determined by their training time as in Strouse et al. (2021). Therefore, we model the context encoder as a classifier in Fig. 3.

After the population is obtained, we generate many policy-ensemble partners of different levels and roll out the partners to collect training data τ{1,..,N}\tau_{\{1,..,N\}} and corresponding level-based labels c{1,..,N}c_{\{1,..,N\}} for the context encoder. The network is updated by gradient descent to predict context c^\hat{c} with the one-hot label cc. And the loss function L\mathcal{L} is the cross-entropy between c^\hat{c} and cc in Eq. 5.

where NN is the batchsize. Each input τ\bm{\tau} in the batch is a set of trajectories with the same label, and c^ij=fj(τi)\hat{c}_{ij}=f_{j}(\bm{\tau}_{i}) is the probability of group jj with input τi\bm{\tau}_{i}.

We predict the context from past trajectories. The sequence of transitions in the trajectory are encoded by a self-attention Vaswani et al. (2017) module followed by multi-layer perceptrons (MLP). The order of transitions and trajectories is irrelevant to the partner’s level. Therefore, we take sum over both encoded transitions within a trajectory and encoded trajectories, so that the prediction result is permutation-invariant w.r.t. transitions in a trajectory and trajectories in the buffer. This architecture has the capacity to represent any permutation-invariant function Zaheer et al. (2017). More implementation details of PECAN are given in the supplementary material.

By conditioning on the predicted context c^\hat{c} indicating the partner’s level, the ego agent’s policy πego(⋅∣s,c^)\pi_{ego}(\cdot|s,\hat{c}) learns level-specific coordination skills, which are more universal than previous partner-specific coordination skills.

The ego agent will learn to coordinate with partners based on their level-based context, and thus reach a level-based common BR to partners of different levels. Different from common BR to specific partners, the level-based common BR is more universal and enables the ego agent to better coordinate with unfamiliar novel partners by analyzing their level of skills and taking actions accordingly.

During zero-shot evaluation with a novel partner, we use the collected trajectories to infer context of the partner. Akin to posterior sampling Strens (2000); Osband et al. (2013); Rakelly et al. (2019), as more trajectories are collected, the ego agent’s belief narrows, and the prediction becomes more accurate. But we would like to note that the process is different from posterior sampling, since we model the context identification as a mapping from trajectories to classes rather than a distribution. Although the context inference uses past trajectories, we do not fine-tune or update any parameters during evaluation. Thus, the evaluation with a novel partner is still zero-shot.

Experiments

In this section, we will first introduce the tasks, baselines and procedures that we adopt to evaluate zero-shot human-AI coordination performance of PECAN. Then, experimental results of coordinating with human proxy models and real human players will be given. Finally, case studies are conducted to show the adaptiveness of PECAN intuitively.

Tasks We follow the evaluation protocol proposed in Carroll et al. (2019) and evaluate the proposed method on a challenging collaborative game Overcooked Games (2016); Carroll et al. (2019). Five layouts (Cramped Room, Asymmetric Advantages, Coordination Ring, Forced Coordination and Counter Circuit) in Overcooked are adopted to evaluate the ego agent’s ability to coordinate with some novel partners. See more details of the layouts in Carroll et al. (2019). Each layout exhibits a unique challenge, which can be overcome if the players coordinate well with each other. The players are required to put three onions in a pot, collect an onion soup from the pot after 20 timesteps and deliver the dish to a counter. The agents will receive 20 points for each dish served. The objective is to serve as many dishes as possible in 1 minute.

Baselines The baseline methods include Self-play PPO (SP) Carroll et al. (2019); Schulman et al. (2017), population-based training (PBT) Jaderberg et al. (2017); Carroll et al. (2019), Fictitious Co-Play (FCP) Strouse et al. (2021), TrajeDi Lupu et al. (2021) and Maximum-Entropy Population-based training (MEP) Zhao et al. (2021).

Procedure First, we pair the agents with a human proxy model, a behavior-cloning agent that mimics human’s behaviors, to test their coordination performance. The effect of each proposed component of PECAN is studied in the ablation study. Besides, we design other experiments to reinforce our claim that (a) policy ensemble is able to improve partners’ diversity and (b) the PECAN agent learns a context-aware policy. Then, we recruit human players to evaluate the human-AI coordination ability of PECAN. The human players are required to give their subjective ratings to the agents. Finally, two case studies are conducted to demonstrate the adaptiveness of PECAN in human-AI coordination. More experiment details are given in the supplementary material.

2. Experiments with Human Proxy Model

Overall Result Fig. 2(a) shows the overall coordination performance of PECAN, MEP, SP, PBT, FCP and TrajeDi agents when paired with a human proxy model. We run PECAN and MEP for 4 times with different random seeds and report their average performance. The results of SP, PBT, FCP and TrajeDi are taken from Zhao et al. (2021). For a fair comparison, MEP and PECAN have the same population size. And we adopt the recommended hyperparameter settings for MEP in their paper Zhao et al. (2021). From the results, we can see that PECAN outperforms baseline methods on all five layouts. Especially in Asymm. Adv., the best score by PECAN agents exceeds the baselines by a very large margin (+26.8%). But in layouts that require less coordination like Cramped Rm., PECAN has relatively marginal performance advantage than the baselines, which indicates that PECAN effectively improves the ego agent’s ability to coordinate with its partner rather than to accomplish the task by itself.

Ablation Study To study the effect of each component, we ablate policy ensemble and context encoder of PECAN respectively. In PECAN-e, policy ensemble is removed, and the partners are chosen the same as in MEP to train the ego agent. In PECAN-c, we remove the context encoder and make the ego agent’s policy no longer condition on context cc.

Fig. 2(b) shows that without policy ensemble and context encoder, PECAN-e and PECAN-c have similar or worse performance than MEP (average performance drop −18.1-18.1 for PECAN-e and −14.4-14.4 for PECAN-c), while PECAN consistently outperforms the baseline. The result validates the effectiveness of the two proposed modules.

Policy Ensemble and Partner Diversity We will empirically validate our previous claim that policy ensemble is able to improve the diversity of partners. We randomly sample some states ss from 5 trajectories and partners with/without policy ensemble. Then, we plot the distribution of partners’ actions πp(⋅∣s)\pi_{p}(\cdot|s) by t-SNE Van der Maaten and Hinton (2008).

Fig. 4(a) gives the visualization results. Each point represents the action distribution πp(⋅∣s)\pi_{p}(\cdot|s) of some partner pp over state ss. There are the same number of data points in the graphs with and without policy ensemble. But because there are limited number of partners without policy ensemble, the points excessively overlap with each other, indicating limited diversity. On the contrary, the action distributions become much more diverse with policy ensemble, which shows that policy ensemble is able to improve partner diversity effectively. The diverse partners allow the ego agent to learn more universal coordination behaviors, and therefore have stronger zero-shot human-AI coordination performance.

Study of the Context-aware Policy We claim that the ego agent’s policy πe(⋅∣s,c^)\pi_{e}(\cdot|s,\hat{c}) is context-aware, which means that the ego agent’s policy conditions on context c^\hat{c}. The context encoder in PECAN will recognize the partner’s context based on past behaviors and help the ego agent take actions accordingly. If incorrect context is fed to the ego agent, it may make improper decisions. There will be a performance gap between the true context and incorrect context, if the ego agent’s policy is context-aware. Thus, we design an experiment to test whether the ego agent’s policy is context-aware by manually creating context mismatch.

Specifically, we feed a random or manually-assigned context to the ego agent and pair them with a human proxy model to compare their performance with PECAN, where the context is the recognized context. Fig. 4(b) gives the results with context mismatch on Asymmetric Advantages. c=RNDc=RND represents replacing the real context with a random context, and c=SPc=SP represents forcing the context to indicate the ego agent is collaborating with a partner from the self-play group G4G_{4}, which is a clear mismatch. The results demonstrates significant performance drop with context mismatch, which means the ego agent’s policy is context-aware and the recognized context is effective. And it is worth noting that c=SPc=SP has very poor performance. This confirms the conclusions from previous studies Carroll et al. (2019); Strouse et al. (2021) that agents with a self-play cooperative pattern have significantly different behaviors from an agent with strong human-AI coordination performance.

Effect of Level-based Grouping Similar to section 5.1.3, we randomly sample some states ss and plot the action distribution πp(⋅∣s)\pi_{p}(\cdot|s) by t-SNE. Fig. 5 gives the visualization results. It can been seen that the policy ensembles with level-based grouping are more diverse and their levels are controllable based on the level of their policy primitives, while the policy of partners without grouping always tends to form two clusters, which may be the result of compromise between high-level and low-level partners and limits the diversity of partners. Therefore, the result confirms that the level-based grouping is able to improve partner diversity and provide partners with more controllable levels. And the controllable level of partners allows the ego agent to learn level-based common BR more easily.

3. Human-AI Coordination

We follow the Human-AI coordination test protocol proposed in Carroll et al. (2019) and recruit 15 human players to participate in the study. We evaluate the average performance across layouts of PECAN and state-of-the-art method MEP and reuse the evaluation results of other baselines in Zhao et al. (2021). The results are compatible because the test procedure is consistent. For the convenience of human-AI coordination experiments on Overcooked, we integrate models from Carroll et al. (2019) with PantheonRL Sarkar et al. (2022), a newly released library for dynamic training. The code is available herehttps://github.com/LxzGordon/pecan_human_AI_coordination.

Fig. 6 gives the result of our human study. PECAN outperform all other baselines, and the recruited human players give higher adaptiveness rating to PECAN and noticeably prefer coordinating with PECAN than MEP.

Case Study To show the adaptiveness of PECAN, we give two case studies in our human-AI coordination experiments. Demo videos are available at https://sites.google.com/view/pecan-overcooked.

Case Study 1 See Fig. 7. In Cramped Room, we (blue chef) intentionally block the agents’(green chef) way to the onions to see how the agents will react. The MEP agent gets stuck and stand still until we move aside, while our PECAN agent makes adjustments immediately and turn around to pick up onions from the other side. It shows the PECAN agents are more adaptive and capable of adjusting its policy according to the human player’s behaviors.

Case Study 2 We record the pot usage by the agent in Asymmetric Advantages (see the layout in Fig. 8). The trick for the agent (green chef) is to pick up dishes from both pots and serve, because it’s much nearer to the serving counter than the human player (blue chef). Table 1 shows that the PECAN agent has no clear preference over pots than the MEP agent, which has very strong preference for Pot 1. The MEP agent adopts a non-adaptive strategy which has poor generalization to typical human behaviors of using both pots, which further leads to worse performance. However, the PECAN agent’s policy is diverse and adaptive, which helps it better coordinate with real humans and serve much more dishes than the MEP agent.

Conclusion and Future Work

In this paper, we propose a new method (PECAN) for zero-shot human-AI coordination. Policy-ensemble partners and a context encoder are proposed to improve diversity of partners and help the ego agent learn more universal coordination behaviors. We evaluate PECAN with a human proxy model on Overcooked and shows that PECAN is able to outperform all baselines. Ablation studies, further studies and visualization experiments are conducted to demonstrate each component in PECAN. We also organize a human study to evaluate the proposed method’s capability of human-AI coordination. The results indicate that PECAN outperforms all other baselines on performance as well as subjective ratings.

Our future work is to study how to analyze and identify the human player’s behavior pattern as the ego agents’ context (rather than the level-based context in the current method) with the population and no human data during training. In this way, the ego agent can take actions accordingly and better coordinate with humans. This is a very challenging research subject because it requires the ego agent to comprehend human behaviors given only the population of AI agents. And since PECAN is a two-stage method, another future direction is to study how to train it in an end-to-end manner.

References

A. Experiment Details

The five layouts in our experiments are given in Fig. 9. There are unique challenges Carroll et al. (2019) for players to overcome in each layout:

Cramped Room: This layout presents low-level coordination challenges. In this shared, confined space it is very easy for the agents to collide;

Asymmetric Advantages: This layout tests whether players can choose high-level strategies that play to their strengths;

Coordination Ring: In this layout, players must coordinate to travel between the bottom left and top right corners of the layout;

Forced Coordination: This layout removes collision coordination problems, and forces players to develop a high-level joint strategy, since neither player can serve a dish by themselves.

Counter Circuit: This layout involves a non-obvious coordination strategy, where onions are passed over the counter to the pot, rather than being carried around.

A.2 Implementation Details

The population in PECAN is trained with population entropy maximization proposed in Zhao et al. (2021). Thus, we remain the recommended agent architecture and hyperparameter settings for population training, such as 5 initial agents with different random seeds, learning rate of 8×10−48\times 10^{-4} for agent training, reward shaping horizon 5×1065\times 10^{6} and so on. As in Strouse et al. (2021), the initial, the middle and the final checkpoints of the 5 agents are saved to form the population. In our level-based grouping, the initial checkpoints form the low-level group, the middle checkpoints form the middle-level group, and the final checkpoints form the high-level group. β\beta for group-level prioritized sampling is 33 in our experiments. The probability of directly selecting a partner from the population rather than providing a policy-ensemble partner decays to 0.10.1 after 50 training iterations.

In the context encoder, each transition vector in the trajectories are first encoded by a fully-connected layer with hidden size 128 . Then, the trajectory sequence are encoded by 5 self-attention Vaswani et al. (2017) modules with dimensions of Q,K,VQ,K,V matrices being 128. The transition-wise sum is taken after the self-attention modules. Then, 2 fully-connected layers with hidden size 128 follows. Before the last fully-connected layer, the trajectory-wise sum is taken. We adopt the LeakyReLULeakyReLU Maas et al. (2013) activation function for most layers of the network and SigmoidSigmoid for the layer before the output layer. The network is optimized by Adam Kingma and Ba (2014) with learning rate of 10−310^{-3}. In the cross-entropy loss function of the context encoder, batch size N=128N=128 and the input set of trajectories τ\bm{\tau} randomly contains 1,21,2 or 3 trajectories to make sure the ego agent can predict the partner’s context from different number of trajectories.

B. Human Study Statement

Below is the statement of our human-AI coordination study, which are shown to the participants. The participants are required to sign this statement after reading carefully.

You have been asked to participate in a research study that studies human-AI coordination. We would like your permission to enroll you as a participant in this research study.

The instruments involved in the experiment are a computer screen and a keyboard. The experimental task consisted of playing 5 layouts of the computer game Overcooked and manipulating the keyboard to coordinate with the AI agent to cook and serve dishes. You will be given specific instructions for the task before it begins.

B.2 Procedure

In this study, you should read the experimental instructions and ensure that you understand the experimental content. The whole experiment process lasts about 40 minutes, and the experiment is divided into the following steps:

1) Read and sign the experimental statement;

2) Test the experimental instrument, and adjust the seat height, sitting posture, and the distance between your eyes and the screen. Please ensure that you are in a comfortable sitting position during the experiment;

3) You will be paired with a dummy agent in a demo layout. You should comprehend the specific instrument operation rules and be familiar with the experimental process in the demo layout;

4) Start the formal experiment. You will be paired with an AI agent and play 8 rounds in each layout. Please cooperate with the AI agent to complete the 5 layouts and get as much scores as possible within 1 minutes. Note that after each layout, you should have a rest;

5) After the experiment, you need to fill in a questionnaire.

B.3 Risks and Discomforts

The only potential risk factor for this experiment is trace electron radiation from the computer. Relevant studies have shown that radiation from computers and related peripherals will not cause harm to the human body.

B.4 Costs

Each participant who completes the experiment will be paid 100 RMB, and the top 5 participants will be paid another 100 RMB.

B.5 Confidentiality

The results of this study may be published in an academic journal/book or used for teaching purposes. However, your name or other identifiers will not be used in any publication or teaching materials without your specific permission. In addition, if photographs, audio tapes or videotapes were taken during the study that would identify you, then you must give special permission for their use.

B.6 Participant Declaration

I confirm that the purpose of the research, the study procedures and the possible risks and discomforts as well as potential benefits that I may experience have been explained to me. All my questions have been satisfactorily answered. I have read this consent form. My signature below indicates my willingness to participate in this study.