Learning Model Predictive Controllers with Real-Time Attention for Real-World Navigation

Xuesu Xiao, Tingnan Zhang, Krzysztof Choromanski, Edward Lee, Anthony Francis, Jake Varley, Stephen Tu, Sumeet Singh, Peng Xu, Fei Xia, Sven Mikael Persson, Dmitry Kalashnikov, Leila Takayama, Roy Frostig, Jie Tan, Carolina Parada, Vikas Sindhwani

Introduction and Related Work

Real-world robot deployment in human-centric environments, such as cluttered homes or crowded offices, remains an unsolved problem . These challenging situations require safe and efficient navigation through tight spaces, such as squeezing between coffee tables and couches, handling tight corners, doorways, untidy rooms, and so on. An equally critical requirement is to navigate in a manner that complies with unwritten social norms around humans.

Classical approaches using model-based control can already move robots from one point to another safely and reliably. However, when deploying these systems in the complex real world , extensive engineering effort is required to construct world representations , model vehicle kinodynamics , hand-craft cost functions , fine-tune system parameters , and design backup planners to recover from stuck scenarios . While providing verifiable guarantees, these cascaded components need to be hand-engineered before deployment based on the roboticist’s best expectations of what would be encountered in the real world, and become cumbersome when in-situ modifications are necessary to enable adaptive behaviors .

In contrast to these classical methods, machine learning enables robots to learn these behaviors directly from data . End-to-end learning is an appealing paradigm to reduce the engineering effort and cascading errors caused by separate components, but it usually requires extensive real-world training data or simulation with inevitably simplified human and environment representations. Most importantly, it lacks safety, optimality, generalizability, interpretability, and explainability, which are crucial for real robots moving around humans . Therefore, researchers have looked at individually learning global planners , local planners , and other navigation components including cost representations , kinodynamic models , and planner parameters to enable both better navigation performance, and also off-road and social navigation .

Both classical and learning-based methods have their merits. Model Predictive Control (MPC) enables synthesis of real-time feedback controllers for robots operating in real-world environments that satisfy given safety constraints, optimality criteria, and kinodynamic models. To get the best of both worlds, we design a class of Learnable-MPC policies enabling robots to learn navigation behaviors in real-world use cases by combining the flexibility of learning from demonstrations with the optimality properties (e.g., collision-free, shortest path) of MPC solutions. Our framework can also been seen as a class of Implicit Behavior Cloning policies that are aware of real-world robot-environment and robot-human interactions. Our contributions in this paper are three-fold:

We augment the cost function of MPC with learnable components parameterized by rich Transformer-based latent embeddings of real-world context. Transformer architectures have produced stunning advances in language modeling , image generation , and multi-modal reasoning . This indisputable success comes at a computational price in proportion to the massive number of parameters learned (e.g., 175175 billion for GPT-3 ) as well as quadratic scaling in input sequence length of the core attention modules of these models. By generating context-dependent quadratic costs using Performers —a low-rank linear-attention Transformer, we demonstrate how we can embed powerful pixel-to-pixel attention mechanisms in MPC while crucially retaining real-time solutions on a CPU onboard a mobile robot.

Using distributed bilevel optimization with implicit differentiation mechanisms, we train navigation policies on expert demonstrations to handle difficult navigation scenarios, with data augmentation strategies to mitigate well-known distribution shift issues that frequently plague behavioral cloning and other imitation learning approaches.

We demonstrate that our Performer-MPC outperforms its counterparts in real-world challenging navigation scenarios, including highly constrained and human-occupied environments. Performer-MPC learns to achieve >>40% better goal reached in highly constrained environments and >>65% better behavior as captured by social metrics defined in the Appendix when moving around humans in a social navigation pilot study.

Performer-MPC: Learnable MPC with Scalable Real-Time Attention

In order to respond to dynamic uncertainty in the environment, the principal challenges of synthesis of model predictive controllers are: (i) to construct cost functions that remain suitable across a wide variety of robot-environment situations, and (ii) to generate reliable solutions to the underlying trajectory optimization problems in real-time. This work focuses on the first challenge above, in particular, by eschewing the classical approach of hand-engineered cost functions, and adopting a learning-based inverse optimal control framework, where the sensory/visual context is used to induce what MPC problem to solve in real-time in order to generate actions. We first provide some details on training and inference of a learnable MPC framework, and then discuss the details of the learnable components; see Fig. 2 for an overview.

Let C0C_{0} denote the current “context”, e.g., such as a list of tensors encoding for example RGB/D streams, force-torque readings, and proprioceptive states over a short time-window. Consider a Learnable-MPC feedback policy implicitly defined by solving the following parametric optimal control problem at each time instant:

Denote the optimizer as u∗(x0;C0,θ)\mathbf{u}^{*}(\mathbf{x}_{0};C_{0},\theta), and the corresponding optimal state sequence as x∗(x0;C0,θ)\mathbf{x}^{*}(\mathbf{x}_{0};C_{0},\theta). Here, θ\theta are learnable parameters for stagewise and terminal cost neural networks {c,cT}\{c,c_{T}\}, ff the dynamics function, and gg current state estimator. While our framework generalizes to learning cost, dynamics, and state estimators simultaneously, in this paper we focus on the inverse optimal control setting: We study how the multi-layer self-attention cores of Transformers may be embedded in the cost networks to handle sensor fusion while retaining real-time speed expectations of MPC. The dynamics function here corresponds to the differential drive dynamics of our robot.

While the MPC-structured policy can be trained using any flavor of Reinforcement Learning, real-world trial-and-error data is prohibitively expensive and a general reward function that captures all intricacies of different real-world scenarios is difficult to design. Therefore, we take an Imitation Learning approach where the robot has access to NN expert demonstrations. The MPC structure provides a form of a strong inductive bias for Imitation Learning, and can lead to improved data efficiency, robustness, and generalization. Denote uˉi=(uˉ0i,…,uˉTi−1i)\bar{\mathbf{u}}^{i}=(\bar{\mathbf{u}}^{i}_{0},\ldots,\bar{\mathbf{u}}^{i}_{T_{i}-1}) and xˉi=(xˉ0i,…,xˉTii)\bar{\mathbf{x}}^{i}=(\bar{\mathbf{x}}^{i}_{0},\ldots,\bar{\mathbf{x}}^{i}_{T_{i}}) as the control and state sequence for the ithi^{th} demonstration snippet, with associated sensor context C0iC^{i}_{0}, which can be extracted from offline planning or human teleoperation. We optimize θ\theta as follows:

where JlJ_{l} denotes total imitation loss that measures discrepancy between MPC-generated and expert state-control trajectories. We assume that JlJ_{l} also admits a stagewise and terminal decomposition using loss functions ll and lTl_{T}:

Above, xi∗,ui∗\mathbf{x}^{i*},\mathbf{u}^{i*} is the MPC solution with cost parameters θ\theta, given context C0iC^{i}_{0} and initial state xˉ0i\bar{\mathbf{x}}_{0}^{i}.

The training optimization problem in (3) has embedded the MPC optimization, (1). Together, the two may be viewed as an instance of bilevel optimization, where the higher-level searches for the best cost-network parameters θ\theta via imitation loss minimization, while the lower-level synthesizes the optimal predictive control sequences given fixed θ\theta. To use stochastic gradient descent for the higher-level problem, we need the gradient of JlJ_{l} with respect to θ\theta evaluated at a control sequence u∗(θk)\mathbf{u}^{*}(\theta_{k}) where θk\theta_{k} denotes the parameters during the current iterate kk during training. This quantity decomposes as a vector-Jacobian product (VJP),

The first term in the product on the right hand side is the gradient of the total imitation loss, which can be efficiently computed using the Adjoint method in Optimal Control thanks to its stagewise structure. The second term is the sensitivity of the MPC solution with respect to parameters, which may be efficiently computed using the Implicit Function Theorem (IFT), as featured in several works exploring differentiable “optimization layers” , see Appendix for details.

MPC Solver:

We use a second order Gauss-Newton trajectory optimizer called Iterative LQR (iLQR) with line searches inspired by Differential Dynamic Programming (DDP) . At each iteration, iLQR quadratizes the cost and linearizes the dynamics to compute the search direction by solving a time-varying LQR (TVLQR) problem. Upon convergence, a single additional LQR solve suffices for computing ∂θu∗(θ)\partial_{\theta}\mathbf{u}^{*}(\theta) in (4).

Policy:

While one may use the first component of the optimal solution for problem (1), i.e., u0∗(x0;C0,θ)\mathbf{u}^{*}_{0}(\mathbf{x}_{0};C_{0},\theta), as the policy map, we noted better performance by leveraging a secondary (non-learnable) MPC problem similar to (1), featuring a “tracking” objective w.r.t. the solution of the learnable MPC problem. The details of this “tracking MPC” problem are provided in the Appendix.

2 Attentive Cost Functions for Learnable MPC

We adopt an inverse optimal control framework for learnable MPC whereby only the cost function is learnable. We structure this cost as the sum of a user-engineered function and a context-dependent quadratic, parameterized by an embedding matrix P\mathbf{P} and vector q\mathbf{q} (described in more detail below):

Here, cˉ\bar{c} refers to the hand-designed cost function (see Appendix), appended to a Transformer-backed cost model that attends to the current context C0C_{0} to generate residual quadratic cost terms for MPC to optimize. This structure removes the computational cost of repeated quadratization of a large network in the iLQR solver. Furthermore, since the residual cost is convex and well-conditioned, crude MPC solutions can generate reliable descent directions for the higher-level optimizer, even though applying IFT in gradient computation assumes the MPC solution is precisely a local minimum. We next describe the details of the embeddings P\mathbf{P} and q\mathbf{q}.

We outline a general Transformer-based backend for learnable-MPCs which leads to Performer-MPCs. The backend maps the current contexts C0C_{0} into a latent embedding which can be reshaped into the matrices P\mathbf{P} and q\mathbf{q} to support the quadratic parameterization of Eqn. 5. For concreteness, let C0C_{0} be an image frame, i.e., the occupancy grid in the robot frame. As in Vision Transformer architectures , each frame is first independently pre-processed by a convolution layer, and then flattened to a sequence. Each element (token) of the sequence corresponds to a different patch of the original frame which is then enriched with positional encodings. The length LL of this sequence is a patch-size hyper-parameter. The preprocessed input is then fed to regular attention and MLP layers. The final embedding of one of the tokens is chosen as a latent representation of the entire context C0C_{0} to parameterize the learnable cost (e.g., via de-vectorization to P\mathbf{P} and q\mathbf{q} as in Eqn. 5). Even though we take the final embedding of a single token, it contains signal from all the tokens since attention mixes information across tokens.

Experiments

Performer-MPC is compared with two baselines, a regular MPC policy (RMPC) without the learned cost components, and an Explicit Policy (EP) that predicts a reference/goal state using the same Performer architecture, but without being coupled to the MPC structure. The control action (i.e., final policy output) is implicitly defined via the solution of the “tracking MPC” problem (see previous section), where the reference trajectory is generated by one of RMPC, EP, or Performer-MPC.

We evaluate our method in four scenarios, one in simulation and three in the real world (Fig. 3). For each scenario, the learned policies (EP and Performer-MPC ) are trained with demonstrations specifically collected for that scenario. To address the distribution shift issue, we not only collect positive examples where the robot is driving smoothly with the intended behavior, but also start the robot in randomly selected “disadvantage locations” (e.g., near-collision situations), and steer the robot to recover from them. For more data collection details please refer to the Appendix. We visualize the planning results of Performer-MPC (green) and RMPC (red) along with expert demonstrations (grey) in the top half of Fig. 4 and the train and test curves in the bottom half.

We first evaluate our method in a simulated doorway traversal scenario (Fig. 3a). 100 start and goal pairs are randomly sampled from opposing sides of the wall. A planner guided by a greedy cost function often leads the robot to a local minimum, i.e., getting stuck at the closest point to the goal on the other side of the wall. Although such a problem can be mitigated by using a global planner, we use this as a test case to showcase the learning results. We generate 2000 expert demonstrations using an off-line iLQR planner , which iteratively solves for intermediate way points provided by a Dijkstra’s global planner. Using these off-line demonstrations, Performer-MPC learns a cost landscape that steers the robot towards the doorway, even if it must veer away from the goal and travel further. Performer-MPC passes the doorway in 86 out of 100 trials while RMPC only passes 24 out of 100.

Learning Highly-Constrained Maneuvers:

We next test our method in a challenging real-world scenario—a cluttered home/office setting where the robot must perform sharp, near-collision maneuvers (Fig. 3b). A global planner provides coarse way points for the robot to follow. Each policy is run ten times and we report Success Rate (SR) and average Completion Percentage (CP) with variance (VAR) of the obstacle course that the robot is able to traverse without collisions or getting stuck (Fig. 5). Performer-MPC outperforms both RMPC and EP in SR and CP.

Learning to Anticipate Pedestrians at Blind Corners:

Going beyond static obstacles, we apply our method to social robot navigation , where robots must respect unwritten social norms for which cost functions are hard to design. One such scenario is blind corner (Fig. 3c), where robots should avoid the inner side of a hallway corner in case a human suddenly appears in this “blind spot”. For blind corner, we collect 30 demonstrations with a human driving the robot from a randomly chosen location on one side of the corner to the other side in a socially compliant manner. After training, we evaluate each policy twenty times in the real world: ten times in the corner where the demonstrations were collected (seen), and ten times in a different corner (unseen). During each run, the robot and a pedestrian human subject will approach the corner from opposite sides, while a third-party observer monitors from a distance. We use a social navigation evaluation protocol in which pedestrians and observers rate the performance of the policy using a standardized questionnaire scored with Likert scales which we combine into a joint social navigation score (see Appendix for further details). Social navigation scores for blind corner seen and unseen scenarios are shown in Fig. 6a. RMPC has the least social compliance: its hand-crafted cost function efficiently cuts the corner, causing uncomfortable near-collisions (Fig. 1). EP performs slightly better than Performer-MPC in the seen environments, but does not generalize well to unseen scenarios, with worse social scores and a 20%20\% failure rate (e.g., safety stops or not reaching the goal).

Learning to Respect Comfort Distance When Obstructed by Pedestrians:

Another common social navigation scenario is pedestrian obstruction, when a human unexpectedly impedes the prescribed path of a robot (Fig. 3d). While static obstacle avoidance is a largely solved problem, pedestrian obstruction is particularly challenging for MPC policies that are guided by waypoints that were valid before the human entered the environment. A hand-crafted cost function may guide the robot too near to the human, causing uncomfortably close interactions or the robot getting stuck right in front of the human. We evaluate policy performance for pedestrian obstruction using the social navigation evaluation protocol . Again, Performer-MPC is the most socially compliant (Fig. 6b), and in a few cases even shows emergent maneuvers unseen in the dataset (i.e., passing the human on the left if there is not enough space on the right, while the demonstrations only include right-side passing). In contrast, RMPC usually gets stuck in front of the human due to a local minimum close to the goal behind the human. While EP does stay away from the human subject in both seen and unseen, it struggles to reach the goal location and sometimes comes close to colliding with nearby walls, leading to a 5%5\% overall failure rate. Please refer to the Appendix for a detailed description of both social scenarios and statistical analysis of the results.

2 Speed Studies Over Various Performer Architectures

Limitations

Currently our Transformer-backend uses spatial attention, but in principle it can leverage the temporal axis. For example, in a face-to-face approach with a walking human in a hallway, motion history may shed light on how the human intends to move in the future, e.g., yielding left or right. A promising future research direction is to add history dependency so that the robot can sequentially reason about potential future interactions and conflicts. Furthermore, exploring richer modalities than the occupancy grid (e.g., RGB images, human traces, language contexts ), to enable robot-environment and robot-human interactions beyond simple geometry is another natural way to extend our approach. Another limitation is while the quadratic cost assures global convexity and training stability, it limits the expressiveness and complexity of the cost function. Furthermore, the cost function is learned individually for each navigation scenario, but it is unclear how one stand-alone learned cost function can handle multiple scenarios. Also, our user study pilot questionnaire could be refined, and our social evaluations could be expanded to a broader set of scenarios.

Conclusions

We present in this paper Performer-MPC, a learnable MPC system utilizing scalable Transformers to learn rich context representations parameterizing trainable cost function. We show that Performer-MPCs can be used as robotic controllers for navigation in challenging real-world environments where regular MPCs struggle, including learning to avoid local minima, to maneuver through highly constrained spaces, and to adhere to unwritten social norms, while maintaining real-time speed, even for nearly Pixel-to-Pixel attention.

References

Appendix

Appendix A Differentiable MPC

We provide some additional details regarding the learnable MPC, including the structure of the static cost and differentiation of the optimal solution w.r.t. the parameters of the cost neural networks.

Tracking MPC:

The final policy action is determined as the solution to a secondary, non-learnable MPC problem (termed “tracking MPC”), taking the form of (1) but with the cost comprising of only the static portion, i.e., cˉ\bar{c}. As defined above, this cost features a goal-reaching penalty. For the “higher-level” MPC problem (i.e., the problem solved by either Performer-MPC or RMPC), this goal is determined from problem context, e.g., a global waypoint. For the “tracking MPC” problem, this goal is set as an intermediate state extracted from the optimal state trajectory solution to the “higher-level” MPC problem. In the EP case, this goal is directly output by the EP policy. Using such a dual/tracking-MPC structure allows a more fair comparison between the three policies since the lowest-level control action is output in an identical manner.

Jacobian of Optimal MPC solution:

The key tool for computing the desired Jacobian is the implicit function theorem, stated below (note: we suppress the dependence of the MPC cost function JcJ_{c} on the context C0C_{0} for readability):

If in the neighborhood of (u⋆,θk)(\mathbf{u}^{\star},\theta_{k}) where ∇uJc(u⋆,θk)=0\nabla_{\mathbf{u}}J_{c}(\mathbf{u}^{\star},\theta_{k})=0, the Hessian ∇u2Jc(u⋆,θk)\nabla^{2}_{\mathbf{u}}J_{c}(\mathbf{u}^{\star},\theta_{k}) is non-singular, then we have:

The Hessian term above need not be explicitly materialized, since by chaining Eqns 4 and 11, the VJP can be efficiently calculated as,

The term inside the large square brackets may be computed as the solution to the following quadratic problem:

which in turn decomposes into a TV-LQR problem (Thm. 11). Finally, the dot product with ∇θ,u2Jc(u⋆,θk)\nabla^{2}_{\theta,\mathbf{u}}J_{c}(\mathbf{u}^{\star},\theta_{k}) may be computed by differentiating the co-state equations associated with the gradient ∇uJc(u⋆,θk)\nabla_{\mathbf{u}}J_{c}(\mathbf{u}^{\star},\theta_{k}).

Appendix B Speed studies over various Performers’ architectures

We run speed studies over various Performer-MPC variants to choose the most suitable one for on-robot deployment characterized by strict latency constraints. The tests were run for 100×100100\times 100 resolution images and two architecture sizes: large and medium (details below).

Tested Performer variants:

Results:

Appendix C Experimental Setup

We further provide details about our experimental setup.

Our robot can move at a linear speed between [−0.8,0.8]m/s[-0.8,0.8]\textrm{m/s}, and can rotate at an angular speed between [−1.2,1.2]rad/s[-1.2,1.2]\textrm{rad/s}. The on-board software stack can perform SLAM (Simultaneous Localization And Mapping) and generate at each time step an occupancy grid with 0.05×0.05m0.05\times 0.05\textrm{m} cells (occupied or free) around the robot within [−7.5,7.5]m[-7.5,7.5]\textrm{m} of range for both xx and yy axis. For most experiments such as in tight spaces and blind corners, we clip the occupancy grid to [−2.5,2.5]m[-2.5,2.5]\textrm{m} (100×100100\times 100 cells) since there is enough local information to make navigation decisions. Only for the pedestrian obstruction scenario we reduce the occupancy grid to [−4.5,4.5]m[-4.5,4.5]\textrm{m} (180×180180\times 180 cells) so the human can be detected early to leave enough decision making time.

C.2 Data Collection for Doorway Traversal

To learn to avoid local minima in the doorway traversal scenario, we use an artificial expert to generate 2000 training episodes. Each episode is generated by (1) randomly sampling a start configuration on one side of the door and a goal position on the other side, (2) running a coarse Dijkstra’s search to generate sparse global way points leading from the start through the doorway to the goal, and (3) feeding these way points on by one (i.e., using the solution of the last way point as the initial condition for the MPC solve for the next one) to an off-line MPC planner with 100 planning horizon and 200 maximum iterations (instead of 20 horizon and 20 iterations of the online MPC deployed on our robot). Such a heavy-duty planner requires too much computation to run onboard our robot. For the 2000 expert trajectories of 100 horizon, we apply a moving window of size 20 and extract 80 trajectories of 20 horizon, but keep the original goal of the 100-horizon trajectory on the other side of the wall, as expert demonstration data for the Performer-MPC to learn (the starting locations of these expert trajectories are shown in Fig. 7, with a few rare failure cases of the offline planner displayed by a few erratic points).

C.3 Data Collection for Highly Constrained Maneuvers

To learn agile maneuvers in the highly constrained obstacle course, we collect data from a human expert via joystick teleoperation. The human expert teleoperates the robot to randomly navigate in a collision-free manner in the cluttered obstacle course (Fig. 3 b). The entire demonstration lasts around 30 minutes. For each data point, we set the goal as the 200-th future state on the demonstrated trajectory. At around 60Hz60\textrm{Hz} state rate and 0.5m/s0.5\textrm{m/s} driving speed, the goal is roughly 2m2\textrm{m} in front of the robot. In addition to this demonstration of desirable navigation behavior, we further collect about 10 minutes of demonstration starting from failure locations, e.g., where the robot is stuck and unable to recover from such situations, to address the distribution shift problem.

C.4 Data Collection for the Two Social Scenarios

We defer the data collection details for the two social scenarios to Appendix D after the social scenarios are formally defined to facilitate understanding.

C.5 Training Time

Our bilevel optimization takes around 0.36 seconds per iteration, i.e., a forward inference pass and a backward training pass (Fig. 2) to update the learnable parameters θ\theta (Eqn. 2) on four TPUs. As shown in Fig. 4, it takes around 15, 60, 1.5, and 10 hours for our Performer-MPC model to converge for the doorway traversal, highly constrained maneuvers, blind corner, and pedestrian obstruction scenarios, respectively.

Appendix D Social Navigation Evaluation

To make evaluation of social navigation policies well-defined, realistic, scalable and repeatable, we use a social navigation benchmark based on a set of human-robot interaction scenarios to be evaluated using user surveys previously designed by Pirk et al. . This benchmark consists of scenarios with well-defined roles and expected behavior for both humans and robots, along with a series of questions for human raters with answers defined on a five-level Likert scale . To provide a concise metric for comparing policies, we average the answers to these questions into a social navigation score reported in the main body of the paper (Fig. 6). This section justifies that choice by describing this social navigation protocol and the results of our experiments in more detail.

Crucially for our purposes, this protocol can be used to generate trajectories and evaluate whether they meet the scenario’s social criteria, enabling us to create curated datasets of expert trajectories which score highly on our benchmark. For training Performer-MPCs, we chose two scenarios:

blind corner (Fig. 8a), in which a robot is expected to apply some strategy (such as slowing down or swinging wide) to reduce the likelihood of collision with a possible unseen pedestrian coming around a corner.

pedestrian obstruction (Fig. 8b), in which the robot is expected to drive around a human obstructing its path, while respecting the human’s comfort distance.

Table 5 details the questions used to evaluate each scenario, along with a set of scenario-independent general questions. For most questions, the Likert scale is implemented as Strongly Disagree, Disagree, Neutral, Agree, Strongly Agree, except for question G3, which evaluates overall performance with Very Poor, Poor, Neutral, Good, Very Good.

D.2 Collecting Expert Demonstrations

We use the scenario definitions to collect a variety of expert demonstration datasets. For both scenarios, we collect around 30 episodes each scenario, which prove sufficient for training Performer-MPC. Note compare to taking visual RGB camera input which requires over 700 episodes each scenario , Performer-MPC takes occupancy grids as input and requires substantially smaller amount of data.

For our social scenario data collection, human expert trajectories are collected by individuals trained as both participants and evaluators of the social scenarios, who attempt in their runs to guide the robot according to the social norms defined in the scenarios. In turn, our evaluation of Performer-MPC against baselines in these scenarios uses the same definitions and questionnaires which guide the trajectories; thus our experimental results gauge how well Performer-MPCs can successfully navigate with respect to social norms whose cost functions are difficult to explicitly design.

For pedestrian obstruction, we discover that both Performer-MPC and EP tend to memorize the building configurations of the training environment (e.g., walls, chairs, and tables) and sometimes do not respond to the human properly during deployment. Therefore, we augment the existing training data by randomly shuffling the background (i.e., randomly removing or adding obstacle pixels to the surrounding area, but keeping the space around the robot-human interaction point intact). We posit that such data augmentation may not be necessary when we scale up our data collection to different building configurations, and more importantly, adding extra information to distinguish humans from obstacles (e.g., human detection, tracking, and prediction). For the blind corner scenario, swinging wide to avoid the inner side of the corner is part of the desired learning process, therefore such augmentation is not necessary.

We also randomly select goal locations between behind the human and the final robot position of each episode for the pedestrian obstruction scenario to improve the model’s robustness against different goals. Again, for blind corner, we find such augmentation not necessary and simply select the 300-th future state on the demonstrated trajectory as goal for each data point. In this way, most data points have a goal behind the corner. We posit that these two data augmentation techniques contribute to the much longer training time for pedestrian obstruction than that for blind corner (Fig. 4).

D.3 Human Evaluation Pilot Study

To evaluate the performance of our social navigation policies, we conduct a pilot study, gathering both participant and observer perspectives using the scenario questionnaires. Due to covid-19 restrictions and limited availability of participants for in-person user study participation (N=9), we ran this pilot study with members of our research team. In this pilot study, we aim to address the research question: How well does our Performer-MPC social navigation policy perform on our social navigation metrics when compared against RMPC and EP?

Beyond that research question, we also aim to explore several other variables. First, we want to see how these policies would perform when comparing previously seen environments (i.e., environments where the training set is collected) vs. unseen environments (i.e., novel environments not in the training set). Second, we want to examine how direct interactants (1st person) vs. bystanders (3rd person) would rate the robot’s social navigation performance because we hypothesize that 1st person responses might be stronger, especially when it comes to perceptions of comfort and safety. Third, we want to test the policies’ performances across at least two different social navigation scenarios, starting with blind corner and pedestrian obstruction.

As this is a large set of variables (navigation policy, performing in seen vs. unseen environments, measuring from 1st vs. 3rd person perspectives, and navigating in two different social navigation scenarios) and we have a very limited set of research study participants, we opt to run this as an exploratory pilot study, not as a fully controlled, counterbalanced, human subjects experiment. Altogether we run 120 sessions, gathering 240 sets of questionnaire responses from 1st and 3rd person perspectives in the course of one day in June 2022 on our campus (N=9). To minimize possible bias, neither pedestrians nor observers are aware of what policy is being tested during each episode, and the ordering of policies is randomized.

D.4 Evaluating Social Navigation Performance Factors

To assess the relationship between people’s responses to our social navigation questionnaire items, we run principal component analyses (using varimax rotations) to see if the variables really do hang together. We also run reliability analyses to see if those items that load onto a single factor are indeed reliable measures of an underlying factor.

blind corner: The three general questions in this set (Questions G1-G3) create a highly reliable factor, Cronbach’s alpha = .92 (N=120). When we run PCA on the questions that are specific to the blind corner scenario (Questions BC1-B4), one of the items does not correlate strongly with the other three (BC2: The robot stopped to let me pass).

pedestrian obstruction: We find that the three general social navigation performance questions (Questions G1-G3) create a highly reliable factor, Cronbach’s alpha = .99 (N=240). For the four questions specific to the pedestrian obstruction, we find that all four questions (Questions PO1-PO4) also create a highly reliable factor, Cronbach’s alpha = .97 (N=120).

Because these results show most these variables are highly correlated, we average them into a “social navigation score” for presenting our results in the main body of the paper concisely. The following section presents the detailed results of the social questionnaire evaluation.

D.5 Pilot Study Results

As we are unable to fully balance the experiment design (e.g., getting each of our participants to try out each of the 24 experiment conditions once), we cannot satisfy the statistical analysis assumptions of repeated measures ANOVAs. As such we are reporting upon the descriptive statistics (means and standard errors) of our pilot study data, but we recommend interpreting these results as pilot study findings, not statistically significant findings that indicate causal relationships.

blind corner (Tab. 6): The best performing policy on blind corner is the EP policy in the seen condition, with a combined score of 4.164.16 and individual questionnaire scores equal to or higher than both other policies. However, EP does not generalize well to the unseen condition, suffering a 20%20\% failure rate and a drop in social score to 3.173.17; differences between the performance of these policies on individual questions are generally greater than the standard errors of these policies, potentially indicating a real difference that could be teased out with a larger study. In contrast, Performer-MPC generalizes well, with scores on seen and unseen of 3.883.88 and 3.873.87 respectively, with individual questions generally showing greater differences. RMPC is the worst performing policy in the overall social score and on most individual questions, except BC2, “The robot stopped to let me pass,” on which it is slightly superior due to stopping for users. However, this illuminated issues on question BC2, which we discuss further below.

pedestrian obstruction (Tab. 7): The best performing policy on pedestrian obstruction is Performer-MPC in the seen condition with overall social score of 4.394.39, with a relatively small drop to 4.104.10 in the unseen condition. EP also performs well with scores in seen and unseen of 4.204.20 and 4.084.08, respectively, though it fails to complete the task 5%5\% of the time in both conditions. While the difference in seen performance of these policies is typically greater than the standard error of their performance, potentially indicating a real difference which could be teased out with a larger study, this is not true in the unseen case. RMPC is the worst performing policy in both social score and individual questions. Results within questions, between questions and within policies under given conditions are generally more consistent for pedestrian obstruction than blind corner.

Overall, we interpret these results to indicate that Performer-MPC has better generalization than EP and is comparable in social navigation performance to EP, and that both of those policies are superior to RMPC at social navigation. The social navigation score results presented in the main body of the paper are consistent with this detailed analysis.

However, the outlier question BC2 is worth further discussion. All other questions in our survey hang together to create highly reliable factors, but BC2 does not, and show lower scores for both Performer-MPC and EP, which otherwise score highly on the social navigation questionnaire. Discussion with the participants and analysis of robot behaviors in the episodes reveal that this question inadvertently prescribes a solution: that the robot should stop at the blind corner. However, our expert training demonstrations incorporate a different solution: swinging wide at the corner to avoid collisions, which our human robot drivers determine is the preferred solution based on the navigation speed and stopping distance of the robot. In contrast, RMPC, while not scoring well on most social questions, nevertheless stop for the user, giving it an artificially high score on this question even though its behavior is not very social due to stopping very close to the user. While we report the results of this question for completeness, for future work we plan to craft questions which evaluate social navigation without prescribing a solution.

D.6 Limitations and Future Work

This paper focuses on whether Performer-MPC can successfully learn from expert demonstrations derived from real-world scenarios and then be successfully deployed in those scenarios on-robot. Therefore, we select a limited set of social scenarios which enables us to evaluate this research question. These scenarios are necessarily limited to those which can be detected from the occupancy grid, which preclude the use of the visual gesture-based scenarios proposed by Pirk et al. . Furthermore, data for these scenarios are collected by a limited number of human experts, and the policies for each scenario are trained separately. In future work, we plan to expand to a wider range of scenarios collected by a broader range of experts, and to train policies to solve sets of scenarios rather than a single scenario.

The user evaluation pilot study is a first step toward developing a more robust user study protocol for evaluating future versions of social navigation policies. Our pilot study is limited by its use of research team members as study participants; their perspectives on social navigation behavior are quite influenced by their experience with operating and running tests on these robots in their work. In the future, we will recruit user study participants from people, who are not part of our research team. Second, our pilot study is not properly balanced so we cannot run the usual statistical analyses necessary to evaluate the statistical significance of the effects we observed. Instead, we report upon the descriptive statistics for this paper and we will run a full user study as a next step in this research project. In future user studies, we will focus upon more targeted research questions so that the experiments are simpler to run, analyze, and interpret, and will ensure these questions focus on quality of social navigation without tying evaluation quality to mimicking a specific solution behavior.