SACSoN: Scalable Autonomous Control for Social Navigation

Noriaki Hirose, Dhruv Shah, Ajay Sridhar, Sergey Levine

I Introduction

Even the simplest forms of interaction between humans, such as how to pass someone in a hallway, are governed by complex non-verbal cues, and may be challenging to script. In order for robots to inhabit the same environments as people, they must also be cognizant of basic social cues and etiquette, even for seemingly simple navigational tasks. While a range of prior works have proposed approaches for modeling human behavior , the complexity of such interactions often defies analytic modeling techniques.

We approach this challenge from a data-driven perspective: acquiring policies for navigation around humans by leveraging data of human-robot interactions to learn how to navigate in socially unobtrusive ways. We propose a definition for such behavior, which is based on the counterfactual perturbation of humans. Specifically, we consider whether humans would have acted in the same way if the robot had not intruded into their space. By minimizing this counterfactual perturbation, we can guide robots to behave in a manner that does not alter the natural behavior of humans in the shared space. To instantiate this principle, we train the SACSoN (Scalable Autonomous Control for Social Navigation) policy to minimize the impact on human behavior for vision-based navigation using a single camera. This requires us to both formalize the notion of counterfactual perturbation into an objective, and to collect a dataset that has the kinds of human-robot interactions that can allow our model to learn to predict human behavior in the presence of robots. Thus, our work focuses on two complementary technical components: the design of a policy learning method that can utilize predictive models of humans for unobtrusive navigation, and the collection of a large dataset of human-robot interactions to train these predictive models.

To collect such a dataset, we propose a data collection system, which we call HuRoN (Human-Robot interaction data collection for vision-based Navigation) system. In contrast to previous social navigation datasets that involve expensive manual tele-operation , or simple scripted policies that fail to capture data diversity . Instead, we devise an intelligent system that can autonomously collect rich interaction data with little-to-no human intervention, and can improve its data collection policy over time as the ever-growing dataset is reused to further train our policy.

We deploy our data collection system to collect the HuRoN dataset, which comprises over 75 hours of robotic navigation in 5 different office environments populated by people. To the best of our knowledge, this represents the largest such dataset of an autonomous mobile robot interacting with humans, with over 4000 individual human-robot interactions. In the process of collecting the HuRoN dataset, our robot traveled for a combined total of about 58.7 km over four months. Since our dataset includes time sequences of camera images, 2D LiDAR, and wheel odometry, our dataset can be useful for visual SLAM tasks including visual odometry estimation and depth estimation.

Our work makes the following contributions: (i) a model-based method for learning a socially compliant SACSoN policy for visual navigation around humans, (ii) an autonomous data collection system, HuRoN, that encourages rich interactions with human pedestrians using a novel training objective, and (iii) the HuRoN dataset, a large and diverse dataset comprising over 4000 human-robot interactions of an autonomous robot operating in a densely populated office-space environment. Please see the project page for the dataset and videos.

II Related Work

Social navigation has been widely studied in the literature . Model-based approaches based on the dynamic pedestrian model have clasically been applied for behavior modeling . These methods determine the robot’s actions in a virtual space with the predicted pedestrians’ behavior , considering social momentum , a maximum entropy model , a model predictive controller , or a classical planner . Social navigation has also been viewed through the lens of model-free data-driven learning such as reinforcement learning .

Our method using the pedestrians’ predictive model belongs to the former. However, different from prior works, including model-based reinforcement learning, we apply the predictive model to estimate the counterfactual perturbation from the pedestrians’ intended trajectory and train the control policy offline by penalizing the perturbations. Hence our control policy enables the robot to navigate to the target position while allowing the pedestrian to walk as intended. Moreover, since our approach is end-to-end learning, the robot actions can be derived from raw images without detecting and predicting pedestrians in inference.

Similar to the data-driven approaches, our training method needs a large dataset. For vision-based navigation, prior works in collecting real-world data tend to use manual teleoperation, which is expensive and scales poorly . Instead, in addition to the training method, our work focuses on autonomous data collection of rich human-robot interactions, aiming to train our control policy.

While there has been prior work on autonomously collecting robot navigation data , our task is particularly challenging due to the dynamic agents (i.e., humans) present in the environment. To autonomously learn an accurate predictive model and socially-compliant behavior around humans, the training data must contain rich human-robot interactions, with humans walking close to the robot, and it must include a wide perceptual and behavioral diversity. The closest prior works are SCAND , which is teleoperated in indoor and outdoor environments, and CoBoT, THÖR , which are autonomous but contain no visual observations; therefore, they have limited utility for learning visual navigation policies. MuSoHu collects a dataset without using real robots. Instead, they have human participants walk in human-occupied spaces. Hence, they do not include any interactions between real robots and humans. Table I summarizes the existing robot navigation datasets, highlighting scene, method, size, and contained sensor signals. In addition to the training method, we propose the HuRoN system that can autonomously collect a large and diverse dataset of rich human interactions, and can be scaled with minimal human effort to multiple environments.

III Preliminaries

We propose a method and dataset for social compliant robotic navigation with a learning-based approach. The design of our method extends ExAug , a control policy for vision-based navigation that optimizes a goal-directed cost function (but does not by itself consider interaction with humans). This system can navigate to user-specified goal images using a combination of a topological graph and a learned low-level control policy, and its design is related to a number of recent works on vision-based navigation with learned policies and topological maps . We build our data collection system, HuRoN, on top of the same visual navigation system.

where JposeJ_{\text{pose}} corresponds to the prediction error in the relative pose estimates, JcolJ_{\text{col}} penalizes collisions, and JregJ_{\text{reg}} is a regularization term for predicted velocities. ExAug uses a geometric and kinematics model to estimate the relevant states of the robot in a virtual space and calculate these objectives, akin to the model predictive control. These objectives enable us to train the policy by minimizing the differentiable cost JnavJ_{\text{nav}} without imitating the ground truth values. Please refer to the original paper for implementation details of this system .

Overview: Section IV introduces our method to train the SACSoN policy, which aims to enable robotic navigation among humans with minimal disruption. In addition to JnavJ_{\text{nav}}, we introduce two new objectives using the counterfactual human trajectories. We pre-train the predictive model of the pedestrians’ future trajectory to estimate the counterfactual human trajectories in training. Section V describes the HuRoN system for autonomously collecting a dataset with human-robot interactions that allows us to effectively train the SACSoN policy. Our data collection system includes two key components. First, we use a policy that is similar to SACSoN, but optimized to encourage rather than avoid interactions with humans, so as to gather the maximum number of human-robot interactions. Second, the HuRoN system is designed for scalable, autonomous data collection, and includes a number of design choices to enable autonomy and continual improvement that we detail in Section V.

IV Learning a Socially Compliant Policy

We posit that a possible way to achieve “social compliance” is for robots to avoid disrupting the intended behavior of pedestrians, i.e., allow humans to carry on with their activities without disruption. In our proposed method, we penalize the counterfactual perturbation of the intended trajectories of the pedestrians. We define the intended trajectory of a pedestrian as the predicted trajectory of the pedestrian from our predictive model conditioned on the robot being stationary and non-intrusive. Our method aims to control the robot so that the humans in the environment do not act differently than they would have if the robot had been stationary. This principle could be further generalized to minimize the difference to other counterfactual situations, such as ones where the robot is absent all together, but we focus on the stationary robot counterfactual as a simple instantiation of the principle. For safety, the complete design of our full objective function also includes a term to penalize the predicted distance between the human and the robot, to encourage the robot to maintain clearance, as well as the standard navigation terms described in the preceding section. Thus, we add two terms to JnavJ_{\text{nav}}, forming our full objective:

where JcpJ_{\text{cp}} is an objective to suppress the counterfactual perturbation (Fig. 2 left) and JpsJ_{\text{ps}} is an objective to penalize the penetration of the personal space of the pedestrians (Fig. 2 right), where wcpw_{cp} and wpsw_{ps} are weights for each objective. Here, our control policy πθ\pi_{\theta} predicts velocity commands {vi,ωi}\{v_{i},\omega_{i}\} from It:t−NpI_{t:t-N_{p}} and IgI_{g}, defined as follows:

Concatenating the past image frames gives the robot additional context that can be useful to avoid obstacles, detect pedestrians in the environment, and reduce partial observability .

JcpJ_{\text{cp}}: To train the policies without distracting pedestrians, we design JcpJ_{\text{cp}} using counterfactual pedestrians trajectories,

where h^t+igw\hat{h}^{gw}_{t+i} is the estimated pedestrian’s 2D trajectory conditioned on the robot virtually stopping at the current position to give way and h^t+i\hat{h}_{t+i} is the estimated pedestrian’s 2D trajectory conditioned on the robot future action. By minimizing JcpJ_{\text{cp}} with the other objectives to train our control policy, the pedestrian can walk a path similar to what they would have taken when the robot stopped and gave way, while allowing the robot itself to move toward the goal position. Here, we estimate h^t+i\hat{h}_{t+i} as

where fψf_{\psi} is a trained predictive model of a pedestrian’s future trajectory, conditioned on their past trajectory ht−α:th_{t-\alpha:t}, as well as the robot’s past trajectory rt−α:tr_{t-\alpha:t} and future trajectory rt+1:t+βr_{t+1:t+\beta}. All trajectories in Eqn. 5 are on the current robot coordinate. The values for rt−α:tr_{t-\alpha:t} are obtained from past wheel odometry, and rt+1:t+βr_{t+1:t+\beta} is derived by integrating the velocity commands {vi,ωi}i=1…Ns\{v_{i},\omega_{i}\}_{i=1\ldots N_{s}} from our control policies. To obtain ht−α:th_{t-\alpha:t}, we use YOLO and DeepSORT to detect and track pedestrians in the images (processed into a panorama) from the recorded observations of the robot , and project these detections in 3D using the depth and scale estimates obtained from the ExAug perception module, as shown in Fig. 3.

For the other counterfactual trajectory, we input a zero vector instead of rt+1:t+βr_{t+1:t+\beta} to estimate h^t+1:t+βgw\hat{h}^{gw}_{t+1:t+\beta} as fψ(ht−α:t,rt−α:t,0)f_{\psi}(h_{t-\alpha:t},r_{t-\alpha:t},{\bf 0}). Giving the zeros vector as the robot future trajectory corresponds to stopping at the current pose. Note that we only consider scenes involving a single pedestrian for simplicity; for scenes with multiple pedestrians, we consider the nearest non-stationary pedestrians for training, since they are most likely to interact with the robot. To obtain an accurate predictive model fψf_{\psi}, we collect an interaction-enriched dataset using the HuRoN system (Section V), and train fψf_{\psi} before training πθ\pi_{\theta}.

JpsJ_{\text{ps}}: We design JpsJ_{\text{ps}} to encourage the robot to avoid the personal space of the pedestrians.

where Rh\mathcal{R}_{h} is the personal space, Rr\mathcal{R}_{r} is the robot radius, did_{i} is the distance on 2D plane between the future pedestrians’ position h^t+i\hat{h}_{t+i} and the future robot position rt+ir_{t+i}, and cc is the function to limit did_{i} between 0 and Rh+Rr\mathcal{R}_{h}+\mathcal{R}_{r} to penalize the robot trajectories only penetrating the personal space. JpsJ_{\text{ps}} may be alternatively defined as the mean of the set \{\left|{\color[rgb]{0,0,0}\mathcal{R}_{h}+\mathcal{R}_{r}}-c(d_{i})\right|\}, but empirically, we found the min formulation of Eqn. 6 to better capture the desired behavior.

V Autonomous Data Collection System

For our counterfactual objective to effectively supervise the robot’s policy, we rely on the predictive model fψf_{\psi} to make accurate predictions about hypothetical human-robot interactions. This requires training fψf_{\psi} on a diverse dataset that contains many interactions between pedestrians and our robot. Therefore, the second major contribution of our work is an autonomous data collection system that can collect such a dataset. During collection, we wish to maximize interactions between the robot and pedestrians, while also maintaining autonomy, to collect high-quality data.

We design a scalable data collection system with the data collection policy πρ\pi_{\rho} that is largely autonomous and can operate in large, indoor environments without any high-fidelity indoor positioning system. Our proposed system (see Fig. 4) builds on top of the existing ExAug navigation system using three key components: (a) help-and-rescue module for collision recovery, (b) long-term anchors for coarse localization in the environment, and (c) continual learning for improving performance over the course of deployment.

Encouraging interactions: In contrast to the SACSoN policy, the data collection policy πρ\pi_{\rho} is trained to collect a dataset with enriched human-robot interactions. We introduce an additional objective JintJ_{\text{int}} to encourage the robot to approach pedestrians while collecting data towards the desired goal.

where wiw_{i} is a scaling factor. Here, we employs the same network structure πρ\pi_{\rho} as πθ\pi_{\theta} of Eqn. 3 for the data collection control policy. JintJ_{\text{int}} is designed to minimize the distance between human and robot trajectories as Jint(ρ)=\mboxmini{∣rt+i−ht+i∣}J_{\text{int}}(\rho)=\mbox{min}_{i}\{\left|r_{t+i}-h_{t+i}\right|\} where R={rt+1,rt+2,…rt+Ns}{\mathcal{R}}=\{r_{t+1},r_{t+2},\ldots r_{t+N_{s}}\} and H={ht+1,ht+2,…ht+Ns}{\mathcal{H}}=\{h_{t+1},h_{t+2},\ldots h_{t+N_{s}}\} are the robot and human trajectories estimated by same approach in Section IV. Similar to JpsJ_{\text{ps}}, we only penalize the smallest ∣ri−hi∣\left|r_{i}-h_{i}\right| by giving min formulation to better capture the desired interaction behavior.

Help-and-rescue module: To make the data collection process as seamless and autonomous as possible, we designed a pipeline for autonomous recovery from collisions and remote help in case of irrecoverable collisions for the challenging obstacles. We built a messaging and remote teleoperation interface, where the robot sends a signal to a remote operator when in need of remote teleoperation.

Long-term anchors: To overcome the limitation of localization in the repetitive environments, we place AR tags throughout the environment at approximately 10 meters apart. Since these tags are located at fixed anchor locations, we can use their coarse positions to find the corresponding nodes of our topological graph. Please see the supplemental material on our project page for more information about Help-and-rescue module and Long-term anchors.

Continual learning: As the data collection system is deployed, it may encounter novel challenges—such as varying environmental lighting throughout the day, new obstacles in the environment etc.—and it must adapt the learned data collection behavior to these changes. To achieve this, our system adopts a continual learning approach, where training data at the end of each day of deployment is used to fine-tune the data collection policy to incorporate new experience. We also use data collected across multiple days, and times of day, to augment the fine-tuning data by chaining diverse trajectory segments between the same subgoals . This allows us to effectively incorporate experience over multiple days of deployment, while also improving robustness to variations across different days and times of day.

V-B Data collection

We use the above HuRoN system with the data collection policy πρ\pi_{\rho} to autonomously collect over 75 hours of robot navigation data in 5 diverse human-occupied environments, capturing over 4000 rich interactions with humans. We describes the robotic system used for data collection environment setup, as well as key characteristics of the dataset.

Robot and Environment Setup: Figure 5 shows an overview of our platform, built on top of an iRobot Roomba base . The robot is equipped with two visual sensors (a spherical camera, and a 170∘170^{\circ} wide-angle RGB camera), and a 2D LiDAR. We use two identical data collection robots with identical sensors, equipped with different onboard computers: an NVIDIA Jetson Xavier AGX, and an Intel i5 NUC, with all computation run onboard without a dedicated GPU. Our system commands angular and linear velocity commands to the base, and has access to the bumper collision sensor for triggering our help-and-rescue module.

During deployment, we instrument the environment with NARN_{\text{AR}} AR tags to coarsely define the robot’s route for data collection (approximately 10 m apart), and collect an example trajectory by teleoperation. This example trajectory is subsampled at a fixed frame rate of 0.5 fps to generate a topological graph of the environment. Additionally, we associate each AR tag with neighboring image nodes by collecting their ID and relative pose estimates in the robot’s local frame. Starting with a base control policy that does not encourage human interactions, HuRoN system autonomously collects data that is used to train a new control policy that can interact with humans (Section V-A) and can improve with increasing environmental experience (Section V). Please see our supplemental materials for further details.

Implementation details: We use the same hyperparameters and architecture for training πθ\pi_{\theta} and πρ\pi_{\rho}. Following ExAug , we set the control horizon Ns=8N_{s}=8 and the past observations Np=5N_{p}=5 (see Eqn. 3). For pedestrian detection and tracking, we use the spherical camera on the robot to allow detection and interactions with pedestrians behind it. For the trajectory chaining procedure described in Section V, we merge multiple trajectories across several days from an environment to enhance robustness to visual distractors. We use a batch size of 80, with the training pair (past observations and subgoal images) sampled from the same trajectory for one half of the batch, and the pair coming from different trajectories in the other half of the batch. We empirically set the weights wi=1.5w_{i}=1.5, wcp=10.0w_{cp}=10.0, and wps=100.0w_{ps}=100.0 for each objective, after analyzing closed-loop navigation performance.

For training πθ\pi_{\theta}, we pre-train fψf_{\psi} with α=Ns−1\alpha=N_{s}-1 and β=Ns\beta=N_{s} by minimizing the MSE loss using supervised learning and frozen fψf_{\psi} while training πθ\pi_{\theta}. We calculate the gradient of πθ\pi_{\theta} by back-propagation via fψf_{\psi} for updating πθ\pi_{\theta}. To train a more accurate predictive model, we generate the human and robot trajectories by social force model and mix them with our real data in the batch. One half of the batch is from our real dataset and the other half of the batch is from the social force model. Please see our supplemental materials for more information on the simulation data. Following , we set the personal space Rh\mathcal{R}_{h} as 0.45 and the robot radius {\color[rgb]{0,0,0}\mathcal{R}_{r}} as 0.25 including a small margin. All other hyperparameters are replicated from ExAug .

Dataset Characteristics: We collected the HuRoN dataset over the course of 24 days in 5 diverse environments, spread across 3 university buildings. The dataset spans 75 hours and 58 kilometers of autonomous robot navigation trajectories, containing over 4000 interactions with humans. The dataset includes visual observations (spherical and fisheye), 2D LiDAR scans, velocity information, and collision signals from the bumper. Figure 6 shows example images of rich human-robot interactions captured in our dataset.

To evaluate the efficacy of the proposed interaction objective JintJ_{\text{int}} (Section V-A), our dataset contains two equal subsets: the interaction-enriched dataset corresponding to data collected by the collection policy with interaction objective (wi=1.5w_{i}=1.5), and the naïve dataset collected without (wi=0w_{i}=0). We have released this dataset publicly on our project page.

VI Evaluation

We design our experiments to evaluate the socially compliant control policy πθ\pi_{\theta} with our proposed objectives JcpJ_{\text{cp}} and JpsJ_{\text{ps}}, as well as the proposed interaction-enriched dataset collected by our autonomous data collection system. Specifically, we study the following questions:

Does our proposed objective lead to better socially unobtrusive behavior?

Does our proposed data collection system lead to more interactions, and does this in turn lead to better predictive models of pedestrians?

How does the navigation capabilities of our policy improve over the course of collecting our dataset?

Towards answering Q1, we train two different policies with and without our proposed objectives JcpJ_{\text{cp}} and JpsJ_{\text{ps}}. Here, the control policy without JcpJ_{\text{cp}} and JpsJ_{\text{ps}} corresponds to the most relevant baseline method, ExAug . In addition, we train different social navigation policy on the naïve dataset without the proposed interaction objective. We conduct fifteen experiments using the real robot across a few days (five experiments in each three difference real environments). In these experiments, we use goal images which were collected over two months ago to evaluate the robustness of the policies to environmental changes. The distance between the start and goal positions ranges from 13.0 to 37.8 meters, which is considered relatively long for vision-based navigation in indoor settings. In order to ensure equivalent experimental conditions, we request during the evaluation that the pedestrians navigate around the robot, creating similar interaction scenarios for each control policy. If the robot collides with a pedestrian or obstacle, we request the pedestrian to distance themselves from the robot’s perimeter, and we allow the robot to continue navigation.

Table II presents the comparison of our method to the above baselines along several metrics: Goal arrival Rate (GR), Success weighted by Path Length (SPL) , Success weighted by Time Length (STL) , Collision count for Pedestrians (CP), Collision count for static Objects (CO), and Personal Space Violation duration (PSV). Our control policy trained on our proposed dataset with JcpJ_{\text{cp}} and JpsJ_{\text{ps}} shows a clear improvement over ExAug and the control policy trained on the naïve dataset without JintJ_{\text{int}} in all metrics. In particular, our method decreases the collision counts for pedestrians by more than 80%\%, reduces PSV by over 30%\%, and successfully leads the robot to the goal position.

Furthermore, we conduct a user study to evaluate the robot’s behavior in real-world environments. We recruited 17 subjects from a university campus, encompassing diverse genders, races, and backgrounds; however, there was a bias with 80% being students. We conduct 3 navigation experiments by three different control policies in Table II for each subject (51 experiments in total). We ask them to walk around the robot without explaining which control policy we are running, and we have them evaluate the social compliance and smoothness of our policy between 4 and 0 (larger is better) for four questions after each experiment. Fig. 7 demonstrates the advantage of our method in human ratings across all questions. The comparison suggests that our proposed objective improves the robot’s ability to navigate unobtrusively in the presence of humans, and our proposed dataset collected via an interaction-seeking policy leads to better performance for our method.

In Fig. 8, we qualitatively observe the robot’s behavior to be significantly more “compliant” when trained with the interaction-enriched dataset (left). Even in the narrow corridors, our control policy makes space for the pedestrians while still maintaining clearance from the walls. The control policy trained on the naïve dataset does not take avoidance action when a pedestrian approaches the robot, so the robot often violates personal space, collides with the pedestrian (top right), or fails to reach the goal (bottom right).

VI-B The Value of Interaction-Rich Data

Modeling Pedestrian Dynamics: While the previous evaluation studies the end-to-end performance of our system, in the next experiment we specifically examine the pedestrian prediction model at the core of our method, and how its predictive accuracy changes based on the composition of the training dataset. Our aim is to understand whether our proposed interaction-seeking data collection scheme actually leads to more accurate pedestrian prediction models. For Q2, we train the predictive model fψf_{\psi} on a combination of three datasets: the interaction-enriched dataset, the naïve dataset, and the simulation dataset from social force model . We report the mean squared error, to capture how close each predicted point is to the true future positions, and the cosine similarity score, that measures the alignment between the vectors corresponding to the predicted and true positions (a scale-invariant metric proposed in GNM ).

Table III shows the evaluation results of the predictive model. We find that a predictive model trained with the interaction-enriched dataset leads to better predictions, both in terms of the direction and scale, suggesting that the proposed objective indeed allows better prediction of future human behavior. In addition, its performance is much better than the trained model solely from the simulator. Since the simulation dataset from social force model can help to improve the predictive performance by mixing with our real dataset in training, we use these models for training the socially compliant control policy in Table II. Fig. 9 illustrates the predictive model in action for two example interactions. Estimated trajectories of our predictive model trained on our dataset with JintJ_{\text{int}} (magenta) coincide well with the ground truth pedestrians trajectories (red), different from the estimated trajectories trained on the naïve dataset (cyan).

Moreover, to investigate the effect of the proposed interaction loss on the quality of the data collected, we conduct controlled experiments with 5 human participants tasked with interacting with the data collection system running two different collection policies: one that encourages interactions and another that does not. While quantifying the amount of human interaction is a challenging problem by itself , we propose three metrics that coarsely capture these interactions: (i) the mean distance of the robot to an observed pedestrian, (ii) bounding box area (in sq. pixels) of the observed pedestrian, as detected by an object detector , and (iii) the offset (in pixels) of the observed pedestrian from the center of the robot’s frame (e.g., this would correspond to the visual servoing error for a follower robot ). Table IV shows the results of this evaluation on the two subsets of our dataset. We observe the explicit difference, which results in better predictive model and socially compliant control policies.

Continual Learning with the HuRoN System: Lastly, we evaluate how the navigation capabilities of the robotic policy improve over the course of collecting our dataset. While this experiment does not directly evaluate the robot’s ability to interact with humans, it does show how our data collection system can enable autonomous improvement, validating the scalability of our data gathering approach for Q3.

We deploy HuRoN system to operate autonomously, with occasional remote assistance, to collect data through the environment. At the end of a collection day, this data is used to fine-tune the policy to incorporate the new experience (as described in Sec. V) . Figure 10(a) shows the average number of remote interventions requested by the data collection system during different times of the day. We notice that at the start, the variability in environmental lighting is significant and the initialized model (red) performs significantly worse as the day progresses. However, HuRoN system is able to quickly incorporate this new experience and improve it’s performance in subsequent data collection days, requiring fewer interventions each time. Over the course of multiple days (b), our system learns near-perfect autonomous navigation in the challenging indoor environment with dynamic obstacles, requesting an average of 0.18 interventions per a 10 minute trajectory, representing a 95.5% improvement over the day 1 baseline.

VII Discussion

In this paper, we proposed a method for training the SACSoN policy for vision-based navigation to build the socially unobstrusive navigation system. In training SACSoN policy, we introduced novel objectives using the predictive model of the pedestrians’ future trajectories to suppress the counterfactual perturbation from the intended human trajctories. To obtain an accurate predictive model for a better SACSoN policy, we proposed the HuRoN system, a scalable data collection system, to autonomously collect a dataset with enriched human-robot interactions. HuRoN system has the data collection control policy to interact with the pedestrians while collecting the dataset. We used this data collection system to collect the HuRoN dataset: the publicly available dataset of visual navigation around humans, spanning over 75 hours of data collected in 5 different environments and comprising over 4000 rich human-robot interactions. Our experiments show that policies trained on the collected dataset enables the real robot to navigate with the socially unobtrusive behavior.

Our SACSoN policy, when trained on a dataset with enriched human-robot interactions, still has some limitations. Our current system only learns simple social interactions such as avoiding a pedestrian’s personal space and giving way to the pedestrians by considering the closest pedestrians’ behavior. To understand more complex scenes, we will need to incorporate better objectives accounting for multiple pedestrians and their grouping in the data collection and deployment policies.

We believe that HuRoN opens up many exciting avenues for socially compliant navigation systems in human inhabited spaces. The possibility of scaling such as system to new environments and platforms and objectives is promising. The limitation of the HuRoN dataset we present is that it lacks complex scenes that include groups of multiple pedestrians. Also, the environments in the dataset are limited to office buildings.

Acknowledgments

This research was supported by Berkeley DeepDrive at the University of California, Berkeley, and Toyota Motor North America. Additionally, partial support for this research was provided by ARL DCIST CRA W911NF-17-2-0181. The authors would like to express their gratitude to Marwa Abdulhai, Qiyang Li, Manan Tomar, Mitsuhiko Nakamoto, Roxana Infante, Ami Katagiri, Katie Kang, Zheyuan Hu, Oier Mees, Jakub Grudzien Kuba, Pranav Atreya, Isadora White, Zhiyuan Zhou, Anjali Thakrar, Niclas Joswig, Kyle Stachowicz, and Catherine Glossop for their valuable assistance in evaluating the SACSoN.

We consulted the Committee for Protection of Human Subjects at our home university (UC Berkeley) and it was determined that the study does not meet the definition of research with human subjects set forth in Federal Regulations at 45 CFR 46.102.

References

Appendix

In order to avoid navigation failure by the localization errors, we placed some AR tags along the topological graph to assist localization. Our idea is simply overriding the estimated node number by the node number associated with the AR tags. When collecting the topological graph, we also save the list of {niar,piar,ninode}i=1…Nar\{n^{ar}_{i},p^{ar}_{i},n^{node}_{i}\}_{i=1\ldots N_{ar}}. Here niarn^{ar}_{i} and piarp^{ar}_{i} are the detected AR tag number and its pose on the robot coordinate, respectively . ninoden^{node}_{i} is the node number on the topological graph, which detects the AR tag of niarn^{ar}_{i}.

We basically override the estimated node number by njnoden^{node}_{j} when detecting AR tag of niarn^{ar}_{i} in the data collection. If the multiple node images detect the same AR tag in the topological graph, we use the closest one to assist moving forward. However, the mobile robot may pass over the node location linked to the AR tag and still detect its AR tag. Such a case causes unnatural movement like stopping abruptly because the subgoal image will be behind the current robot pose. To avoid unnatural behavior, the robot compares the estimated pose of AR tag with piarp^{ar}_{i} to detect whether the its is passing by the tag. If the robot is passing by, it is overwritten with the next node number niarn^{ar}_{i} + 1.

VII-B Help-and-rescue module

In the help-and-rescue module, we implement a pipeline for autonomous recovery from collisions and seeking remote help in case of irrecoverable collisions for the challenging obstacles (e.g., that may be shorter than the camera height, made of glass etc.). When a collision is detected by the robot’s collision detector sensor (e.g., a mechanical contact sensor), an automatic backup maneuver is executed. This maneuver drives the robot away from the obstacle for a short distance along the normal vector corresponding to the point of contact. Specifically, the robot moves back about 0.5 meter and rotates about 45 degrees at a point. The direction of rotation is determined by the detection of two bumper sensors in right and left. If the left sensor detects a collision, the robot rotates to the right; if the right sensor detects a collision, the robot rotates to the left.

This allows the robot to automatically recover from 70% of the simple collisions where the robot accidentally runs into challenging obstacles (e.g., that may be shorter than the camera height, made of glass etc.). Complete autonomy, however, may not be possible to achieve. The robot may drive itself into a convex hull of multiple obstacles, leading to repeated collisions, or get it’s wheels stuck (e.g., on an air vent) and be unable to rescue itself. Only for these accidental cases, we use a messaging and remote teleoperation interface in Fig. 11 to recover and continues the data collection without any physical interventions.

VII-C Trajectory Chaining for Continual Learning

To chain the different sequences in training, we need to take TgtT_{gt} between current and subgoal image from different sequences. Fig. 12 visualizes how to obtain TgtT_{gt} from different sequences. Since we place AR tags along the topological graph to assist the localization module, some frames in our dataset detect AR tag and estimate the relative pose for each AR tag. In Fig. 12, TmcT_{mc} and TmgT_{mg} indicate the estimated relative pose against same AR tag from different sequence scs_{c} and sgs_{g}. Here, ncn_{c} and ngn_{g} are corresponding node number on scs_{c} and sgs_{g}.

To take various pairs of current and subgoal images, we randomly select two step numbers within Nm=18N_{m}=18 as ncrn_{cr} and ngrn_{gr} and decide the node number of the current image as ncn_{c} - ncrn_{cr} on scs_{c} and the node number of the subgoal image as ngn_{g} + ngrn_{gr} on sgs_{g}, respectively. The sign of ncrn_{cr} and ngrn_{gr} are decided so that the subgoal image position is forward with respect to the current image position. Note that we assume that the dataset can be collected with a positive linear velocity. Since NmN_{m} is not large number, we can have accurate relative pose TocT_{oc} between nc−ncrn_{c}-n_{cr} and ncn_{c}, and an accurate relative pose TogT_{og} between ngn_{g} and ng+ngrn_{g}+n_{gr} from the odometry. As the result, we calculate TgtT_{gt} between nc−ncrn_{c}-n_{cr} and ng+ngrn_{g}+n_{gr} as Tgt=Toc⋅Tmc⋅Tmg−1⋅TogT_{gt}=T_{oc}\cdot T_{mc}\cdot T_{mg}^{-1}\cdot T_{og}.

VII-D Network structures

Figure 13 describes the neural network architecture of πθ\pi_{\theta}. An 8-layer CNN is used to extract the image features zz from the image history It:t−NpI_{t:t-N_{p}} and the subgoal image IgI_{g}, with each layer using BatchNorm and ReLU activations. Following our previous work, ExAug, the predicted velocity commands {vi,ωi}i=1…Ns\{v_{i},\omega_{i}\}_{i=1\ldots N_{s}} from 3 fully-connected layers “FCv” are conditioned on the robot size {rs,vl}\{r_{s},v_{l}\} and zz. A scaled tanh activation is given to limit the output velocities as per the specified constraints. We can control the robot by giving v1v_{1} and ω1\omega_{1} as the actual robot velocity command.

In addition to the core part of our control policy, we can implement ”FCt” to estimate traversability {ti}i=1…Ns\{t_{i}\}_{i=1\ldots N_{s}}, following ExAug. We integrate the velocities to obtain waypoints predictions and feed them to a set of fully-connected layers “FCt” along with the observation embedding zz and target robot size rs′r_{s}^{\prime}, followed by a sigmoid function to limit ti∈(0,1)t_{i}\in(0,1). Although rs=rs′r_{s}=r_{s}^{\prime} in training, we found the flexibility of an independent rs′≠rsr_{s}^{\prime}\neq r_{s} crucial to the collision-avoidance performance of our system in inference. Note that we can remove the gray color part to construct πθ\pi_{\theta} for the minimum implementation.

In our evaluation section, we train the predictive model fθpf_{\theta_{p}} for the pedestrians dynamics. Fig. 14 is the network structure of fθpf_{\theta_{p}}. At first, we feed the concatenated past human trajectory ht−α:t−1h_{t-\alpha:t-1} and the past robot trajectory rt−α:t−1r_{t-\alpha:t-1} into “FC1” with the three fully connected layers using BatchNorm and ReLU activations to extract the features zpz_{p}. Then, we predict the human future trajectory condition on the robot actions (=future trajectories) by giving zpz_{p} with rt−1:t+βr_{t-1:t+\beta}. Here, the last layer of “FC2” with three fully connected layers has the tanh activation to limit the human velocity within ±\pm 1.5 m/s.

VII-E Simulation dataset from social force model

To evaluate the effectiveness of SACSoN dataset, we generate the pedestrian trajectories and the robot trajectories from the social force model . In addition, we mix this simulation dataset with the real data from the SACSoN dataset in training to improve the accuracy of the predictive model for the pedestrians’ future trajectories. In this appendix, we show the implementation details to generate the simulation dataset.

We set two agents (the robot and the pedestrian) with different initial velocity: 0.8 m/s for the pedestrian and 0.3 m/s for the robot toward the goal position, because the pedestrian is much faster than our robot in our case. Note that the initial velocity decides a nominal velocity for each agent, not a maximum velocity. To simulate these agents in each scenario, we randomly place these agents on their own circle with varying radii, such that the two circles centered at the origin. We decide the robot’s and pedestrian’s goal position as the opposite side of their respective circles. However, we randomly shorten the goal position for the robot to stop before arriving at the original goal to emulate giving way to the pedestrian. We decide the radius for the robot’s circle as 2.0 m and the radius for the pedestrian’s circle as 5.3 m. Since the center of these circles is the origin, the robot and the pedestrian often has the interaction around the origin, because the radius for the pedestrian 5.3 m is calculated as 0.80.3×\frac{0.8}{0.3}\times 2.0 m. We run the social force model with these hyperparameters for 80 steps and collect 10000 scenarios.

In training, we randomly choose the scenario to make the batch. Since our predictive model estimates the pedestrians trajectory on the robot local coordinate, we transform the sampled trajectories before making batch. In Fig. 15, we show the examples of the simulation dataset from the social force model.