ChatGPT for Robotics: Design Principles and Model Abilities

Sai Vemprala, Rogerio Bonatti, Arthur Bucker, Ashish Kapoor

Experimental Details

MuSHR: We adapt the codebase available from the MuSHR open-sourced project (\citetsrinivasa2019mushr, available at \urlhttps://mushr.io/tutorials/quickstart/) towards designing our simulator setup. The data collection procedure used a MPC controller to generate a trajectory library with 2727 candidate trajectories at each time step, and the lowest-cost trajectory (considering obstacle avoidance and control effort minimization) was selected at a rate of 5050Hz during re-planning. The simulation environment consisted of a real office floor plan of approximately 74m×30m74m\times 30m which was mapped using the gmapping library \citepgrisetti2007improved, and for each episode we sampled a random valid goal location. In total we collected approximately 1.51.5 million perception-action pairs of the vehicle in action. The data consisted of 2D LiDAR measurements with angular resolution of 0.50.5 degrees (720720 returns per scan), and vehicle wheel angle (limited to a motion amplitude of of 43.5 degrees).

Habitat: We use the Habitat simulator \citephabitat19iccv and sample random valid goal locations for the agent accross 1010 environments. Use use Habitat’s built-in shortest path function to generate the agent’s actions, and record a total of 800800K perception-action pairs consisting of RGB images of size 224×224224\times 224 with their respective discrete actions (left turn, right turn, go forward, stop).

2 Tokenizer Network Architectures

RGB images. We use a ResNet-18 backbone [he2016deep] to compute features for RGB images, which are then converted into a token of length 128.

PointNet LiDAR scans. We use a 2D LiDAR that returns a sequence of range values. We convert these values into XY locations in a vehicle-oriented bird’s eye view, and then use the PointNet \citepqi2017pointnet to compute a feature for each scan, which is then converted into a token. We remove PointNet’s transform blocks so that the resulting token is not agnostic to the point cloud’s orientation relative to the vehicle.

BEV LiDAR scans. For real-world experiments only we found that using a ResNet-18 backbone was more robust as a tokenizer for LiDAR data. There, we converted LiDAR return values into a bird’s eye view image of size 200×200200\times 200, which is processed through a ResNet-18 backbone, results and finally converted into a token of length of size 128128.

Discrete actions. In our experiment on Habitat, we have a 4-D discrete action space, i.e. ‘left’, ‘right’, ‘forward’, and ‘stop’. To tokenize such a discrete action space, we use a simple linear embedding to map the 4-D actions to a token.

Continuous actions. In our experiments with a continuous action space we are use a 2-layer MLP to map actions into a 128-D token embedding.

3 Training Parameters

The training and network parameters used for the main paper experiments are described in the table below, unless where noted differently.

Additional Training Results

We analyze the pre-training model performance for MuSHR as a function of the number of tokens used for training and as a function of model capacity, expressed as the number of layers of the transformer architecture. We evaluated 4 model sizes (3, 6, 12, 24 layers), as shown in Fig 1. Performance is measured in terms of average number of meters traversed over 150 model deployments in a realistic floor plan.

In general see an improvement in model performance as we increase the number of training tokens. Interestingly, larger models did not necessarily result in better performance for robot navigation. Even though larger models consistently presented better loss values for action prediction on a static dataset (Fig. 2), when it comes to real-time deployment the larger network capacity introduces inference delays that become a disadvantage and lead to earlier crashes. For example, while LiDAR perception measurements arrive to the vehicle every 0.077s (1313Hz), the largest model of 24 layers takes on average 0.023s for inference with a RTX3090 GPU, roughly 40% longer the 3 layer model (0.016s). These time differences can amount to even larger performance gaps in small embedded systems, and further emphasize the importance of multiple downstream task architectures sharing a common representation branch for real-time robotics applications.

2 Attention Maps

We visualize the attention maps for MuSHR (Fig. 3) and Habitat (Fig. 4). For both maps we have states and actions intercalated in time order (s0,a0,s1,a1,...s_{0},a_{0},s_{1},a_{1},...). To interpret the attention maps one can consider that the embedding in row ii pays attention to the embedding in column jj. Matrices are lower-diagonal because of the causal transformer architecture, where tokens at time tt can only attend to tokens from the beginning of the sequence up until that step.

3 Sequence length and accuracy

We evaluate the impact of longer transformer sequence lengths on the accuracy of action prediction in the pre-trained PACT model. As seen in Figure 5, longer sequences lead to lower mean absolute errors of action prediction. In practice one must find a good trade-off point because longer sequences lead to longer model training time, and signific longer delays in real-time deployments. For our main paper experiments we used a sequence length of size 1616, which presented a good trade-off between accuracy and real-time performance.

4 Habitat downstream tasks

We present additional visualizations of Habitat’s downstream tasks of mapping and localization in Figure 6.