What Matters in Language Conditioned Robotic Imitation Learning over Unstructured Data
Oier Mees, Lukas Hermann, Wolfram Burgard
I Introduction
One of the grand challenges in robotics is to create a generalist robot: a single agent capable of performing a wide variety of tasks in everyday settings based on arbitrary user commands. Doing so requires the robot to acquire a diverse repertoire of general-purpose skills and non-expert users to be able to effectively specify tasks for the robot to solve. This stands in contrast to most current end-to-end models, which typically learn individual tasks one at a time from manually-specified rewards and assume tasks being specified via goal images or one-hot skill selectors , which are not practical for untrained users to instruct robots. Not only is this inefficient, but also limits the versatility and adaptivity of the systems that can be built. How can we design learning systems that can efficiently acquire a diverse repertoire of useful skills that allows them to solve many different tasks based on arbitrary user commands?
To address this problem, we must resolve two questions. (1) How can untrained users direct the robot to perform specific tasks? Natural language presents a promising alternative form of specification, providing an intuitive and flexible way for humans to communicate tasks and refer to abstract concepts. However, learning to follow language instructions involves addressing a difficult symbol grounding problem , relating a language instruction to a robot’s onboard perception and actions. (2) How can the robot efficiently learn general-purpose skills from offline data, without hand-specified rewards? A simple and versatile choice is to define skills as being continuous instead of discrete, endowing the agent of task-agnostic control: the ability to reach any reachable goal state from any current state . These forms of task specification can in principle enable a robot to solve multi-stage tasks by following several language instructions in a row.
Recent advances have been made at learning language conditioned policies for continuous visuomotor-control in 3D environments via imitation learning or reinforcement learning . These approaches typically require offline data sources of robotic interaction together with post-hoc crowd-sourced natural language labels. Although all methods share the basic idea of leveraging instructions that are grounded in the agent’s high-dimensional observation space, their details vary greatly. Moreover, evaluating published methods and their components in language conditioned policy learning is difficult due to incomparable setups or subjective task definitions. In this work we systematically compare, improve, and integrate key components by leveraging the recently proposed CALVIN benchmark to further our understanding and provide a unified framework for long-horizon language conditioned policy learning. We build upon relabeled imitation learning to distill many reusable behaviors into a goal-directed policy, as seen in Fig. 1. Our approach consists of only standard supervised learning subroutines, and learns perceptual and linguistic understanding, together with task-agnostic control end-to-end as a single neural network. Our contributions are:
We systematically compare key components of language conditioned imitation learning over unstructured data, such as observation and action spaces, losses for aligning visuo-lingual representations, language models and latent plan representations, and we analyze the effect of other choices, such as data augmentation and optimization.
We propose four improvements to these key components: a multimodal transformer encoder to learn to recognize and organize behaviors during robotic interaction into a global categorical latent plan, a hierarchical division of the robot control learning that learns local policies in the gripper camera frame conditioned on the global plan, balancing terms within the KL loss and a self-supervised contrastive visual-language alignment loss.
We integrate the best performing improved components in a unified framework, Hierarchical Universal Language Conditioned Policies (HULC). Our model sets a new state of the art on the challenging CALVIN benchmark , on learning a single 7-DoF policy that can perform long-horizon manipulation tasks in a 3D environment, directly from images, and only specified with natural language.
II Related Work
Natural language processing has recently received much attention in the field of robotics , following the advances made towards learning groundings between vision and language and grounding behaviors in language . Early works have approached instruction following by designing interactive fetching systems to localize objects mentioned in referring expressions or by grounding not only objects, but also spatial relations to follow language expressions characterizing pick-and-place commands . Unlike these approaches, we directly learn robotic control from images and natural language instructions, and do not assume any predefined motion primitives.
More recently, end-to-end deep learning has been used to condition agents on natural language instructions , which are then trained under an imitation or reinforcement learning objective. These works have pushed the state of the art and generated a range of ideas for language conditioned policy learning, such as losses for aligning visual observations and language instructions. However, each work evaluates a different combination of ideas and uses different setups or task definitions, making it unclear how individual ideas compare to each other and which ideas combine well together. For example, the methods BC-Z and MIA use both behavior cloning, but different actions spaces and multi-modal alignment losses, such as regressing the language embedding from visual observations or cross-modality matching . Moreover, BC-Z leverages expert trajectories and task labels, and MIA includes mobile navigation, making them difficult to implement directly in CALVIN, which contains unlabeled play data on different tabletop environments. Nair et. al. learn a reward classifier which predicts if a change in state completes a language instruction and leverage it for offline multi-task RL given four camera views. Similar to BC-Z they rely on discrete task labels and do not focus on solving long-horizon language-specified tasks. Most related to our approach is multi-context imitiation learning (MCIL) , which also uses relabeled imitation learning to distill reusable behaviors into a goal-reaching policy. Besides different action and observation spaces, these works leverage different language models to encode the raw text instructions into a semantic pre-trained vector space, making it difficult to analyze which language models are best suited for language conditioned policy learning. The ablation studies presented in these papers show that each novel contribution of each work does indeed improve the performance of their model, but due to incomparable setups and evaluation protocols, it is difficult to asses what matters for language conditioned policy learning. Our work addresses this problem by systematically comparing and combining different observation and action spaces, auxiliary losses and latent representations and integrating the best performing components in a unified framework.
III Problem Formulation and Method Overview
We consider the problem of learning a goal-conditioned policy that outputs action , conditioned on the current state and free-form language instruction , under environment dynamics . We note that the agent does not have access to the true state of the environment, but to visual observations. In CALVIN the action space consists of the 7-DoF control of a Franka Emika Panda robot arm with a parallel gripper.
We model the interactive agent with a general-purpose goal-reaching policy based on multi-context imitation learning (MCIL) from play data . To learn from unstructured “play” we assume access to an unsegmented teleoperated play dataset of semantically meaningful behaviors provided by users, without a set of predefined tasks in mind. To learn control, this long temporal state-action stream is relabeled , treating each visited state in the dataset as a “reached goal state”, with the preceding states and actions treated as optimal behavior for reaching that goal. Relabeling yields a dataset of where each goal state has a trajectory demonstration solving for the goal. These short horizon goal image conditioned demonstrations can be fed to a simple maximum likelihood goal conditioned imitation objective:
to learn a goal-reaching policy . We address the inherent multi-modality in free-form imitation datasets by auto-encoding contextual demonstrations through a latent “plan” space with an sequence-to-sequence conditional variational auto-encoder (seq2seq CVAE) . Conditioning the policy on the latent plan frees up the policy to use the entirety of its capacity for learning uni-modal behavior. To generate latent plans we make use of the variational inference framework . The objective of the latent plan sampler is to model the full distribution over all high-level behaviors that might connect the current and goal state, to provide multi-modal plans at inference time. This distribution is learned with a CVAE by maximizing the marginal log likelihood of the observed behaviors in the dataset , where are sampled state-action trajectories from . The Evidence Lower Bound (ELBO) for the CVAE can be written as:
The decoder is a policy trained to reconstruct input actions, conditioned on state , goal , and an inferred plan for how to get from to . At test time, it takes a goal as input, and infers and follows plan in closed-loop.
However, when learning language conditioned policies it is not possible to relabel any visited state to a natural language goal as the goal space is no longer equivalent to the observation space. Lynch et al. showed that pairing a small number of random windows with language after-the-fact instructions enables learning a single language conditioned visuomotor policy that can perform a wide variety of robotic manipulation tasks. The key insight here is that solving a single imitation learning policy for either goal image or language goals, allows for learning control mostly from unlabeled play data and reduces the burden of language annotation to less than 1% of the total data. Concretely, given multiple contextual imitation datasets , with a different way of describing tasks, MCIL trains a single latent goal conditioned policy over all datasets simultaneously, as well as one parameterized encoder per dataset.
IV Key Components of Language Conditioned Imitation Learning over Unstructured Data
This section compares and improves key components of language conditioned imitation learning over unstructured data. We base our model on MCIL and improve it by decomposing control into a hierarchical approach of generating global plans with a static camera and learning local policies with a gripper camera conditioned on the plan. Then we go through different components that have a large impact on performance: architectures to encode sequences in relabeled imitation learning, the representation of the latent distributions, how to best align language and visual representations, data augmentation and optimization. We visualize the full architecture in Fig. 2.
How to best represent motion skills is an age-old question in robotics. From a learning perspective, generating the action sequences to solve diverse manipulation tasks with a single network from high-dimensional observations is challenging, because the distribution is multi-modal, discontinuous and imbalanced. For these reasons, finding an efficient representation is crucial to perform this non-trivial reasoning using learning-based methods. MCIL uses global actions learned from a single static RGB camera. We observe that predicting 7-DoF global actions leads to the network primarily solving static element tasks, such as pushing a button, but failing to generalize to dynamic tasks, such as manipulating colored blocks. To alleviate this problem, we propose generating global plans that correspond to reusable common behavior seen in the play data, but learning local policies conditioned on the plan. This results in a hierarchical approach that frees up the network from having to memorize all locations in the scene were the behaviors were performed. Concretely, we encode RGB images from both the static and a gripper camera to learn a compact representation of all the different high-level plans that take an agent from a current state to a goal state, learning . Inspired by a recent line of work that aims to learn hierarchies of controllers based on static and gripper cameras , we use the encoded gripper camera representations in the policy network, the global contextualized latent plan, and perform control in the gripper frame with relative actions for an efficient robot control learning. The action space consists of delta XYZ position, delta euler angles and the gripper action. Our proposed formulation has several advantages: a) local policies based on the gripper camera generalize better to different locations of the objects to be manipulated b) the policy has a prior in the form of a global contextualized latent plan, but is free to discover the exact strategy on how to interact with the objects.
IV-B Latent Plan Encoding
IV-C Semantic Alignment of Video and Language
Learning to follow language instructions involves addressing a difficult symbol grounding problem , relating a language instruction to a robots onboard perception and actions. Although instructions and visual observations are aligned in CALVIN, learning to manipulate the colored blocks is a challenging problem. This is due to the fact that the robot needs to learn a wide variety of diverse behaviors to manipulate the blocks, but also needs to understand which colored block the user is referring to. Thus, the block related instructions are very similar, for the exception of a word that might disambiguate the instruction by indicating a color. Therefore, most pre-trained language models struggle to learn such semantics from text only and the policy needs to learn referring expression comprehension via the imitation loss. There have been a number of multi-modal alignment losses proposed, such as regressing the language embedding from the visual observation or cross-modality matching . We maximize the cosine similarity between the visual features of the sequence and the corresponding language features while, at the same time, minimizing the cosine similarity between the current visual features and other language instructions in the same batch. We define our loss the same way as the contrastive loss for pairing images and captions in CLIP . However, ideally our model should use the time-dependent representation of the sequence visual observations in order to capture the meaning of a language instruction. This can be appreciated only after the sequence of actions have been executed for several timesteps. The usage of in-batch negatives enables re-use of computation both in the forward and the backward pass making training highly efficient. The logits for one batch is a matrix, where each entry is given by where is a trainable temperature parameter. Only entries on the diagonal of the matrix are considered positive examples. The final loss is the sum of the cross entropy losses on the row and the column direction.
IV-D Action Decoder
A challenge in learning control from free-form imitation data, in which different ways of executing the same skill are shown, is that a standard unimodal predictor, such as a Gaussian distribution, will average out dissimilar motions. To address this multimodality, we follow the solution proposed by Lynch et. al. of discretizing the action space and then parameterizing the policy as a discretized logistic mixture distribution . Each of the predicted logistic distributions have a separate mean and scale, and are weighed with to form the mixture distribution. The imitation loss is the negative log-likelihood for this distribution:
Where, and is the logistic CDF. Additionally, we use a cross-entropy loss to model the binary gripper open/close action.
IV-E Optimization and Implementation Details
Our full training objective for the 1% of the total data that is annotated with after-the-fact language instructions is given by . The windows without annotations are trained with the same imitation learning objective, but the language goals are replaced by the last visual frame of the sampled window to learn control in a fully self-supervised manner. A common problem in training VAEs is finding the right balance in the weight of the KL loss. A high value can result in an over-regularized model in which the decoder ignores the latent plans from the prior, also known as a “posterior collapse” . On the other hand, setting too low results in the plan sampler network being unable to catch up to plan over the latent space created by the posterior, and as a result at test time the plans generated by the plan sampler network will be unfamiliar inputs for the decoder. Orthogonal to this, as the KL loss is bidirectional, we want to avoid regularizing the plans generated by the posterior toward a poorly trained prior. To solve this problem, we minimize the KL loss faster with respect to the prior than the posterior by using different learning rates, for the prior and for the posterior, similar to Hafner et. al. . We set and for all experiments and train with the Adam optimizer with a learning rate of . During training, we randomly sample windows between length 20 and 32 and pad them until the max length of 32. For the latent plan representation we use 32 categoricals with 32 classes each. To better compare the differences between approaches, we use the same convolutional encoders as the MCIL baseline available in CALVIN for processing the images of the static and gripper camera. Our multimodal transformer encoder has 2 blocks, 8 self-attention heads, and a hidden size of 2048. In order to encode raw text into a semantic pre-trained vector space, we leverage the paraphrase-MiniLM-L3-v2 model , which distills a large Transformer based language model and is trained on paraphrase language corpora that is mainly derived from Wikipedia. It has a vocabulary size of 30,522 words and maps a sentence of any length into a vector of size 384.
IV-F Data Augmentation
To aid learning we apply data augmentation to image observations, both in our method and across all baselines. During training, we apply stochastic image shifts of 0-4 pixels to the gripper camera images and of 0-10 pixels to the static camera images as in Yarats et. al. . Additionally, a bilinear interpolation is applied on top of the shifted image by replacing each pixel with the average of the nearest pixels.
V Experiments
We evaluate our model in an extensive comparison and ablation study, to determine which components matter for language conditioned imitation learning over unstructured data. We ablate single components of our full approach to study the influence of each component. We then compare our resulting model to the best published methods on the CALVIN benchmark, and show that it outperforms all previous methods.
The goal of the agent in CALVIN is to solve sequences of up to 5 language instructions in a row using only onboard sensors. This setting is very challenging as it requires agents to be able to transition between different subgoals. CALVIN has a total of 34 different subtasks and evaluates 1000 unique sequence instruction chains. The robot is set to a neutral position after every sequence to avoid biasing the policies through the robot’s initial pose. This neutral initialization breaks correlation between initial state and task, forcing the agent to rely entirely on language to infer and solve the task. For each subtask in a row the policy is conditioned on the current subgoal instruction and transitions to the next subgoal only if the agent successfully completes the current task. We perform the ablation studies on the environment D of CALVIN and additionally report numbers of our approach for the other two CALVIN splits, the multi environment and zero-shot multi environment splits. We emphasize that the CALVIN dataset for each of the four environment consists of 6 hours of teleoperated undirected play data that might contain suboptimal behavior. To simulate a real-world scenario, only 1% of that data contains crow-sourced language annotations.
V-B Results and Ablations of Key Components
Observation and Actions Spaces: We compare our approach of dividing the robot control learning into generating global contextualized plans and conditioning a local policy that receives only the observations of the the gripper camera on the global plan against a “No Local Policy” baseline. Unlike our approach, which performs control in the gripper camera frame, the baseline’s policy receives both cameras images and performs control in the robot’s base frame, as is usual in most published approaches. We observe in Fig. 3, that despite the baseline’s decoder having more perceptual information, the performance for completing 5 chains of language instructions sequentially drops from 28.3% to 20.1%. In order to analyze the big performance difference with respect to the original MCIL baseline, we train a MCIL baseline with relative actions and observe that its performance improves significantly from the original MCIL baseline with absolute actions, but performs worse than our models. We speculate that using relative actions with a local policy is easier for the agent to learn instead of memorizing all the locations where interactions have been performed with global actions and a global observation space. By decoupling the control into a hierarchical structure, we show that performance increases significantly. Additionally, we analyze the influence of using the 7-DoF proprioceptive information as input for both the plan encodings and conditioning the policy, as many works report improved performance from it . We observe that the performance drops significantly and the agent relies too much on the robot’s initial position, rather than learning to disentangle initial states and tasks. We hypothesize this might be due to a causal confusion between the proprioceptive information and the target actions . We also analyze the effect of modeling the full action space, including the binary gripper action dimension, with the mixture of logistics distribution instead of using the log loss for the open/close gripper action and observe that the average sequence length drops from 2.64 to 2.45. Finally, we note that applying stochastic image shifts to the input images increases the performance significantly.
Latent Plan Encoding: In our CVAE framework the latent plan represents valid ways of connecting the actual state and the goal state and thus, frees up the policy to use the entirety of its capacity for learning uni-modal behavior. As language is inherently discrete and discrete representations are a natural fit for complex reasoning and planning, we represent latent plans as a vector of multiple categorical latent variables and and optimize them using straight-through gradients . We observe that the performance for 5-chain evaluation drops from 28.3% to 23.6% when we train our model with a diagonal Gaussian distribution as in MCIL. While it is difficult to judge why categorical latents work better than continuous latent variables, we hypothesize that categorical latents could be a better inductive bias for non-smooth aspects of the CALVIN benchmark, such as when a block is hidden behind the sliding door. Besides, the sparsity level enforced by a categorical distribution could be beneficial for generalization. Additionally, we compare against a goal-conditioned Behavior Cloning (GCBC) baseline which does not condition the policy on a latent plan, and observe that it performs worse than MCIL with relative actions, highlighting the importance of modeling latent behaviors in free-form imitation datasets. We also observe that balancing the KL loss is beneficial in the CVAE training. By scaling up the prior cross entropy relative to the posterior entropy, the agent is encouraged to minimize the KL loss by improving its prior toward the more informed posterior, as opposed to reducing the KL by increasing the posterior entropy. We visualize a t-SNE plot of our learned discrete latent space in Figure 5 and that see that even for unseen language instructions it appears to organize the latent space functionally. Additionally, we report degraded performance for an over-regularized model which learns to ignore the latent plans, in which we weight the KL divergence with . Finally, we evaluate replacing the transformer encoder in the posterior with a GRU bidirectional recurrent network of the same hidden dimension of 2048, similar to MCIL. The results suggest that besides an improved performance, the multimodal transformer encoder is significantly more efficient both memory and model size wise (5.9 M vs 106 M parameters for the posterior network) and overall training wall clock time. For comparison, with the transformer encoder, our full approach contains 47.1 M trainable parameters.
Semantic Alignment of Video and Language: One of the main challenges for language conditioned continuous visuomotor-control is solving a difficult symbol grounding problem , relating a language instruction to a robots onboard perception and actions. An agent in CALVIN needs to learn a wide variety of diverse behaviors to manipulate blocks with different shapes, but also needs to understand which colored block the user is instructing it to manipulate. We compare commonly used auxiliary losses for aligning visual and language representations. Concretely, we compare our contrastive loss against predicting the language embedding from the sequence’s visual observations with a cosine loss , cross-modality matching and not having an auxiliary visuo-lingual alignment loss. We observe that using an auxiliary loss to semantically align the sampled video sequences and the language instructions helps, but both baselines perform similarly. We hypothesize that our contrastive loss works best because it leverages a larger number of in-batch negatives than the cross-modality matching loss. Concretely, we maximize the cosine similarity for real pairs in the batch while minimizing the cosine similarity of the multimodal embeddings of the incorrect pairings. The cross-modality matching loss implements a discriminator that produces a binary predictor of whether the embeddings match or not. The batch is shuffled only once to produce the negative samples, contrasting only negative samples.
Language Models: Despite steady progress in language conditioned policy learning, a fundamental, but less considered aspect is the choice of the pre-trained language model to encode raw text into a semantic pre-trained vector space. We compare the lightweight paraphrase-MiniLM-L3-v2 language embeddings from our full model against several popular alternatives, such as the larger BERT , Distilroberta and MPNet , which double the embedding size from 384 to 768. Besides the architecture of the language model, we analyze the impact of the loss functions the language models are trained on, by comparing the original embeddings of MPNet and Distilroberta against versions that have been finetuned with contrastive losses at the sentence level to map semantically similar sentences into the same latent space . We observe that the SBERT models that have been finetuned on sentence semantic similarity achieve significantly better results than the original language models trained on masked language modeling. Concretely, the original Distilroberta model achieves an average sequence length of 2.21, while the SBERT Distilroberta model achieves an average sequence length of 2.50. Finally, we also compare against a model conditioned on visual (ResNet-50) and language-goal features from a pre-trained CLIP model , which has been trained to align visual and language features from millions of image-caption pairs from the internet. Surprisingly, we find that performance is slightly worse than our best performing model. We hypothesize that this might be due to a domain gap between the natural images that CLIP has been trained on and the simulated images from CALVIN. The results suggest that for complex semantics, the choice of the pre-trained language model has a large impact and models finetuned on sentence level semantic similarity should be preferred. While in this paper, we do not finetune the language models with the action loss, we anticipate this might lead to better performance, specially in order to ground instructions referring to the colored blocks.
Multi Environment and Zero-Shot Generalization: Finally, we investigate the performance of our approach on the larger multi environment splits of CALVIN on Fig. 4. On the zero-shot split, which consists on training on three environments and testing on an unseen environment with unseen instructions, we observe that despite modest improvements over the MCIL baseline, the policy achieves just an average sequence length of 0.67. We hypothesize that in order to achieve better zero-shot performance, additional techniques from the domain adaptation literature, such as adversarial skill-transfer losses might be helpful. On the split that trains on all four environments and evaluates on one of them, we observe that HULC benefits from the larger dataset size and sets a new state of the art with an average sequence length of 3.06, which is higher than our best performing model trained and tested on environment D (2.64). The results suggest that increasing the number of collected language pairs aids addressing the complicated perceptual grounding problem.
VI Conclusion
We have presented a study into what matters in language conditioned robotic imitation learning over unstructured data that systematically analyzes, compares, and improves a set of key components. This study results in a range of novel observations about these components and their interactions, from which we integrate the best components and improvements into a state-of-the-art approach. Our resulting hierarchical HULC model learns a single policy from unstructured imitation data that substantially surpasses the state of the art on the challenging language conditioned long-horizon robot manipulation CALVIN benchmark. We hope it will be useful as a starting point for further research and will bring us closer towards general-purpose robots that can relate human language to their perception and actions.
Acknowledgments
This work has been supported partly by the German Federal Ministry of Education and Research under contract 01IS18040B-OML.