Offline Reinforcement Learning from Images with Latent Space Models
Rafael Rafailov, Tianhe Yu, Aravind Rajeswaran, Chelsea Finn
I Introduction
For robots and artificial agents to be competent in a wide variety of dynamic and uncertain environments, they require the ability to perceive the world and act based on rich sensory observations like vision. In most real-world scenarios like homes or disaster management, it is difficult to hand-design state representations or simulators, let alone instrument the world to estimate the states. This suggests the need for an end-to-end integration of sensing and control. Despite recent advances \citepDreamer2020Hafner,kostrikov2020image,laskin2020reinforcement, the interactive (or online) sample complexity for learning control policies from vision is prohibitively high. Furthermore, interactive reinforcement learning (RL) with physical systems in the real world is fraught with safety challenges that limit widespread applicability. Our goal in this work is to develop approaches for overcoming these challenges by utilizing offline datasets.
Offline RL \citepLangeGR12 involves the learning of control policies from static pre-collected data. By using previously collected data, we alleviate the safety challenges associated with online exploration. Large offline datasets are also already available in domains like autonomous driving \citepcaesar2020nuscenes, recommendation systems \citepharper2015movielens, and robotic manipulation \citepfinn2017deep,sharma2018multiple,dasari2019robonet,mandlekar2019scaling, typically in image (or video) format. Prior work in offline RL (see Section II) has typically focused on environments with compact state representations, which are not representative of challenges faced in real-world applications. In this work, we focus on control from pixels using offline datasets. We believe the ability to train vision based policies using offline data can reduce interactive sample complexity, enhance safety, and greatly expand the applicability of RL.
Our work builds on recent advances in model-based offline RL \citepkidambi2020morel, yu2020mopo. Model-based RL algorithms have demonstrated impressive sample efficiency in interactive RL \citepjanner2019trust, RajeswaranGameMBRL, Dreamer2020Hafner. For offline RL, model-based algorithms \citepkidambi2020morel, yu2020mopo have been shown to be mimimax optimal, obtain state of the art results in a variety of benchmark tasks, as well as generalize to new out-of-distribution tasks. Uncertainty quantification and pessimism have emerged as key principles and requirements for successful offline RL, with both model-based and model-free approaches. This represents a unique challenge for control from pixels, since directly translating successful approaches in state based domains (e.g. ensembles of models) to image spaces can be computationally prohibitive.
The main contribution of our work is an algorithm, latent offline model-based policy optimization (LOMPO), which enables learning of visuomotor policies using offline datasets. LOMPO can be summarized as follows. (i) Using the available offline data, we learn a variational model with an image encoder, an image decoder, and an ensemble of latent dynamics models. (ii) We construct an uncertainty penalized MDP in the latent state space (which induces a corresponding uncertainty-penalized POMDP in observation space), where we quantify uncertainty based on disagreement between forward models in the latent state space. (iii) We learn a control policy in the learned latent space using the offline dataset by optimizing an uncertainty-penalized objective. The learned uncertainty penalized MDP provides a pessimistic regularizing effect for policy learning and guards against major challenges like distributional shift and model exploitation. We evaluate our algorithm on four simulated visuomotor control tasks and one real-world robotic manipulation task. We find that LOMPO outperforms or matches prior model-based and model-free methods across the board on these challenging tasks.
II Related Work
Our work is at the intersection of offline RL and control from high-dimensional inputs (i.e. images). We review related work from these fields below.
Offline RL has recently emerged as a prominent paradigm for learning control policies \citepLangeGR12, levine2020offline. Most offline RL algorithms augment well known RL algorithms with various forms of regularization. These include regularized variants of importance sampling based algorithms \citepLiuSAB19, SwaminathanJ15, nachum2019algaedice, zhanggendice, actor-critic algorithms \citepwu2019behavior, jaques2019way, siegel2020keep, peng2019advantage, approximate dynamic programming algorithms \citepfujimoto2018off, kumar2019stabilizing, kumar2020conservative, Liu2020ProvablyGB, agarwal2019striving, and model-based RL algorithms \citepkidambi2020morel, yu2020mopo, argenson2020model, matsushima2020deployment, swazinna2020overcoming. However, most of these prior works focus on problems with low-dimensional compact state information, except that a few that learn to play Atari games \citepfujimoto2018off, agarwal2019striving,kumar2020conservative or propose benchmarks that include simulated locomotion tasks with pixel inputs \citepgulcehre2020rl; our focus in this work is on continuous robot control from high-dimensional perceptual inputs in both simulation and the real world.
Control from Pixels.
Control from high-dimensional observation inputs has become an important problem within control and robotics as it makes real world applications more practical. Prior works have tackled this problem with model-free RL methods by learning policies either from pixel inputs end-to-end \citephaarnoja2018soft,gelada2019deepmdp, singh2019end, kostrikov2020image, laskin2020reinforcement, VRM2020Han or on top of unsupervised visual representations [lange2010deep, ghadirzadeh2017deep, nair2018visual]. Alternatively \citepSLAC2020Lee, Merlin2018Wayne train a variational latent space model, but use it only as a filter and train a separate policy on top of the learned latent representation. Model-based RL learns a dynamics model either in the pixel space \citepfinn2017deep, ebert2018visual or in a latent space \citeplevine2016end, finn2016deep, watter2015embed, banijamali2018robust, zhang2019solar, Hafner2019PlanNet, Dreamer2020Hafner, ha2018world, kipf2019contrastive, suraj2020bee and can either learn a policy within the model or deploy shooting-based planning methods. However, most of those prior works rely critically on online data collection to be successful. Visual foresight algorithms \citepfinn2017deep,ebert2018visual,suh2020surprising,yen2019experience,suraj2020bee handle control from pixels in a fully offline setting, but do not explicitly tackle the distributional shift issue that arises; meanwhile, our method is designed to specifically address this. As a result, we find in Section V-B that our approach significantly outperforms visual foresight.
III Preliminaries
where and is the prior action distribution.
Offline RL.
In the offline RL problem, the agent must learn a policy using only a fixed dataset of interactions. In our work, we focus on offline RL in high-dimensional POMDPs, where the agent has access to the fixed dataset of trajectories, consisting of high-dimensional observations, actions and rewards. The dataset is collected by a behavior policy , which may correspond to a mixture of policies. No additional environment interaction is possible. We call the distribution induced by as the behavioral distribution.
Model-based offline RL.
Model-based RL in conjunction with the key idea of pessimism or conservatism, has emerged as a promising paradigm for offline RL \citepkidambi2020morel, yu2020mopo, matsushima2020deployment, argenson2020model. In this work, we build on the MOPO framework \citepyu2020mopo. Given a dataset from an MDP , MOPO learns a dynamics model and reward model , and uses these models to construct an uncertainty-penalized MDP , with dynamics and a modified reward function , where is an estimate of model uncertainty. To account for distribution shift, the policy is optimized in this uncertainty penalized MDP. With an admissible uncertainty estimator such that , \citepyu2020mopo theoretically show that optimizing a policy under the uncertainty-penalized MDP is equivalent to optimizing a lower bound of the return under the learned policy in the true MDP . While MOPO achieves impressive results in tasks with low-dimensional observation spaces, it is hard to scale to realistic environments with image observations. In the next section, we present LOMPO, which aims to solve the offline RL problem in a model-based way using high-dimensional observation spaces.
IV LOMPO: Latent Offline Model-Based Policy Optimization
Our goal is to design an offline model-based RL method that handles high-dimensional observations. Since we need to learn a model from a fixed dataset without further interaction with the environment, the model predictions will become less trustworthy as the policy rollouts move further from the behavioral distribution. Such inaccurate model predictions would generate observations that could negatively impact the policy optimization (the model exploitation phenomenon). Therefore, quantifying the uncertainty of the observations generated by the learned model is important for offline model-based RL to avoid large extrapolation error on out-of-distribution observations. However, estimating model uncertainty in high-dimensional spaces is challenging: the common approach to uncertainty quantification of learning an ensemble of models is notably memory-intensive and computationally-expensive when applied to visual dynamics models.
In this section, we present our offline visual model-based RL algorithm that address the above challenges by learning a latent dynamics model and estimating the model uncertainty in the compact latent space. Specifically, under the assumption that the control problem in the latent space is an MDP, we construct an uncertainty-penalized latent MDP with the reward penalized by the uncertainty of the latent dynamics model (Section IV-A). Next, we construct the corresponding uncertainty-penalized POMDP and optimize the policy and the latent dynamics model in the control as inference framework by maximizing the ELBO in the uncertainty-penalized POMDP, which is a lower bound of the ELBO in the true POMDP (Section IV-B). Finally, we discuss our overall practical algorithm LOMPO (Section IV-C).
where is the policy learned via maximizing the return under . In practice, we do not have access to the uncertainty quantification oracle that upper bounds the latent model error, and we estimate the uncertainty of the latent model using heuristics as discussed in Section IV-C.
IV-B Latent Model Training and Policy Optimization with Uncertainty-Penalized ELBO
With the uncertainty-penalized latent MDP defined, we can construct the corresponding uncertainty-penalized POMDP . In the following subsection, we will derive the objectives of the policy learning and latent model learning under the uncertainty-penalized POMDP.
We define the learned variational distribution as a product of inference terms , learned latent dynamics terms and policy terms as follows
With and Eq. 1, we can bound the expected return term in Eq. 1 from below as follows:
where Eq. 4 follows from the definition of and the discounted state-action distribution with the initial state distribution being and the inequality in Eq. 5 follows from Eq. 2 and the latent MDP assumption. Now we can derive the ELBO in the uncertainty-penalized POMDP from the ELBO in the original POMDP defined in Equation 1 as follows:
IV-C Practical Implementation of LOMPO
We now present our practical LOMPO algorithm, outlined in Algorithm 1 in Appendix D, and visualize the training framework in Figure 2.
Variational Model Training. To estimate the uncertainty term in the uncertainty-penalized POMDP, we train an ensemble of latent transition models and use model-disagreement as a proxy. In designing the optimization approach, we have two main considerations: (1) we need all of the members of the ensemble to be grounded within the same latent space, (2) we should minimize additional model complexity and training overhead. Considering these, we optimize the following objective:
Here at each step we sample a random learned forward transition model from a fixed set of models . The inference distribution is modeled via the standard mean-field approximation as a uni-modal Gaussian distribution and is shared across all time steps. This explicitly grounds all forward models within the same latent space state representation induced by the inference distribution addressing concern (1) above. Moreover since we use a single forward model at each time step this procedure has the same computational overhead as the regular training of a single variational model, which addresses concern (2).
Model-based Policy Optimization. We train a policy and a critic with components and on top of the latent space representation similar to \citepSLAC2020Lee, however we deploy model inference for the policy as well as the critic. We made this design choice as it’s faster and more efficient to carry out rollouts in the latent space, rather than sampling from the observation model. We maintain two replay buffers , . The real data replay buffer contains transition tuples from the latent MDP, where states are sampled from the inference distribution over trajectories from the real dataset . The latent data buffer contains transitions from rolling-out the policy in the model latent space utilizing the ensemble of learned forward models. During rollouts transitions are carried out at each step by picking a a random forward model from the ensemble. As discussed in Section IV-A, the final rewards for the rollout use ensemble estimates and the uncertainty penalty term and are computed as:
where are sampled from each forward model and is sampled from . Here is an estimate of model uncertainty and is a penalty parameter. In particular, we pick as disagreement of the latent model predictions in the ensemble, i.e. the variance of log-likelihoods under the ensemble since ensembles have been shown to capture the epistemic uncertainty \citepbickel1981some and also work well for model-based RL in practice \citepyu2020mopo, kidambi2020morel. We used this heuristic to estimate uncertainty as it estimates disagreement across the means of the forward models, as well as the variances. Finally the actor and critic are trained using standard off-policy training algorithm using batches of equal data mixed from the real and sampled replay buffers, as we find that maintaining a fixed sampling proportion is important in preventing distributional shift in the actor-critic training.
V Experiments
The goal of our experimentation evaluation is to answer the following questions. (1) Can offline RL reliably scale to realistic robot environments with complex dynamics and interactions? (2) How does LOMPO compare to prior offline model-free RL algorithms and online model-based RL algorithms when learning vision-based control tasks from offline data? (3) How does the quality and size of the dataset affect performance? (4) Can LOMPO be applied to an offline RL task on a real robot with raw camera image observations? We answer questions (1), (2) and (3) in Section V-A and address question (4) in Section V-B. All implementation details such as model architectures and hyperparameter choices are included in Appendix C.
Previous offline RL benchmarks are not well-suited for answering questions (1), (2) and (3) above as they largely lack image-based robot control problems. Thus, to answer those three questions, we design a suite of four simulated image-based offline RL problems, focusing on robotics applications, described in Appendix A and visualized in the four pictures on the left in Figure 3. We also include the visualization of the samples generated from our learned variational model on the four environments in Appendix E. All of the environments and datasets will be open-sourced to allow future work to also study this problem and make direct comparisons.
Comparisons. We compare our proposed method to both model-free and model-based learning algorithms. Our first comparison is direct behavior cloning (BC) from raw image observations, which has proved to be a strong baseline in the past ([fu2020d4rl]). We also benchmark to Conservative Q-Learning ([kumar2020conservative]), which is a state of the art offline learning algorithm in the low-dimensional case ([fu2020d4rl]); however, we again train it from raw image observations. We evaluate an MBPO based model ([janner2019trust]), which also carries out policy rollouts in latent space similar to LOMPO, but does not apply an uncertainty penalty. The performance of this method is indicative of online model-based methods \citepHafner2019PlanNet, Dreamer2020Hafner. We also train the Stochastic Latent Actor Critic (SLAC) model ([SLAC2020Lee]), a state of the art online learning algorithm from images, however we train fully online. The goal of this benchmark is to evaluate the need for representation learning in offline RL.
Results. Results are reported in Table I. We see that LOMPO achieves high-scores across most high-fidelity simulation environments, using raw observations. Moreover, our proposed model outperforms other model-based learning algorithms across the board and is the only model-based learning algorithm that achieves any success on several environments. Comparing to model-free algorithms, LOMPO still outperforms CQL and behaviour cloning across most environments, with the exception of learning on the expert dataset on the Door Open task. This is a well-known phenomenon when learning dynamics models from narrow expert data. On the other hand, given the thin data distribution and relatively simple dynamics of the task, direct behavior cloning from images performs well on both the medium-expert and expert dataset. We hypothesize that LMOPO performs well on the D’Claw and Adroit expert datasets, as these environments are relatively stationary, as compared to a robot arm manipulation task, and even actions from a stochastic expert cover a wide range of the environment dynamics. The question of dataset size (question (3)) has not been extensively studied in offline RL, however we found that this almost as important as the data distribution itself. We believe this is an important question as collecting large robot dataset can still be complex and expensive task. We carried additional ablation experiments (Appendix B) and discover that performance for regular model-based and model-free methods can decline drastically as data size decrease, while LOMPO performance remains relatively stable.
V-B Real Robot Experiments
To answer question (4), we deploy LOMPO on a real Franka Emika Panda robot arm.
Task. The environment consists of a Panda arm mounted in front of an Ikea desk cluttered with random distractor objects. The robot arm is initialized randomly above the desk and the drawer is initialized randomly in a open position. The goal of the robot is to navigate to the handle, hook it, and close the drawer. Observations are raw RGB images from a single overhead camera. The complete setup is shown in the rightmost picture in Figure 3.
Dataset. We use a pre-existing dataset of 1000 trajectories that was collected using a semi-supervised batch exploration algorithm \citepsuraj2020bee. A small balanced dataset of 200 images (0.2% of the full dataset) is manually labeled with whether the drawer is open or closed. Following the set up by \citepsuraj2020bee, we use this dataset to train a classifier to predict whether the drawer is open or closed. We use the classifier probability as a reward for RL, which leads to a sparse, noisy, and unstable reward signal, which is reflective of one challenge of real-world RL. Since the dataset was collected and labeled in the context of a different paper \citepsuraj2020bee, this experiment evaluates the ability to reuse existing offline datasets, which further exemplifies real-world problems.
Comparisons. We compare LOMPO, LMBRL, Offline SLAC, and visual foresight \citepfinn2017deep using an SV2P model \citepbabaeizadeh2018stochastic and a CEM planner. As CQL did not achieve competitive performance on the simulated environments, we did not deploy it on the real robot. Moreover the offline dataset has high variance and consists of mostly non-task centric exploration, which is not suitable for imitation; hence we also did not evaluate behavioral cloning.
Results. We carry out 25 evaluation rollouts on the real robotEvaluation videos are available at https://sites.google.com/view/lompo/. and summarize the results in Table II. Overall, 24/25 of the LOMPO agent rollouts successfully navigate to the drawer handle, hook it, and push the drawer in; however the agent fully closes the drawer in only 19 of the rollouts for a final success rate of 76%. We hypothesize that the agent does not always close the door as the classifier reward incorrectly predicts the drawer as closed when the drawer is slightly open. In contrast, the LMBRL, Offline SLAC, and visual foresight agents do not manage to successfully navigate to the correct handle location, hence achieving a success rate of 0%. These experiments suggest that LOMPO’s uncertainty estimation and pessimism are critical for good offline RL performance. Finally, we note that \citepsuraj2020bee evaluate visual foresight in the same environment but on an easier version of this task, where the robot arm is initialized near the drawer handle. In this shorter-horizon problem, visual foresight achieves a success rate of 65% (Figure 8 of [suraj2020bee]), which is still lower than LOMPO’s success rate in the more difficult setting. Hence, this suggests that LOMPO is better able to solve problems with a longer time horizons by incorporating pessimism and a learned value function.
VI Conclusion
We present the first offline model-based RL algorithm that handles high-dimensional observations. We noted that learning a visual dynamics model is challenging and quantifying uncertainty in the pixel space is extremely costly. We address such challenges by learning a latent space and extend the theoretical findings in prior work [yu2020mopo] to quantify the model uncertainty in the latent space. Our algorithm, LOMPO, penalizes latent states with latent model uncertainty implemented as the latent model ensemble disagreement. LOMPO empirically outperforms previous latent model-based and model-free models in the offline setting on both the standard locomotion tasks and complex robotic manipulation domains.
Appendix A Environments and Datasets
We describe the simulated environments and the dataset collection in more detail below.
Environments and Tasks. We evaluate our method in four simulated environments.
The first task is the standard walker task from the DeepMind Control suite \citeptassa2020dmcontrol. The observations consist of raw images, similiar to previous works. We apply an action repeat of 2 over the base environment. This is a standard practice \citepSLAC2020Lee, Dreamer2020Hafner as timestep in these environments is relatively short and leaves low visual footprint, which makes model learning hard.
The second experiment consists of a modified version of the D’Claw screw task from the Robel benchmark \citepahn2019robel, where the goal of the robot is to continuously turn the valve as fast as possible. The agent receives a dense negative penalty for positioning the fingers and a sparse reward whenever it turns the valve. The observation space consist of robot proprioception and raw images. We apply an action repeat of 2 over the base environment.
The third environment is based on the Adroit pen task \citepRajeswaran-RSS-18 with a fixed goal, which requires the agent to flip the pen around and catch it at certain angle. The observation space consist of robot proprioception and raw images. We apply an action repeat of 4 over the base environment.
The final simulation experiment is based on a Sawyer manipulation task \citepyu2020meta, which requires opening a door. The agent received a sparse reward when the door is fully opened and no reward otherwise. The goal of this environment is to test learning in a realistic multi-stage robot arm environment with a sparse reward. This is a hard environment, not solvable with online RL. The observation space consists of images without access to the robot state. We apply an action repeat of 4 over the base environment.
We provide visualizations of all simulated tasks in Figure 3. Each of these environments presents different challenges. In the Walker task the agent needs to learn a forward dynamics model completely from images, which are generated by a non-stationary camera. The D’Claw and Adroit environments are both quite challenging as the model needs to merge proprioception and visual information in order to estimate a hard contact model with realistic physics, as well as forward dynamics under a high-dimensional action space (9 and 26 respectively) and a sparse reward function (for the calw environment). The Sawyer arm environment requires learning a 3D dynamics model without access to depth estimation and a sparse reward function, both of which are common in real world RL.
Datasets. We construct new sets of offline datasets with image observations. Similar to the protocol by \citetfu2020d4rl, we create three types of datasets for each task, which are obtained by training an agent using soft-actor critic \citephaarnoja2018soft from the ground-truth state and recording the corresponding image observations, actions, and rewards. The medium-replay datasets consist of data from the training replay buffer up to the point where the policy reaches performance of about half the expert level performance. The goal of these datasets is to test learning on incomplete training data. The medium-expert datasets consist of the second half of the replay buffer after the agent reaches medium-level performance. The goal of these datasets is to test learning on a mixture of data from sub-optimal policies. The expert dataset consist of data sampled from the stochastic SAC expert policy. The goal of these datasets is to test learning on a thin data distribution. The dataset sizes for Walker, D’Claw, Adroit and the Sawyer environments are 100K, 180K, 1M and 50K transitions respectively.
Appendix B Ablation Studies on Varying Offline Dataset Size
We test the effect of dataset size, given that the data is sampled from the same distribution. We create a medium-expert dataset of size 1M on the D’Claw Screw environment by mixing data from 3 separate policy training runs. We then created two more datasets by sub-sampling by a factor of 5 and 25 respectively. We observe that LMOPO still performs well in the low-data regime, as compared to regular latent model-based RL and Offline SLAC.
Appendix C Implementation details
The latent dynamics model and the observation model consist of the following components as in \citepDreamer2020Hafner:
where are the environment observation, latent state, action at time respectively. We use to denote the concatenation of all the parameters involved in the latent space dynamics model. The latent dynamics model is represented by a RSSM ([Hafner2019PlanNet]). Specifically, we adopt the latent space representation , which consists of a deterministic and a sampled stochastic representation . With such a latent space representation, we use the following components:
where are observation features as defined in Eq. 10. The deterministic representation is implemented as a single GRU cell and is shared between the forward and inference models. All the learned forward models in the ensemble share the same deterministic model but separate stochastic transition models , which are implemented as MLPs. Finally the stochastic inference model is also implemented as an MLP network.
The encoder network is modeled as a convolutional neural network. For the DeepMind Control Walker task the network has 4 layers with channels respectively. For the D’Claw Screw, Adroit Pen and Door Open environments the convolutional model has 5 layers with channels. All kernels have size 4 and stride 2. The reconstruction model for DeepMind Control walker task has 4 layers with channels with kernel size and stride 2. For the other environments the reconstruction network has 5 layers with channels with kernel size and stride 2. The reward reconstruction network is a two-layer fully connected network. The deterministic path of the RSSM is modeled as a GRU cell with 256 units for all environments, except the Adroit Pen task, which uses 512 units. All the forward models and inference models are 3-layer fully-connected networks with 256 units, except for the Adroit Pen task, which uses 512. The variational model is trained with the Adam optimizer with .
Both the actor and the critic are modeled as fully-connected networks with 3 layers and 256 units. We use the Adam optimizer with .
Appendix D Main Algorithm
We present the full LOMPO algorithm in Algorithm 1.
Appendix E Variational Latent Model Samples
For samples generated by our variational latent models, see Figure 5.