RvS: What is Essential for Offline RL via Supervised Learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, Sergey Levine
Introduction
Offline and off-policy reinforcement learning (RL) are typically addressed using value-based methods. While theoretically appealing because they include performance guarantees under certain assumptions , such methods can be difficult to apply in practice; they tend to require complex tricks to stabilize learning and delicate tuning of many hyperparameters. Recent work has explored an alternative approach: convert the RL problem into a conditional, filtered, or weighted imitation learning problem. This typically uses a simple insight: suboptimal experience for one task may be optimal for another task. By conditioning on some piece of information, such as a goal, reward function parameterization, or reward value, such experience can be used for simple behavior cloning . We refer to this set of approaches as RL via supervised learning (RvS). These approaches commonly condition on goals or reward values , but they can also involve reweighting or filtering .
RvS methods are appealing because of their algorithmic simplicity. However, prior work has put forward conflicting hypotheses about which factors are essential for their good performance, including online data , advantage weighting , or large Transformer sequence models . The first question we study is: what elements are essential for effective RvS learning? Beyond this, it also remains unclear on which tasks and datasets such methods work well. For example, prior work has argued that temporal compositionality (dubbed “subtrajectory stitching”) is an important component for solving offline RL when there are few near-optimal trajectories present in the data (e.g., the Franka Kitchen and AntMaze tasks in D4RL ). A priori, one might expect that dynamic programming via TD learning is needed for these tasks. So we also ask: what are the limits of RvS learning, and does it scale to settings with few near-optimal trajectories?
The main contribution of this paper is a study of RvS methods on a wide range of offline RL problems, as well as a set of analyses about which design factors matter most for such methods. First, we show that pure supervised learning (maximizing the likelihood of actions observed in the data) performs as well as conservative TD learning across a diverse set of environments. Second, simple feedforward models can match the performance of more complex sequence models from prior work across a wide range of tasks. Finally, choosing to condition on reward values versus goals can have a large effect on performance, with different choices working better in different domains. These simple results contradict the narrative put forward in many prior works that argue for more complex design decisions . To the best of our knowledge, our results match or exceed those reported by any prior RvS method. We believe that these findings will be useful to the RL community because they help to understand the essential ingredients and limitations of RvS methods, providing a foundation for future work on simple and performant offline RL algorithms.
Related Work
Recent offline RL methods use many techniques, including value functions , dynamics estimation alongisde value functions , dynamics estimation alone , and uncertainty quantification . In this paper, we focus on offline RL methods based on conditional behavior cloning that avoid value functions. The most common instantiation of these methods are goal-conditioned behavior cloning and reward-conditioned behavior cloning . Prior work has also looked at conditioning on different information, such as many previous timesteps , tasks inferred by inverse RL , and other task information . While these methods can be directly applied to the offline RL setting, some prior works combine these methods with iterative data collection , which is not permitted in the typical offline RL setting. Kumar et al. study reward-conditioned policies and do present experiments in the offline RL setting. However, in contrast with our work, the results from Kumar et al. suggest that good performance cannot be achieved through RvS learning but only through additional advantage-weighting . More recent work presents a conditional imitation learning method for the offline RL setting that avoids advantage weighting by introducing a higher-capacity, autoregressive Transformer policy. In contrast to these prior works, we show that simple conditioning with standard feedforward networks can attain state-of-the-art results. While these results do not require advantage weighting or high-capacity sequence models, they do require careful tuning of the policy capacity.
Reinforcement Learning via Supervised Learning
In this section, we describe a formulation of Reinforcement Learning via Supervised Learning. We do not propose a new method but rather place many existing methods under a common framework. After this, we will investigate what design decisions are important to make such methods work well.
While the basic formulation of this conditional policy is simple, instantiating RvS methods that attain excellent results has proven challenging . In the remainder of this paper, we present an empirical analysis of the design choices in RvS methods. We aim to understand which design choices are needed to make RvS methods perform well on diverse benchmark tasks, including how two different choices for (goals and rewards) compare in practice.
Tasks and Datasets
To provide a comprehensive empirical study of RvS methods, we selected a broad range of tasks; the state-based tasks in prior work on various types of RvS or conditional imitation methods typically include only a subset of the domains we consider . Our goal will be to include: (1) domains and datasets that are appropriate for different types of conditioning, including goals and rewards; (2) datasets that include different proportions of near-optimal data, ranging from near-expert datasets to ones with very few or no optimal trajectories at all; (3) datasets that run the gamut in terms of task dimensionality; (4) domains that have been studied by prior value-based offline RL methods.
GCSL is a suite of goal-conditioned environments used by Ghosh et al. to evaluate GCSL, a goal-conditioned RvS method with online data collection. We adapt these tasks for offline RL by using a random policy to collect training data, which results in suboptimal trajectories. The tasks include 2D navigation with obstacles (FourRooms, Eysenbach et al. ); two end-effector controlled Sawyer robotic arm tasks (Door and Pusher, Nair et al. ); the Lunar Lander video game, which requires controlling thrusters to land a simulated Lunar Excursion Module (Lander); and a manipulation task that requires rotating a valve with a robotic claw (Claw, Ahn et al. ).
Gym Locomotion v2 tasks consist of the HalfCheetah, Hopper, and Walker datasets from the D4RL offline RL benchmark . We use the random, medium, medium-expert, and medium-replay datasets in our evaluations, which consist of (a mixture of different) policies with varying levels of optimality. This requires learning from mixed and suboptimal data, and we will see that TD learning methods perform comparatively well on the random data.
Franka Kitchen v0 is a 9-DoF robotic manipulation task paired with datasets of human demonstrations. This task originates from Gupta et al. and was formalized as an offline RL task in D4RL . Solving this task requires composing multi-step behaviors (e.g., open the microwave, then flip a switch) from component skills. This task includes three datasets: complete, where all trajectories solve all tasks in sequence; partial, where only a subset of trajectories perform the desired tasks in sequence; and mixed, which contains various subtasks but never all in sequence, requiring generalization and temporal composition.
AntMaze v2 involves controlling an 8-DoF quadruped to navigate to a particular goal state. This benchmark task, from D4RL , uses a non-Markovian demonstrator policy and was intended to test an agent’s ability to learn temporal compositionality by combining subtrajectories of different demonstrations. There are three mazes of increasing size: umaze, medium, and large, and there are two datasets types: diverse and play. The D4RL paper claims that “diverse” data navigates from random start locations to random goal locations while “play” navigates between hand-picked start-goal pairs. Prior work has proposed that, for tasks requiring temporal compositionality, dynamic programming with value-based methods should be particularly important . However, we find in AntMaze that RvS outperforms all the dynamic programming methods we consider.
When applying RvS learning methods, one must choose a type of outcome, such as rewards or goals. In Kitchen and AntMaze, the reward function is implemented in terms of goals. The kitchen reward is for completing multiple subtasks, so we condition RvS-G on a state in which all the subtasks have been achieved. For AntMaze, we condition RvS-G on the goal location in the maze. This choice makes the additional assumption that we know how the reward function is defined, rather than just receiving samples from the reward function. Because GCSL is inherently a multitask, goal-conditioned setup, we omit reward conditioning and only report RvS-G. Conversely, because there is no clear way to define performance in Gym w.r.t. a goal, we omit goal conditioning and only report RvS-R. For all tasks, we report scores in the range $100\cdot\frac{\text{return}-\text{random}}{\text{expert}-\text{random}}$. (Images from .)
Architecture, Capacity, and Regularization
We instantiate the RvS framework in Section 3 using a simple feedforward, fully-connected neural network. We condition on either the reward-to-go or a goal state by merely concatenating with the input state (see Figure 2). Our aim is to identify the essential components of such methods; we eschew more complex architectures such as Transformer sequence models . On the tasks we consider, we show that our simple implementation achieves performance competitive with prior work that uses these more complex components. However, seemingly minor choices in the architecture do matter. These architectural choices include tuning capacity and regularization, suggesting that overfitting and underfitting are major challenges for RvS. Additionally, the choice of what to condition on (goals or rewards) also has important and domain-specific repercussions.
In Figure 3, we compare different architecture sizes and regularization settings on a single task from each task suite. The best-performing architectures are generally larger than the architectures used in standard online RL and imitation learning . This conclusion is intuitive, since RvS policies must represent both the optimal policy and policies for other conditioning values (e.g., other goals or suboptimal rewards). This result may help explain the contradictory conclusions from prior work regarding the importance of more complex design decisions . The results in Figure 3 also highlight the importance of regularization: while dropout sometimes has no effect (pusher), it improves performance in some tasks (kitchen-complete) and worsens performance in other tasks (hopper-medium-expert and antmaze-medium-play).
What can we say about these results on network capacity and regularization? The hopper-medium-expert dataset is large, and it contains only two modes: one from a medium-quality policy, and one from an expert policy. So, it is not surprising that this data doesn’t need regularization. In contrast, kitchen-complete is a small dataset of human demonstrations solving many different subtasks. Prior work has found that human demonstrations are harder to fit than artificial demonstrations , and this may explain why kitchen-complete generally benefits from dropout. In antmaze-medium-play, relatively few trajectories are successful, so we are surprised to see that the best-performing hyperparameter setting does not use regularization. In pusher, regularization neither helps nor hurts, underscoring that the (lack of) impact of regularization depends on the task at hand. Overall, we speculate that regularization balances a tension between two competing demands on the policy. First, even if the behavior policy is simple, the conditional action distribution may be complex. Second, the policy must be sufficiently well regularized to generalize well to new goals or conditioning variables.
Output distributions.
The form of the policy’s output distribution also influences the policy’s capacity. While unimodal Gaussians are a common choice with continuous action spaces , a categorical distribution over a discretized action space allows the policy to represent more complex, multi-modal distributions . To study the importance of the policy output distribution, we use the low-dimensional GCSL environments, where we can discretize the full action space. Figure 4 (left) shows that across the GCSL suite, categorical distributions either match or outperform Gaussian distributions. Similar to the results from the previous section, these results highlight how increasing the policy’s model capacity can improve performance. Again, our finding stands in contrast to standard RL methods, which work well with unimodal Gaussians We conjecture that upweighting better data, as done by previous methods , may make it easier for low-capacity policies to fit the data. While both upweighting data and increasing model capacity allow the policy to better fit the data, increasing model capacity may be a simpler approach.
Tuning with validation loss.
Hyperparameter optimization is especially important in the offline RL setting, as evaluating different hyperparameter configurations requires interacting with the environment . Since RvS methods reduce the RL problem to a supervised learning problem, we hypothesize that the validation loss of this supervised learning problem might be an effective metric for tuning important hyperparameters, such as policy capacity and regularization. To test this, we train models on 80% of each Franka Kitchen dataset with two different hyperparameter settings: a regularized setting with dropout and small batch size ; and an unregularized setting without dropout and a large batch size of . We show results in Figure 4 (right). For all three datasets, validation set error does correlate with performance, but the strength of this correlation varies significantly. In general, the validation loss does not provide a reliable approach for hyperparameter tuning. Fully automated tuning of hyperparameters remains an open question .
A recipe for practitioners.
We suggest the following process for online hyperparameter tuning: incrementally increase network width until performance saturates, and then try adding a bit of dropout regularization (e.g., dropout ). If validation loss indicates that under/overfitting is an issue, one can also try increasing/decreasing the batch size, respectively.
Table 2 in the Appendix summarizes the hyperparameters that we found to work best for each task.
Comparing RvS with Prior Offline RL Methods
Having identified key design decisions for implementing RvS learning, we now compare RvS to prior offline RL methods. We evaluate both goal-conditioned behavioral cloning (RvS-G) and reward-conditioned behavioral cloning (RvS-R) using the domains in Section 4. We then discuss what these results imply about the performance, design parameters, and limitations of RvS methods.
On the D4RL benchmarks, we compare RvS learning to (i) value-based methods and (ii) prior supervised learning methods and behavioral cloning baselines. For (i), we include CQL as well as the more recent TD3+BC and Onestep RL methods. TD3+BC and Onestep RL both involve elements of behavior cloning. We use CQL-p to denote the CQL numbers published in and CQL-r to denote our best attempt to replicate these results using open-source code and hyperparameters from the CQL authors. For (ii), we include behavioral cloning (BC), which does not perform any conditioning; Filtered (“Filt.”) BC, a baseline appearing in which performs BC after filtering for the trajectories with highest cumulative reward; and Decision Transformer (DT) , which conditions on rewards and uses a large Transformer sequence model. Both BC and Filtered BC are our own optimized implementations , tuned thoroughly in a similar manner as RvS-R and RvS-G. In Kitchen and Gym locomotion, Filtered BC clones the top 10% of trajectories based on cumulative reward. In AntMaze, Filtered BC clones all successful trajectories (those with a reward of 1) and ignores the failing trajectories (those with a reward of 0). On the GCSL benchmarks, we compare to GCSL using numbers reported by the authors . This gives the GCSL baseline the advantage of online data, whereas RvS uses only offline data. We do not run RvS-G on the Gym locomotion tasks, which are typically expressed as reward-maximization tasks, not goal-reaching tasks. For additional details about the baselines, see Appendix A.
Overall performance and comparison to prior methods.
Table 1 shows the results of this comparison, leading to several conclusions about how our RvS implementations compare to prior value-based and supervised methods. In all suites, either RvS-G or RvS-R attains results that are comparable to the best prior method. The finding that conditional imitation with standard fully connected networks attains competitive results stands in contrast to prior work, which emphasizes advantage weighting and Transformer sequence models . As we discussed in Section 5, good performance requires a careful combination of high capacity and regularization, and it may simply be that this seemingly contradictory combination of components was not adequately explored in prior work.
Subtrajectory stitching.
The mixed dataset in Kitchen and all of the AntMaze tasks consist of suboptimal trajectories. Attaining expert performance requires recombining parts of these trajectories. Such stitching is usually viewed as a key benefit of dynamic programming methods, so it is surprising that RvS-G performs as well as prior methods based on dynamic programming (TD3+BC and CQL). In these same tasks, RvS-G outperforms RvS-R. We speculate that the spatial information encoded by goal information helps RvS-G generalize. BC works well in the Kitchen environment but poorly in the AntMaze tasks, perhaps becaue the Kitchen task contains experience that is easier to imitate. As a direction for future work, we propose studying if conditioning on goals can help provide compositionality in space just as the Bellman backup provides compositionality in time.
Reaching arbitrary goals.
An important capability of RvS-G, compared with more standard offline RL methods, is that it learns a policy that can reach many goals. To test how well RvS-G can learn to reach arbitrary goals from random offline data, we compare to GCSL, an RvS learning method that performs iterative online data collection . We use the same tasks used by GCSL. Each of the tasks in the GCSL suite requires that the agent reach a goal drawn uniformly at random from the set of possible goals. The results shown in Figure 5 indicate that RvS-G can successfully reach many different goals entirely from randomly collected offline data. Although it may still be that iterative online collection, as proposed in prior work , may be helpful in some domains, in this case simple offline training appears to suffice.
Random datasets.
While on average RvS-R is competitive with prior methods on the Gym Locomotion tasks, TD learning performs better on the random datasets; indeed, CQL performs especially well on halfcheetah-random (Table 1). This suggests that RvS methods may more generally perform worse on random datasets in comparison to TD learning methods.
Analysis of reward conditioning.
Next, we analyze what RvS-R actually learns and how it uses the conditioning variable. First, we use the walker2d-medium-expert task to analyze the behavior of RvS-R for target rewards. The walker2d-medium-expert dataset contains two modes (“medium” and “expert”, as shown in Figure 6). Conditioning the network on a range of reward targets (X-axis), Figure 6 shows that the policy’s achieved return (Y-axis) corresponds only to the two modes in the dataset. The policy cannot interpolate between the two modes. In effect, it appears that RvS is mimicking a subset of the demonstrations, but doing so without explicit filtering (as is done in Filtered BC).
This analysis highlights the importance of the conditioning variable. It is not obvious a priori that a reward target of 100 will lead to a return of 70, whereas a reward target of 110 will lead to a return of 105. Following prior work , we tune the reward target for every task (see Table 2). Choosing this hyperparameter in a truly offline setting is an important problem for future work.
Discussion and Future Work
We presented an empirical study of offline RL via supervised learning (RvS) methods, which solve RL problems via conditional imitation learning. While prior work is divided on which elements of RvS are crucial for good performance, or even how well such methods actually work, our results provide three conclusions. First, if capacity and regularization are chosen correctly, RvS methods with simple fully connected architectures can match or outperform the best prior methods. Second, properly choosing the conditioning variable (e.g., goal or reward) is critical to the performance of RvS methods. Third, RvS methods remain competitive in tasks such as Franka Kitchen and AntMaze, where there is little optimal data. These conclusions suggest multiple directions for future work. Because policy capacity and regularization are critical to good performance, is there a method for automatically tuning these hyperparameters? Validation error is not a reliable metric. Also, the choice of conditioning variable is important: how might we automate this choice? Addressing these questions may increase the performance and applicability of RvS learning methods.
This work was supported in part by the DOE CSGF under grant number DE-SC0020347.
We thank Michael Janner for discussions about architectures, and we thank Aviral Kumar and Justin Fu for discussions about datasets. We thank Sam Toyer, Micah Carroll, and anonymous reviewers for giving feedback on initial drafts of this work.
References
Appendix A Experiment Details
In this section, we provide more details about our experiments in each environment suite. For every task, we use 5 random training seeds and 200 evaluation rollouts for RvS.
GCSL Following the GCSL protocol , we collect offline data for RvS by performing random rollouts in each environment. As in GCSL, we set the environment time limit (i.e., max episode duration) to 50 actions, and the rollout policy uses a discretized action space. We collect 50,000 timesteps of experience in FourRooms, 100,000 timesteps of experience in Door and Lander, and 500,000 timesteps of experience in Pusher and Claw. We take our GCSL baseline numbers from those reported in the GCSL paper at the corresponding number of timesteps for each environment .
Gym Locomotion We use the v2 versions of the datasets for all algorithms. Because the v2 datasets fix issues with the v0 datasets, we omit the published CQL-p numbers, which use v0 and are no longer comparable. We instead report our replication of CQL, denoted CQL-r, using v2. We report the DT , TD3+BC , and Onestep numbers from the corresponding papers. We use the reverse KL regularization variant of Onestep because this variant attains the best performance. Because DT did not run experiments with the random datasets, we run them ourselves using the default hyperparameters in the open-source codebase released by the DT authors. We omit numbers for RvS-G because there is no clear way to define the Gym tasks in terms of goals.
Franka Kitchen We use the v0 versions of the datasets and omit TD3+BC, Onestep, and DT numbers, which are not provided by the authors. We include the previously published CQL-p numbers from D4RL as well as CQL-r, our own replication of CQL, and BC and Filtered BC, our own baselines.
AntMaze We use the v2 versions of the datasets for RvS, CQL-r, BC, Filtered BC, and DT. Because the DT authors do not test in AntMaze, we ran these experiments ourselves using the default hyperparameters in the open-source codebase released by the DT authors. We include previously published numbers for CQL-p and TD3+BC , as well as Onestep numbers we received via email correspondence with the Onestep authors. These numbers for CQL-p, TD3+BC, and Onestep use the v0 versions of the datasets, which are the same as the v2 versions except that the v0 timeout flags indicating the ends of episodes contain errors. Because our initial experiments suggest that TD learning methods are fairly robust to errors in the timeout flags, we include these numbers for the sake of completeness.
Codebases Here are links to all of the codebases we use in our experiments:
CQL-r: https://github.com/scottemmons/youngs-cql
DT: https://github.com/scottemmons/decision-transformer
BC and Filtered BC: https://github.com/ikostrikov/jaxrl
Appendix B The Impact of Model Capacity and Regularization
As we find in Section 5 that model capacity and regularization are important for performance in seemingly contradictory ways, we hypothesize that this can be explained by validation loss. So, for 5 random seeds in each of the 3 Franka Kitchen datasets (complete, partial, and mixed), we measure validation loss and return throughout training for three different settings of policy hyperparameters: network width 1024 with dropout p = 0.1, network width 1024 with dropout p = 0 (no dropout), and network width 256 with dropout p = 0 (no dropout). We expect that a larger network width leads to better validation loss, explaining its better return. We expect adding regularization also has this pattern of better validation loss and therefore better return.
In Figure 7, we plot the results: validation loss is nearly identical across all of the hyperparameter settings, but return can vary by as much as 1.4x between different hyperparameter settings. We also observe that evaluation return steadily increases even after validation loss has mostly leveled off. This is surprising and refutes our hypothesis; furthermore, it provides the insight that model capacity and dropout are having effects beyond just the validation loss. A similar phenomenon occurs in Neural Machine Translation, where cross entropy loss can decline while the quality of the translation (measured by BLEU score) remains constant , and in contrastive representation learning, where lower capacity models have worse losses on the test set but still learn better representations for downstream tasks . It is an important open problem for the field to understand why architecture decisions can impact performance beyond validation loss.
Appendix C Comparison of Conditioning Strategies
To further study how different goal and reward conditioning strategies compare, in Figure 8(a), we show the performance of four strategies: Reward Scalar, Reward Goal, Length Goal, and Optimized Goal. In Reward Scalar, we condition RvS-R on a return that is various fractions of expert performance. In Reward Goal, we condition RvS-G on states from randomly selected timesteps near the end of the highest-reward demonstration trajectories. In Length Goal, we condition RvS-G on states from randomly selected timesteps near the end of the longest demonstration trajectories. In Optimized Goal, we do a random search over 200 Length Goals and condition RvS-G on the highest performing one. (Note that Optimized Goal requires access to the environment, so it is not strictly offline RL; we are including it here as an oracle experiment.) We see that across all the goal selection strategies we consider, the performance of RvS-G cannot match the performance of RvS-R, since the Gym tasks are not inherently goal-conditioned.
In Figure 8(b), we compare the performance of RvS-G in kitchen-complete when commanded to reach different goal states. The kitchen-complete task requires the agent to open a microwave, move a kettle, flip a light switch, and slide open a cabinet. For “all,” we command the policy to reach a state where all the subtasks are achieved. For “random,” we randomly select one of the tasks and command the policy to reach a state where that task is complete. For “microwave,” “kettle,” “light switch,” and “slide cabinet,” we command the agent to reach a state where that one subtask is achieved. For “dynamic,” rather than leaving the goal fixed throughout the trajectory, we initially command the agent to complete the first subtask. Then, once the first subtask is complete, we update the command to include the second subtask (and so on for all remaining subtasks). As expected, we see that the oracle dynamic strategy, as well as conditioning on all subtasks completed, perform the best. Surprisingly, commanding the microwave to be opened outperforms commanding the other individual subtasks. We speculate that the microwave subtask might appear most often in the demonstrations.