Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
Tianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman, Sergey Levine, Chelsea Finn
Introduction
Recent advances in offline reinforcement learning (RL) make it possible to train policies for real-world scenarios, such as robotics and healthcare , entirely from previously collected data. Many realistic settings where we might want to apply offline RL are inherently multi-task problems, where we want to solve multiple tasks using all of the data available. For example, if our goal is to enable robots to acquire a range of different behaviors, it is more practical to collect a modest amount of data for each desired behavior, resulting in a large but heterogeneous dataset, rather than requiring a large dataset for every individual skill. Indeed, many existing datasets in robotics and offline RL include data collected in precisely this way. Unfortunately, leveraging such heterogeneous datasets leaves us with two unenviable choices. We could train each task only on data collected for that task, but such small datasets may be inadequate for good performance. Alternatively, we could combine all of the data together and use data relabeled from other tasks to improve offline training, but this naïve data sharing approach can actually often degrade performance over simple single-task training in practice . In this paper, we aim to understand how data sharing affects RL performance in the offline setting and develop a reliable and effective method for selectively sharing data across tasks.
A number of prior works have studied multi-task RL in the online setting, confirming that multi-tasking can often lead to performance that is worse than training tasks individually . These prior works focus on mitigating optimization challenges that are aggravated by the online data generation process . As we will find in Section 4, multi-task RL remains a challenging problem in the offline setting when sharing data across tasks, even when exploration is not an issue. While prior works have developed heuristic methods for reweighting and relabeling data , they do not yet provide a principled explanation for why data sharing can hurt performance in the offline setting, nor do they provide a robust and general approach for selective data sharing that alleviates these issues while preserving the efficiency benefits of sharing experience across tasks.
In this paper, we hypothesize that data sharing can be harmful or brittle in the offline setting because it can exacerbate the distribution shift between the policy represented in the data and the policy being learned. We analyze the effect of data sharing in the offline multi-task RL setting, and present evidence to support this hypothesis. Based on this analysis, we then propose an approach for selective data sharing that aims to minimize distributional shift, by sharing only data that is particularly relevant to each task. Instantiating a method based on this principle requires some care, since we do not know a priori which data is most relevant for a given task before we’ve learned a good policy for that task. To provide a practical instantiation, we propose the conservative data sharing (CDS) algorithm. CDS reduces distributional shift by sharing data based on a learned conservative estimate of the Q-values that penalizes Q-values on out-of-distribution actions. Specifically, CDS relabels transitions when the conservative Q-value of the added transitions exceeds the expected conservative Q-values on the target task data. We visualize how CDS works in Figure 1.
The main contributions of this work are an analysis of data sharing in offline multi-task RL and a new algorithm, conservative data sharing (CDS), for multi-task offline RL problems. CDS relabels a transition into a given task only when it is expected to improve performance based on a conservative estimate of the Q-function. After data sharing, similarly to prior offline RL methods, CDS applies a standard conservative offline RL algorithm, such as CQL , that learns a conservative value function or BRAC , a policy-constraint offline RL algorithm. Further, we theoretically analyze CDS and characterize scenarios under which it provides safe policy improvement guarantees. Finally, we conduct extensive empirical analysis of CDS on multi-task locomotion, multi-task robotic manipulation with sparse rewards, multi-task navigation, and multi-task imaged-based robotic manipulation. We compare CDS to vanilla offline multi-task RL without sharing data, to naïvely sharing data for all tasks, and to existing data relabeling schemes for multi-task RL. CDS is the only method to attain good performance across all of these benchmarks, often significantly outperforming the best domain-specific method, improving over the next best method on each domain by 17.5% on average.
Related Work
Offline RL. Offline RL has shown promise in domains such as robotic manipulation , NLP , recommender systems & advertising , and healthcare . The major challenge in offline RL is distribution shift , where the learned policy might generate out-of-distribution actions, resulting in erroneous value backups. Prior offline RL methods address this issue by regularizing the learned policy to be “close“ to the behavior policy , through variants of importance sampling , via uncertainty quantification on Q-values , by learning conservative Q-functions , and with model-based training with a penalty on out-of-distribution states . While current benchmarks in offline RL contain datasets that involve multi-task structure, existing offline RL methods do not leverage the shared structure of multiple tasks and instead train each individual task from scratch. In this paper, we exploit the shared structure in the offline multi-task setting and train a general policy that can acquire multiple skills.
Multi-task RL algorithms. Multi-task RL algorithms focus on solving multiple tasks jointly in an efficient way. While multi-task RL methods seem to provide a promising way to build general-purpose agents , prior works have observed major challenges in multi-task RL, in particular, the optimization challenge . Beyond the optimization challenge, how to perform effective representation learning via weight sharing is another major challenge in multi-task RL. Prior works have considered distilling per-task policies into a single policy that solves all tasks , separate shared and task-specific modules with theoretical guarantees , and incorporating additional supervision . Finally, sharing data across tasks emerges as a challenge in multi-task RL, especially in the off-policy setting, as naïvely sharing data across all tasks turns out to hurt performance in certain scenarios . Unlike most of these prior works, we focus on the offline setting where the challenges in data sharing are most relevant. Methods that study optimization and representation learning issues are complementary and can be readily combined with our approach.
Data sharing in multi-task RL. Prior works have found it effective to reuse data across tasks by recomputing the rewards of data collected for one task and using such relabeled data for other tasks, which effectively augments the amount of data available for learning each task and boosts performance. These methods perform relabeling either uniformly or based on metrics such as estimated Q-values , domain knowledge , the distance to states or images in goal-conditioned settings , and metric learning for robust inference in the offline meta-RL setting . All of these methods either require online data collection and do not consider data sharing in a fully offline setting, or only consider offline goal-conditioned or meta-RL problems . While these prior works empirically find that data sharing helps, we believe that our analysis in Section 4 provides the first analytical understanding of why and when data sharing can help in multi-task offline RL and why it hurts in some cases. Specifically, our analysis reveals the effect of distributional shift introduced during data sharing, which is not taken into account by these prior works. Our proposed approach, CDS, tackles the challenge of distributional shift in data sharing by intelligently sharing data across tasks and improves multi-task performance by effectively trading off between the benefits of data sharing and the harms of excessive distributional shift.
Preliminaries and Problem Statement
Standard offline RL is concerned with learning policies using only a given static dataset of transitions , collected by a behavior policy , without any additional environment interaction. In the multi-task offline RL setting, the dataset is partitioned into per-task subsets, , where consists of experience from task . While algorithms can choose to train the policy for task (i.e., ) only on , in this paper, we are interested in data-sharing schemes that correspond to relabeling data from a different task, with the reward function , and learn on the combined data. To be able to do so, we assume access to the functional form of the reward , a common assumption in goal-conditioned RL , and which often holds in robotics applications through the use of learned classifiers , and discriminators .
Offline RL algorithms. A central challenge in offline RL is distributional shift: differences between the learned policy and the behavior policy can lead to erroneous target values, where the Q-function is queried at actions that are far from the actions it is trained on, leading to massive overestimation . A number of offline RL algorithms use some kind of regularization on either the policy or on the learned Q-function to ensure that the learned policy does not deviate too far from the behavior policy. For our analysis in this work, we will abstract these algorithms into a generic constrained policy optimization problem :
We will utilize this generic optimization problem to motivate our method in Section 5.
When Does Data Sharing Actually Help in Offline Multi-Task RL?
Our goal is to leverage experience from all tasks to learn a policy for a particular task of interest. Perhaps the simplest approach to leveraging experience across tasks is to train the task policy on not just the data coming from that task, but also relabeled data from all other tasks . Is this naïve data sharing strategy sufficient for learning effective behaviors from multi-task offline data? In this section, we aim to answer this question via empirical analysis on a relatively simple domain, which will reveal interesting aspects of data sharing. We first describe the experimental setup and then discuss the results and possible explanations for the observed behavior. Using insights obtained from this analysis, we will then derive a simple and effective data sharing strategy in Section 5.
Experimental analysis setup. To assess the efficacy of data sharing, we experimentally analyze various multi-task RL scenarios created with the walker2d environment in Gym . We construct different test scenarios on this environment that mimic practical situations, including settings where different amounts of data of varied quality are available for different tasks . In all these scenarios, the agent attempts three tasks: run forward, run backward, and jump, which we visualize in Figure 3. Following the problem statement in Section 3, these tasks share the same state-action space and transition dynamics, differing only in the reward function that the agent is trying to optimize. Different scenarios are generated with varying size offline datasets, each collected with policies that have different degrees of suboptimality. This might include, for each task, a single policy with mediocre or expert performance, or a mixture of policies given by the initial part of the replay buffer trained with online SAC . We refer to these three types of offline datasets as medium, expert and medium-replay, respectively, following Fu et al. .
Analysis of results in Table 1. To begin, note that even naïvely sharing data is better than not sharing any data at all on 5/9 tasks considered (compare the performance across No Sharing and Sharing All in Table 1). However, a closer look at Table 1 suggests that data-sharing can significantly degrade performance on certain tasks, especially in scenarios where the amount of data available for the original task is limited, and where the distribution of this data is narrow. For example, when using expert data for jumping in conjunction with more than 25 times as much lower-quality (mediocre & random) data for running forward and backward, we find that the agent performs poorly on the jumping task despite access to near-optimal jumping data.
Why does naïve data sharing degrade performance on certain tasks despite near-optimal behavior for these tasks in the original task dataset? We argue that the primary reason that naïve data sharing can actually hurt performance in such cases is because it exacerbates the distributional shift issues that afflict offline RL. Many offline RL methods combat distribution shift by implicitly or explicitly constraining the learned policy to stay close to the training data. Then, when the training data is changed by adding relabeled data from another task, the constraint causes the learned policy to change as well. When the added data is of low quality for that task, it will correspondingly lead to a lower quality learned policy for that task, unless the constraint is somehow modified. This effect is evident from the higher divergence values between the learned policy without any data-sharing and the effective behavior policy for that task after relabeling (e.g., expert+jump) in Table 1. Although these results are only for CQL, we expect that any offline RL method would, insofar as it combats distributional shift by staying close to the data, would exhibit a similar problem.
To mathematically quantify the effects of data-sharing in multi-task offline RL, we appeal to safe policy improvement bounds and discuss cases where data-sharing between tasks and can degrade the amount of worst-case guaranteed improvement over the behavior policy. Prior work has shown that the generic offline RL algorithm in Equation 1 enjoys the following guarantees of policy improvement on the actual MDP, beyond the behavior policy:
To conclude, our analysis reveals that while data sharing is often helpful in multi-task offline RL, it can lead to substantially poor performance on certain tasks as a result of exacerbated distributional shift between the optimal policy and the effective behavior policy induced after sharing data.
CDS: Reducing Distributional Shift in Multi-Task Data Sharing
The analysis in Section 4 shows that naïve data sharing may be highly sub-optimal in some cases, and although it often does improve over no data sharing at all in practice, it can also lead to exceedingly poor performance. Can we devise a conservative approach that shares data intelligently to not exacerbate distributional shift as a result of relabeling?
(4) The scheme presented in Equation 4 would guarantee that distributional shift (i.e., second term in Equation 2) is reduced. Moreover, since sharing data can only increase the size of the dataset and not reduce it, this scheme is guaranteed to not increase the sampling error term in Equation 3. We refer to this scheme as the basic variant of conservative data sharing (CDS (basic)).
2 The Complete Version of Conservative Data Sharing (CDS)
Let be the policy obtained by optimizing Equation 5, and let be the behavior policy for . Then, w.h.p. , is a -safe policy improvement over , i.e., , where is given by:
3 Practical implementation of CDS
The pseudocode of CDS is summarized in Algorithm 1. The complete variant of CDS can be directly implemented using the rule in Equation 6 with conservative Q-value estimates obtained via any offline RL method that constrains the learned policy to the behavior policy. For implementing CDS (basic), we reparameterize the divergence in Equation 4 to use the learned conservative Q-values. This is especially useful for our implementation since we utilize CQL as the base offline RL method, and hence we do not have access to an explicit divergence. In this case, can be redefined as,
Equation 7 can be viewed as the difference between the CQL regularization term on a given and the original dataset for task , . This CQL regularization term is equal to the divergence between the learned policy and the behavior policy , therefore Equation 7 practically computes Equation 4.
Experimental Evaluation
We conduct experiments to answer six main questions: (1) can CDS prevent performance degradation when sharing data as observed in Section 4?, (2) how does CDS compare to vanilla multi-task offline RL methods and prior data sharing methods? (3) can CDS handle sparse reward settings, where data sharing is particularly important due to scarce supervision signal? (4) can CDS handle goal-conditioned offline RL settings where the offline dataset is undirected and highly suboptimal? (5) Can CDS scale to complex visual observations? (6) Can CDS be combined with any offline RL algorithms? Besides these questions, we visualize CDS weights for better interpretation of the data sharing scheme learned by CDS in Figure 4 in Appendix C.2.
Comparisons. To answer these questions, we consider the following prior methods. On tasks with low dimensional state spaces, we compare with the online multi-task relabeling approach HIPI , which uses inverse RL to infer for which tasks the datapoints are optimal and in practice routes a transition to task with the highest Q-value. We adapt HIPI to the offline setting by applying its data routing strategy to a conservative offline RL algorithm. We also compare to naïvely sharing data across all tasks (denoted as Sharing All) and vanilla multi-task offline RL method without any data sharing (denoted as No Sharing). On image-based domains, we compare CDS to the data sharing strategy based on human-defined skills (denoted as Skill), which manually groups tasks into different skills (e.g. skill “pick” and skill “place”) and only routes an episode to target tasks that belongs to the same skill. In these domains, we also compare to HIPI, Sharing All and No Sharing. Beyond these multi-task RL approaches with data sharing, to assess the importance of data sharing in offline RL, we perform an additional comparison to other alternatives to data sharing in multi-task offline RL settings. One traditionally considered approach is to use data from other tasks for some form of “pre-training” before learning to solve the actual task. We instantiate this idea by considering a method from Yang and Nachum that conducts contrastive representation learning on the multi-task datasets to extract shared representation between tasks and then runs multi-task RL on the learned representations. We discuss this comparison in detail in Table 7 in Appendix C.3. To answer question (6), we use CQL (a Q-function regularization method) and BRAC (a policy-constraint method) as the base offline RL algorithms for all methods. We discuss evaluations of CDS with CQL in the main text and include the results of CDS with BRAC in Table 5 in Appendix C.1. For more details on setup and hyperparameters, see Appendix B.
Multi-task environments. We consider a number of multi-task reinforcement learning problems on environments visualized in Figure 3. To answer questions (1) and (2), we consider the walker2d locomotion environment from OpenAI Gym with dense rewards. We use three tasks, run forward, run backward and jump, as proposed in prior offline RL work . To answer question (3), we also evaluate on robotic manipulation domains using environments from the Meta-World benchmark . We consider four tasks: door open, door close, drawer open and drawer close. Meaningful data sharing requires a consistent state representation across tasks, so we put both the door and the drawer on the same table, as shown in Figure 3. Each task has a sparse reward of 1 when the success condition is met and 0 otherwise. To answer question (4), we consider maze navigation tasks where the temporal “stitching” ability of an offline RL algorithm is crucial to obtain good performance. We create goal reaching tasks using the ant robot in the medium and hard mazes from D4RL . The set of goals is a fixed discrete set of size 7 and 3 for large and medium mazes, respectively. Following Fu et al. , a reward of +1 is given and the episode terminates if the state is within a threshold radius of the goal. Finally, to explore how CDS scales to image-based manipulation tasks (question (5)), we utilize a simulation environment similar to the real-world setup presented in . This environment, which was utilized by Kalashnikov et al. as a representative and realistic simulation of a real-world robotic manipulation problem, consists of 10 image-based manipulation tasks that involve different combinations of picking specific objects (banana, bottle, sausage, milk box, food box, can and carrot) and placing them in one of the three fixtures (bowl, plate and divider plate) (see example task images in Fig. 3). We pick these tasks due to their similarity to the real-world setup introduced by Kalashnikov et al. , which utilized a skill-based data-sharing heuristic strategy (Skill) for data-sharing that significantly outperformed simple data-sharing alternatives, which we use as a point of comparison. More environment details are in the appendix. We report the average return for locomotion tasks and success rate for AntMaze and both manipluation environments, averaged over 6 and 3 random seeds for environments with low-dimensional inputs and image inputs respectively.
Multi-task datasets. Following the analysis in Section 4, we intentionally construct datasets with a variety of heterogeneous behavior policies to test if CDS can provide effective data sharing to improve performance while avoiding harmful data sharing that exacerbates distributional shift. For the locomotion domain, we use a large, diverse dataset (medium-replay) for run forward, a medium-sized dataset for run backward, and an expert dataset with limited data for run jump. For Meta-World, we consider medium-replay datasets with 152K transitions for task door open and drawer close and expert datasets with only 2K transitions for task door close and drawer open. For AntMaze, we modify the D4RL datasets for antmaze-*-play environments to construct two kinds of multi-task datasets: an “undirected” dataset, where data is equally divided between different tasks and the rewards are correspondingly relabeled, and a “directed” dataset, where a trajectory is associated with the goal closest to the final state of the trajectory. This means that the per-task data in the undirected setting may not be relevant to reaching the goal of interest. Thus, data-sharing is crucial for good performance: methods that do not effectively perform data sharing and train on largely task-irrelevant data are expected to perform worse. Finally, for image-based manipulation tasks, we collect datasets for all the tasks individually by running online RL until the task reaches medium-level performance (40% for picking tasks and 80% placing tasks). At that point, we merge the entire replay buffers from different tasks creating a final dataset of 100K RL episodes with 25 transitions for each episode.
Results on domains with low-dimensional states. We present the results on all non-vision environments in Table 3. CDS achieves the best average performance across all environments except that on walker2d, it achieves the second best performance, obtaining slightly worse the other variant CDS (basic). On the locomotion domain, we observe the most significant improvement on task jump on all three environments. We interpret this as strength of conservative data sharing, which mitigates the distribution shift that can be introduced by routing large amount of other task data to the task with limited data and narrow distribution. We also validate this by measuring the in Table 2 where is the behavior policy after we perform CDS to share data. As shown in Table 2, CDS achieves lower KL divergence between the single-task optimal policy and the behavior policy after data sharing on task jump with limited expert data, whereas Sharing All results in much higher KL divergence compared to No Sharing as discussed in Section 4 and Table 1. Hence, CDS is able to mitigate distribution shift when sharing data and result in performance boost.
On the Meta-World tasks, we find that the agent without data sharing completely fails to solve most of the tasks due to the low quality of the medium replay datasets and the insufficient data for the expert datasets. Sharing All improves performance since in the sparse reward settings, data sharing can introduce more supervision signal and help training. CDS further improves over Sharing All, suggesting that CDS can not only prevent harmful data sharing, but also lead to more effective multi-task learning compared to Sharing All in scenarios where data sharing is imperative. It’s worth noting that CDS (basic) performs worse than CDS and Sharing All, indicating that relabeling data that only mitigates distributional shift is too pessimistic and might not be sufficient to discover the shared structure across tasks.
In the AntMaze tasks, we observe that CDS performs better than Sharing All and significantly outperforms HIPI in all four settings. Perhaps surprisingly, No Sharing is a strong baseline, however, is outperformed by CDS with the harder undirected data. Moreover, CDS performs on-par or better in the undirected setting compared to the directed setting, indicating the effectiveness of CDS in routing data in challenging settings.
Results on image-based robotic manipulation domains. Here, we compare CDS to the hand-designed Skill sharing strategy, in addition to the other methods. Given that CDS achieves significantly better performance than CDS (basic) on low-dimensional robotic manipulation tasks in Meta-World, we only evaluate CDS in the vision-based robotic manipulation domains. Since CDS is applicable to any offline multi-task RL algorithm, we employ it as a separate data-sharing strategy in while keeping the model architecture and all the other hyperparameters constant, which allows us to carefully evaluate the influence of data sharing in isolation. The results are reported in Table 4. CDS outperforms both Skill and other approaches, indicating that CDS is able to scale to high-dimensional observation inputs and can effectively remove the need for manual curation of data sharing strategies.
Conclusion
In this paper, we study the multi-task offline RL setting, focusing on the problem of sharing offline data across tasks for better multi-task learning. Through empirical analysis, we identify that naïvely sharing data across tasks generally helps learning but can significantly hurt performance in scenarios where excessive distribution shift is introduced. To address this challenge, we present conservative data sharing (CDS), which relabels data to a task when the conservative Q-value of the given transition is better than the expected conservative Q-value of the target task. On multitask locomotion, manipulation, navigation, and vision-based manipulation domains, CDS consistently outperforms or achieves comparable performance to existing data sharing approaches. While CDS attains superior results, it is not able to handle data sharing in settings where dynamics vary across tasks and requires functional forms of rewards. We leave these as future work.
Acknowledgements
We thank Kanishka Rao, Xinyang Geng, Avi Singh, other members of RAIL at UC Berkeley, IRIS at Stanford and Robotics at Google and anonymous reviewers for valuable and constructive feedback on an early version of this manuscript. This research was funded in part by Google, ONR grant N00014-20-1-2675, Intel Corporation and the DARPA Assured Autonomy Program. CF is a CIFAR Fellow in the Learning in Machines and Brains program.
References
Appendices
In this section, we will analyze the key idea behind our method CDS (Section 5) and show that the abstract version of our method (Equation 5) provides better policy improvement guarantees than naïve data sharing and that the practical version of our method (Equation 6) approximates Equation 5 resulting in an effective practical algorithm.
We begin with analyzing Equation 5, which is used to derive the practical variant of our method, CDS. We build on the analysis of safe-policy improvement guarantees of conventional offline RL algorithms and show that data sharing using CDS attains better guarantees in the worst case. To begin the analysis, we introduce some notation and prior results that we will directly compare to.
Notation and prior results. Let denote the behavior policy for task (note that index was dropped from for brevity). The dataset, is generated from the marginal state-action distribution of , i.e., . We define as the state marginal distribution introduced by the dataset under . Let denote the following distance between two distributions and with equal support :
The first term in Equation 8 corresponds to the decrease in performance due to sampling error and this term is high when the single-task optimal policy visits rarely observed states in the dataset and/or when the divergence from the behavior policy is higher under the states visited by the single-task policy .
Let denote the return of a policy in the empirical MDP induced by the transitions in the dataset . Further, let us assume that optimizing Equation 5 gives us the following policies:
We now show the following result for CDS:
Let be the policy obtained by optimizing Equation 5, and let be the behavior policy for . Then, w.h.p. , is a -safe policy improvement over , i.e., , where is given by:
where the notation hides constants depending upon the concentration properties of the MDP and , the probability with which the statement holds. Next, we provide guarantees on policy improvement in the empirical MDP. To see this, note that the following statements on are true:
where ignores sampling error terms that do not depend on distributional shift measures like because and are behavior policies which generated the complete and part of the dataset, and hence these terms are dominated by and subsumed into the sampling error for . Combining Equations 10 (by setting ) and 15, we obtain the following safe-policy improvement guarantee for : , where is given by:
Finally, we show that the sampling error term is controlled when utilizing Equation 5. We will show in Lemma A.2 that the sampling error in Proposition A.1 is controlled to be not much bigger than the error just due to variance, since distributional shift is bounded with Equation 5.
If and obtained from Equation 5 satisfy, , then:
This lemma can be proved via a simple application of the Cauchy-Schwarz inequality. We can partition the first term as a sum over dot products of two vectors such that:
A.2 From Equation 5 to Practical CDS (Equation 6)
Appendix B Experimental details
In this section, we provide the training details of CDS in Appendix B.1 and also include the details on the environment and datasets that we use for the evaluation in Appendix B.2. Finally, we include the discussion on the compute information in Appendix B.3. We also compare CDS to an offline RL with pretrained representations from multi-task datasets method .
Our practical implementation of CDS optimizes the following objectives for training the critic and the policy:
where is the coefficient of the CQL penalty on distribution shift, is a wide sampling distribution as in CQL and is the sample-based Bellman operator.
For state-based experiments, we use a stratified batch with transitions for each task for the critic and policy learning. For each task , we sample transitions from and another transitions from , i.e. the relabeled datasets of all the other tasks. When computing , we only apply the weight to relabeled data on multi-task Meta-World environments and multi-task vision-based robotic manipulation tasks while also applying the weight to the original data drawn from with 50% chance for each task in the remaining domains.
We use CQL as the base offline RL algorithm. On state-based experiments, we mostly follow the hyperparameters provided in prior work . One exception is that on the multi-task ant domain, we set and on the other two locomotion environments and the multi-task Meta-World domain, we use . On multi-task AntMaze, we use the Lagrange version of CQL, where the multiplier is automatically tuned against a pre-specific constraint value on the CQL loss equal to . We use a policy learning rate and a critic learning rate as in . On the vision-based environment, instead of using the direct CQL algorithm, we follow and sample unseen actions according to the soft-max distritbution of the Q-values and set its Q target value to . This algorithm can be viewed the version of CQL with in Eq.1 in , i.e. removing the term of negative expected Q-values on the dataset. We follow the other hyperparameters from prior work .
For the choice architectures, in the domains with low-dimensional state inputs, we use 3-layer feedforward neural networks with hidden units for both the Q-networks and the policy. We append a one-hot task vector to the state of each environment. For the vision-based experiment, our Q-network architecture follows from multi-headed convolutional networks used in MT-Opt . For the observation input, we use images with dimension along with additional state features as well as the one-hot task vector as in . For the action input, we use Cartesian space control of the end-effector of the robot in 4D space (3D position and azimuth angle) along with two discrete actions for opening/closing the gripper and terminating the episode respectively. More details can be found in .
B.2 Environment and dataset details
In this subsection, we discuss the details of how we set up the multi-task environment and how we collect the offline datasets. We want to acknowledge that all datasets with state inputs use the MIT License.
Multi-task locomotion domains. We construct the environment by changing the reward function in . On the halfcheetah environment, we follow and set the reward functions of task run forward, run backward and jump as , and respectively where denotes the velocity along the x-axis and denotes the z-position of the half-cheetah and init z denotes the initial z-position. Similarly, on walker2d, the reward functions of the three tasks are , and respectively. Finally, on ant, the reward functions of the three tasks are , and .
On each of the multi-task locomotion environment, we train each task with SAC for 500 epochs. For medium-replay datasets, we take the whole replay buffer after the online SAC is trained for 100 epochs. For medium datasets, we take the online single-task SAC policy after 100 epochs and collect 500 trajectories with the medium-level policy. For expert datasets, we take the final online SAC policy and collect 5 trajectories with it for walker2d and halfcheetah and 20 trajectories for ant.
Meta-World domains. We take the door open, door close, drawer open and drawer close environments from the open-sourced Meta-World repoThe Meta-World environment can be found at the public repo https://github.com/rlworkgroup/metaworld. We put both the door and the drawer on the same scene to make sure the state space of all four tasks are shared. For offline training, we use sparse rewards for each task by replacing the dense reward defined in Meta-World with the success condition defined in the public repo. Therefore, each task gets a reward of 1 if the task is fully completed and 0 otherwise.
For generating the offline datasets, we train each task with online SAC using the dense reward defined in Meta-World for 500 epochs. For medium-replay datasets, we take the whole replay buffer of the online SAC until 150 epochs. For the expert datasets, we run the final online SAC policy to collect 10 trajectories.
AntMaze domains. We take the antmaze-medium-play and antmaze-large-play datasets from D4RL and convert the datasets into multi-task datasets in two ways. In the undirected version of these tasks, we split the dataset randomly into equal sized partitions, and then assign each partition to a particular randomly chosen task. Thus, the task data observed in the data for each task is largely unsuccessful for the particular task it is assigned to and effective data sharing is essential for obtaining good performance. The second setting is the directed data setting where a trajectory in the dataset is marked to belong to the task corresponding to the actual end goal of the trajectory. A sparse reward equal to +1 is provided to an agent when the current state reaches within a 0.5 radius of the task goal as was used default by Fu et al. .
Vision-based robotic manipulation domains. Following MT-Opt , we use sparse rewards for each task, i.e. reward 1 for success episodes and 0 otherwise. We define successes using the success detectors defined in . To collect data for vision-based experiments, we train a policy for each task individually by running QT-Opt with default hyperparameters until the task reaches 40% success rate for picking skills and 80% success rate for placing skills. We take the whole replay buffer of each task and combine all of such replay buffers to form the multi-task offline dataset with total 100K episodes where each episode has 25 transitions.
B.3 Computation Complexity
For all the state-based experiments, we train CDS on a single NVIDIA GeForce RTX 2080 Ti for one day. For the image-based robotic manipulation experiments, we train it on 16 TPUs for three days.
Appendix C Visualizations, Comparisons and Additional Experiments
In this section, we perform diagnostic and ablation experiments to: (1) understand the efficacy of CDS when applied with other base offline RL algorithms, such as BRAC , (2) visualize the weights learned by CDS to understand if the weighting scheme induced by CDS corresponds to what we would intuitively expect on different tasks, and (3) compare CDS to a prior approach that performs representation learning from offline multi-task datasets and then runs vanilla multi-task RL algorithm on top of the learned representations. We discuss these experiments next.
We evaluated BRAC + CDS on the Meta-World tasks and compared it to BRAC + Sharing All and BRAC + No Sharing. We present the results in Table 5. We use to denote the 95%-confidence interval. As observed below, BRAC + CDS significantly outperforms BRAC with Sharing All and BRAC with No sharing. This indicates that CDS is effective on top of BRAC.
C.2 Analyzing CDS weights for Different Scenarios
Next, to understand if the weights assigned by CDS align with our expectation for which transitions should be shared between tasks, we perform diagnostic analysis studies on the weights learned by CDS on the Meta-World and Antmaze domains.
On the Meta-World environment, we would expect that for a given target task, say Drawer Close, transitions from a task that involves a different object (door) and a different skill (open) would not be as useful for learning. To understand if CDS weights reflect this expectation, we compare the average CDS weights on transitions from all the other tasks to two target tasks, Door Open and Drawer Close, respectively and present the results in Table 6. We sort the CDS weights in the descending order. As shown, indeed CDS assigns higher weights to more related tasks and thus shares data from those tasks. In particular, the CDS weights for relabeling data from the task that handles the same object as the target task are much higher than the weights for tasks that consider a different object.
For example, when relabeling to the target task Door Open, datapoints from task Door Close are assigned with much higher weights than those from either task Drawer Open or task Drawer Close. This suggests that CDS filters the irrelevant transitions for learning a given task.
On the AntMaze-large environment, with undirected data, we visualize the CDS weight for the various tasks (goals) in the form of a heatmap and present the results in Figure 4. To generate this plot, we sample a set of state-action pairs from the entire dataset for all tasks, and then plot the weights assigned by CDS as the color of the point marker at the locations of these state-action pairs in the maze. Each plot computes the CDS weight corresponding to the target task (goal) indicated by the red in the plot. As can be seen in Figure 4, CDS assigns higher weights to transitions from nearby goals as compared to transitions from farther away goals. This matches our expectation: transitions from nearby locations are likely to be the most useful in learning a particular target task and CDS chooses to share these transitions to the target task.
Finally, we aim to empirically verify how other alternatives to data sharing perform on multi-task offline RL problems. One simple approach to utilize data from other tasks is to use this data to learn low-dimensional representations that capture meaningful information about the environment initially in a pre-training phase and then utilize these representations for improved multi-task RL without any specialized data sharing schemes. To assess the efficacy of this alternate approach of using multi-task offline data, in Table 7, we performed an experiment on the Meta-World domain that first utilizes the data from all the tasks to learn a shared representation using the best method, ACL and then runs standard offline multi-task RL on top of this representation. We denote the method as Offline Pretraining. We include the average task success rates of all tasks in the table below. While the representation learning approach improves over standard multi-task RL without representation learning (No Sharing) consistent with the findings in , we still find that CDS with no representation learning outperforms this representation learning approach by a large margin on multi-task performance, which suggests that conservative data sharing is more important than pure pretrained representation from multi-task datasets in the offline multi-task setting. We finally remark that in principle, we could also utilize representation learning approaches in conjunction with data sharing strategies and systematically characterizing this class of hybrid approaches is a topic of future work.