Rethinking Goal-conditioned Supervised Learning and Its Connection to Offline RL
Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, Chongjie Zhang
Introduction
Reinforcement learning (RL) enables automatic skill learning and has achieved great success in various tasks (Mnih et al., 2015; Lillicrap et al., 2016; Haarnoja et al., 2018; Vinyals et al., 2019). Recently, goal-conditioned RL is gaining attention from the community, as it encourages agents to reach multiple goals and learn general policies (Schaul et al., 2015). Previous advanced works for goal-conditioned RL consist of a variety of approaches based on hindsight experience replay (Andrychowicz et al., 2017; Fang et al., 2019), exploration (Florensa et al., 2018; Ren et al., 2019; Pitis et al., 2020), and imitation learning (Sun et al., 2019; Ding et al., 2019; Sun et al., 2020; Ghosh et al., 2021). However, these goal-conditioned methods need intense online interaction with the environment, which could be costly and dangerous for real-world applications.
A direct solution to learn goal-conditioned policy from offline data is through imitation learning. For instance, goal-conditioned supervised learning (GCSL) (Ghosh et al., 2021) iteratively relabels collected trajectories and imitates them directly. It is substantially simple and stable, as it does not require expert demonstrations or value estimation. Theoretically, GCSL guarantees to optimize over a lower bound of the objective for goal reaching problem. Although hindsight relabeling (Andrychowicz et al., 2017) with future reached states can be optimal under certain conditions (Eysenbach et al., 2020), it would generate non-optimal experiences in more general offline goal-conditioned RL setting, as discussed in Appendix B.1. As a result, GCSL suffers from the same issue as other behavior cloning methods by assigning uniform weights for all experiences and results in suboptimal policies.
To this end, leveraging the simplicity and stability of GCSL, we generalize it to the offline goal-conditioned RL setting and propose an effective and theoretically grounded method, named Weighted GCSL (WGCSL). We first revisit the theoretical foundation of GCSL by additionally considering discounted rewards, and this allows us to obtain a weighted supervised learning objective. To learn a better policy from the offline dataset and promote learning efficiency, we introduce a more general weighting scheme by considering the importance of different relabeling goals and the expected return estimated with a value function. Theoretically, the discounted weight for relabeling goals contributes to optimizing a tighter lower bound than GCSL. Based on the discounted weight, we show that additionally re-weighting with an exponential function over advantage value guarantees monotonic policy improvement for goal-conditioned RL. Moreover, the introduced weights build a natural connection between WGCSL and offline RL, making WGCSL available for both online and offline setting. Another major challenge in goal-conditioned RL is the multi-modality problem (Lynch et al., 2020), i.e., there are generally many valid trajectories from a state to a goal, which can present multiple counteracting action labels and even impede learning. To tackle the challenge, we further introduce the best-advantage weight under the general weighting scheme, which contributes to the asymptotic performance when the data is multi-modal. Although WGCSL has the cost of learning the value function, we empirically show that it is worth for its remarkable improvement. For the evaluation of offline goal-conditioned RL algorithms, we provide a public benchmark and offline datasets including a set of challenging multi-goal manipulation tasks with a robotics arm or an anthropomorphic hand. ExperimentsCode and offline dataset are available at https://github.com/YangRui2015/AWGCSL conducted in the introduced benchmark show that WGCSL significantly outperforms other state-of-the-art baselines in the fully offline goal-conditioned settings, especially in the difficult anthropomorphic hand task and when learning from randomly collected datasets with sparse rewards.
Preliminaries
Goal-conditioned RL
Goal-conditioned Supervised Learning
Weighted Goal-Conditioned Supervised Learning
In this section, we will revisit GCSL and introduce the general weighting scheme which generalizes GCSL to offline goal-conditioned RL.
As an imitation learning method, GCSL sticks the agent’s policy to the relabeled data distribution, therefore it naturally alleviates the problem of out-of-distribution actions. Besides, it has the potential to reach any goal in the offline dataset with hindsight relabeling and the generalization ability of neural networks. Despite its advantages, GCSL has a major disadvantage for offline goal-conditioned RL, i.e., it only considers the last step reward and generally results in suboptimal policies.
We provide training results of GCSL in the PointReach task in Figure 2. The objective of the task is to move a point from the starting position to the desired goal as quickly as possible. The offline dataset is collected using a random policy. As shown in Figure 2, GCSL learns a suboptimal policy which detours to reach goals. This is because GCSL only considers the last step reward . As long as the trajectories reaching the goal at the end of the episode, they are all considered equally to GCSL. To improve the learned policy, a straightforward way is to evaluate the importance of samples or trajectories for policy learning using importance weights. As a comparison, Weighted GCSL (WGCSL), which uses a novel weighting scheme, learns the optimal policy in the PointReach task. We will derive the formulation of WGCSL in the following analysis.
Connection between Goal-Conditioned RL and SL
GCSL has been proved to optimize a lower bound on the goal-reaching objective (Ghosh et al., 2021), which is different from the goal-conditioned RL objective that optimizes the cumulative return. In this section, we will introduce a surrogate function for goal-conditioned RL and then derive the connection between goal-conditioned RL and weighted goal-conditioned supervised learning (WGCSL). For overall consistency, we provide the formulation of WGCSL as below:
Assume a finite-horizon discrete MDP, a stochastic discrete policy which selects actions with non-zero probability and a sparse reward function , where is the state-to-goal mapping and is an indicator function. Given trajectories and discount factor , let the weight , then the following bounds hold:
We defer the proof to Appendix B.2, where we also show that under mild conditions, (1) is a lower bound of , and (2) shares the same gradient direction with at . Theorem 1 reveals the connection between the goal-conditioned RL objective and the WGCSL/GCSL objective. Meanwhile, it suggests that GCSL with the discount weight is a tighter lower bound compared to the unweighted version. We name this weight as discounted relabeling weight (DRW) as it intuitively assigns smaller weights on longer trajectories reaching the same relabeled goal.
2 A More General Weighting Scheme
DRW can be viewed as a weight evaluating the importance of relabeled goals. In this subsection, we consider a more general weighting scheme using a weight function measuring the importance of state-action-goal combination, revealed by the following corollary.
Suppose function over the state-action-goal combination. Given trajectory , let , then the following bound holds:
The proof are provided in Appendix B.3. Naturally, the Q-value in RL with some constant shift is a candidate of the function family . Inspired by prior offline RL approaches (Wang et al., 2018; Peng et al., 2019; Nair et al., 2020), we choose , an exponential function over the goal-conditioned advantage, which we refer to as goal-conditioned exponential advantage weight (GEAW). The constant is used to ensure that . The intuition is that GEAW assigns larger weight for samples with higher values. In addition, the exponential advantage weighted formulation has been demonstrated to be a closed-form solution of an offline RL problem, where the learned policy is constrained to stay close to the behavior policy (Wang et al., 2018). Compared to DRW, GEAW evaluates the importance of state-action-goal samples, taking advantage of the universal value function (Schaul et al., 2015).
To tackle the multi-modality challenge in goal-conditioned RL (Lynch et al., 2020), we further introduce the best-advantage weight (BAW) based on the learned value function. BAW has the following form:
where is a threshold, and is a small positive value. Implementation details of BAW can be found in Section 4. In our implementation, gradually increases. Therefore, BAW gradually leads the policy toward the modal with the highest return rather than sticks to a position between multiple modals, especially when fitting using a Gaussian policy. Combining all three introduced weights, we have the weighting form: . The constant only introduces a coefficient independent of for all transitions, therefore, we simply omit it for further analysis.
3 Policy Improvement via WGCSL
We formally show that combining the three introduced weights, WGCSL can consistently improve the policy learned from the offline dataset. First, we assume there exists a policy that can produce the relabeled experiences . GCSL is essentially equivalent to imitating :
The proof can be found in Appendix Proposition. Proposition 1 implies that WGCSL can generate a uniformly non-worse policy using the relabeled data . To be rigorous, we also need to discuss the relationship between the relabeled return and the original return. In fact, there is a monotonic improvement for the relabeled return over the original one when relabeling with inverse RL (Eysenbach et al., 2020). Below, we provide a simpler relabeling strategy which also offers monotonic value improvement.
Assume the state space can be perfectly mapped to the goal space, , and the dataset has sufficient coverage. We define a relabeling strategy for as
We also assume that when , the strategy only accepts relabeled trajectories with higher future returns than any trajectory with goal in the original dataset . Then, is uniformly as good as or better than , i.e., , where is the behavior policy forming the dataset .
The proof is provided in Appendix Proposition. In Proposition 2, the relabeling strategy is slightly different from the strategy of relabeling with random future states (Andrychowicz et al., 2017). When there is no successful future states, the relabeling strategy in Proposition 2 randomly samples from the future visited states for relabeling, otherwise it keeps the original goal. For random datasets, there are rare successful states and the relabeling strategy in Proposition 2 acts similarly to relabeling with random future states. Combining Propositions 1 and 2, we know that under certain assumptions, after performing relabeling and weighted supervised learning, the policy can be monotonically improved. In fully offline settings, WGCSL is able to learn a better policy than the policy learned by GCSL and the behavior policy generating the offline dataset.
Algorithm
In this section, we summarize our proposed algorithm. Denote the relabeled dataset as , and we maximize the following WGCSL objective based on the relabeled data
where the weight is composed of three parts conforming to the general weighting scheme
(1) is the discounted relabeling weight (DRW) as introduced in Section 3.1.
(3) is the best-advantage weight (BAW) with the form in Eq. 2. In our experiments, is set as percentile of recent advantage values and is set as 0.05. Motivated by curriculum learning (Bengio et al., 2009), gradually increases from to , i.e., we learn from all the experiences in the beginning stage and converge to learning from samples with the top 20% advantage values.
The advantage value in Eq. 3 can be estimated by , and we learn a Q-value function by minimizing the TD error
where we relabel original goals as with a probability of and remain the original goal unchanged for the rest 20% time. For simplicity, in our experiments, once it falls into the probability of goal relabeling, we use purely relabeling with random future states, no matter whether there exist successful future states in the trajectory or not, to avoid parsing all the future states as discussed in Appendix B.5. refers to the target network which is slowly updated to stabilize training. The same value function training method is adopted in (Andrychowicz et al., 2017; Fang et al., 2019). The value function is estimated using Q-value function as . The entire algorithm is provided in Appendix C.
Experiments
In our experiments, we demonstrate our proposed method on a diverse collection of continuous sparse-reward goal-conditioned tasks.
As shown in Figure 3, the introduced benchmark includes two point-based environments, and eight simulated robot environments. In all tasks, the rewards are sparse and binary: the agent receives a reward of if it achieves a desired goal and a reward of otherwise. More details about those environments are provided in Appendix F. For the evaluation of offline goal-conditioned algorithms, we collect two types of offline dataset, namely ’random’ and ’expert’, following prior offline RL work (Fu et al., 2020). Each dataset has the sample size of for hard tasks (FetchPush, FetchSlide, FetchPick and HandReach) and for other relatively easy tasks. The ’random’ dataset is collected by a uniform random policy, while the ’expert’ dataset is collected by the final policy trained using online HER. We also add Gaussian noise with zero mean and standard deviation to increase the diversity of the ’expert’ dataset. This may lead to the consequence that the deterministic policy learned by behavior cloning achieves a higher average return than that in the dataset, which has also been observed in previous works (Kumar et al., 2019; Fujimoto et al., 2019). More details about the two datasets are provided in Appendix D. Beyond the offline setting, we also evaluate our method in the online setting and the results are reported in Appendix E.8.
As for the implementation of GCSL and WGCSL, we utilize the Diagonal Gaussian policy with a mean vector and a constant variance for continuous control. In the offline setting, does not need to be defined explicitly because we evaluate the policy without randomness. The loss function of GCSL can be rewritten as:
Our implemented WGCSL learns a policy to minimize the following loss
where is defined in Eq 4. The policy networks of GCSL and WGCSL are 3-layer MLP with 256 units each layer and relu activation. Besides, WGCSL learns a value network with the same structure except for the input and output layers. The batch size is set as 128 for the first 6 tasks, and 512 for 4 harder tasks from FetchPush to HandReach. We use Adam optimizer with a learning rate of . For all the experiments we repeat for 5 different random seeds and report the average evaluation performance with standard deviation.
2 Experimental Results
In the offline goal-conditioned setting, we compare WGCSL with GCSL, HER and modified versions of most related offline works such as goal-conditioned MARWIL (Wang et al., 2018) (Goal MARWIL), goal-conditioned behavior cloning (Bain & Sammut, 1995) (Goal BC), goal-conditioned Batch-Constrained Q-Learning (Fujimoto et al., 2019) (Goal BCQ), and goal-conditioned Conservative Q-learning (Kumar et al., 2020) (Goal CQL). The results are presented in Table 1 and Figure 4. Additional comparison results with more baselines and results on 10 dataset are provided in Appendix E.
The top row of Figure 4 reports the performance of different algorithms using the expert datasets of four hard tasks. In all these tasks, WGCSL outperforms other baselines in terms of both learning efficiency and final performance. GCSL converges very slowly but can finally reach beyond the average return of expert dataset in FetchSlide and FetchPick tasks. Goal MARWIL is more efficient than GCSL and achieves the second best performance. However, none of the baseline algorithms succeeds in the HandReach task, which might be due to high-dimensional states, goals, actions and the noise used to collect the dataset. In addition, HER itself has the ability to handle offline dataset as the relabeled data is far off-policy (Plappert et al., 2018). But we can conclude from Table 1 that the performance of HER is unstable and inconsistent across different tasks. In contrast, learning curves of WGCSL are consistently stable. It is also verified that techniques used in WGCSL are effective to generate better policies than original and relabeled datasets in offline goal-conditioned RL setting.
Random Datasets
The bottom row of Figure 4 shows the training results of different methods on random datasets. We can observe that WGCSL outperforms other baselines by a large margin. Baselines such as Goal MARWIL, Goal BC, Goal BCQ and Goal CQL can hardly learn to reach goals with the random datasets. GCSL and HER can learn relatively good policies in some tasks, which proves the importance of goal relabeling. Interestingly, in tasks such as Reacher, SawyerReach and FetchReach, results of WGCSL and HER on the random dataset exceeds that of the expert dataset. The possible reason is that the data coverage of random dataset could be larger than the expert dataset in these tasks, and the agent can learn to reach more goals via goal relabeling.
Ablation Studies
We also include ablation studies for offline experiments in Figure 5. In our ablation studies, we investigate the effect of the three proposed weights, i.e., DRW, GEAW and BAW, by applying them to GCSL respectively. The results demonstrate that all three weights are effective compared to the plain GCSL, while GEAW and BAW show larger benefits. This is reasonable since GEAW and BAW learn additional value networks while DRW does not. Moreover, the results suggest that the learned policy can be improved by combining these weights.
Value Estimation
Off-policy RL algorithms are prone to distributional shift and suffer from value overestimation problem on out-of-distribution actions. Such problem can be exacerbated in tasks with high-dimensional states and actions. As shown in Figure 6, DDPG and HER exhibit large estimated values during offline training. In contrast, WGCSL has a more robust value approximation through weighted supervised learning, thus achieving higher returns in the difficult HandReach task.
Related Work
Solving goal-conditioned RL task is an important yet challenging domain in reinforcement learning. In those tasks, agents are required to achieve multiple goals rather than a single task (Schaul et al., 2015). Such setting puts much burden on the generalization ability of the learned policy when dealing with unseen goals. Another challenge arises in goal-conditioned RL is the sparse reward (Plappert et al., 2018). HER (Andrychowicz et al., 2017) tackles the sparse reward issue via relabeling the failed rollouts as successful ones to increase the proportion of successful trails during learning. HER is then extended to deal with demonstration (Nair et al., 2018; Ding et al., 2019) and dynamic goals (Fang et al., 2018), and combined with curriculum learning for exploration-exploitation trade-off (Fang et al., 2019). Eysenbach et al. (2020) and Li et al. (2020) further improved the goal sampling strategy from the perspective of inverse RL. Zhao et al. (2021) encouraged agents to control their states to reach goals by optimizing a mutual information objective without external rewards. Different from prior works, we consider the offline setting, where the agent cannot access the environment to collect data in the training phase.
Our work has strong connections to imitation learning (IL). Behavior cloning is one of the simplest IL methods that learns a policy predicting the expert actions (Bain & Sammut, 1995; Bojarski et al., 2016). Inverse RL (Ng et al., 2000) extracts a reward function from demonstrations and then uses the reward function to train policies. Several works also leverage self-imitation to improve the stability and efficiency of RL without expert experience. Self-imitation learning (Oh et al., 2018) imitates the agent’s past advantageous decisions and indirectly drives exploration. Sun et al. (2019) introduced a dynamic programming method to learn policies progressively via self-imitation. Lynch et al. (2020) employed VAE to jointly learn a latent plan representation and goal-conditioned policy from replay without the need for rewards. Our work is most relevant to GCSL (Ghosh et al., 2021), which iteratively relabels and imitates its own collected experiences. The difference is that we further generalize GCSL to the offline goal-conditioned RL setting via weighted supervised learning.
Our work belongs to the scope of offline RL. Learning policies from static offline datasets is challenging for general off-policy RL algorithms due to error accumulation with distributional shift (Fujimoto et al., 2019; Kumar et al., 2019). To address such issue, a series of algorithms are proposed to constrain the policy updates to avoid excessive deviation from the data distribution (Wang et al., 2018; Fujimoto et al., 2019; Wu et al., 2019; Yang et al., 2021). MARWIL (Wang et al., 2018) can learn a better policy than the behavior policy of the dataset with theoretical guarantees. In addition to policy regularization, other works leverage techniques such as value underestimation (Kumar et al., 2020; Yu et al., 2020), robust value estimation (Agarwal et al., 2020), data augmentation (Sinha & Garg, 2021; Wang et al., 2021), and Expectile -Learning (Ma et al., 2021) to handle distribution shift. As offline RL can easily overfit to one task, generalizing offline RL to unseen tasks can be even more challenging (Li et al., 2021). Li et al. (2019) tackled the challenge by leveraging the triplet loss for robust task inference. Our method has similar exponential advantage weighted form as MARWIL, however, we tackle goal-conditioned tasks and introduces a general weighting scheme for WGCSL. A most recent work (Chebotar et al., 2021) also addresses the offline goal-conditioned problem through underestimating the value of unseen actions. The difference is that WGCSL is a weighted supervised learning method without explicitly value underestimation. Therefore, WGCSL is substantially simpler, more stable, and more applicable for real-world tasks.
Conclusion
In this paper, we have proposed a novel algorithm, Weighted Goal-Conditioned Supervised Learning (WGCSL), to solve offline goal-conditioned RL problems with sparse rewards, by first revisiting GCSL and then generalizing it to offline goal-conditioned RL. We further derive theoretical guarantees to show that WGCSL can generate monotonically improved policies under certain conditions. Experiments on an introduced benchmark demonstrate that WGCSL outperforms current offline goal-conditioned approaches by a great margin in terms of learning efficiency and convergent performance. For future directions, we are interested in building more stable and efficient offline goal-conditioned RL algorithms based on WGCSL for real-world applications.
Acknowledgments
This work was partly supported by the Science and Technology Innovation 2030-Key Project under Grant 2021ZD0201404, in part by Science and Technology Innovation 2030 – “New Generation Artificial Intelligence” Major Project (No. 2018AAA0100904) and National Natural Science Foundation of China (No. 20211300509). Part of this work was done when Rui Yang and Hao Sun worked as interns at Tencent Robotics X Lab.
Ethics and Reproducibility Statements
Our work provides a way to learn general policy from offline data. Learning policy for general purpose is one of the fundamental challenges in artificial general intelligence (AGI). The proposed offline goal-conditioned algorithm can provide new insights into training agents with diverse skills, but currently we cannot foresee any negative social impact of our work. The algorithm is evaluated on simulated robot environments, thus our experiments would not suffer from discrimination/bias/fairness concerns. The details of experimental settings and additional results are also provided in the Appendix. All the implementation code and offline dataset are available in our code link: https://github.com/YangRui2015/AWGCSL.
References
Appendix A A Comprehensive Summary for WGCSL
WGCSL is strongly efficient in the offline goal-conditioned RL due to the following advantages:
WGCSL learns a universal value function (Schaul et al., 2015) which is more generalizable than the value function learned in a single task, thus alleviating the overfitting problem in offline RL.
WGCSL utilizes goal relabeling which contributes to augment huge amount of data ( times) for offline training and meanwhile alleviates the sparse reward issue.
WGCSL leverages weighted supervised learning to reweight the action distribution in the dataset, avoiding the out-of-distribution action problem in offline RL.
In the three weights, GEAW keeps the learned policy close to the offline dataset implicitly and BAW handles the multi-modality problem in offline goal-conditioned RL. WGCSL utilizes the three weights to learn a better policy with theoretical guarantees.
Appendix B Theoretical Results
In our analysis, we will assume finite-horizon MDPs with discrete state and goal spaces.
Given a trajectory , we show that under general offline goal-conditioned RL setting with the sparse reward function , relabeling with future states is not optimal. This is a direct result from Eq. 7 of (Eysenbach et al., 2020), which has proven that the optimal relabelling strategy is
This is different from simply relabeling with the last state as in Section 4.1 of (Eysenbach et al., 2020) and relabeling with future states as .
B.2 Proof of Theorem 1
Assume a finite-horizon discrete MDP, a stochastic discrete policy which selects actions with non-zero probability and a sparse reward function , where is the state-to-goal mapping and is an indicator function. Given trajectory and discount factor , let the weight , then the following bounds hold:
Eq. 5 holds because of the non-positive logarithmic function over . Compared to the surrogate function used in Appendix B.1 of (Ghosh et al., 2021), further considers the discounted cumulative return. The theorem shows that WGCSL optimizes over a tighter lower bound of compared to GCSL.
While Theorem 1 bounds the surrogate function with and , here we bridge the gap between and the goal-conditioned RL objective .
We first show that is an upper bound of in addition to KL divergence and a constant. Let and , and assume . Since practically is sampled from offline dataset, we can find . Similarly to (Kober & Peters, 2014), we have following results using Jensen’s inequality and absorbing constants into ’constant’:
Naturally, in offline setting, we explicitly or implicitly bound the KL divergence between and via policy constraints or (weighted) supervised learning. Therefore, we can regard as a lower bound of .
Furthermore, We show that shares the same gradient with when after scaling with .
Therefore, taking a gradient update with a proper step size on can also improve the policy on the objective . In our algorithm, the exponential weight poses an implicit constraint on the policy update similar to (Wang et al., 2018; Nair et al., 2020). With a proper constraint, optimizing on the is equivalent as optimizing on the goal-conditioned objective . Intuitively, maximizes the likelihood for state-action pairs with large returns, and therefore it can lead to policy improvement based on the offline dataset.
B.3 Proof of Corollary 1
Suppose function over the state-action-goal combination. Given trajectory , let , then the following bounds hold:
Similarly to the proof of Theorem 1, we have
B.4 Proof of Proposition 1
We assume the advantage value is for the behavior policy of the relabeled data, i.e., . We consider the compound state from the relabeled distribution . When is fixed (the index of are fixed), is a constant and we have:
The above proposition assumes the advantage value is for . However, we empirically find that using advantage value of the learned policy (the value function is learned to minimize in Section 4) works better in offline goal-conditioned RL setting. As analyzed in prior works (Nair et al., 2020; Yang et al., 2021), using the value function of is equivalent to solving the following problem:
In addition, for offline goal-conditioned RL, learning policies with inter-trajectory information is important for reaching diverse goals. Therefore, using to compute TD target helps value estimation across trajectories compared with . Moreover, is trained via weighted supervised learning to avoid the OOD action problem.
B.5 Proof of Proposition 2
Assume the state space can be perfectly mapped to the goal space, , and the dataset has sufficient coverage. We define a relabeling strategy for as
We also assume that when , the strategy only accepts relabeled trajectories with higher future returns than any trajectory with goal in the original dataset . Then, is uniformly as good as or better than , i.e., , where is the behavior policy forming the dataset .
Let . We assume that the data of and are sufficient and cover the full state and goal space, i.e., . Note that every trajectory has only a single goal, and it does not have to reach the goal within the trajectory. Therefore, we can partition into two subsets (positive set) and (negative set):
where contains trajectories that can reach the goal and contains trajectories that cannot reach the goal. Let , and similarly we define the partition
Hence, the value can be written as
Given , we relabel the future trajectory as for value estimation following the strategy
and when , the strategy only accepts with higher return than any trajectory in . Let be the relabeled trajectories with as goal. Hence, the value can be written as
For any , we enumerate the situation of its counterpart without relabeling.
. Following the relabeling strategy, we have , such that .
The relabeling strategy of Proposition 2 requires a complete scan of future states of trajectories, which is costly. Empirically we find that relabeling with random future states as HER has comparable performance, and is very simple and efficient to implement. We provide the comparison between the relabeling strategy of Proposition 2 (called WGCSL-slow-relabel) and the relabeling strategy that we used in practice in Table 2.
Appendix C Algorithm
We summarize our proposed approach WGCSL into Algorithm 1.
Appendix D Offline Settings
In the offline setting, we only use collected datasets to train agents, without additional interaction with the environment. The difference between the offline goal-conditioned RL dataset and the general offline RL dataset is that it also needs to save the original desired goals and the achieved goals . We follow previous works (Fujimoto et al., 2019; Fu et al., 2020) to introduce two different types of datasets:
Random dataset: using a uniformly random policy to collect trajectories.
Expert dataset: using the final online policy trained by HER with additional Gaussian noise (zero mean and standard deviation) to collect data.
For the dataset of six easy tasks (PointReach, PointRooms, Reacher, SawyerReach, SawyerDoor, FetchReach), we save a dict of transitions (2000 trajectories) containing observations, actions, desired goals and achieved goals. Regarding the other four hard tasks (FetchPush, FetchSlide, FetchPick, HandReach), we store transitions (40000 trajectories) in the offline dataset. The rewards can be computed by leveraging the corresponding interface (namely env.compute_reward) of the environment. The average returns of the two types of datasets are shown in Table 3. Generally, average returns of the expert datasets are close to that of the optimal policy, while average returns of the random datasets are close to 0.
Baseline Introduction
In the offline goal-conditioned setting, we denote the offline dataset as and consider the following baselines:
GCSL: GCSL (Ghosh et al., 2021) without online interaction, only using offline samples after relabeling (denoted as ) to maximize
Goal MARWIL: goal-conditioned MARWIL (Wang et al., 2018), using offline samples to maximize
To fairly compare with our method, Goal MARWIL is implemented with actor-critic style as AWAC (Nair et al., 2020) and WGCSL, therefore it can also be treated as ’Goal AWAC’. Besides, we also clip the exponential advantage value as WGCSL for numerical stability.
Goal AWR: goal-conditioned AWR (Peng et al., 2019), using offline samples to maximize
where . Different from the original AWR that uses TD() or Monte Carlo return to train the value function, we use TD loss to learn value function as WGCSL. Empirically, we find TD target works better than Monte Carlo in our setting.
Goal BC: goal-conditioned behavior cloning (Bain & Sammut, 1995), using offline samples to maximize
Goal BCQ: goal-conditioned Batch-Constrained Q-Learning (Fujimoto et al., 2019), which we use the official codebase and concatenate observations, desired goals and achieved goals as states. The objective function of Goal BCQ is as follows:
where is the goal-conditioned perturbation model, which outputs an adjustment to an action within the range . is the VAE fitted on the behavior policy of the offline dataset. The policy is optimized through optimizing . The function is learned using the Clipped Double Q-learning in (Fujimoto et al., 2019), and the only difference is that the function also includes goal as input.
Goal CQL: goal-conditioned Conservative Q-Learning (Kumar et al., 2020), which we also use the official codebase and concatenate observations, desired goals and achieved goals as states. The objective function of Goal CQL is as follows:
which jointly optimizes estimated return and policy entropy. The Q function of CQL is learned by minimizing the following loss:
where is the behavior policy for the offline dataset and is the Bellman operator.
HER: Hindsight Experience Replay (Andrychowicz et al., 2017), using offline samples after relabeling to maximize
where and the function is learned by minimizing TD error:
Implementation of Baselines
For fair comparison, all the implementation of WGCSL, GCSL, Goal MARWIL, Goal AWR, Goal BC, and HER use the Diagnal Gaussian policy with a mean vector and a non-zero constant variance , where is set as 0.2. WGCSL, Goal MARWIL, Goal AWR, and HER share the same network structure, i.e., the value function and the policy network (along with their target networks) are both 3-layer MLPs with 256 units each layer and relu activation. Besides, GCSL and Goal BC both keep a single policy network without target network and value function. We also normalize the observations and goals with estimated mean and standard deviation for WGCSL and other baselines, which is helpful for offline multi-goal manipulation tasks. The parameter in Goal MARWIL and Goal AWR are set as by default. As regards to Goal BCQ and Goal CQL, we use the same batch size as other algorithms and other parameters are set as default (we find that default parameters perform equally well, if not better than others, in the grid search of critical parameters). All other hyper-parameters of different algorithms are kept the same for each task, such as batch size (128 for the first 6 tasks and 512 for 4 harder taks), discount factor , and Adam optimizer with learning rate .
Implementation of WGCSL
The basic implementation is introduced as above. To compute the best-advantage weight for WGCSL, we keep a First-In-First-Out queue of size to store recent calculated advantage values, and the percentile threshold gradually increases from to . For each training step, increases by 0.01 (for 4 harder tasks from FetchPush to HandReach) or 0.15 (for other tasks). The maximum clipping upper bound for WGCSL is set as 10. Though DRW brings improvements for most of the tasks on expert and random datasets, we find DRW slightly affects the performance of harder tasks on random datasets, which may be because in the random datasets we need more timesteps to obtain enough coverage of achieved goals, especially for harder tasks. Tuning for DRW trade-offs between the value and the effective coverage of relabeling goals, and excluding DRW is equivalent to setting for DRW. Therefore, we only use GEAW and BAW for these tasks on random datasets.
Evaluation Settings
For evaluation, we evaluate all the algorithms for 100 episodes without randomness. We repeat for 5 different random seeds and report the average evaluation performance with standard deviation.
Appendix E Additional Experimental Results
In this section, we provide more empirical results in both offline and online goal-conditioned RL setting, including average return, final distance of the offline setting, results on small () dataset, and the results of online experiments.
We provide the average return curves of all tasks in Figure 7. In these results, WGCSL performs consistently well across ten environments, especially when handling hard tasks (FetchPush, FetchSlide, FetchPick, HandReach) and learning from the random offline dataset. Other conclusions have been discussed in Section 5.2. Comparing AWR with MARWIL, the major difference is using Monte Carlo return to replace the TD estimation for learning the policy. Hence, we can find that TD bootstrap learns more efficiently in weighted supervised learning, which is similar to the conclusion of prior work (Nair et al., 2020). In addition, we present all final performance of different algorithms in Table 4.
E.2 Comparison Results with Actionable Models, HER and DDPG
In this subsection, we make comparison with Actionable Models (AM) (Chebotar et al., 2021), which applies the conservative estimation idea similar to (Kumar et al., 2020) into offline goal-conditioned RL. There are some different settings between AM and ours. For example, we consider cumulative return of the maximum horizon, i.e., the agent doesn’t terminate the interaction after reaching the goal. To fairly compare with AM, we implement AM based on HER. Specifically, we revise the value update as:
In the formula, AM tries to minimize TD error and minimize unseen Q values at the same time. The action is sampled according to exponential Q values. In our implementation, we sample 20 actions from the action space and compute their exponential Q value in order to sample .
The results are shown in Figure 8. We can observe that DDPG and HER are unstable during training, with large variance and performance drops at the end. In both expert and random dataset, HER outperforms DDPG, which demonstrates the benefits of goal relabeling for offline goal-conditioned RL. Besides, Actionable Model (AM) achieves higher final performance than HER and DDPG especially in expert dataset, which indicates the conservative estimation technique is useful for offline goal-conditioned RL. Moreover, WGCSL is more efficient than AM, especially in harder tasks with large action space. Intuitively, WGCSL promotes the probability of visited actions, while AM restrains the probability of unvisited actions. For tasks with large action space, conservative strategy can be ineffective compared to WGCSL.
E.3 Normalized Score Across 10 Tasks
For clarity, we report the normalized score across 10 tasks to compare different algorithms in Figure 9. We can observe the significant improvement of WGCSL over other baselines. We take an insight into the impact of the goal relabeling technique. For expert datasets, goal relabeling is not a dominant factor, e.g., the learning curve of Goal BC and GCSL are very close. In random datasets, goal relabeling is essential to achieve a high performance. In addition to goal relabeling, WGCSL contains other techniques to jointly promote the performance in offline goal-conditioned RL tasks. Other conclusions are the same as prior subsections (Appendix E.1 and Appendix E.2).
E.4 Comparison of Training Time
We further analyze the computation cost of the four most effective algorithms in the random offline setting, i.e., WGCSL, GCSL, HER and Actionable Models (AM). The results are shown in Figure 10, from which we can conclude that GCSL needs the minimal training time, while WGCSL and HER require similar cost. It needs to be emphasized that WGCSL has only a small amount of additional training time over GCSL, but it achieves significantly higher performance. Therefore, we believe that the cost of learning the value function on top of GCSL is worthwhile. Moreover, AM requires the largest computation among these algorithms, which is mainly because it needs to sample several actions and compute their exponential Q value for constraining the large Q value. These procedures make AM the most time-consuming algorithm. In contrast, WGCSL can achieve higher performance with less training time. The training time would vary across different platforms, and we use one single GPU (Tesla P100 PCIe 16GB) and 5 cpu cores (Intel Xeon E5-2680 v4 @ 2.40GHz) to run 5 random seeds parallelly.
E.5 Average Probability of Improvement
In this subsection, we adopt the average probability of improvement (Agarwal et al., 2021), a robust metric to measure how likely it is for one algorithm to outperform another on a randomly selected task. The results are reported in Figure 11. As shown in the results, WGCSL robustly outperforms other baselines in both expert and random datasets. For instance, WGCSL is better than 4 baselines on the expert dataset, and 6 baselines on the random dataset. Note that the most two effective and robust algorithms on both random and expert datasets are WGCSL and Actionable Models (AM), which are specifically designed for offline goal-conditioned RL. Comparing the two algorithms, WGCSL outperforms AM with a probability of on the expert dataset and on the random dataset.
E.6 Final Distances in the Offline Setting
GCSL (Ghosh et al., 2021) uses the average final distance to the desired goal as a measure for goal-reaching problems. We also provide average final distances in the offline setting to evaluate the learned goal-conditioned policy in Figure 12. From the results we can conclude that WGCSL substantially outperforms other baselines in both random and expert dataset. Although GCSL learns a suboptimal policy, it also achieves relatively consistent final distance across different environments and offline datasets. Besides, Goal MARWIL and Goal BC are two effective baselines when learning from the expert dataset, but they do not perform well on the random dataset. In addition, HER and DDPG learn unstable results mainly because of bootstrapping from out-of-distribution actions.
E.7 Results on 10%percent1010\% Offline Dataset
To evaluate the impact of dataset size on the performance, we select of samples (i.e., transitions) from the offline dataset of 4 hard tasks (FetchPush, FetchSlide, FetchPick, HandReach). The results are reported in Figure 13 and Figure 14. From the results, we can conclude that with only offline data WGCSL consistently achieves the best performance out of these algorithms. In Figure 14, we can see that WGCSL is very robust to the dataset size on the expert dataset, while the performance is affected more by the small data size on the random dataset. The probable reason is that in the offline setting, the number of samples influences the data coverage greatly and learning generalizable value function from a small set of data is very difficult.
E.8 Online Experiments
We introduce implementation details and hyper-parameters of the online setting in this subsection. As GCSL is originally proposed for online goal-conditioned RL, WGCSL can also be applied to online goal-conditioned setting. In the online training setting, the agent iterates between collecting trajectories from the environment and training with samples from the replay buffer. We conduct experiments on the first six taks, i.e., PointReach (here we call Point2DLargeEnv), PointRooms (here we call Point2D-FourRooms), SawyerReach, FetchReach, Reacher and SawyerDoor. The buffer size for first four tasks is set as , and for the other two tasks. Before training starts, we first randomly sample 20 trajectories to initialize the buffer. The policies of WGCSL and GCSL are both Diagonal Gaussian policy with a mean vector parameterized as and a non-zero constant variance . The standard deviation is set as 0.2. All the policy networks and value networks are the same as that used in the offline setting. When interacting with the environments, we use additional random actions for exploration with a probability of . After data collection, we sample a mini-batch from the replay buffer and perform goal relabeling for WGCSL and GCSL. Then we train the agent for 1 mini-batch (Point2DLarge, Point2D-FourRooms) or 4 mini-batches for other tasks. The batch size, optimizer, learning rate keeps the same as the offline setting. The maximum clipping upper bound for WGCSL is set as 10. We also utilize target nets for WGCSL to stabilize training, and the polyak average coefficient for target nets are 0.9. With respect to the evaluation phase, we use the deterministic policy without randomness. After each epoch, we evaluate the policy for 100 episodes and compute the average success rate. For all the experiments we repeat for 10 different random seeds and report the mean performance with standard deviation. In this subsection, we report the results of WGCSL with the discounted relabeling weight (DRW) and the goal-conditioned exponential advantage weight (GEAW), though the performance can be improved with the best-advantage weight (BAW).
The empirical results regarding the average success rate of the online setting are shown in Figure 15. From Figure 15 we can conclude that WGCSL achieves better success rate and converges faster than other baselines in most of the environments except SawyerReach. Therefore, WGCSL is verified to be highly efficient in both online and offline goal-conditioend RL, and we hope it could be a stepping stone for future real-world RL applications.
E.9 The Impact of Clipping for Exponential Advantage Weights
We study the impact of clipping in WGCSL’s exponential advantage weights by comparing with the following variants:
WGCSL w/o clip: no clipping weights performed.
WGCSL clip 5: clipping the exponential advantage weights to .
WGCSL clip 10: clipping the exponential advantage weights to , which is used in our paper by default.
From the results in Figure 16 we can observe that without clipping the advantage estimation can harm the performance in Reacher and SawyerDoor. Therefore, we perform weight clipping by default.
Appendix F Details of Environments
All of the ten tasks have continuous state space, action space and goal space. The maximum episode horizon is set as 50. In this section, we introduce these tasks in detail.
PointReach is taken from the open-source environment multi-worldhttps://github.com/vitchyr/multiworld and it requires the blue point to reach the green circle. The state space has two dimensions representing Cartesian coordinates of the blue point, and the action space also has two dimensions meaning the horizontal and vertical displacement. The goal space is the same as state space, which means . The bule point and the green circle are randomly initialized in the state space. The allowable error of reaching goals is the radius of the target circle and is set as . The reward function is defined as:
PointRooms
The PointRooms environment is built on multi-world. The state space, the action space, the goal space, and the reward function are the same as PointReach. The difference is that there are four rooms separated by walls.
Reacher
The Reacher environment is revised from OpenAI Gym (Brockman et al., 2016). States are 11-dimensional, which indicate the angles, the positions, and the velocity of the joints. Actions are -dimensional and they control the movement of two joints. Goals are -dimensional vectors representing the expected XY position. And the state-to-goal mapping is , where the last three dimensions are the XYZ position of the end-effect. The reward function is defined as:
where the allowable error is set as .
SawyerReach
The SawyerReach environment is taken from multi-world. The Sawyer robot aims to reach a target position with its end-effector. The observation space is 3-dimensional, representing the 3D Cartesian position of the end-effector. Correspondingly, the goal space is 3-dimensional and describes the expected position, and the state-to-goal mapping is . Besides, the action space has 3 dimensions describing the next position of the end-effector. The reward function is defined as:
where the allowable error .
SawyerDoor
The SawyerDoor environment is revised from multi-world. The Sawyer robot is required to open the door to a desired angle. The state space (4-dimensional) consists of the Cartesian coordinates of the Sawyer end-effector and the door’s angle. As in the SawyerReach task, the action space is -dimensional and controls the position of the end-effector. The desired goals are uniform from 0 to 0.83 radians. And the reward function is defined as:
where the allowable error is set as .
FetchReach
The FetchReach environment is taken from OpenAI Gym (Brockman et al., 2016; Plappert et al., 2018). In this environment, a 7-DoF robotic arm is expected to touch a desired location with its two-finger gripper. The state space is 10-dimensional, including the gripper’s position and linear velocities. The action space is 4-dimensional, which represents the gripper’s movements and its status about the opening and closing. Moreover, the goals are 3-dimensional vectors representing the target places of the gripper. The state-to-goal mapping is , because the first dimensions of the state describe the position of the gripper. The allowable error in FetchReach is . The reward function is defined as:
FetchPush
Similar to FetchReach, the FetchPush environment is taken from OpenAI Gym where a 7-DoF robotic arm is controlled to move a box to the desired location. The state space is 25-dimensional, including the gripper’s position, linear velocities, and the box’s position, rotation, linear and angular velocities. The action space is also 4-dimensional as in FetchReach. Different from FetchReach, the 3-dimensional goals and achieved goals represent the target and current position of the box. The state-to-goal mapping is , because the dimensions of state describe the position of the box. The allowable error in FetchPush is . The reward function is defined as:
FetchSlide
The FetchSlide environment is like FetchPush where a 7-DoF robotic arm is controlled to slide a box to the desired location. The state space, the action space, the state-to-goal mapping, the reward function and the allowable error are all the same as FetchPush. The difference is that the target position is outside of the robot’s reach thus it has to hit the box and makes the box slide and then stop at the desired position.
FetchPick
Like FetchPush and FetchSlide environments, the FetchPick environment controls a 7-DoF robotic arm to pick up a box to the desired location. The state space, the action space, the state-to-goal mapping, the reward function and the allowable error are all the same as FetchPush. The difference is that the target position may be on the table or in the air, thus the agent needs to pick the box with its gripper.
HandReach
In HandReach environment, the agent is required to control a 24-DoF anthropomorphic hand and use its fingers to reach the target place. Observations is 63-dimensional containing two 24-dimensional vectors about positions and velocities of the joints, and additional 15 dimensions indicating current state of fingertips. Actions are 20-dimensional vectors that control the non-coupled joints of the hand. Goals have 15 dimensions for HandReach representing the target Cartesian positions of each fingertip. The reward function is the same as Fetch tasks except the state-to-goal-mapping () and the allowable threshold ().