Skew-Fit: State-Covering Self-Supervised Reinforcement Learning

Vitchyr H. Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, Sergey Levine

Introduction

Reinforcement learning (RL) provides an appealing formalism for automated learning of behavioral skills, but separately learning every potentially useful skill becomes prohibitively time consuming, both in terms of the experience required for the agent and the effort required for the user to design reward functions for each behavior. What if we could instead design an unsupervised RL algorithm that automatically explores the environment and iteratively distills this experience into general-purpose policies that can accomplish new user-specified tasks at test time?

In the absence of any prior knowledge, an effective exploration scheme is one that visits as many states as possible, allowing a policy to autonomously prepare for user-specified tasks that it might see at test time. We can formalize the problem of visiting as many states as possible as one of maximizing the state entropy H(S){\mathcal{H}}({\mathbf{S}}) under the current policy.We consider the distribution over terminal states in a finite horizon task and believe this work can be extended to infinite horizon stationary distributions. Unfortunately, optimizing this objective alone does not result in a policy that can solve new tasks: it only knows how to maximize state entropy. In other words, to develop principled unsupervised RL algorithms that result in useful policies, maximizing H(S){\mathcal{H}}({\mathbf{S}}) is not enough. We need a mechanism that allows us to reuse the resulting policy to achieve new tasks at test-time.

We argue that this can be accomplished by performing goal-directed exploration: a policy should autonomously visit as many states as possible, but after autonomous exploration, a user should be able to reuse this policy by giving it a goal G\mathbf{G} that corresponds to a state that it must reach. While not all test-time tasks can be expressed as reaching a goal state, a wide range of tasks can be represented in this way. Mathematically, the goal-conditioned policy should minimize the conditional entropy over the states given a goal, H(S∣G){\mathcal{H}}({\mathbf{S}}\mid\mathbf{G}), so that there is little uncertainty over its state given a commanded goal. This objective provides us with a principled way to train a policy to explore all states (maximize H(S){\mathcal{H}}({\mathbf{S}})) such that the state that is reached can be determined by commanding goals (minimize H(S∣G){\mathcal{H}}({\mathbf{S}}\mid\mathbf{G})).

Directly optimizing this objective is in general intractable, since it requires optimizing the entropy of the marginal state distribution, H(S){\mathcal{H}}({\mathbf{S}}). However, we can sidestep this issue by noting that the objective is the mutual information between the state and the goal, I(S;G)I({\mathbf{S}};\mathbf{G}), which can be written as:

Equation 1 thus gives an equivalent objective for an unsupervised RL algorithm: the agent should set diverse goals, maximizing H(G){\mathcal{H}}(\mathbf{G}), and learn how to reach them, minimizing H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}).

While learning to reach goals is the typical objective studied in goal-conditioned RL (Kaelbling, 1993; Andrychowicz et al., 2017), setting goals that have maximum diversity is crucial for effectively learning to reach all possible states. Acquiring such a maximum-entropy goal distribution is challenging in environments with complex, high-dimensional state spaces, where even knowing which states are valid presents a major challenge. For example, in image-based domains, a uniform goal distribution requires sampling uniformly from the set of realistic images, which in general is unknown a priori.

Our paper makes the following contributions. First, we propose a principled objective for unsupervised RL, based on Equation 1. While a number of prior works ignore the H(G){\mathcal{H}}(\mathbf{G}) term, we argue that jointly optimizing the entire quantity is needed to develop effective exploration. Second, we present a general algorithm called Skew-Fit and prove that under regularity conditions Skew-Fit learns a sequence of generative models that converges to a uniform distribution over the goal space, even when the set of valid states is unknown (e.g., as in the case of images). Third, we describe a concrete implementation of Skew-Fit and empirically demonstrate that this method achieves state of the art results compared to a large number of prior methods for goal reaching with visually indicated goals, including a real-world manipulation task, which requires a robot to learn to open a door from scratch in about five hours, directly from images, and without any manually-designed reward function.

Problem Formulation

To ensure that an unsupervised reinforcement learning agent learns to reach all possible states in a controllable way, we maximize the mutual information between the state S{\mathbf{S}} and the goal G\mathbf{G}, I(S;G)I({\mathbf{S}};\mathbf{G}), as stated in Equation 1. This section discusses how to optimize Equation 1 by splitting the optimization into two parts: minimizing H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}) and maximizing H(G){\mathcal{H}}(\mathbf{G}).

Standard RL considers a Markov decision process (MDP), which has a state space S{\mathcal{S}}, action space A{\mathcal{A}}, and unknown dynamics p(st+1∣st,at):S×S×A↦[0,+∞)p(\mathbf{s}_{t+1}\mid\mathbf{s}_{t},\mathbf{a}_{t}):{\mathcal{S}}\times{\mathcal{S}}\times{\mathcal{A}}\mapsto[0,+\infty). Goal-conditioned RL also includes a goal space G{\mathcal{G}}. For simplicity, we will assume in our derivation that the goal space matches the state space, such that G=S{\mathcal{G}}={\mathcal{S}}, though the approach extends trivially to the case where G{\mathcal{G}} is a hand-specified subset of S{\mathcal{S}}, such as the global XY position of a robot. A goal-conditioned policy π(a∣s,g)\pi(\mathbf{a}\mid\mathbf{s},\mathbf{g}) maps a state s∈S\mathbf{s}\in{\mathcal{S}} and goal g∈S\mathbf{g}\in{\mathcal{S}} to a distribution over actions a∈A\mathbf{a}\in{\mathcal{A}}, and its objective is to reach the goal, i.e., to make the current state equal to the goal.

Goal-reaching can be formulated as minimizing H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}), and many practical goal-reaching algorithms (Kaelbling, 1993; Lillicrap et al., 2016; Schaul et al., 2015; Andrychowicz et al., 2017; Nair et al., 2018; Pong et al., 2018; Florensa et al., 2018a) can be viewed as approximations to this objective by observing that the optimal goal-conditioned policy will deterministically reach the goal, resulting in a conditional entropy of zero: H(G∣S)=0{\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}})=0. See Appendix E for more details. Our method may thus be used in conjunction with any of these prior goal-conditioned RL methods in order to jointly minimize H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}) and maximize H(G){\mathcal{H}}(\mathbf{G}).

2 Maximizing ℋ​(𝐆)ℋ𝐆{\mathcal{H}}(\mathbf{G}): Setting Diverse Goals

Skew-Fit: Learning a Maximum Entropy Goal Distribution

Our method, Skew-Fit, learns a maximum entropy goal distribution qϕG{q^{G}_{{\phi}}} using samples collected from a goal-conditioned policy. We analyze the algorithm and show that Skew-Fit maximizes the goal distribution entropy, and present a practical instantiation for unsupervised deep RL.

To learn a uniform distribution over valid goal states, we present a method that iteratively increases the entropy of a generative model qϕG{q^{G}_{{\phi}}}. In particular, given a generative model qϕtG{q^{G}_{{\phi}_{t}}} at iteration tt, we want to train a new generative model, qϕt+1G{q^{G}_{{\phi}_{t+1}}} that has higher entropy. While we do not know the set of valid states S{\mathcal{S}}, we could sample states sn∼iidpϕtS\mathbf{s}_{n}\overset{\text{iid}}{\sim}{p^{S}_{{\phi}_{t}}} using the goal-conditioned policy, and use the samples to train qϕt+1G{q^{G}_{{\phi}_{t+1}}}. However, there is no guarantee that this would increase the entropy of qϕt+1G{q^{G}_{{\phi}_{t+1}}}.

The intuition behind our method is simple: rather than fitting a generative model to these samples sn\mathbf{s}_{n}, we skew the samples so that rarely visited states are given more weight. See Figure 2 for a visualization of this process. How should we skew the samples if we want to maximize the entropy of qϕt+1G{q^{G}_{{\phi}_{t+1}}}? If we had access to the density of each state, pϕtS(S){p^{S}_{{\phi}_{t}}}({\mathbf{S}}), then we could simply weight each state by 1/pϕtS(S)1/{p^{S}_{{\phi}_{t}}}({\mathbf{S}}). We could then perform maximum likelihood estimation (MLE) for the uniform distribution by using the following importance sampling (IS) loss to train ϕt+1{{\phi}_{t+1}}:

where we use the fact that the uniform distribution US(S){U_{\mathcal{S}}}({\mathbf{S}}) has constant density for all states in S{\mathcal{S}}. However, computing this density pϕtS(S){p^{S}_{{\phi}_{t}}}({\mathbf{S}}) requires marginalizing out the MDP dynamics, which requires an accurate model of both the dynamics and the goal-conditioned policy.

We avoid needing to model the entire MDP process by approximating pϕtS(S){p^{S}_{{\phi}_{t}}}({\mathbf{S}}) with our previous learned generative model: pϕtS(S)≈qϕtG(S){p^{S}_{{\phi}_{t}}}({\mathbf{S}})\approx{q^{G}_{{\phi}_{t}}}({\mathbf{S}}). We therefore weight each state by the following weight function

where α\alpha is a hyperparameter that controls how heavily we weight each state. If our approximation qϕtG{q^{G}_{{\phi}_{t}}} is exact, we can choose α=−1\alpha=-1 and recover the exact IS procedure described above. If α=0\alpha=0, then this skew step has no effect. By choosing intermediate values of α\alpha, we trade off the reliability of our estimate qϕtG(S){q^{G}_{{\phi}_{t}}}({\mathbf{S}}) with the speed at which we want to increase the goal distribution entropy.

As described, this procedure relies on IS, which can have high variance, particularly if qϕtG(S)≈0{q^{G}_{{\phi}_{t}}}({\mathbf{S}})\approx 0. We therefore choose a class of generative models where the probabilities are prevented from collapsing to zero, as we will describe in Section 4 where we provide generative model details. To further reduce the variance, we train qϕt+1G{q^{G}_{{\phi}_{t+1}}} with sampling importance resampling (SIR) (Rubin, 1988) rather than IS. Rather than sampling from pϕtS{p^{S}_{{\phi}_{t}}} and weighting the update from each sample by wt,α{w_{t,\alpha}}, SIR explicitly defines a skewed empirical distribution as

where δ\delta is the indicator function and ZαZ_{\alpha} is the normalizing coefficient. We note that computing ZαZ_{\alpha} adds little computational overhead, since all of the weights already need to be computed. We then fit the generative model at the next iteration qϕt+1G{q^{G}_{{\phi}_{t+1}}} to pskewedt{p_{\text{skewed}_{t}}} using standard MLE. We found that using SIR resulted in significantly lower variance than IS. See Section B.2 for this comparision.

Goal Sampling Alternative

Because qϕt+1G≈pskewedt{q^{G}_{{\phi}_{t+1}}}\approx{p_{\text{skewed}_{t}}}, at iteration t+1t+1, one can sample goals from either qϕt+1G{q^{G}_{{\phi}_{t+1}}} or pskewedt{p_{\text{skewed}_{t}}}. Sampling goals from pskewedt{p_{\text{skewed}_{t}}} may be preferred if sampling from the learned generative model qϕt+1G{q^{G}_{{\phi}_{t+1}}} is computationally or otherwise challenging. In either case, one still needs to train the generative model qϕtG{q^{G}_{{\phi}_{t}}} to create pskewedt{p_{\text{skewed}_{t}}}. In our experiments, we found that both methods perform well.

Summary

Overall, Skew-Fit collects states from the environment and resamples each state in proportion to Equation 2 so that low-density states are resampled more often. Skew-Fit is shown in Figure 2 and summarized in Algorithm 1. We now discuss conditions under which Skew-Fit converges to the uniform distribution.

2 Skew-Fit Analysis

This section provides conditions under which qϕtG{q^{G}_{{\phi}_{t}}} converges in the limit to the uniform distribution over the state space S{\mathcal{S}}. We consider the case where N→∞N\rightarrow\infty, which allows us to study the limit behavior of the goal distribution pskewedt{p_{\text{skewed}_{t}}}. Our most general result is stated as follows:

Let S{\mathcal{S}} be a compact set. Define the set of distributions Q={p:support(p)⊆S}{\mathcal{Q}}=\{p:{\text{support}}(p)\subseteq{\mathcal{S}}\}. Let F:Q↦Q{\mathcal{F}}:{\mathcal{Q}}\mapsto{\mathcal{Q}} be continuous with respect to the pseudometric dH(p,q)≜∣H(p)−H(q)∣{d_{\mathcal{H}}}(p,q)\triangleq|{\mathcal{H}}(p)-{\mathcal{H}}(q)| and H(F(p))≥H(p){\mathcal{H}}({\mathcal{F}}(p))\geq{\mathcal{H}}(p) with equality if and only if pp is the uniform probability distribution on S{\mathcal{S}}, denoted as US{U_{\mathcal{S}}}. Define the sequence of distributions P=(p1,p2,… )P=(p_{1},p_{2},\dots) by starting with any p1∈Qp_{1}\in{\mathcal{Q}} and recursively defining pt+1=F(pt)p_{t+1}={\mathcal{F}}(p_{t}). The sequence PP converges to US{U_{\mathcal{S}}} with respect to dH{d_{\mathcal{H}}}. In other words, lim⁡t→0∣H(pt)−H(US)∣→0\lim_{t\rightarrow 0}|{\mathcal{H}}(p_{t})-{\mathcal{H}}({U_{\mathcal{S}}})|\rightarrow 0.

We will apply Lemma 3.1 to be the map from pskewedt{p_{\text{skewed}_{t}}} to pskewedt+1{p_{\text{skewed}_{t+1}}} to show that pskewedt{p_{\text{skewed}_{t}}} converges to US{U_{\mathcal{S}}}. If we assume that the goal-conditioned policy and generative model learning procedure are well behaved (i.e., the maps from qϕtG{q^{G}_{{\phi}_{t}}} to pϕtS{p^{S}_{{\phi}_{t}}} and from pskewedt{p_{\text{skewed}_{t}}} to qϕt+1G{q^{G}_{{\phi}_{t+1}}} are continuous), then to apply Lemma 3.1, we only need to show that H(pskewedt)≥H(pϕtS){\mathcal{H}}({p_{\text{skewed}_{t}}})\geq{\mathcal{H}}({p^{S}_{{\phi}_{t}}}) with equality if and only if pϕtS=US{p^{S}_{{\phi}_{t}}}={U_{\mathcal{S}}}. For the simple case when qϕtG=pϕtS{q^{G}_{{\phi}_{t}}}={p^{S}_{{\phi}_{t}}} identically at each iteration, we prove the convergence of Skew-Fit true for any value of α∈[−1,0)\alpha\in[-1,0) in Section A.3. However, in practice, qϕtG{q^{G}_{{\phi}_{t}}} only approximates pϕtS{p^{S}_{{\phi}_{t}}}. To address this more realistic situation, we prove the following result:

Given two distribution pϕtS{p^{S}_{{\phi}_{t}}} and qϕtG{q^{G}_{{\phi}_{t}}} where pϕtS≪qϕtG{p^{S}_{{\phi}_{t}}}\ll{q^{G}_{{\phi}_{t}}} p≪qp\ll q means that pp is absolutely continuous with respect to qq, i.e. p(s)=0  ⟹  q(s)=0p(\mathbf{s})=0\implies q(\mathbf{s})=0. and

define the pskewedt{p_{\text{skewed}_{t}}} as in Equation 3 and take N→∞N\rightarrow\infty. Let Hα(α){\mathcal{H}}_{\alpha}(\alpha) be the entropy of pskewedt{p_{\text{skewed}_{t}}} for a fixed α\alpha. Then there exists a constant a<0a<0 such that for all α∈[a,0)\alpha\in[a,0),

This lemma tells us that our generative model qϕtG{q^{G}_{{\phi}_{t}}} does not need to exactly fit the sampled states. Rather, we merely need the log densities of qϕtG{q^{G}_{{\phi}_{t}}} and pϕtS{p^{S}_{{\phi}_{t}}} to be correlated, which we expect to happen frequently with an accurate goal-conditioned policy, since pϕtS{p^{S}_{{\phi}_{t}}} is the set of states seen when trying to reach goals from qϕtG{q^{G}_{{\phi}_{t}}}. In this case, if we choose negative values of α\alpha that are small enough, then the entropy of pskewedt{p_{\text{skewed}_{t}}} will be higher than that of pϕtS{p^{S}_{{\phi}_{t}}}. Empirically, we found that α\alpha values as low as α=−1\alpha=-1 performed well.

In summary, pskewedt{p_{\text{skewed}_{t}}} converges to US{U_{\mathcal{S}}} under certain assumptions. Since we train each generative model qϕt+1G{q^{G}_{{\phi}_{t+1}}} by fitting it to pskewedt{p_{\text{skewed}_{t}}} with MLE, qϕtG{q^{G}_{{\phi}_{t}}} will also converge to US{U_{\mathcal{S}}}.

Training Goal-Conditioned Policies with Skew-Fit

Thus far, we have presented Skew-Fit assuming that we have access to a goal-reaching policy, allowing us to separately analyze how we can maximize H(G){\mathcal{H}}(\mathbf{G}). However, in practice we do not have access to such a policy, and this section discusses how we concurrently train a goal-reaching policy.

Maximizing I(S;G)I({\mathbf{S}};\mathbf{G}) can be done by simultaneously performing Skew-Fit and training a goal-conditioned policy to minimize H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}), or, equivalently, maximize −H(G∣S)-{\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}). Maximizing −H(G∣S)-{\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}) requires computing the density log⁡p(G∣S)\log p(\mathbf{G}\mid{\mathbf{S}}), which may be difficult to compute without strong modeling assumptions. However, for any distribution qq, the following lower bound on −H(G∣S)-{\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}):

The RL algorithm we use is reinforcement learning with imagined goals (RIG) (Nair et al., 2018), though in principle any goal-conditioned method could be used. RIG is an efficient off-policy goal-conditioned method that solves vision-based RL problems in a learned latent space. In particular, RIG fits a β\beta-VAE (Higgins et al., 2017) and uses it to encode observations and goals into a latent space, which it uses as the state representation. RIG also uses the β\beta-VAE to compute rewards, log⁡q(G∣S)\log q(\mathbf{G}\mid{\mathbf{S}}). Unlike RIG, we use the goal distribution from Skew-Fit to sample goals for exploration and for relabeling goals during training (Andrychowicz et al., 2017). Since RIG already trains a generative model over states, we reuse this β\beta-VAE for the generative model qϕG{q^{G}_{{\phi}}} of Skew-Fit. To make the most use of the data, we train qϕG{q^{G}_{{\phi}}} on all visited state rather than only the terminal states, which we found to work well in practice. To prevent the estimated state likelihoods from collapsing to zero, we model the posterior of the β\beta-VAE as a multivariate Gaussian distribution with a fixed variance and only learn the mean. We summarize RIG and provide details for how we combine Skew-Fit and RIG in Section C.4 and describe how we estimate the likelihoods given the β\beta-VAE in Section C.1.

Related Work

Many prior methods in the goal-conditioned reinforcement learning literature focus on training goal-conditioned policies and assume that a goal distribution is available to sample from during exploration (Kaelbling, 1993; Schaul et al., 2015; Andrychowicz et al., 2017; Pong et al., 2018), or use a heuristic to design a non-parametric (Colas et al., 2018b; Warde-Farley et al., 2018; Florensa et al., 2018a) or parametric (Péré et al., 2018; Nair et al., 2018) goal distribution based on previously visited states. These methods are largely complementary to our work: rather than proposing a better method for training goal-reaching policies, we propose a principled method for maximizing the entropy of a goal sampling distribution, H(G){\mathcal{H}}(\mathbf{G}), such that these policies cover a wide range of states.

Our method learns without any task rewards, directly acquiring a policy that can be reused to reach user-specified goals. This stands in contrast to exploration methods that modify the reward based on state visitation frequency (Bellemare et al., 2016; Ostrovski et al., 2017; Tang et al., 2017; Chentanez et al., 2005; Lopes et al., 2012; Stadie et al., 2016; Pathak et al., 2017; Burda et al., 2018, 2019; Mohamed & Rezende, 2015; Tang et al., 2017; Fu et al., 2017). While these methods can also be used without a task reward, they provide no mechanism for distilling the knowledge gained from visiting diverse states into flexible policies that can be applied to accomplish new goals at test-time: their policies visit novel states, and they quickly forget about them as other states become more novel. Similarly, methods that provably maximize state entropy without using goal-directed exploration (Hazan et al., 2019) or methods that define new rewards to capture measures of intrinsic motivation (Mohamed & Rezende, 2015) and reachability (Savinov et al., 2018) do not produce reusable policies.

Other prior methods extract reusable skills in the form of latent-variable-conditioned policies, where latent variables are interpreted as options (Sutton et al., 1999) or abstract skills (Hausman et al., 2018; Gupta et al., 2018b; Eysenbach et al., 2019; Gupta et al., 2018a; Florensa et al., 2017). The resulting skills are diverse, but have no grounded interpretation, while Skew-Fit policies can be used immediately after unsupervised training to reach diverse user-specified goals.

Some prior methods propose to choose goals based on heuristics such as learning progress (Baranes & Oudeyer, 2012; Veeriah et al., 2018; Colas et al., 2018a), how off-policy the goal is (Nachum et al., 2018), level of difficulty (Florensa et al., 2018b), or likelihood ranking (Zhao & Tresp, 2019). In contrast, our approach provides a principled framework for optimizing a concrete and well-motivated exploration objective, can provably maximize this objective under regularity assumptions, and empirically outperforms many of these prior work (see Section 6).

Experiments

Our experiments study the following questions: (1) Does Skew-Fit empirically result in a goal distribution with increasing entropy? (2) Does Skew-Fit improve exploration for goal-conditioned RL? (3) How does Skew-Fit compare to prior work on choosing goals for vision-based, goal-conditioned RL? (4) Can Skew-Fit be applied to a real-world, vision-based robot task?

To see the effects of Skew-Fit on goal distribution entropy in isolation of learning a goal-reaching policy, we study an idealized example where the policy is a near-perfect goal-reaching policy. The environment consists of four rooms (Sutton et al., 1999). At the beginning of an episode, the agent begins in the bottom-right room and samples a goal from the goal distribution qϕtG{q^{G}_{{\phi}_{t}}}. To simulate stochasticity of the policy and environment, we add a Gaussian noise with standard deviation of 0.060.06 units to this goal, where the entire environment is 11×1111\times 11 units. The policy reaches the state that is closest to this noisy goal and inside the rooms, giving us a state sample sn\mathbf{s}_{n} for training qϕtG{q^{G}_{{\phi}_{t}}}. Due to the relatively small noise, the agent cannot rely on this stochasticity to explore the different rooms and must instead learn to set goals that are progressively farther and farther from the initial state. We compare multiple values of α\alpha, where α=0\alpha=0 corresponds to not using Skew-Fit. The β\beta-VAE hyperparameters used to train qϕtG{q^{G}_{{\phi}_{t}}} are given in Section C.2. As seen in Figure 3, sampling uniformly from previous experience (α=0\alpha=0) to set goals results in a policy that primarily sets goal near the initial state distribution. In contrast, Skew-Fit results in quickly learning a high entropy, near-uniform distribution over the state space.

Exploration with Skew-Fit

We next evaluate Skew-Fit while concurrently learning a goal-conditioned policy on a task with state inputs, which enables us study exploration performance independently of the challenges with image observations. We evaluate on a task that requires training a simulated quadruped “ant” robot to navigate to different XY positions in a labyrinth, as shown in Figure 4. The reward is the negative distance to the goal XY-position, and additional environment details are provided in Appendix D. This task presents a challenge for goal-directed exploration: the set of valid goals is unknown due to the walls, and random actions do not result in exploring locations far from the start. Thus, Skew-Fit must set goals that meaningfully explore the space while simultaneously learning to reach those goals.

We use this domain to compare Skew-Fit to a number of existing goal-sampling methods. We compare to the relabeling scheme described in the hindsight experience replay (labeled HER). We compare to curiosity-driven prioritization (Ranked-Based Priority) (Zhao et al., 2019), a variant of HER that samples goals for relabeling based on their ranked likelihoods. Florensa et al. (2018b) samples goals from a GAN based on the difficulty of reaching the goal. We compare against this method by replacing qϕG{q^{G}_{{\phi}}} with the GAN and label it AutoGoal GAN. We also compare to the non-parametric goal proposal mechanism proposed by (Warde-Farley et al., 2018), which we label DISCERN-g. Lastly, to demonstrate the difficulty of the exploration challenge in these domains, we compare to #-Exploration (Tang et al., 2017), an exploration method that assigns bonus rewards based on the novelty of new states. We train the goal-conditioned policy for each method using soft actor critic (SAC) (Haarnoja et al., 2018). Implementation details of SAC and the prior works are given in Section C.3.

We see in Figure 4 that Skew-Fit is the only method that makes significant progress on this challenging labyrinth locomotion task. The prior methods on goal-sampling primarily set goals close to the start location, while the extrinsic exploration reward in #-Exploration dominated the goal-reaching reward. These results demonstrate that Skew-Fit accelerates exploration by setting diverse goals in tasks with unknown goal spaces.

Vision-Based Continuous Control Tasks

We now evaluate Skew-Fit on a variety of image-based continuous control tasks, where the policy must control a robot arm using only image observations, there is no state-based or task-specific reward, and Skew-Fit must directly set image goals. We test our method on three different image-based simulated continuous control tasks released by the authors of RIG (Nair et al., 2018): Visual Door, Visual Pusher, and Visual Pickup. These environments contain a robot that can open a door, push a puck, and lift up a ball to different configurations, respectively. To our knowledge, these are the only goal-conditioned, vision-based continuous control environments that are publicly available and experimentally evaluated in prior work, making them a good point of comparison. See Figure 5 for visuals and Appendix C for environment details. The policies are trained in a completely unsupervised manner, without access to any prior information about the image-space or any pre-defined goal-sampling distribution. To evaluate their performance, we sample goal images from a uniform distribution over valid states and report the agent’s final distance to the corresponding simulator states (e.g., distance of the object to the target object location), but the agent never has access to this true uniform distribution nor the ground-truth state information during training. While this evaluation method is only practical in simulation, it provides us with a quantitative measure of a policy’s ability to reach a broad coverage of goals in a vision-based setting.

We compare Skew-Fit to a number of existing methods on this domain. First, we compare to the methods described in the previous experiment (HER, Rank-Based Priority, #-Exploration, Autogoal GAN, and DISCERN-g). These methods that we compare to were developed in non-vision, state-based environments. To ensure a fair comparison across methods, we combine these prior methods with a policy trained using RIG. We additionally compare to Hazan et al. (2019), an exploration method that assigns bonus rewards based on the likelihood of a state (labeled Hazan et al.). Next, we compare to RIG without Skew-Fit. Lastly, we compare to DISCERN (Warde-Farley et al., 2018), a vision-based method which uses a non-parametric clustering approach to sample goals and an image discriminator to compute rewards.

We see in Figure 6 that Skew-Fit significantly outperforms prior methods both in terms of task performance and sample complexity. The most common failure mode for prior methods is that the goal distributions collapse, resulting in the agent learning to reach only a fraction of the state space, as shown in Figure 1. For comparison, additional samples of qϕG{q^{G}_{{\phi}}} when trained with and without Skew-Fit are shown in Section B.3. Those images show that without Skew-Fit, qϕG{q^{G}_{{\phi}}} produces a small, non-diverse distribution for each environment: the object is in the same place for pickup, the puck is often in the starting position for pushing, and the door is always closed. In contrast, Skew-Fit proposes goals where the object is in the air and on the ground, where the puck positions are varied, and the door angle changes.

We can see the effect of these goal choices by visualizing more example rollouts for RIG and Skew-Fit. These visuals, shown in Figure 14 in Section B.3, show that RIG only learns to reach states close to the initial position, while Skew-Fit learns to reach the entire state space. For a quantitative comparison, Figure 7 shows the cumulative total exploration pickups for each method. From the graph, we see that many methods have a near-constant rate of object lifts throughout all of training. Skew-Fit is the only method that significantly increases the rate at which the policy picks up the object during exploration, suggesting that only Skew-Fit sets goals that encourage the policy to interact with the object.

Real-World Vision-Based Robotic Manipulation

We also demonstrate that Skew-Fit scales well to the real world with a door opening task, Real World Visual Door, as shown in Figure 5. While a number of prior works have studied RL-based learning of door opening (Kalakrishnan et al., 2011; Chebotar et al., 2017), we demonstrate the first method for autonomous learning of door opening without a user-provided, task-specific reward function. As in simulation, we do not provide any goals to the agent and simply let it interact with the door, without any human guidance or reward signal. We train two agents using RIG and RIG with Skew-Fit. Every seven and a half minutes of interaction time, we evaluate on 55 goals and plot the cumulative successes for each method. Unlike in simulation, we cannot easily measure the difference between the policy’s achieved and desired door angle. Instead, we visually denote a binary success/failure for each goal based on whether the last state in the trajectory achieves the target angle. As Figure 8 shows, standard RIG only starts to open the door after five hours of training. In contrast, Skew-Fit learns to occasionally open the door after three hours of training and achieves a near-perfect success rate after five and a half hours of interaction. Figure 8 also shows examples of successful trajectories from the Skew-Fit policy, where we see that the policy can reach a variety of user-specified goals. These results demonstrate that Skew-Fit is a promising technique for solving real world tasks without any human-provided reward function. Videos of Skew-Fit solving this task and the simulated tasks can be viewed on our website. https://sites.google.com/view/skew-fit

Additional Experiments

To study the sensitivity of Skew-Fit to the hyperparameter α\alpha, we sweep α\alpha across the values [−1,−0.75,−0.5,−0.25,0][-1,-0.75,-0.5,-0.25,0] on the simulated image-based tasks. The results are in Appendix B and demonstrate that Skew-Fit works across a large range of values for α\alpha, and α=−1\alpha=-1 consistently outperform α=0\alpha=0 (i.e. outperforms no Skew-Fit). Additionally, Appendix C provides a complete description our method hyperparameters, including network architecture and RL algorithm hyperparameters.

Conclusion

We presented a formal objective for self-supervised goal-directed exploration, allowing researchers to quantify and compare progress when designing algorithms that enable agents to autonomously learn. We also presented Skew-Fit, an algorithm for training a generative model to approximate a uniform distribution over an initially unknown set of valid states, using data obtained via goal-conditioned reinforcement learning, and our theoretical analysis gives conditions under which Skew-Fit converges to the uniform distribution. When such a model is used to choose goals for exploration and to relabeling goals for training, the resulting method results in much better coverage of the state space, enabling our method to explore effectively. Our experiments show that when we concurrently train a goal-reaching policy using self-generated goals, Skew-Fit produces quantifiable improvements on simulated robotic manipulation tasks, and can be used to learn a door opening skill to reach a 95%95\% success rate directly on a real-world robot, without any human-provided reward supervision.

Acknowledgement

This research was supported by Berkeley DeepDrive, Huawei, ARL DCIST CRA W911NF-17-2-0181, NSF IIS-1651843, and the Office of Naval Research, as well as Amazon, Google, and NVIDIA. We thank Aviral Kumar, Carlos Florensa, Aurick Zhou, Nilesh Tripuraneni, Vickie Ye, Dibya Ghosh, Coline Devin, Rowan McAllister, John D. Co-Reyes, various members of the Berkeley Robotic AI & Learning (RAIL) lab, and anonymous reviewers for their insightful discussions and feedback.

References

Appendix A Proofs

The definitions of continuity and convergence for pseudo-metrics are similar to those for metrics, and we state them below.

A function f:Q↦Qf:{\mathcal{Q}}\mapsto{\mathcal{Q}} is continuous with respect to a pseudo-metric dd if for any p∈Qp\in{\mathcal{Q}} and any scalar ϵ>0\epsilon>0, there exists a δ\delta such that for all q∈Qq\in{\mathcal{Q}},

An infinite sequence p1,p2…p_{1},p_{2}\dots converges to a value pp with respect to a pseudo-metric dd, which we write as

Let S{\mathcal{S}} be a compact set. Define the set of distributions Q={p:support(p)⊆S}{\mathcal{Q}}=\{p:{\text{support}}(p)\subseteq{\mathcal{S}}\}. Let F:Q↦Q{\mathcal{F}}:{\mathcal{Q}}\mapsto{\mathcal{Q}} be continuous with respect to the pseudometric dH(p,q)≜∣H(p)−H(q)∣{d_{\mathcal{H}}}(p,q)\triangleq|{\mathcal{H}}(p)-{\mathcal{H}}(q)| and H(F(p))≥H(p){\mathcal{H}}({\mathcal{F}}(p))\geq{\mathcal{H}}(p) with equality if and only if pp is the uniform probability distribution on S{\mathcal{S}}, denoted as US{U_{\mathcal{S}}}. Define the sequence of distributions P=(p1,p2,… )P=(p_{1},p_{2},\dots) by starting with any p1∈Qp_{1}\in{\mathcal{Q}} and recursively defining pt+1=F(pt)p_{t+1}={\mathcal{F}}(p_{t}). The sequence PP converges to US{U_{\mathcal{S}}} with respect to dH{d_{\mathcal{H}}}. In other words, lim⁡t→0∣H(pt)−H(US)∣→0\lim_{t\rightarrow 0}|{\mathcal{H}}(p_{t})-{\mathcal{H}}({U_{\mathcal{S}}})|\rightarrow 0.

The idea of the proof is to show that the distance (with respect to dH{d_{\mathcal{H}}}) between ptp_{t} and US{U_{\mathcal{S}}} converges to a value. If this value is , then the proof is complete since US{U_{\mathcal{S}}} uniquely has zero distance to itself. Otherwise, we will show that this implies that F{\mathcal{F}} is not continuous, which a contradiction.

For shorthand, define dtd_{t} to be the dH{d_{\mathcal{H}}}-distance to the uniform distribution, as in

First we prove that dtd_{t} converges. Since the entropies of the sequence (p1,… )(p_{1},\dots) monotonically increase, we have that

We also know that dtd_{t} is lower bounded by , and so by the monotonic convergence theorem, we have that

To prove the lemma, we want to show that d∗=0d^{*}=0. Suppose, for contradiction, that d∗≠0d^{*}\neq 0. Then consider any distribution, q∗q^{*}, such that dH(q∗,US)=d∗{d_{\mathcal{H}}}(q^{*},{U_{\mathcal{S}}})=d^{*}. Such a distribution always exists since we can continuously interpolate entropy values between H(p1){\mathcal{H}}(p_{1}) and H(US){\mathcal{H}}({U_{\mathcal{S}}}) with a mixture distribution. Note that q∗≠USq^{*}\neq{U_{\mathcal{S}}} since dH(US,US)=0{d_{\mathcal{H}}}({U_{\mathcal{S}}},{U_{\mathcal{S}}})=0. Since lim⁡t→∞dt→d∗\lim_{t\rightarrow\infty}d_{t}\rightarrow d^{*}, we have that

Because the function F{\mathcal{F}} is continuous with respect to dH{d_{\mathcal{H}}}, Equation 5 implies that

However, since F(pt)=pt+1{\mathcal{F}}(p_{t})=p_{t+1} we can equivalently write the above equation as

which, through a change of index variables, implies that

Since q∗q^{*} is not the uniform distribution, we have that H(F(q∗))>H(q∗){\mathcal{H}}({\mathcal{F}}(q^{*}))>{\mathcal{H}}(q^{*}), which implies that F(q∗){\mathcal{F}}(q^{*}) and q∗q^{*} are unique distributions. So, ptp_{t} converges to two distinct values, q∗q^{*} and F(q∗){\mathcal{F}}(q^{*}), which is a contradiction. Thus, it must be the case that d∗=0d^{*}=0, completing the proof. ∎

A.2 Proof of Lemma 3.2

Given two distribution p(x)p(x) and q(x)q(x) where p≪qp\ll q and

Observe that {pα:α∈}\{p_{\alpha}:\alpha\in\} is a one-dimensional exponential family

with log carrier density k(x)=log⁡p(x)k(x)=\log p(x), natural parameter α\alpha, sufficient statistic T(x)=log⁡q(x)T(x)=\log q(x), and log-normalizer A(α)=∫XeαT(x)+k(x)dxA(\alpha)=\int_{{\mathcal{X}}}e^{\alpha T(x)+k(x)}dx. As shown in (Nielsen & Nock, 2010), the entropy of a distribution from a one-dimensional exponential family with parameter α\alpha is given by:

The derivative with respect to α\alpha is then

which is negative by assumption. Because the derivative at α=0\alpha=0 is negative, then there exists a constant a>0a>0 such that for all α∈[−a,0]\alpha\in[-a,0], Hα(α)>Hα(0)=H(p){\mathcal{H}}_{\alpha}(\alpha)>{\mathcal{H}}_{\alpha}(0)={\mathcal{H}}(p). ∎

The paper applies to the case where q=qϕGq={q^{G}_{{\phi}}} and p=pϕSp={p^{S}_{{\phi}}}. When we take N→∞N\rightarrow\infty, we have that pskewed{p_{\text{skewed}}} corresponds to pαp_{\alpha} above.

A.3 Simple Case Proof

We prove the convergence directly for the (even more) simplified case when pθ=p(S∣qϕtG){p_{\theta}}=p({\mathbf{S}}\mid{q^{G}_{{\phi}_{t}}}) using a similar technique:

Assume the set S{\mathcal{S}} has finite volume so that its uniform distribution US{U_{\mathcal{S}}} is well defined and has finite entropy. Given any distribution p(s)p(\mathbf{s}) whose support is S{\mathcal{S}}, recursively define ptp_{t} with p1=pp_{1}=p and

where ZαtZ_{\alpha}^{t} is the normalizing constant and α∈[0,1)\alpha\in[0,1).

The sequence (p1,p2,… )(p_{1},p_{2},\dots) converges to US{U_{\mathcal{S}}}, the uniform distribution S{\mathcal{S}}.

If α=0\alpha=0, then p2p_{2} (and all subsequent distributions) will clearly be the uniform distribution. We now study the case where α∈(0,1)\alpha\in(0,1).

At each iteration tt, define the one-dimensional exponential family {pθt:θ∈}\{{p_{\theta}^{t}}:\theta\in\} where pθt{p_{\theta}^{t}} is

with log carrier density k(s)=0k(\mathbf{s})=0, natural parameter θ\theta, sufficient statistic T(s)=log⁡pt(s)T(\mathbf{s})=\log p_{t}(\mathbf{s}), and log-normalizer A(θ)=∫SeθT(s)dsA(\theta)=\int_{{\mathcal{S}}}e^{\theta T(\mathbf{s})}d\mathbf{s}. As shown in (Nielsen & Nock, 2010), the entropy of a distribution from a one-dimensional exponential family with parameter θ\theta is given by:

The derivative with respect to θ\theta is then

This monotonically increasing sequence is upper bounded by the entropy of the uniform distribution, and so this sequence must converge.

The sequence can only converge if ddθHθt(θ)\frac{d}{d\theta}{{\mathcal{H}}_{\theta}^{t}}(\theta) converges to zero. However, because α\alpha is bounded away from , Equation 8 states that this can only happen if

Because ptp_{t} has full support, then so does pθt{p_{\theta}^{t}}. Thus, Equation 9 is only true if log⁡pt(s)\log p_{t}(\mathbf{s}) converges to a constant, i.e. ptp_{t} converges to the uniform distribution. ∎

Appendix B Additional Experiments

In our experiments, we combined Skew-Fit with soft actor critic (SAC) (Haarnoja et al., 2018). We conduct a set of experiments to test whether Skew-Fit may be used with other RL algorithms for training the goal-conditioned policy. To that end, we replaced SAC with twin delayed deep deterministic policy gradient (TD3) (Fujimoto et al., 2018) and ran the same Skew-Fit experiments on Visual Door, Visual Pusher, and Visual Pickup. In Figure 9, we see that Skew-Fit performs consistently well with both SAC and TD3, demonstrating that Skew-Fit is beneficial across multiple RL algorithms.

Sensitivity to α𝛼\alpha Hyperparameter

We study the sensitivity of the α\alpha hyperparameter by testing values of α∈[−1,−0.75,−0.5,−0.25,0]\alpha\in[-1,-0.75,-0.5,-0.25,0] on the Visual Door and Visual Pusher task. The results are included in Figure 10 and shows that our method is robust to different parameters of α\alpha, particularly for the more challenging Visual Pusher task. Also, the method consistently outperform α=0\alpha=0, which is equivalent to sampling uniformly from the replay buffer.

B.2 Variance Ablation

We measure the gradient variance of training a VAE on an unbalanced Visual Door image dataset with Skew-Fit vs Skew-Fit with importance sampling (IS) vs no Skew-Fit (labeled MLE). We construct the imbalanced dataset by rolling out a random policy in the environment and collecting the visual observations. Most of the images contained the door in a closed position; in a few, the door was opened. In Figure 11, we see that the gradient variance for Skew-Fit with IS is catastrophically large for large values of α\alpha. In contrast, for Skew-Fit with SIR, which is what we use in practice, the variance is relatively similar to that of MLE. Additionally we trained three VAE’s, one with MLE on a uniform dataset of valid door opening images, one with Skew-Fit on the unbalanced dataset from above, and one with MLE on the same unbalanced dataset. As expected, the VAE that has access to the uniform dataset gets the lowest negative log likelihood score. This is the oracle method, since in practice we would only have access to imbalanced data. As shown in Table 1, Skew-Fit considerably outperforms MLE, getting a much closer to oracle log likelihood score.

B.3 Goal and Performance Visualization

We visualize the goals sampled from Skew-Fit as well as those sampled when using the prior method, RIG (Nair et al., 2018). As shown in Figure 12 and Figure 13, the generative model qϕG{q^{G}_{{\phi}}} results in much more diverse samples when trained with Skew-Fit. We we see in Figure 14, this results in a policy that more consistently reaches the goal image.

Appendix C Implementation Details

We estimate the density under the VAE by using a sample-wise approximation to the marginal over xx estimated using importance sampling:

where qθq_{\theta} is the encoder, pψp_{\psi} is the decoder, and p(z)p(z) is the prior, which in this case is unit Gaussian. We found that sampling N=10N=10 latents for estimating the density worked well in practice.

C.2 Oracle 2D Navigation Experiments

We initialize the VAE to the bottom left corner of the environment for Four Rooms. Both the encoder and decoder have 2 hidden layers with units, ReLU hidden activations, and no output activations. The VAE has a latent dimension of 88 and a Gaussian decoder trained with a fixed variance, batch size of 256256, and 10001000 batches at each iteration. The VAE is trained on the exploration data buffer every 1000 rollouts.

C.3 Implementation of SAC and Prior Work

For all experiments, we trained the goal-conditioned policy using soft actor critic (SAC) (Haarnoja et al., 2018). To make the method goal-conditioned, we concatenate the target XY-goal to the state vector. During training, we retroactively relabel the goals (Kaelbling, 1993; Andrychowicz et al., 2017) by sampling from the goal distribution with probabilty 0.50.5. Note that the original RIG (Nair et al., 2018) paper used TD3 (Fujimoto et al., 2018), which we also replaced with SAC in our implementation of RIG. We found that maximum entropy policies in general improved the performance of RIG, and that we did not need to add noise on top of the stochastic policy’s noise. In the prior RIG method, the VAE was pre-trained on a uniform sampling of images from the state space of each environment. In order to ensure a fair comparison to Skew-Fit, we forego pre-training and instead train the VAE alongside RL, using the variant described in the RIG paper. For our RL network architectures and training scheme, we use fully connected networks for the policy, Q-function and value networks with two hidden layers of size 400400 and 300300 each. We also delay training any of these networks for 1000010000 time steps in order to collect sufficient data for the replay buffer as well as to ensure the latent space of the VAE is relatively stable (since we continuously train the VAE concurrently with RL training). As in RIG, we train a goal-conditioned value functions (Schaul et al., 2015) using hindsight experience replay (Andrychowicz et al., 2017), relabelling 50%50\% of exploration goals as goals sampled from the VAE prior N(0,1)\mathcal{N}(0,1) and 30%30\% from future goals in the trajectory.

For our implementation of (Hazan et al., 2019), we trained the policies with the reward

For rHazan et al.r_{\text{Hazan et al.}}, we use the reward described in Section 5.2 of Hazan et al. (2019), which requires an estimated likelihood of the state. To compute these likelihood, we use the same method as in Skew-Fit (see Section C.1). With 3 seeds each, we tuned λ\lambda across values [100,10,1,0.1,0.01,0.001][100,10,1,0.1,0.01,0.001] for the door task, but all values performed poorly. For the pushing and picking tasks, we tested values across [1,0.1,0.01,0.001,0.0001][1,0.1,0.01,0.001,0.0001] and found that 0.1 and 0.01 performed best for each task, respectively.

C.4 RIG with Skew-Fit Summary

Algorithm 2 provides detailed pseudo-code for how we combined our method with RIG. Steps that were removed from the base RIG algorithm are highlighted in blue and steps that were added are highlighted in red. The main differences between the two are (1) not needing to pre-train the β\beta-VAE, (2) sampling exploration goals from the buffer using pskewed{p_{\text{skewed}}} instead of the VAE prior, (3) relabeling with replay buffer goals sampled using pskewed{p_{\text{skewed}}} instead of from the VAE prior, and (4) training the VAE on replay buffer data data sampled using pskewed{p_{\text{skewed}}} instead of uniformly.

C.5 Vision-Based Continuous Control Experiments

In our experiments, we use an image size of 48x48. For our VAE architecture, we use a modified version of the architecture used in the original RIG paper (Nair et al., 2018). Our VAE has three convolutional layers with kernel sizes: 5x5, 3x3, and 3x3, number of output filters: 16, 32, and 64 and strides: 3, 2, and 2. We then have a fully connected layer with the latent dimension number of units, and then reverse the architecture with de-convolution layers. We vary the latent dimension of the VAE, the β\beta term of the VAE and the α\alpha term for Skew-Fit based on the environment. Additionally, we vary the training schedule of the VAE based on the environment. See the table at the end of the appendix for more details. Our VAE has a Gaussian decoder with identity variance, meaning that we train the decoder with a mean-squared error loss.

When training the VAE alongside RL, we found the following three schedules to be effective for different environments:

For first 5K5K steps: Train VAE using standard MLE training every 500500 time steps for 10001000 batches. After that, train VAE using Skew-Fit every 500500 time steps for 200200 batches.

For first 5K5K steps: Train VAE using standard MLE training every 500500 time steps for 10001000 batches. For the next 45K45K steps, train VAE using Skew-Fit every 500500 steps for 200200 batches. After that, train VAE using Skew-Fit every 10001000 time steps for 200200 batches.

For first 40K40K steps: Train VAE using standard MLE training every 40004000 time steps for 10001000 batches. Afterwards, train VAE using Skew-Fit every 40004000 time steps for 200200 batches.

We found that initially training the VAE without Skew-Fit improved the stability of the algorithm. This is due to the fact that density estimates under the VAE are constantly changing and inaccurate during the early phases of training. Therefore, it made little sense to use those estimates to prioritize goals early on in training. Instead, we simply train using MLE training for the first 5K5K timesteps, and after that we perform Skew-Fit according to the VAE schedules above. Table 2 lists the hyper-parameters that were shared across the continuous control experiments. Table 3 lists hyper-parameters specific to each environment. Additionally, Section C.4 discusses the combined RIG + Skew-Fit algorithm.

Appendix D Environment Details

Four Rooms: A 20 x 20 2D pointmass environment in the shape of four rooms (Sutton et al., 1999). The observation is the 2D position of the agent, and the agent must specify a target 2D position as the action. The dynamics of the environment are the following: first, the agent is teleported to the target position, specified by the action. Then a Gaussian change in position with mean and standard deviation 0.06050.0605 is appliedIn the main paper, we rounded this to 0.060.06, but this difference does not matter.. If the action would result in the agent moving through or into a wall, then the agent will be stopped at the wall instead.

Ant: A MuJoCo (Todorov et al., 2012) ant environment. The observation is a 3D position and velocities, orientation, joint angles, and velocity of the joint angles of the ant (8 total). The observation space is 29 dimensions. The agent controls the ant through the joints, which is 8 dimensions. The goal is a target 2D position, and the reward is the negative Euclidean distance between the achieved 2D position and target 2D position.

Visual Pusher: A MuJoCo environment with a 7-DoF Sawyer arm and a small puck on a table that the arm must push to a target position. The agent controls the arm by commanding x,yx,y position for the end effector (EE). The underlying state is the EE position, ee and puck position pp. The evaluation metric is the distance between the goal and final puck positions. The hand goal/state space is a 10x10 cm2 box and the puck goal/state space is a 30x20 cm2 box. Both the hand and puck spaces are centered around the origin. The action space ranges in the interval $$ in the x and y dimensions.

Visual Door: A MuJoCo environment with a 7-DoF Sawyer arm and a door on a table that the arm must pull open to a target angle. Control is the same as in Visual Pusher. The evaluation metric is the distance between the goal and final door angle, measured in radians. In this environment, we do not reset the position of the hand or door at the end of each trajectory. The state/goal space is a 5x20x15 cm3 box in the x,y,zx,y,z dimension respectively for the arm and an angle between [0,.83][0,.83] radians. The action space ranges in the interval $$ in the x, y and z dimensions.

Visual Pickup: A MuJoCo environment with the same robot as Visual Pusher, but now with a different object. The object is cube-shaped, but a larger intangible sphere is overlaid on top so that it is easier for the agent to see. Moreover, the robot is constrained to move in 2 dimension: it only controls the y,zy,z arm positions. The xx position of both the arm and the object is fixed. The evaluation metric is the distance between the goal and final object position. For the purpose of evaluation, 75%75\% of the goals have the object in the air and 25%25\% have the object on the ground. The state/goal space for both the object and the arm is 10cm in the yy dimension and 13cm in the zz dimension. The action space ranges in the interval $inthein theyandandz$ dimensions.

Real World Visual Door: A Rethink Sawyer Robot with a door on a table. The arm must pull the door open to a target angle. The agent controls the arm by commanding the x,y,zx,y,z velocity of the EE. Our controller commands actions at a rate of up to 10Hz with the scale of actions ranging up to 1cm in magnitude. The underlying state and goal is the same as in Visual Door. Again we do not reset the position of the hand or door at the end of each trajectory. We obtain images using a Kinect Sensor. The state/goal space for the environment is a 10x10x10 cm3 box. The action space ranges in the interval (incm)inthex,yandzdimensions.Thedoorangleliesintherange(in cm) in the x, y and z dimensions. The door angle lies in the range degrees.

Some goal-conditioned RL methods such as Warde-Farley et al. (2018); Nair et al. (2018) present methods for minimizing a lower bound for H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}), by approximating log⁡p(G∣S)\log p(\mathbf{G}\mid{\mathbf{S}}) and using it as the reward. Other goal-conditioned RL methods (Kaelbling, 1993; Lillicrap et al., 2016; Schaul et al., 2015; Andrychowicz et al., 2017; Pong et al., 2018; Florensa et al., 2018a) are not developed with the intention of minimizing the conditional entropy H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}). Nevertheless, one can see that goal-conditioned RL generally minimizes H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}) by noting that the optimal goal-conditioned policy will deterministically reach the goal. The corresponding conditional entropy of the goal given the state, H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}), would be zero, since given the current state, there would be no uncertainty over the goal (the goal must have been the current state since the policy is optimal). So, the objective of goal-conditioned RL can be interpreted as finding a policy such that H(G∣S)=0{\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}})=0. Since zero is the minimum value of H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}), then goal-conditioned RL can be interpreted as minimizing H(G∣S){\mathcal{H}}(\mathbf{G}\mid{\mathbf{S}}).