Off-Policy Deep Reinforcement Learning without Exploration

Scott Fujimoto, David Meger, Doina Precup

Introduction

Batch reinforcement learning, the task of learning from a fixed dataset without further interactions with the environment, is a crucial requirement for scaling reinforcement learning to tasks where the data collection procedure is costly, risky, or time-consuming. Off-policy batch reinforcement learning has important implications for many practical applications. It is often preferable for data collection to be performed by some secondary controlled process, such as a human operator or a carefully monitored program. If assumptions on the quality of the behavioral policy can be made, imitation learning can be used to produce strong policies. However, most imitation learning algorithms are known to fail when exposed to suboptimal trajectories, or require further interactions with the environment to compensate (Hester et al., 2017; Sun et al., 2018; Cheng et al., 2018). On the other hand, batch reinforcement learning offers a mechanism for learning from a fixed dataset without restrictions on the quality of the data.

Most modern off-policy deep reinforcement learning algorithms fall into the category of growing batch learning (Lange et al., 2012), in which data is collected and stored into an experience replay dataset (Lin, 1992), which is used to train the agent before further data collection occurs. However, we find that these “off-policy” algorithms can fail in the batch setting, becoming unsuccessful if the dataset is uncorrelated to the true distribution under the current policy. Our most surprising result shows that off-policy agents perform dramatically worse than the behavioral agent when trained with the same algorithm on the same dataset.

This inability to learn truly off-policy is due to a fundamental problem with off-policy reinforcement learning we denote extrapolation error, a phenomenon in which unseen state-action pairs are erroneously estimated to have unrealistic values. Extrapolation error can be attributed to a mismatch in the distribution of data induced by the policy and the distribution of data contained in the batch. As a result, it may be impossible to learn a value function for a policy which selects actions not contained in the batch.

To overcome extrapolation error in off-policy learning, we introduce batch-constrained reinforcement learning, where agents are trained to maximize reward while minimizing the mismatch between the state-action visitation of the policy and the state-action pairs contained in the batch. Our deep reinforcement learning algorithm, Batch-Constrained deep Q-learning (BCQ), uses a state-conditioned generative model to produce only previously seen actions. This generative model is combined with a Q-network, to select the highest valued action which is similar to the data in the batch. Under mild assumptions, we prove this batch-constrained paradigm is necessary for unbiased value estimation from incomplete datasets for finite deterministic MDPs.

Unlike any previous continuous control deep reinforcement learning algorithms, BCQ is able to learn successfully without interacting with the environment by considering extrapolation error. Our algorithm is evaluated on a series of batch reinforcement learning tasks in MuJoCo environments (Todorov et al., 2012; Brockman et al., 2016), where extrapolation error is particularly problematic due to the high-dimensional continuous action space, which is impossible to sample exhaustively. Our algorithm offers a unified view on imitation and off-policy learning, and is capable of learning from purely expert demonstrations, as well as from finite batches of suboptimal data, without further exploration. We remark that BCQ is only one way to approach batch-constrained reinforcement learning in a deep setting, and we hope that it will be serve as a foundation for future algorithms. To ensure reproducibility, we provide precise experimental and implementation details, and our code is made available (https://github.com/sfujim/BCQ).

Background

The Bellman operator is a contraction for γ∈[0,1)\gamma\in[0,1) with unique fixed point Qπ(s,a)Q^{\pi}(s,a) (Bertsekas & Tsitsiklis, 1996). Q∗(s,a)=max⁡πQπ(s,a)Q^{*}(s,a)=\max_{\pi}Q^{\pi}(s,a) is known as the optimal value function, which has a corresponding optimal policy obtained through greedy action choices. For large or continuous state and action spaces, the value can be approximated with neural networks, e.g. using the Deep Q-Network algorithm (DQN) (Mnih et al., 2015). In DQN, the value function QθQ_{\theta} is updated using the target:

Q-learning is an off-policy algorithm (Sutton & Barto, 1998), meaning the target can be computed without consideration of how the experience was generated. In principle, off-policy reinforcement learning algorithms are able to learn from data collected by any behavioral policy. Typically, the loss is minimized over mini-batches of tuples of the agent’s past data, (s,a,r,s′)∈B(s,a,r,s^{\prime})\in\mathcal{B}, sampled from an experience replay dataset B\mathcal{B} (Lin, 1992). For shorthand, we often write s∈Bs\in\mathcal{B} if there exists a transition tuple containing ss in the batch B\mathcal{B}, and similarly for (s,a)(s,a) or (s,a,s′)∈B(s,a,s^{\prime})\in\mathcal{B}. In batch reinforcement learning, we assume B\mathcal{B} is fixed and no further interaction with the environment occurs. To further stabilize learning, a target network with frozen parameters Qθ′Q_{\theta^{\prime}}, is used in the learning target. The parameters of the target network θ′\theta^{\prime} are updated to the current network parameters θ\theta after a fixed number of time steps, or by averaging θ′←τθ+(1−τ)θ′\theta^{\prime}\leftarrow\tau\theta+(1-\tau)\theta^{\prime} for some small τ\tau (Lillicrap et al., 2015).

In a continuous action space, the analytic maximum of Equation (2) is intractable. In this case, actor-critic methods are commonly used, where action selection is performed through a separate policy network πϕ\pi_{\phi}, known as the actor, and updated with respect to a value estimate, known as the critic (Sutton & Barto, 1998; Konda & Tsitsiklis, 2003). This policy can be updated following the deterministic policy gradient theorem (Silver et al., 2014):

which corresponds to learning an approximation to the maximum of QθQ_{\theta}, by propagating the gradient through both πϕ\pi_{\phi} and QθQ_{\theta}. When combined with off-policy deep Q-learning to learn QθQ_{\theta}, this algorithm is referred to as Deep Deterministic Policy Gradients (DDPG) (Lillicrap et al., 2015).

Extrapolation Error

Extrapolation error is an error in off-policy value learning which is introduced by the mismatch between the dataset and true state-action visitation of the current policy. The value estimate Q(s,a)Q(s,a) is affected by extrapolation error during a value update where the target policy selects an unfamiliar action a′a^{\prime} at the next state s′s^{\prime} in the backed-up value estimate, such that (s′,a′)(s^{\prime},a^{\prime}) is unlikely, or not contained, in the dataset. The cause of extrapolation error can be attributed to several related factors:

Absent Data. If any state-action pair (s,a)(s,a) is unavailable, then error is introduced as some function of the amount of similar data and approximation error. This means that the estimate of Qθ(s′,π(s′))Q_{\theta}(s^{\prime},\pi(s^{\prime})) may be arbitrarily bad without sufficient data near (s′,π(s′))(s^{\prime},\pi(s^{\prime})).

Model Bias. When performing off-policy Q-learning with a batch B\mathcal{B}, the Bellman operator Tπ\mathcal{T}^{\pi} is approximated by sampling transitions tuples (s,a,r,s′)(s,a,r,s^{\prime}) from B\mathcal{B} to estimate the expectation over s′s^{\prime}. However, for a stochastic MDP, without infinite state-action visitation, this produces a biased estimate of the transition dynamics:

where the expectation is with respect to transitions in the batch B\mathcal{B}, rather than the true MDP.

Training Mismatch. Even with sufficient data, in deep Q-learning systems, transitions are sampled uniformly from the dataset, giving a loss weighted with respect to the likelihood of data in the batch:

If the distribution of data in the batch does not correspond with the distribution under the current policy, the value function may be a poor estimate of actions selected by the current policy, due to the mismatch in training.

We remark that re-weighting the loss in Equation (5) with respect to the likelihood under the current policy can still result in poor estimates if state-action pairs with high likelihood under the current policy are not found in the batch. This means only a subset of possible policies can be evaluated accurately. As a result, learning a value estimate with off-policy data can result in large amounts of extrapolation error if the policy selects actions which are not similar to the data found in the batch. In the following section, we discuss how state of the art off-policy deep reinforcement learning algorithms fail to address the concern of extrapolation error, and demonstrate the implications in practical examples.

Deep Q-learning algorithms (Mnih et al., 2015) have been labeled as off-policy due to their connection to off-policy Q-learning (Watkins, 1989). However, these algorithms tend to use near-on-policy exploratory policies, such as ϵ\epsilon-greedy, in conjunction with a replay buffer (Lin, 1992). As a result, the generated dataset tends to be heavily correlated to the current policy. In this section, we examine how these off-policy algorithms perform when learning with uncorrelated datasets. Our results demonstrate that the performance of a state of the art deep actor-critic algorithm, DDPG (Lillicrap et al., 2015), deteriorates rapidly when the data is uncorrelated and the value estimate produced by the deep Q-network diverges. These results suggest that off-policy deep reinforcement learning algorithms are ineffective when learning truly off-policy.

Our practical experiments examine three different batch settings in OpenAI gym’s Hopper-v1 environment (Todorov et al., 2012; Brockman et al., 2016), which we use to train an off-policy DDPG agent with no interaction with the environment. Experiments with additional environments and specific details can be found in the Supplementary Material.

Batch 1 (Final buffer). We train a DDPG agent for 1 million time steps, adding N(0,0.5)\mathcal{N}(0,0.5) Gaussian noise to actions for high exploration, and store all experienced transitions. This collection procedure creates a dataset with a diverse set of states and actions, with the aim of sufficient coverage.

Batch 2 (Concurrent). We concurrently train the off-policy and behavioral DDPG agents, for 1 million time steps. To ensure sufficient exploration, a standard N(0,0.1)\mathcal{N}(0,0.1) Gaussian noise is added to actions taken by the behavioral policy. Each transition experienced by the behavioral policy is stored in a buffer replay, which both agents learn from. As a result, both agents are trained with the identical dataset.

Batch 3 (Imitation). A trained DDPG agent acts as an expert, and is used to collect a dataset of 1 million transitions.

In Figure 1, we graph the performance of the agents as they train with each batch, as well as their value estimates. Straight lines represent the average return of episodes contained in the batch. Additionally, we graph the learning performance of the behavioral agent for the relevant tasks.

Our experiments demonstrate several surprising facts about off-policy deep reinforcement learning agents. In each task, the off-policy agent performances significantly worse than the behavioral agent. Even in the concurrent experiment, where both agents are trained with the same dataset, there is a large gap in performance in every single trial. This result suggests that differences in the state distribution under the initial policies is enough for extrapolation error to drastically offset the performance of the off-policy agent. Additionally, the corresponding value estimate exhibits divergent behavior, while the value estimate of the behavioral agent is highly stable. In the final buffer experiment, the off-policy agent is provided with a large and diverse dataset, with the aim of providing sufficient coverage of the initial policy. Even in this instance, the value estimate is highly unstable, and the performance suffers. In the imitation setting, the agent is provided with expert data. However, the agent quickly learns to take non-expert actions, under the guise of optimistic extrapolation. As a result, the value estimates rapidly diverge and the agent fails to learn.

Although extrapolation error is not necessarily positively biased, when combined with maximization in reinforcement learning algorithms, extrapolation error provides a source of noise that can induce a persistent overestimation bias (Thrun & Schwartz, 1993; Van Hasselt et al., 2016; Fujimoto et al., 2018). In an on-policy setting, extrapolation error may be a source of beneficial exploration through an implicit “optimism in the face of uncertainty” strategy (Lai & Robbins, 1985; Jaksch et al., 2010). In this case, if the value function overestimates an unknown state-action pair, the policy will collect data in the region of uncertainty, and the value estimate will be corrected. However, when learning off-policy, or in a batch setting, extrapolation error will never be corrected due to the inability to collect new data.

These experiments show extrapolation error can be highly detrimental to learning off-policy in a batch reinforcement learning setting. While the continuous state space and multi-dimensional action space in MuJoCo environments are contributing factors to extrapolation error, the scale of these tasks is small compared to real world settings. As a result, even with a sufficient amount of data collection, extrapolation error may still occur due to the concern of catastrophic forgetting (McCloskey & Cohen, 1989; Goodfellow et al., 2013). Consequently, off-policy reinforcement learning algorithms used in the real-world will require practical guarantees without exhaustive amounts of data.

Batch-Constrained Reinforcement Learning

Current off-policy deep reinforcement learning algorithms fail to address extrapolation error by selecting actions with respect to a learned value estimate, without consideration of the accuracy of the estimate. As a result, certain out-of-distribution actions can be erroneously extrapolated to higher values. However, the value of an off-policy agent can be accurately evaluated in regions where data is available. We propose a conceptually simple idea: to avoid extrapolation error a policy should induce a similar state-action visitation to the batch. We denote policies which satisfy this notion as batch-constrained. To optimize off-policy learning for a given batch, batch-constrained policies are trained to select actions with respect to three objectives:

Minimize the distance of selected actions to the data in the batch.

Lead to states where familiar data can be observed.

We note the importance of objective (1) above the others, as the value function and estimates of future states may be arbitrarily poor without access to the corresponding transitions. That is, we cannot correctly estimate (2) and (3) unless (1) is sufficiently satisfied. As a result, we propose optimizing the value function, along with some measure of future certainty, with a constraint limiting the distance of selected actions to the batch. This is achieved in our deep reinforcement learning algorithm through a state-conditioned generative model, to produce likely actions under the batch. This generative model is combined with a network which aims to optimally perturb the generated actions in a small range, along with a Q-network, used to select the highest valued action. Finally, we train a pair of Q-networks, and take the minimum of their estimates during the value update. This update penalizes states which are unfamiliar, and pushes the policy to select actions which lead to certain data.

We begin by analyzing the theoretical properties of batch-constrained policies in a finite MDP setting, where we are able to quantify extrapolation error precisely. We then introduce our deep reinforcement learning algorithm in detail, Batch-Constrained deep Q-learning (BCQ) by drawing inspiration from the tabular analogue.

In the finite MDP setting, extrapolation error can be described by the bias from the mismatch between the transitions contained in the buffer and the true MDP. We find that by inducing a data distribution that is contained entirely within the batch, batch-constrained policies can eliminate extrapolation error entirely for deterministic MDPs. In addition, we show that the batch-constrained variant of Q-learning converges to the optimal policy under the same conditions as the standard form of Q-learning. Moreover, we prove that for a deterministic MDP, batch-constrained Q-learning is guaranteed to match, or outperform, the behavioral policy when starting from any state contained in the batch. All of the proofs for this section can be found in the Supplementary Material.

A value estimate QQ can be learned using an experience replay buffer B\mathcal{B}. This involves sampling transition tuples (s,a,r,s′)(s,a,r,s^{\prime}) with uniform probability, and applying the temporal difference update (Sutton, 1988; Watkins, 1989):

If π(s′)=argmax⁡a′Q(s′,a′)\pi(s^{\prime})=\operatorname*{argmax}_{a^{\prime}}Q(s^{\prime},a^{\prime}), this is known as Q-learning. Assuming a non-zero probability of sampling any possible transition tuple from the buffer and infinite updates, Q-learning converges to the optimal value function.

Theorem 1. Performing Q-learning by sampling from a batch B\mathcal{B} converges to the optimal value function under the MDP MBM_{\mathcal{B}}.

We define ϵMDP\epsilon_{\text{MDP}} as the tabular extrapolation error, which accounts for the discrepancy between the value function QBπQ^{\pi}_{\mathcal{B}} computed with the batch B\mathcal{B} and the value function QπQ^{\pi} computed with the true MDP MM:

For any policy π\pi, the exact form of ϵMDP(s,a)\epsilon_{\text{MDP}}(s,a) can be computed through a Bellman-like equation:

This means extrapolation error is a function of divergence in the transition distributions, weighted by value, along with the error at succeeding states. If the policy is chosen carefully, the error between value functions can be minimized by visiting regions where the transition distributions are similar. For simplicity, we denote

To evaluate a policy π\pi exactly at relevant state-action pairs, only ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 is required. We can then determine the condition required to evaluate the exact expected return of a policy without extrapolation error.

Lemma 1. For all reward functions, ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 if and only if pB(s′∣s,a)=pM(s′∣s,a)p_{\mathcal{B}}(s^{\prime}|s,a)=p_{M}(s^{\prime}|s,a) for all s′∈Ss^{\prime}\in\mathcal{S} and (s,a)(s,a) such that μπ(s)>0\mu_{\pi}(s)>0 and π(a∣s)>0\pi(a|s)>0.

Lemma 1 states that if MBM_{\mathcal{B}} and MM exhibit the same transition probabilities in regions of relevance, the policy can be accurately evaluated. For a stochastic MDP this may require an infinite number of samples to converge to the true distribution, however, for a deterministic MDP this requires only a single transition. This means a policy which only traverses transitions contained in the batch, can be evaluated without error. More formally, we denote a policy π∈ΠB\pi\in\Pi_{\mathcal{B}} as batch-constrained if for all (s,a)(s,a) where μπ(s)>0\mu_{\pi}(s)>0 and π(a∣s)>0\pi(a|s)>0 then (s,a)∈B(s,a)\in\mathcal{B}. Additionally, we define a batch B\mathcal{B} as coherent if for all (s,a,s′)∈B(s,a,s^{\prime})\in\mathcal{B} then s′∈Bs^{\prime}\in\mathcal{B} unless s′s^{\prime} is a terminal state. This condition is trivially satisfied if the data is collected in trajectories, or if all possible states are contained in the batch. With a coherent batch, we can guarantee the existence of a batch-constrained policy.

Theorem 2. For a deterministic MDP and all reward functions, ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 if and only if the policy π\pi is batch-constrained. Furthermore, if B\mathcal{B} is coherent, then such a policy must exist if the start state s0∈Bs_{0}\in\mathcal{B}.

Batch-constrained policies can be used in conjunction with Q-learning to form batch-constrained Q-learning (BCQL), which follows the standard tabular Q-learning update while constraining the possible actions with respect to the batch:

BCQL converges under the same conditions as the standard form of Q-learning, noting the batch-constraint is nonrestrictive given infinite state-action visitation.

Theorem 3. Given the Robbins-Monro stochastic convergence conditions on the learning rate α\alpha, and standard sampling requirements from the environment, BCQL converges to the optimal value function Q∗Q^{*}.

The more interesting property of BCQL is that for a deterministic MDP and any coherent batch B\mathcal{B}, BCQL converges to the optimal batch-constrained policy π∗∈ΠB\pi^{*}\in\Pi_{\mathcal{B}} such that Qπ∗(s,a)≥Qπ(s,a)Q^{\pi^{*}}(s,a)\geq Q^{\pi}(s,a) for all π∈ΠB\pi\in\Pi_{\mathcal{B}} and (s,a)∈B(s,a)\in\mathcal{B}.

Theorem 4. Given a deterministic MDP and coherent batch B\mathcal{B}, along with the Robbins-Monro stochastic convergence conditions on the learning rate α\alpha and standard sampling requirements on the batch B\mathcal{B}, BCQL converges to QBπ(s,a)Q^{\pi}_{\mathcal{B}}(s,a) where π∗(s)=argmax⁡a s.t.(s,a)∈BQBπ(s,a)\pi^{*}(s)=\operatorname*{argmax}_{a\text{ s.t.}(s,a)\in\mathcal{B}}Q^{\pi}_{\mathcal{B}}(s,a) is the optimal batch-constrained policy.

This means that BCQL is guaranteed to outperform any behavioral policy when starting from any state contained in the batch, effectively outperforming imitation learning. Unlike standard Q-learning, there is no condition on state-action visitation, other than coherency in the batch.

2 Batch-Constrained Deep Reinforcement Learning

We introduce our approach to off-policy batch reinforcement learning, Batch-Constrained deep Q-learning (BCQ). BCQ approaches the notion of batch-constrained through a generative model. For a given state, BCQ generates plausible candidate actions with high similarity to the batch, and then selects the highest valued action through a learned Q-network. Furthermore, we bias this value estimate to penalize rare, or unseen, states through a modification to Clipped Double Q-learning (Fujimoto et al., 2018). As a result, BCQ learns a policy with a similar state-action visitation to the data in the batch, as inspired by the theoretical benefits of its tabular counterpart.

To maintain the notion of batch-constraint, we define a similarity metric by making the assumption that for a given state ss, the similarity between (s,a)(s,a) and the state-action pairs in the batch B\mathcal{B} can be modelled using a learned state-conditioned marginal likelihood PBG(a∣s)P_{\mathcal{B}}^{G}(a|s). In this case, it follows that the policy maximizing PBG(a∣s)P_{\mathcal{B}}^{G}(a|s) would minimize the error induced by extrapolation from distant, or unseen, state-action pairs, by only selecting the most likely actions in the batch with respect to a given state. Given the difficulty of estimating PBG(a∣s)P_{\mathcal{B}}^{G}(a|s) in high-dimensional continuous spaces, we instead train a parametric generative model of the batch Gω(s)G_{\omega}(s), which we can sample actions from, as a reasonable approximation to argmax⁡aPBG(a∣s)\operatorname*{argmax}_{a}P_{\mathcal{B}}^{G}(a|s).

For our generative model we use a conditional variational auto-encoder (VAE) (Kingma & Welling, 2013; Sohn et al., 2015), which models the distribution by transforming an underlying latent spaceSee the Supplementary Material for an introduction to VAEs.. The generative model GωG_{\omega}, alongside the value function QθQ_{\theta}, can be used as a policy by sampling nn actions from GωG_{\omega} and selecting the highest valued action according to the value estimate QθQ_{\theta}. To increase the diversity of seen actions, we introduce a perturbation model ξϕ(s,a,Φ)\xi_{\phi}(s,a,\Phi), which outputs an adjustment to an action aa in the range [−Φ,Φ][-\Phi,\Phi]. This enables access to actions in a constrained region, without having to sample from the generative model a prohibitive number of times. This results in the policy π\pi:

The choice of nn and Φ\Phi creates a trade-off between an imitation learning and reinforcement learning algorithm. If Φ=0\Phi=0, and the number of sampled actions n=1n=1, then the policy resembles behavioral cloning and as Φ→amax−amin\Phi\rightarrow a_{\text{max}}-a_{\text{min}} and n→∞n\rightarrow\infty, then the algorithm approaches Q-learning, as the policy begins to greedily maximize the value function over the entire action space.

The perturbation model ξϕ\xi_{\phi} can be trained to maximize Qθ(s,a)Q_{\theta}(s,a) through the deterministic policy gradient algorithm (Silver et al., 2014) by sampling a∼Gω(s)a\sim G_{\omega}(s):

To penalize uncertainty over future states, we modify Clipped Double Q-learning (Fujimoto et al., 2018), which estimates the value by taking the minimum between two Q-networks {Qθ1,Qθ2}\{Q_{\theta_{1}},Q_{\theta_{2}}\}. Although originally used as a countermeasure to overestimation bias (Thrun & Schwartz, 1993; Van Hasselt, 2010), the minimum operator also penalizes high variance estimates in regions of uncertainty, and pushes the policy to favor actions which lead to states contained in the batch. In particular, we take a convex combination of the two values, with a higher weight on the minimum, to form a learning target which is used by both Q-networks:

where aia_{i} corresponds to the perturbed actions, sampled from the generative model. If we set λ=1\lambda=1, this update corresponds to Clipped Double Q-learning. We use this weighted minimum as the constrained updates produces less overestimation bias than a purely greedy policy update, and enables control over how heavily uncertainty at future time steps is penalized through the choice of λ\lambda.

This forms Batch-Constrained deep Q-learning (BCQ), which maintains four parametrized networks: a generative model Gω(s)G_{\omega}(s), a perturbation model ξϕ(s,a)\xi_{\phi}(s,a), and two Q-networks Qθ1(s,a),Qθ2(s,a)Q_{\theta_{1}}(s,a),Q_{\theta_{2}}(s,a). We summarize BCQ in Algorithm 1. In the following section, we demonstrate BCQ results in stable value learning and a strong performance in the batch setting. Furthermore, we find that only a single choice of hyper-parameters is necessary for a wide range of tasks and environments.

Experiments

To evaluate the effectiveness of Batch-Constrained deep Q-learning (BCQ) in a high-dimensional setting, we focus on MuJoCo environments in OpenAI gym (Todorov et al., 2012; Brockman et al., 2016). For reproducibility, we make no modifications to the original environments or reward functions. We compare our method with DDPG (Lillicrap et al., 2015), DQN (Mnih et al., 2015) using an independently discretized action space, a feed-forward behavioral cloning method (BC), and a variant with a VAE (VAE-BC), using Gω(s)G_{\omega}(s) from BCQ. Exact implementation and experimental details are provided in the Supplementary Material.

We evaluate each method following the three experiments defined in Section 3.1. In final buffer the off-policy agents learn from the final replay buffer gathered by training a DDPG agent over a million time steps. In concurrent the off-policy agents learn concurrently, with the same replay buffer, as the behavioral DDPG policy, and in imitation, the agents learn from a dataset collected by an expert policy. Additionally, to study the robustness of BCQ to noisy and multi-modal data, we include an imperfect demonstrations task, in which the agents are trained with a batch of 100k transitions collected by an expert policy, with two sources of noise. The behavioral policy selects actions randomly with probability 0.30.3 and with high exploratory noise N(0,0.3)\mathcal{N}(0,0.3) added to the remaining actions. The experimental results for these tasks are reported in Figure 2. Furthermore, the estimated values of BCQ, DDPG and DQN, and the true value of BCQ are displayed in Figure 3.

Our approach, BCQ, is the only algorithm which succeeds at all tasks, matching or outperforming the behavioral policy in each instance, and outperforming all other agents, besides in the imitation learning task where behavioral cloning unsurprisingly performs the best. These results demonstrate that our algorithm can be used as a single approach for both imitation learning and off-policy reinforcement learning, with a single set of fixed hyper-parameters. Furthermore, unlike the deep reinforcement learning algorithms, DDPG and DQN, BCQ exhibits a highly stable value function in the presence of off-policy samples, suggesting extrapolation error has been successfully mitigated through the batch-constraint. In the imperfect demonstrations task, we find that both deep reinforcement learning and imitation learning algorithms perform poorly. BCQ, however, is able to strongly outperform the noisy demonstrator, disentangling poor and expert actions. Furthermore, compared to current deep reinforcement learning algorithms, which can require millions of time steps (Duan et al., 2016; Henderson et al., 2017), BCQ attains a high performance in remarkably few iterations. This suggests our approach effectively leverages expert transitions, even in the presence of noise.

Related Work

Batch Reinforcement Learning. While batch reinforcement learning algorithms have been shown to be convergent with non-parametric function approximators such as averagers (Gordon, 1995) and kernel methods (Ormoneit & Sen, 2002), they make no guarantees on the quality of the policy without infinite data. Other batch algorithms, such as fitted Q-iteration, have used other function approximators, including decision trees (Ernst et al., 2005) and neural networks (Riedmiller, 2005), but come without convergence guarantees. Unlike many previous approaches to off-policy policy evaluation (Peshkin & Shelton, 2002; Thomas et al., 2015; Liu et al., 2018), our work focuses on constraining the policy to a subset of policies which can be adequately evaluated, rather than the process of evaluation itself. Additionally, off-policy algorithms which rely on importance sampling (Precup et al., 2001; Jiang & Li, 2016; Munos et al., 2016) may not be applicable in a batch setting, requiring access to the action probabilities under the behavioral policy, and scale poorly to multi-dimensional action spaces. Reinforcement learning with a replay buffer (Lin, 1992) can be considered a form of batch reinforcement learning, and is a standard tool for off-policy deep reinforcement learning algorithms (Mnih et al., 2015). It has been observed that a large replay buffer can be detrimental to performance (de Bruin et al., 2015; Zhang & Sutton, 2017) and the diversity of states in the buffer is an important factor for performance (de Bruin et al., 2016). Isele & Cosgun (2018) observed the performance of an agent was strongest when the distribution of data in the replay buffer matched the test distribution. These results defend the notion that extrapolation error is an important factor in the performance off-policy reinforcement learning.

Imitation Learning. Imitation learning and its variants are well studied problems (Schaal, 1999; Argall et al., 2009; Hussein et al., 2017). Imitation has been combined with reinforcement learning, via learning from demonstrations methods (Kim et al., 2013; Piot et al., 2014; Chemali & Lazaric, 2015), with deep reinforcement learning extensions (Hester et al., 2017; Večerík et al., 2017), and modified policy gradient approaches (Ho et al., 2016; Sun et al., 2017; Cheng et al., 2018; Sun et al., 2018). While effective, these interactive methods are inadequate for batch reinforcement learning as they require either an explicit distinction between expert and non-expert data, further on-policy data collection or access to an oracle. Research in imitation, and inverse reinforcement learning, with robustness to noise is an emerging area (Evans, 2016; Nair et al., 2018), but relies on some form of expert data. Gao et al. (2018) introduced an imitation learning algorithm which learned from imperfect demonstrations, by favoring seen actions, but is limited to discrete actions. Our work also connects to residual policy learning (Johannink et al., 2018; Silver et al., 2018), where the initial policy is the generative model, rather than an expert or feedback controller.

Uncertainty in Reinforcement Learning. Uncertainty estimates in deep reinforcement learning have generally been used to encourage exploration (Dearden et al., 1998; Strehl & Littman, 2008; O’Donoghue et al., 2018; Azizzadenesheli et al., 2018). Other methods have examined approximating the Bayesian posterior of the value function (Osband et al., 2016, 2018; Touati et al., 2018), again using the variance to encourage exploration to unseen regions of the state space. In model-based reinforcement learning, uncertainty has been used for exploration, but also for the opposite effect–to push the policy towards regions of certainty in the model. This is used to combat the well-known problems with compounding model errors, and is present in policy search methods (Deisenroth & Rasmussen, 2011; Gal et al., 2016; Higuera et al., 2018; Xu et al., 2018), or combined with trajectory optimization (Chua et al., 2018) or value-based methods (Buckman et al., 2018). Our work connects to policy methods with conservative updates (Kakade & Langford, 2002), such as trust region (Schulman et al., 2015; Achiam et al., 2017; Pham et al., 2018) and information-theoretic methods (Peters & Mülling, 2010; Van Hoof et al., 2017), which aim to keep the updated policy similar to the previous policy. These methods avoid explicit uncertainty estimates, and rather force policy updates into a constrained range before collecting new data, limiting errors introduced by large changes in the policy. Similarly, our approach can be thought of as an off-policy variant, where the policy aims to be kept close, in output space, to any combination of the previous policies which performed data collection.

Conclusion

In this work, we demonstrate a critical problem in off-policy reinforcement learning with finite data, where the value target introduces error by including an estimate of unseen state-action pairs. This phenomenon, which we denote extrapolation error, has important implications for off-policy and batch reinforcement learning, as it is generally implausible to have complete state-action coverage in any practical setting. We present batch-constrained reinforcement learning–acting close to on-policy with respect to the available data, as an answer to extrapolation error. When extended to a deep reinforcement learning setting, our algorithm, Batch-Constrained deep Q-learning (BCQ), is the first continuous control algorithm capable of learning from arbitrary batch data, without exploration. Due to the importance of batch reinforcement learning for practical applications, we believe BCQ will be a strong foothold for future algorithms to build on, while furthering our understanding of the systematic risks in Q-learning (Thrun & Schwartz, 1993; Lu et al., 2018).

References

Appendix A Missing Proofs

Definition 1. We define a coherent batch B\mathcal{B} as a batch such that if (s,a,s′)∈B(s,a,s^{\prime})\in\mathcal{B} then s′∈Bs^{\prime}\in\mathcal{B} unless s′s^{\prime} is a terminal state.

Definition 2. We define ϵMDP(s,a)=Qπ(s,a)−QBπ(s,a)\epsilon_{\text{MDP}}(s,a)=Q^{\pi}(s,a)-Q^{\pi}_{\mathcal{B}}(s,a) as the error between the true value of a policy π\pi in the MDP MM and the value of π\pi when learned with a batch B\mathcal{B}.

Definition 3. For simplicity in notation, we denote

To evaluate a policy π\pi exactly at relevant state-action pairs, only ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 is required.

Definition 4. We define the optimal batch-constrained policy π∗∈ΠB\pi^{*}\in\Pi_{\mathcal{B}} such that Qπ∗(s,a)≥Qπ(s,a)Q^{\pi^{*}}(s,a)\geq Q^{\pi}(s,a) for all π∈ΠB\pi\in\Pi_{\mathcal{B}} and (s,a)∈B(s,a)\in\mathcal{B}.

Algorithm 1. Batch-Constrained Q-learning (BCQL) maintains a tabular value function Q(s,a)Q(s,a) for each possible state-action pair (s,a)(s,a). A transition tuple (s,a,r,s′)(s,a,r,s^{\prime}) is sampled from the batch B\mathcal{B} with uniform probability and the following update rule is applied, with learning rate α\alpha:

Theorem 1. Performing Q-learning by sampling from a batch B\mathcal{B} converges to the optimal value function under the MDP MBM_{\mathcal{B}}.

Remark 1. For any policy π\pi and state-action pair (s,a)(s,a), the error term ϵMDP(s,a)\epsilon_{\text{MDP}}(s,a) satisfies the following Bellman-like equation:

Proof. Proof follows by expanding each QQ, rearranging terms and then simplifying the expression.

Lemma 1. For all reward functions, ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 if and only if pB(s′∣s,a)=pM(s′∣s,a)p_{\mathcal{B}}(s^{\prime}|s,a)=p_{M}(s^{\prime}|s,a) for all s′∈Ss^{\prime}\in\mathcal{S} and (s,a)(s,a) such that μπ(s)>0\mu_{\pi}(s)>0 and π(a∣s)>0\pi(a|s)>0.

Proof. From Remark 1, we note that the form of ϵMDP(s,a)\epsilon_{\text{MDP}}(s,a), since no assumptions can be made on the reward function and therefore the expression r(s,a,s′)+γ∑a′π(a′∣s′)QBπ(s′,a′)r(s,a,s^{\prime})+\gamma\sum_{a^{\prime}}\pi(a^{\prime}|s^{\prime})Q^{\pi}_{\mathcal{B}}(s^{\prime},a^{\prime}), we have that ϵMDP(s,a)=0\epsilon_{\text{MDP}}(s,a)=0 if and only if pB(s′∣s,a)=pM(s′∣s,a)p_{\mathcal{B}}(s^{\prime}|s,a)=p_{M}(s^{\prime}|s,a) for all s′∈Ss^{\prime}\in\mathcal{S} and pM(s′∣s,a)γ∑a′π(a′∣s′)ϵMDP(s′,a′)=0p_{M}(s^{\prime}|s,a)\gamma\sum_{a^{\prime}}\pi(a^{\prime}|s^{\prime})\epsilon_{\text{MDP}}(s^{\prime},a^{\prime})=0.

(⇒)(\Rightarrow) Now we note that if ϵMDP(s,a)=0\epsilon_{\text{MDP}}(s,a)=0 then pM(s′∣s,a)γ∑a′π(a′∣s′)ϵMDP(s′,a′)=0p_{M}(s^{\prime}|s,a)\gamma\sum_{a^{\prime}}\pi(a^{\prime}|s^{\prime})\epsilon_{\text{MDP}}(s^{\prime},a^{\prime})=0 by the relationship defined by Remark 1 and the condition on the reward function. It follows that we must have pB(s′∣s,a)=pM(s′∣s,a)p_{\mathcal{B}}(s^{\prime}|s,a)=p_{M}(s^{\prime}|s,a) for all s′∈Ss^{\prime}\in\mathcal{S}.

(⇐)(\Leftarrow) If we have ∑s′∣pM(s′∣s,a)−pB(s′∣s,a)∣=0\sum_{s^{\prime}}|p_{M}(s^{\prime}|s,a)-p_{\mathcal{B}}(s^{\prime}|s,a)|=0 for all (s,a)(s,a) such that μπ(s)>0\mu_{\pi}(s)>0 and π(a∣s)>0\pi(a|s)>0, then for any (s,a)(s,a) under the given conditions, we have ϵ(s,a)=∑s′pM(s′∣s,a)γ∑a′π(a′∣s′)ϵ(s′,a′)\epsilon(s,a)=\sum_{s^{\prime}}p_{M}(s^{\prime}|s,a)\gamma\sum_{a^{\prime}}\pi(a^{\prime}|s^{\prime})\epsilon(s^{\prime},a^{\prime}). Recursively expanding the ϵ\epsilon term, we arrive at ϵ(s,a)=0+γ0+γ20+...=0\epsilon(s,a)=0+\gamma 0+\gamma^{2}0+...=0.

Theorem 2. For a deterministic MDP and all reward functions, ϵMDPπ=0\epsilon^{\pi}_{\text{MDP}}=0 if and only if the policy π\pi is batch-constrained. Furthermore, if B\mathcal{B} is coherent, then such a policy must exist if the start state s0∈Bs_{0}\in\mathcal{B}.

Proof. The first part of the Theorem follows from Lemma 1, noting that for a deterministic policy π\pi, if (s,a)∈B(s,a)\in\mathcal{B} then we must have pB(s′∣s,a)=pM(s′∣s,a)p_{\mathcal{B}}(s^{\prime}|s,a)=p_{M}(s^{\prime}|s,a) for all s′∈Ss^{\prime}\in\mathcal{S}.

We can construct the batch-constrained policy by selecting aa in the state s∈Bs\in\mathcal{B}, such that (s,a)∈B(s,a)\in\mathcal{B}. Since the MDP is deterministic and the batch is coherent, when starting from s0s_{0}, we must be able to follow at least one trajectory until termination.

Theorem 3. Given the Robbins-Monro stochastic convergence conditions on the learning rate α\alpha, and standard sampling requirements from the environment, BCQL converges to the optimal value function Q∗Q^{*}.

Proof. Follows from proof of convergence of Q-learning (see Section A.2), noting the batch-constraint is non-restrictive with a batch which contains all possible transitions.

Theorem 4. Given a deterministic MDP and coherent batch B\mathcal{B}, along with the Robbins-Monro stochastic convergence conditions on the learning rate α\alpha and standard sampling requirements on the batch B\mathcal{B}, BCQL converges to QBπ(s,a)Q^{\pi}_{\mathcal{B}}(s,a) where π∗(s)=argmax⁡a s.t.(s,a)∈BQBπ(s,a)\pi^{*}(s)=\operatorname*{argmax}_{a\text{ s.t.}(s,a)\in\mathcal{B}}Q^{\pi}_{\mathcal{B}}(s,a) is the optimal batch-constrained policy.

Proof. Results follows from Theorem 1, which states Q-learning learns the optimal value for the MDP MBM_{\mathcal{B}} for state-action pairs in (s,a)(s,a). However, for a deterministic MDP MBM_{\mathcal{B}} corresponds to the true MDP in all seen state-action pairs. Noting that batch-constrained policies operate only on state-action pairs where MBM_{\mathcal{B}} corresponds to the true MDP, it follows that π∗\pi^{*} will be the optimal batch-constrained policy from the optimality of Q-learning.

A.2 Sketch of the Proof of Convergence of Q-Learning

The proof of convergence of Q-learning relies large on the following lemma (Singh et al., 2000):

where xt∈Xx_{t}\in X and t=0,1,2,...t=0,1,2,.... Let PtP_{t} be a sequence of increasing σ\sigma-fields such that ζ0\zeta_{0} and Δ0\Delta_{0} are P0P_{0}-measurable and ζt,Δt\zeta_{t},\Delta_{t} and Ft−1F_{t-1} are PtP_{t}-measurable, t=1,2,...t=1,2,.... Assume that the following hold:

ζt(xt)∈,∑tζt(xt)=∞,∑t(ζt(xt))2<∞\zeta_{t}(x_{t})\in,\sum_{t}\zeta_{t}(x_{t})=\infty,\sum_{t}(\zeta_{t}(x_{t}))^{2}<\infty with probability 11 and ∀x≠xt:ζ(x)=0\forall x\neq x_{t}:\zeta(x)=0.

Var[Ft(xt)∣Pt]≤K(1+κ∣∣Δt∣∣)2\left[F_{t}(x_{t})|P_{t}\right]\leq K(1+\kappa||\Delta_{t}||)^{2}, where KK is some constant

Where ∣∣⋅∣∣||\cdot|| denotes the maximum norm. Then Δt\Delta_{t} converges to with probability 11.

Sketch of Proof of Convergence of Q-Learning. We set Δt=Qt(s,a)−Q∗(s,a)\Delta_{t}=Q_{t}(s,a)-Q^{*}(s,a). Then convergence follows by satisfying the conditions of Lemma 2. Condition 1 is satisfied by the finite MDP, setting X=S×AX=\mathcal{S}\times\mathcal{A}. Condition 2 is satisfied by the assumption of Robbins-Monro stochastic convergence conditions on the learning rate αt\alpha_{t}, setting ζt=αt\zeta_{t}=\alpha_{t}. Condition 4 is satisfied by the bounded reward function, where Ft(s,a)=r(s,a,s′)+γmax⁡a′Q(s′,a′)−Q∗(s,a)F_{t}(s,a)=r(s,a,s^{\prime})+\gamma\max_{a^{\prime}}Q(s^{\prime},a^{\prime})-Q^{*}(s,a), and the sequence Pt={Q0,s0,a0,α0,r1,s1,...st,at}P_{t}=\{Q_{0},s_{0},a_{0},\alpha_{0},r_{1},s_{1},...s_{t},a_{t}\}. Finally, Condition 3 follows from the contraction of the Bellman Operator T\mathcal{T}, requiring infinite state-action visitation, infinite updates and γ<1\gamma<1.

Additional and more complete details can be found in numerous resources (Dayan & Watkins, 1992; Singh et al., 2000; Melo, 2001).

Appendix B Missing Graphs

B.2 Complete Experimental Results

We present the complete set of results across each task and environment in Figure 5. These results show that BCQ successful mitigates extrapolation error and learns from a variety of fixed batch settings. Although BCQ with our current hyper-parameters was never found to fail, we noted with slight changes to hyper-parameters, BCQ failed periodically on the concurrent learning task in the HalfCheetah-v1 environment, exhibiting instability in the value function after 750,000750,000 or more iterations on some seeds. We hypothesize that this instability could occur if the generative model failed to output in-distribution actions, and could be corrected through additional training or improvements to the vanilla VAE. Interestingly, BCQ still performs well in these instances, due to the behavioral cloning-like elements in the algorithm.

Appendix C Extrapolation Error in Kernel-Based Reinforcement Learning

where sBa∈Ss_{\mathcal{B}}^{a}\in\mathcal{S} represents states corresponding to the action aa for some tuple (sB,a)∈B(s_{\mathcal{B}},a)\in\mathcal{B}, and V(sB′)=max⁡a s.t.(sB′,a)∈BQ(sB′,a)V(s_{\mathcal{B}}^{\prime})=\max_{a\text{ s.t.}(s_{\mathcal{B}}^{\prime},a)\in\mathcal{B}}Q(s_{\mathcal{B}}^{\prime},a). At each iteration, KBRL updates the estimates of Q(sB,aB)Q(s_{\mathcal{B}},a_{\mathcal{B}}) for all (sB,aB)∈B(s_{\mathcal{B}},a_{\mathcal{B}})\in\mathcal{B} following Equation (19), then updates V(sB′)V(s_{\mathcal{B}}^{\prime}) by evaluating Q(sB′,a)Q(s_{\mathcal{B}}^{\prime},a) for all sB∈Bs_{\mathcal{B}}\in\mathcal{B} and a∈Aa\in\mathcal{A}.

Given access to the entire deterministic MDP, KBRL will provable converge to the optimal value, however when limited to only a subset, we find the value estimation susceptible to extrapolation. In Figure 6, we provide a deterministic two state, two action MDP in which KBRL fails to learn the optimal policy when provided with state-action pairs from the optimal policy. Given the batch {(s0,a1,r=1,s1),(s1,a0,r=0,s0)}\{(s_{0},a_{1},r=1,s_{1}),(s_{1},a_{0},r=0,s_{0})\}, corresponding to the optimal behavior, and noting that there is only one example of each action, Equation (19) provides the following:

After sufficient iterations KBRL will converge correctly to Q(s0,a1)=11−γ2Q(s_{0},a_{1})=\frac{1}{1-\gamma^{2}}, Q(s1,a0)=γ1−γ2Q(s_{1},a_{0})=\frac{\gamma}{1-\gamma^{2}}. However, when evaluating actions, KBRL erroneously extrapolates the values of each action Q(⋅,a1)=11−γ2Q(\cdot,a_{1})=\frac{1}{1-\gamma^{2}}, Q(⋅,a0)=γ1−γ2Q(\cdot,a_{0})=\frac{\gamma}{1-\gamma^{2}}, and its behavior, argmax⁡aQ(s,a)\operatorname*{argmax}_{a}Q(s,a), will result in the degenerate policy of continually selecting a1a_{1}. KBRL fails this example by estimating the values of unseen state-action pairs. In methods where the extrapolated estimates can be included into the learning update, such fitted Q-iteration or DQN (Ernst et al., 2005; Mnih et al., 2015), this could cause an unbounded sequence of value estimates, as demonstrated by our results in Section 3.1.

Appendix D Additional Experiments

BCQ includes a perturbation model ξθ(s,a,Φ)\xi_{\theta}(s,a,\Phi) which outputs a small residual update to the actions sampled by the generative model in the range [−Φ,Φ][-\Phi,\Phi]. This enables the policy to select actions which may not have been sampled by the generative model. If Φ=amax−amin\Phi=a_{\text{max}}-a_{\text{min}}, then all actions can be plausibly selected by the model, similar to standard deep reinforcement learning algorithms, such as DQN and DDPG (Mnih et al., 2015; Lillicrap et al., 2015). In Figure 7 we examine the performance and value estimates of BCQ when varying the hyper-parameter Φ\Phi, which corresponds to how much the model is able to move away from the actions sampled by the generative model.

We observe a clear drop in performance with the increase of Φ\Phi, along with an increase in instability in the value function. Given that the data is based on expert performance, this is consistent with our understanding of extrapolation error. With larger Φ\Phi the agent learns to take actions that are further away from the data in the batch after erroneously overestimating the value of suboptimal actions. This suggests the ideal value of Φ\Phi should be small enough to stay close to the generated actions, but large enough such that learning can be performed when exploratory actions are included in the dataset.

D.2 Uncertainty Estimation for Batch-Constrained Reinforcement Learning

Section 4.2 proposes using generation as a method for constraining the output of the policy π\pi to eliminate actions which are unlikely under the batch. However, a more natural approach would be through approximate uncertainty-based methods (Osband et al., 2016; Gal et al., 2016; Azizzadenesheli et al., 2018). These methods are well-known to be effective for exploration, however we examine their properties for the exploitation where we would like to avoid uncertain actions.

To measure the uncertainty of the value network, we use ensemble-based methods, with an ensemble of size 44 and 1010 to mimic the models used by Buckman et al. (2018) and Osband et al. (2016) respectively. Each network is trained with separate mini-batches, following the standard deep Q-learning update with a target network. These networks use the default architecture and hyper-parameter choices as defined in Section G. The policy πϕ\pi_{\phi} is trained to minimize the standard deviation σ\sigma across the ensemble:

If the ensembles were a perfect estimate of the uncertainty, the policy would learn to select the most certain action for a given state, minimizing the extrapolation error and effectively imitating the data in the batch.

To test these uncertainty-based methods, we examine their performance on the imitation task in the Hopper-v1 environment (Todorov et al., 2012; Brockman et al., 2016). In which a dataset of 1 million expert transitions are provided to the agents. Additional experimental details can be found in Section F. The performance, alongside the value estimates of the agents are displayed in Figure 8.

We find that neither ensemble method is sufficient to constrain the action space to only the expert actions. However, the value function is stabilized, suggesting that ensembles are an effective strategy for eliminating outliers or large deviations in the value from erroneous extrapolation. Unsurprisingly, the large ensemble provides a more accurate estimate of the uncertainty. While scaling the size of the ensemble to larger values could possibly enable an effective batch-constraint, increasing the size of the ensemble induces a large computational cost. Finally, in this task, where only expert data is provided, the policy can attempt to imitate the data without consideration of the value, however in other tasks, a weighting between value and uncertainty would need to be carefully tuned. On the other hand, BCQ offers a computational cheap approach without requirements for difficult hyper-parameter tuning.

D.3 Random Behavioral Policy Study

The experiments in Section 3.1 and 5 use a learned, or partially learned, behavioral policy for data collection. This is a necessary requirement for learning meaningful behavior, as a random policy generally fails to provide sufficient coverage over the state space. However, in simple toy problems, such as the pendulum swing-up task and the reaching task with a two-joint arm from OpenAI gym (Brockman et al., 2016), a random policy can sufficiently explore the environment, enabling us to examine the properties of algorithms with entirely non-expert data.

In Figure 9, we examine the performance of our algorithm, BCQ, as well as DDPG (Lillicrap et al., 2015), on these two toy problems, when learning off-policy from a small batch of 5000 time steps, collected entirely by a random policy. We find that both BCQ and DDPG are able to learn successfully in these off-policy tasks. These results suggest that BCQ is less restrictive than imitation learning algorithms, which require expert data to learn. We also find that unlike previous environments, given the small scale of the state and action space, the random policy is able to provide sufficient coverage for DDPG to learn successfully.

Appendix E Missing Background

A variational auto-encoder (VAE) (Kingma & Welling, 2013) is a generative model which aims to maximize the marginal log-likelihood log⁡p(X)=∑i=1Nlog⁡p(xi)\log p(X)=\sum_{i=1}^{N}\log p(x_{i}) where X={x1,...,xN}X=\{x_{1},...,x_{N}\}, the dataset. While computing the marginal likelihood is intractable in nature, we can instead optimize the variational lower-bound:

where p(z)p(z) is chosen a prior, generally the multivariate normal distribution N(0,I)\mathcal{N}(0,I). We define the posterior q(z∣X)=N(z∣μ(X),σ2(X)I)q(z|X)=\mathcal{N}(z|\mu(X),\sigma^{2}(X)I) as the encoder and p(X∣z)p(X|z) as the decoder. Simply put, this means a VAE is an auto-encoder, where a given sample xx is passed through the encoder to produce a random latent vector zz, which is given to the decoder to reconstruct the original sample xx. The VAE is trained on a reconstruction loss, and a KL-divergence term according to the distribution of the latent vectors. To perform gradient descent on the variational lower bound we can use the re-parametrization trick (Kingma & Welling, 2013; Rezende et al., 2014):

This formulation allows for back-propagation through stochastic nodes, by noting μ\mu and σ\sigma can be represented by deterministic functions. During inference, random values of zz are sampled from the multivariate normal and passed through the decoder to produce samples xx.

Appendix F Experimental Details

Each environment is run for 1 million time steps, unless stated otherwise, with evaluations every 5000 time steps, where an evaluation measures the average reward from 10 episodes with no exploration noise. Our results are reported over 5 random seeds of the behavioral policy, OpenAI Gym simulator and network initialization. Value estimates are averaged over mini-batches of 100100 and sampled every 25002500 iterations. The true value is estimated by sampling 100100 state-action pairs from the buffer replay and computing the discounted return by running the episode until completion while following the current policy.

Each agent is trained after each episode by applying one training iteration per each time step in the episode. The agent is trained with transition tuples (s,a,r,s′)(s,a,r,s^{\prime}) sampled from an experience replay that is defined by each experiment. We define four possible experiments. Unless stated otherwise, default implementation settings, as defined in Section G, are used.

Batch 1 (Final buffer). We train a DDPG (Lillicrap et al., 2015) agent for 1 million time steps, adding large amounts of Gaussian noise (N(0,0.5))(\mathcal{N}(0,0.5)) to induce exploration, and store all experienced transitions in a buffer replay. This training procedure creates a buffer replay with a diverse set of states and actions. A second, randomly initialized agent is trained using the 1 million stored transitions.

Batch 2 (Concurrent learning). We simultaneously train two agents for 1 million time steps, the first DDPG agent, performs data collection and each transition is stored in a buffer replay which both agents learn from. This means the behavioral agent learns from the standard training regime for most off-policy deep reinforcement learning algorithms, and the second agent is learning off-policy, as the data is collected without direct relationship to its current policy. Batch 2 differs from Batch 1 as the agents are trained with the same version of the buffer, while in Batch 1 the agent learns from the final buffer replay after the behavioral agent has finished training.

Batch 3 (Imitation). A DDPG agent is trained for 1 million time steps. The trained agent then acts as an expert policy, and is used to collect a dataset of 1 million transitions. This dataset is used to train a second, randomly initialized agent. In particular, we train DDPG across 15 seeds, and select the 5 top performing seeds as the expert policies.

Batch 4 (Imperfect demonstrations). The expert policies from Batch 3 are used to collect a dataset of 100k transitions, while selecting actions randomly with probability 0.30.3 and adding Gaussian noise N(0,0.3)\mathcal{N}(0,0.3) to the remaining actions. This dataset is used to train a second, randomly initialized agent.

Appendix G Implementation Details

Across all methods and experiments, for fair comparison, each network generally uses the same hyper-parameters and architecture, which are defined in Table 1 and Figure 10 respectively. The value functions follow the standard practice (Mnih et al., 2015) in which the Bellman update differs for terminal transitions. When the episode ends by reaching some terminal state, the value is set to in the learning target yy:

Where the termination signal from time-limited environments is ignored, thus we only consider a state sts_{t} terminal if t<t< max horizon.

BCQ. BCQ uses four main networks: a perturbation model ξϕ(s,a)\xi_{\phi}(s,a), a state-conditioned VAE Gω(s)G_{\omega}(s) and a pair of value networks Qθ1(s,a),Qθ2(s,a)Q_{\theta_{1}}(s,a),Q_{\theta_{2}}(s,a). Along with three corresponding target networks ξϕ′(s,a),Qθ1′(s,a),Qθ2′(s,a)\xi_{\phi^{\prime}}(s,a),Q_{\theta^{\prime}_{1}}(s,a),Q_{\theta^{\prime}_{2}}(s,a). Each network, other than the VAE follows the default architecture (Figure 10) and the default hyper-parameters (Table 1). For ξϕ(s,a,Φ)\xi_{\phi}(s,a,\Phi), the constraint Φ\Phi is implemented through a tanh activation multiplied by I⋅ΦI\cdot\Phi following the final layer.

As suggested by Fujimoto et al. (2018), the perturbation model is trained only with respect to Qθ1Q_{\theta_{1}}, following the deterministic policy gradient algorithm (Silver et al., 2014), by performing a residual update on a single action sampled from the generative model:

To penalize uncertainty over future states, we train a pair of value estimates {Qθ1,Qθ2}\{Q_{\theta_{1}},Q_{\theta_{2}}\} and take a weighted minimum between the two values as a learning target yy for both networks. First nn actions are sampled with respect to the generative model, and then adjusted by the target perturbation model, before passed to each target Q-network:

Both networks are trained with the same target yy, where λ=0.75\lambda=0.75.

The VAE GωG_{\omega} is defined by two networks, an encoder Eω1(s,a)E_{\omega_{1}}(s,a) and decoder Dω2(s,z)D_{\omega_{2}}(s,z), where ω={ω1,ω2}\omega=\{\omega_{1},\omega_{2}\}. The encoder takes a state-action pair and outputs the mean μ\mu and standard deviation σ\sigma of a Gaussian distribution N(μ,σ)\mathcal{N}(\mu,\sigma). The state ss, along with a latent vector zz is sampled from the Gaussian, is passed to the decoder Dω2(s,z)D_{\omega_{2}}(s,z) which outputs an action. Each network follows the default architecture (Figure 10), with two hidden layers of size 750750, rather than 400400 and 300300. The VAE is trained with respect to the mean squared error of the reconstruction along with a KL regularization term:

Noting the Gaussian form of both distributions, the KL divergence term can be simplified (Kingma & Welling, 2013):

where JJ denotes the dimensionality of zz. For each experiment, JJ is set to twice the dimensionality of the action space. The KL divergence term in LVAE\mathcal{L}_{\text{VAE}} is normalized across experiments by setting λ=12J\lambda=\frac{1}{2J}. During inference with the VAE, the latent vector zz is clipped to a range of [−0.5,0.5][-0.5,0.5] to limit generalization beyond previously seen actions. For the small scale experiments in Supplementary Material D.3, L2L_{2} regularization with weight 10−310^{-3} was used for the VAE to compensate for the small number of samples. The other networks remain unchanged. In the value network update (Equation 27), the VAE is sampled multiple times and passed to both the perturbation model and each Q-network. For each state in the mini-batch the VAE is sampled n=10n=10 times. This can be implemented efficiently by passing a latent vector with batch size 1010 ⋅\cdot batch size, effectively 10001000, to the VAE and treating the output as a new mini-batch for the perturbation model and each Q-network. When running the agent in the environment, we sample from the VAE 1010 times, perturb each action with ξϕ\xi_{\phi} and sample the highest valued action with respect to Qθ1Q_{\theta_{1}}.

DDPG. Our DDPG implementation deviates from some of the default architecture and hyper-parameters to mimic the original implementation more closely (Lillicrap et al., 2015). In particular, the action is only passed to the critic at the second layer (Figure 11), the critic uses L2L_{2} regularization with weight 10−210^{-2}, and the actor uses a reduced learning rate of 10−410^{-4}.

As done by Fujimoto et al. (2018), our DDPG agent randomly selects actions for the first 10k time steps for HalfCheetah-v1, and 1k time steps for Hopper-v1 and Walker2d-v1. This was found to improve performance and reduce the likelihood of local minima in the policy during early iterations.

DQN. Given the high dimensional nature of the action space of the experimental environments, our DQN implementation selects actions over an independently discretized action space. Each action dimension is discretized separately into 1010 possible actions, giving 10J10J possible actions, where JJ is the dimensionality of the action-space. A given state-action pair (s,a)(s,a) then corresponds to a set of state-sub-action pairs (s,aij)(s,a_{ij}), where i∈{1,...,10}i\in\{1,...,10\} bins and j={1,...,J}j=\{1,...,J\} dimensions. In each DQN update, all state-sub-action pairs (s,aij)(s,a_{ij}) are updated with respect to the average value of the target state-sub-action pairs (s′,aij′)(s^{\prime},a^{\prime}_{ij}). The learning update of the discretized DQN is as follows:

Where aija_{ij} is chosen by selecting the corresponding bin ii in the discretized action space for each dimension jj. For clarity, we provide the exact DQN network architecture in Figure 12.

Behavioral Cloning. We use two behavioral cloning methods, VAE-BC and BC. VAE-BC is implemented and trained exactly as Gω(s)G_{\omega}(s) defined for BCQ. BC uses a feed-forward network with the default architecture and hyper-parameters, and trained with a mean-squared error reconstruction loss.