Is Conditional Generative Modeling all you need for Decision-Making?

Anurag Ajay, Yilun Du, Abhi Gupta, Joshua Tenenbaum, Tommi Jaakkola, Pulkit Agrawal

Introduction

Over the last few years, conditional generative modeling has yielded impressive results in a range of domains, including high-resolution image generation from text descriptions (DALL-E, ImageGen) (Ramesh et al., 2022; Saharia et al., 2022), language generation (GPT) (Brown et al., 2020), and step-by-step solutions to math problems (Minerva) (Lewkowycz et al., 2022). The success of generative models in countless domains motivates us to apply them to decision-making.

Conveniently, there exists a wide body of research on recovering high-performing policies from data logged by already operational systems (Kostrikov et al., 2022; Kumar et al., 2020; Walke et al., 2022). This is particularly useful in real-world settings where interacting with the environment is not always possible, and exploratory decisions can have fatal consequences (Dulac-Arnold et al., 2021). With access to such offline datasets, the problem of decision-making reduces to learning a probabilistic model of trajectories, a setting where generative models have already found success.

In offline decision-making, we aim to recover optimal reward-maximizing trajectories by stitching together sub-optimal reward-labeled trajectories in the training dataset. Prior works (Kumar et al., 2020; Kostrikov et al., 2022; Wu et al., 2019; Kostrikov et al., 2021; Dadashi et al., 2021; Ajay et al., 2020; Ghosh et al., 2022) have tackled this problem with reinforcement learning (RL) that uses dynamic programming for trajectory stitching. To enable dynamic programming, these works learn a value function that estimates the discounted sum of rewards from a given state. However, value function estimation is prone to instabilities due to function approximation, off-policy learning, and bootstrapping together, together known as the deadly triad (Sutton & Barto, 2018). Furthermore, to stabilize value estimation in offline regime, these works rely on heuristics to keep the policy within the dataset distribution. These challenges make it difficult to scale existing offline RL algorithms.

In this paper, we ask if we can perform dynamic programming to stitch together sub-optimal trajectories to obtain an optimal trajectory without relying on value estimation. Since conditional diffusion generative models can generate novel data points by composing training data (Saharia et al., 2022; Ramesh et al., 2022), we leverage it for trajectory stitching in offline decision-making. Given a fixed dataset of reward-labeled trajectories, we adapt diffusion models (Sohl-Dickstein et al., 2015) to learn a return-conditional model of the trajectory. During inference, we use classifier-free guidance with low-temperature sampling, which we hypothesize to implicitly perform dynamics programming to capture the best behaviors in the dataset and glean return maximizing trajectories (detailed in Appendix A). Our straightforward conditional generative modeling formulation outperforms existing approaches on standard D4RL tasks (Fu et al., 2020).

Viewing offline decision-making through the lens of conditional generative modeling allows going beyond conditioning on returns (Figure 1). Consider an example (detailed in Appendix A) where a robot with linear dynamics navigates an environment containing two concentric circles (Figure 2). We are given a dataset of state-action trajectories of the robot, each satisfying one of two constraints: (i) the final position of the robot is within the larger circle, and (ii) the final position of the robot is outside the smaller circle. With conditional diffusion modeling, we can use the datasets to learn a constraint-conditioned model that can generate trajectories satisfying any set of constraints. During inference, the learned trajectory model can merge constraints from the dataset and generate trajectories that satisfy the combined constraint. Figure 2 shows that the constraint-conditioned model can generate trajectories such that the final position of the robot lies between the concentric circles.

Here, we demonstrate the benefits of modeling policies as conditional generative models. First, conditioning on constraints allows policies to not only generate behaviors satisfying individual constraints but also generate novel behaviors by flexibly combining constraints at test time. Further, conditioning on skills allows policies to not only imitate individual skills but also generate novel behaviors by composing those skills. We instantiate this idea with a state-sequence based diffusion probabilistic model (Ho et al., 2020) called Decision Diffuser, visualized in Figure 1. In summary, our contributions include (i) illustrating conditional generative modeling as an effective tool in offline decision making, (ii) using classifier-free guidance with low-temperature sampling, instead of dynamic programming, to get return-maximizing trajectories and, (iii) leveraging the framework of conditional generative modeling to combine constraints and compose skills during inference flexibly.

Background

Continuous action spaces further require learning a parametric policy πϕ(a∣s)\pi_{\phi}(a|s) that plays the role of the maximizing action in equation 1. This results in a policy objective that must be maximized:

Here, the dataset of transitions D\mathcal{D} evolves as the agent interacts with the environment and both QθQ_{\theta} and πϕ\pi_{\phi} are trained together. These methods make use of function approximation, off-policy learning, and bootstrapping, leading to several instabilities in practice (Sutton, 1988; Van Hasselt et al., 2018).

Offline RL

In this setting, we must find a return-maximizing policy from a fixed dataset of transitions collected by an unknown behavior policy μ\mu (Levine et al., 2020). Using TD-learning naively causes the state visitation distribution dπϕ(s)d^{\pi_{\phi}}(s) to move away from the distribution of the dataset dμ(s)d^{\mu}(s). In turn, the policy πϕ\pi_{\phi} begins to take actions that are substantially different from those already seen in the data. Offline RL algorithms resolve this distribution-shift by imposing a constraint of the form D(dπϕ∣∣dμ)D(d^{\pi_{\phi}}||d^{\mu}), where DD is some divergence metric, directly in the TD-learning procedure. The constrained optimization problem now demands additional hyper-parameter tuning and implementation heuristics to achieve any reasonable performance (Kumar et al., 2021). The Decision Diffuser, in comparison, doesn’t have any of these disadvantages. It does not require estimating any kind of QQ-function, thereby sidestepping TD methods altogether. It also does not face the risk of distribution-shift as generative models are trained with maximum-likelihood estimation.

2 Diffusion Probabilistic Models

Although a tractable variational lower-bound on log⁡pθ\log p_{\theta} can be optimized to train diffusion models, Ho et al. (2020) propose a simplified surrogate loss:

The predicted noise ϵθ(xk,k)\epsilon_{\theta}(\boldsymbol{x}_{k},k), parameterized with a deep neural network, estimates the noise ϵ∼N(0,I)\epsilon\sim{\mathcal{N}}(0,I) added to the dataset sample x0\boldsymbol{x}_{0} to produce noisy xk\boldsymbol{x}_{k}. This is equivalent to predicting the mean of pθ(xk−1∣xk)p_{\theta}(\boldsymbol{x}_{k-1}|\boldsymbol{x}_{k}) since μθ(xk,k)\mu_{\theta}(\boldsymbol{x}_{k},k) can be calculated as a function of ϵθ(xk,k)\epsilon_{\theta}(\boldsymbol{x}_{k},k) (Ho et al., 2020).

Modelling the conditional data distribution q(x∣y)q(\boldsymbol{x}|\boldsymbol{y}) makes it possible to generate samples with attributes of the label y\boldsymbol{y}. The equivalence between diffusion models and score-matching (Song et al., 2021), which shows ϵθ(xk,k)∝∇xklog⁡p(xk)\epsilon_{\theta}(\boldsymbol{x}_{k},k)\propto\nabla_{\boldsymbol{x}_{k}}\log p(\boldsymbol{x}_{k}), leads to two kinds of methods for conditioning: classifier-guided (Nichol & Dhariwal, 2021) and classifier-free (Ho & Salimans, 2022). The former requires training an additional classifier pϕ(y∣xk)p_{\phi}(\boldsymbol{y}|\boldsymbol{x}_{k}) on noisy data so that samples may be generated at test-time with the perturbed noise ϵθ(xk,k)−ω1−αˉk∇xklog⁡p(y∣xk)\epsilon_{\theta}(\boldsymbol{x}_{k},k)-\omega\sqrt{1-\bar{\alpha}_{k}}\nabla_{\boldsymbol{x}_{k}}\log p(\boldsymbol{y}|\boldsymbol{x}_{k}), where ω\omega is referred to as the guidance scale. The latter does not separately train a classifier but modifies the original training setup to learn both a conditional ϵθ(xk,y,k)\epsilon_{\theta}(\boldsymbol{x}_{k},\boldsymbol{y},k) and an unconditional ϵθ(xk,k)\epsilon_{\theta}(\boldsymbol{x}_{k},k) model for the noise. The unconditional noise is represented, in practice, as the conditional noise ϵθ(xk,\O,k)\epsilon_{\theta}(\boldsymbol{x}_{k},\O,k) where a dummy value \O\O takes the place of y\boldsymbol{y}. The perturbed noise ϵθ(xk,k)+ω(ϵθ(xk,y,k)−ϵθ(xk,k))\epsilon_{\theta}(\boldsymbol{x}_{k},k)+\omega(\epsilon_{\theta}(\boldsymbol{x}_{k},\boldsymbol{y},k)-\epsilon_{\theta}(\boldsymbol{x}_{k},k)) is used to later generate samples.

Generative Modeling with the Decision Diffuser

It is useful to solve RL from offline data, both without relying on TD-learning and without risking distribution-shift. To this end, we formulate sequential decision-making as the standard problem of conditional generative modeling:

Our goal is to estimate the conditional data distribution with pθp_{\theta} so we can later generate portions of a trajectory x0(τ)\boldsymbol{x}_{0}(\tau) from information y(τ)\boldsymbol{y}(\tau) about it. Examples of y\boldsymbol{y} could include the return under the trajectory, the constraints satisfied by the trajectory, or the skill demonstrated in the trajectory. We construct our generative model according to the conditional diffusion process:

As usual, qq represents the forward noising process while pθp_{\theta} the reverse denoising process. In the following, we discuss how we may use diffusion for decision making. First, we discuss the modeling choices for diffusion in Section 3.1. Next, we discuss how we may utilize classifier-free guidance to capture the best aspects of trajectories in Section 3.2. We then discuss the different behaviors that may be implemented with conditional diffusion models in Section 3.3. Finally, we discuss practical training details of our approach in Section 3.4.

In images, the diffusion process is applied across all pixel values in an image. Naïvely, it would therefore be natural to apply a similar process to model the state and actions of a trajectory. However, in the reinforcement learning setting, directly modeling actions using a diffusion process has several practical issues. First, while states are typically continuous in nature in RL, actions are more varied, and are often discrete in nature. Furthermore, sequences over actions, which are often represented as joint torques, tend to be more high-frequency and less smooth, making them much harder to predict and model (Tedrake, 2022). Due to these practical issues, we choose to diffuse only over states, as defined below:

Here, kk denotes the timestep in the forward process and tt denotes the time at which a state was visited in trajectory τ\tau. Moving forward, we will view xk(τ)\boldsymbol{x}_{k}(\tau) as a noisy sequence of states from a trajectory of length HH. We represent xk(τ)\boldsymbol{x}_{k}(\tau) as a two-dimensional array with one column for each timestep of the sequence.

Acting with Inverse-Dynamics. Sampling states from a diffusion model is not enough for defining a controller. A policy can, however, be inferred from estimating the action ata_{t} that led the state sts_{t} to st+1s_{t+1} for any timestep tt in x0(τ)\boldsymbol{x}_{0}(\tau). Given two consecutive states, we generate an action according to the inverse dynamics model (Agrawal et al., 2016; Pathak et al., 2018):

Note that the same offline data used to train the reverse process pθp_{\theta} can also be used to learn fϕf_{\phi}. We illustrate in Table 2 how the design choice of directly diffusing state distributions, with an inverse dynamics model to predict action, significantly improves performance over diffusing across both states and actions jointly. Furthermore, we empirically compare and analyze when to use inverse dynamics and when to diffuse over actions in Appendix F.

2 Planning with Classifier-Free Guidance

Given a diffusion model representing the different trajectories in a dataset, we next discuss how we may utilize the diffusion model for planning. To use the model for planning, it is necessary to additionally condition the diffusion process on characteristics y(τ)\boldsymbol{y}(\tau). One approach could be to train a classifier pϕ(y(τ)∣xk(τ))p_{\phi}(\boldsymbol{y}(\tau)|\boldsymbol{x}_{k}(\tau)) to predict y(τ)\boldsymbol{y}(\tau) from noisy trajectories xk(τ)\boldsymbol{x}_{k}(\tau). In the case that y(τ)\boldsymbol{y}(\tau) represents the return under a trajectory, this would require estimating a QQ-function, which requires a separate, complex dynamic programming procedure.

One approach to avoid dynamic programming is to directly train a conditional diffusion model conditioned on the returns y(τ)\boldsymbol{y}(\tau) in the offline dataset. However, as our dataset consists of a set of sub-optimal trajectories, the conditional diffusion model will be polluted by such sub-optimal behaviors. To circumvent this issue, we utilize classifier-free guidance (Ho & Salimans, 2022) with low-temperature sampling, to extract high-likelihood trajectories in the dataset. We find that such trajectories correspond to the best set of behaviors in the dataset. For a detailed discussion comparing Q-function guidance and classifier-free guidance, please refer to Appendix K. Formally, to implement classifier free guidance, a x0(τ)\boldsymbol{x}_{0}(\tau) is sampled by starting with Gaussian noise xK(τ)\boldsymbol{x}_{K}(\tau) and refining xk(τ)\boldsymbol{x}_{k}(\tau) into xk−1(τ)\boldsymbol{x}_{k-1}(\tau) at each intermediate timestep with the perturbed noise:

where the scalar ω\omega applied to (ϵθ(xk(τ),y(τ),k)−ϵθ(xk(τ),\O,k))(\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\boldsymbol{y}(\tau),k)-\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\O,k)) seeks to augment and extract the best portions of trajectories in the dataset that exhibit y(τ)\boldsymbol{y}(\tau). With these ingredients, sampling from the Decision Diffuser becomes similar to planning in RL. First, we observe a state in the environment. Next, we sample states later into the horizon with our diffusion process conditioned on y\boldsymbol{y} and history of last CC states observed. Finally, we identify the action that should be taken to reach the most immediate predicted state with our inverse dynamics model. This procedure repeats in a standard receding-horizon control loop described in Algorithm 1 and visualized in Figure 3.

3 Conditioning beyond Returns

So far we have not explicitly defined the conditioning variable y(τ)\boldsymbol{y}(\tau). Though we have mentioned that it can be the return under a trajectory, we may also consider guiding our diffusion process towards sequences of states that satisfy relevant constraints or demonstrate specific behavior.

To generate trajectories that maximize return, we condition the noise model on the return of a trajectory so ϵθ(xk(τ),y(τ),k)≔ϵθ(xk(τ),R(τ),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\boldsymbol{y}(\tau),k)\coloneqq\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),R(\tau),k). These returns are normalized to keep R(τ)∈R(\tau)\in. Sampling a high return trajectory amounts to conditioning on R(τ)=1R(\tau)=1. Note that we do not make use of any QQ-values, which would then require dynamic programming.

Satisfying Constraints

Composing Skills

Assuming we have learned the data distributions q(x0(τ)∣y1(τ)),…,q(x0(τ)∣yn(τ))q(\boldsymbol{x}_{0}(\tau)|\boldsymbol{y}^{1}(\tau)),\ldots,q(\boldsymbol{x}_{0}(\tau)|\boldsymbol{y}^{n}(\tau)) for nn different conditioning variables, we can sample from the composed data distribution q(x0(τ)∣y1(τ),…,yn(τ))q(\boldsymbol{x}_{0}(\tau)|\boldsymbol{y}^{1}(\tau),\ldots,\boldsymbol{y}^{n}(\tau)) using the perturbed noise (Liu et al., 2022):

This property assumes that {yi(τ)}i=1n\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n} are conditionally independent given the state trajectory x0(τ)\boldsymbol{x}_{0}(\tau). However, we empirically observe that this assumption doesn’t have to be strictly satisfied as long as the composition of conditioning variables is feasible. For more detailed discussion, please refer to Appendix D. We use this property to compose more than one constraint or skill together at test-time. We also show how Decision Diffuser can avoid particular constraint or skill (NOT) in Appendix J.

4 Training the Decision Diffuser

The Decision Diffuser, our conditional generative model for decision-making, is trained in a supervised manner. Given a dataset D\mathcal{D} of trajectories, each labeled with the return it achieves, the constraint that it satisfies, or the skill that it demonstrates, we simultaneously train the reverse diffusion process pθp_{\theta}, parameterized through the noise model ϵθ\epsilon_{\theta}, and the inverse dynamics model fϕf_{\phi} with the following loss:

For each trajectory τ\tau, we first sample noise ϵ∼N(0,I)\epsilon\sim\mathcal{N}(\boldsymbol{0},\boldsymbol{I}) and a timestep k∼U{1,…,K}k\sim{\mathcal{U}}\{1,\ldots,K\}. Then, we construct a noisy array of states xk(τ)\boldsymbol{x}_{k}(\tau) and finally predict the noise as ϵ^θ≔ϵθ(xk(τ),y(τ),k)\hat{\epsilon}_{\theta}\coloneqq\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\boldsymbol{y}(\tau),k). Note that with probability pp we ignore the conditioning information and the inverse dynamics is trained with individual transitions rather than trajectories.

Low-temperature Sampling

In the denoising step of Algorithm 1, we compute μk−1\mu_{k-1} and Σk−1\Sigma_{k-1} from a noisy sequence of states and a predicted noise. We find that sampling xk−1∼N(μk−1,αΣk−1)x_{k-1}\sim{\mathcal{N}}(\mu_{k-1},\alpha\Sigma_{k-1}) where the variance is scaled by α∈[0,1)\alpha\in[0,1) leads to better quality sequences (corresponding to sampling lower temperature samples). For a proper ablation study, please refer to Appendix C.

Experiments

In our experiments section, we explore the efficacy of the Decision Diffuser on a variety of different decision making tasks (performance illustrated in Figure 4). In particular, we evaluate (1) the ability to recover effective RL policies from offline data, (2) the ability to generate behavior that satisfies multiple sets of constraints, (3) the ability compose multiple different skills together. In addition, we empirically justify use of classifier-free guidance, low-temperature sampling (Appendix C), and inverse dynamics (Appendix F) and test the robustness of Decision Diffuser to stochastic dynamics (Appendix G).

We first test whether the Decision Diffuser can generate return-maximizing trajectories. To test this, we train a state diffusion process and inverse dynamics model on publicly available D4RL datasets (Fu et al., 2020). We compare with existing offline RL methods, including model-free algorithms like CQL (Kumar et al., 2020) and IQL (Kostrikov et al., 2022), and model-based algorithms such as trajectory transformer (TT, Janner et al. (2021)) and MoReL (Kidambi et al., 2020). We also compare with sequence-models like the Decision Transformer (DT) (Chen et al. (2021) and diffusion models like Diffuser (Janner et al., 2022).

Results

Across a broad suite of different offline reinforcement learning tasks, we find that the Decision Diffuser is either competitive or outperforms many of our offline RL baselines (Table 1). It also outperforms Diffuser and sequence modeling approaches, such as Decision Transformer and Trajectory Transformer. The difference between Decision Diffuser and other methods becomes even more significant on harder D4RL Kitchen tasks which require long-term credit assignment.

To convey the importance of classifier-free guidance, we also compare with the baseline CondDiffuser, which diffuses over both state and action sequences as in Diffuser without classifier-guidance. In Table 2, we observe that CondDiffuser improves over Diffuser in 22 out of 33 environments. Decision Diffuser further improves over CondDiffuser, performing better across all 33 environments. We conclude that learning the inverse dynamics is a good alternative to diffusing over actions. We further empirically analyze when to use inverse dynamics and when to diffuse over actions in Appendix F. We also compare against CondMLPDiffuser, a policy where the current action is denoised according to a diffusion process conditioned on both the state and return. We see that CondMLPDiffuser performs the worst amongst diffusion models. Till now, we mainly tested on offline RL tasks that have deterministic (or near deterministic) environment dynamics. Hence, we test the robustness of Decision Diffuser to stochastic dynamics and compare it to Diffuser and CQL as we vary the stochasticity in environment dynamics, in Appendix G. Finally, we analyze the runtime characteristics of Decision Diffuser in Appendix E.

2 Constraint Satisfaction

Setup We next evaluate how well we can generate trajectories that satisfy a set of constraints using the Kuka Block Stacking environment (Janner et al., 2022) visualized in Figure LABEL:fig:kuka. In this domain, there are four blocks which can be stacked as a single tower or rearranged into several towers. A constraint like BlockHeight(i)>BlockHeight(j)\texttt{BlockHeight}(i)>\texttt{BlockHeight}(j) requires that block ii be placed above block jj. We train the Decision Diffuser from 10,00010,000 expert demonstrations each satisfying one of these constraints. We randomize the positions of these blocks and consider two tasks at inference: sampling trajectories that satisfy a single constraint seen before in the dataset or satisfy a group of constraints for which demonstrations were never provided. In the latter, we ask the Decision Diffuser to generate trajectories so BlockHeight(i)>BlockHeight(j)>BlockHeight(k)\texttt{BlockHeight}(i)>\texttt{BlockHeight}(j)>\texttt{BlockHeight}(k) for three of the four blocks i,j,ki,j,k. For more details, please refer to Appendix H.

In both the stacking and rearrangement settings, Decision Diffuser satisfies single constraints with greater success rate than Diffuser (Table LABEL:table:kuka_quant). We also compare with BCQ (Fujimoto et al., 2019) and CQL (Kumar et al., 2020), but they consistently fail to stack or rearrange the blocks leading to a 0.00.0 success rate. Unlike these baselines, our method can just as effectively satisfy several constraints together according to Equation 9. For a visualization of these generated trajectories, please see the website https://anuragajay.github.io/decision-diffuser/.

3 Skill Composition

Finally, we look at how to compose different skills together. We consider the Unitree-go-running environment (Margolis & Agrawal, 2022), where a quadruped robot can be found running with various gaits, like bounding, pacing, and trotting. We explore if it is possible to generate trajectories that transition between these gaits after only training on individual gaits. For each gait, we collect a dataset of 25002500 demonstrations on which we train Decision Diffuser.

Results

During testing, we use the noise model of our reverse diffusion process according to equation 9 to sample trajectories of the quadruped robot with entirely new running behavior. Figure 6 shows a trajectory that begins with bounding but ends with pacing. Appendix I provides additional visualizations of running gaits being composed together. Although it visually appears that trajectories generated with the Decision Diffuser contain more than one gait, we would like to quantify exactly how well different gaits can be composed. To this end, we train a classifier to predict at every time-step or frame in a trajectory the running gait of the quadruped (i.e. bound, pace, or trott). We reuse the demonstrations collected for training the Decision Diffuser to also train this classifier, where our inputs are defined as robot joint states over a fixed period of time (i.e. state sub-sequences of length 1010) and the label is the gait demonstrated in this sequence. The complete details of our gait classification procedure can be found in Appendix I.

We use our running gait classifier in two ways: to evaluate how the behavior of the quadruped changes over the course of a single, generated trajectory and to measure how often each gait emerges over several generated trajectories. In the former, we first sample three trajectories from the Decision Diffuser conditioned either on the bounding gait, the pacing gait, or both. For every trajectory, we separately plot the classification probability of each gait over the length of the sequence. As shown in the plots of Figure 7, the classifier predicts bound and pace respectively to be the most likely running gait in trajectories sampled with this condition. When the trajectory is generated by conditioning on both gaits, the classifier transitions between predicting one gait with largest probability to the other. In fact, there are several instances where the behavior of the quadruped switches between bounding and pacing according to the classifier. This is consistent with the visualizations reported in Figure 6. In the table depicted in Figure 7, we consider 10001000 trajectories generated with the Decision Diffuser when conditioned on one or both of the gaits as listed. We record the fraction of time that the quadruped’s running gait was classified as either trott, pace, or bound. It turns out that the classifier identifies the behavior as bounding for 38.5%38.5\% of the time and as pacing for the other 60.1%60.1\% when trajectories are sampled by composing both gaits. This corroborates the fact that the Decision Diffuser can indeed compose running behaviors despite only being trained on individual gaits.

Related Work

Diffusion Models have shown great promise in learning generative models of image and text data (Saharia et al., 2022; Nichol et al., 2021; Nichol & Dhariwal, 2021). It formulates the data sampling process as an iterative denoising procedure (Sohl-Dickstein et al., 2015; Ho et al., 2020). The denoising procedure can be alternatively interpreted as parameterizing the gradients of the data distribution (Song et al., 2021) optimizing the score matching objective (Hyvärinen, 2005) and thus as a Energy-Based Model (Du & Mordatch, 2019; Nijkamp et al., 2019; Grathwohl et al., 2020). To generate data samples (eg: images) conditioned on some additional information (eg:text), prior works (Nichol & Dhariwal, 2021) have learned a classifier to facilitate the conditional sampling. More recent works (Ho & Salimans, 2022) have argued to leverage gradients of an implicit classifier, formed by the difference in score functions of a conditional and an unconditional model, to facilitate conditional sampling. The resulting classifier-free guidance has been shown to generate better conditional samples than classifier-based guidance. All these above mentioned works have mostly focused on generation of text or images. Recent works have also used diffusion models to imitate human behavior (Pearce et al., 2023) and to parameterize policy in offline RL (Wang et al., 2022). Janner et al. (2022) generate trajectories consisting of states and actions with an unconditional diffusion model, therefore requiring a trained reward function on noisy state-action pairs. At inference, the estimated reward function guides the reverse diffusion process towards samples of high-return trajectories. In contrast, we do not train reward functions or diffusion processes separately, but rather model the trajectories in our dataset with a single, conditional generative model instead. This ensures that the sampling procedure of the learned diffusion process is the same at inference as it is during training.

Reward Conditioned Policies

Prior works (Kumar et al., 2019; Schmidhuber, 2019; Srivastava et al., 2019; Emmons et al., 2021; Chen et al., 2021) have studied learning of reward conditioned policies via reward conditioned behavioral cloning. Chen et al. (2021) used a transformer (Vaswani et al., 2017) to model the reward conditioned policies and obtained a performance competitive with offline RL approaches. Emmons et al. (2021) obtained similar performance as Chen et al. (2021) without using a transformer policy but relied on careful capacity tuning of MLP policy. In contrast to these works, in addition to modeling returns, Decision Diffuser can also model constraints or skills and generate novel behaviors by flexibly combining multiple constraints or skills during test time.

Discussion

We propose Decision Diffuser, a conditional generative model for sequential decision making. It frames offline sequential decision making as conditional generative modeling and sidesteps the need of reinforcement learning, thereby making the decision making pipeline simpler. By sampling for high returns, it is able to capture the best behaviors in the dataset and outperforms existing offline RL approaches on standard benchmarks (such as D4RL). In addition to returns, it can also be conditioned on constraints or skills and can generate novel behaviors by flexibly combining constraints or composing skills during test time.

In this work, we focused on offline sequential decision making, thus circumventing the need for exploration. Using ideas from Zheng et al. (2022), future works could look into online fine-tuning of Decision Diffuser by leveraging entropy of the state-sequence model for exploration. While our work focused on state based environments, it can be extended to image based environments by performing the diffusion in latent space, rather than observation space, as done in Rombach et al. (2022). For a detailed discussion on limitations of Decision Diffuser, please refer to Appendix M.

Acknowledgements

The authors would like to thank Ofir Nachum, Anthony Simeonov and Richard Li for their helpful feedback on an earlier draft of the work; Jay Whang and Ge Yang for discussions on classifier-free guidance; Gabe Margolis for helping with unitree experiments; Micheal Janner for providing visualization code for Kuka block stacking; and the members of Improbable AI Lab for discussions and helpful feedback. We thank MIT Supercloud and the Lincoln Laboratory Supercomputing Center for providing compute resources. This research was supported by an NSF graduate fellowship, a DARPA Machine Common Sense grant, ARO MURI Grant Number W911NF-21-1-0328 and an MIT-IBM grant.

This research was also partly sponsored by the United States Air Force Research Laboratory and the United States Air Force Artificial Intelligence Accelerator and was accomplished under Cooperative Agreement Number FA8750-19- 2-1000. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the United States Air Force or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes, notwithstanding any copyright notation herein.

Author Contributions

Anurag Ajay conceived the framework of viewing decision-making as conditional diffusion generative modeling, implemented the Decision Diffuser algorithm, ran experiments on Offline RL and Skill Composition, and helped in paper writing.

Yilun Du helped in conceiving the framework of viewing decision-making as conditional diffusion generative modeling, ran experiments on Constraint Satisfaction, helped in paper writing and advised Anurag.

Abhi Gupta helped in running experiments on Offline RL and Skill Composition, participated in research discussions, and played the leading role in paper writing and making figures.

Joshua Tenenbaum participated in research discussions.

Tommi Jaakkola participated in research discussions and suggested the experiment of classifying running gaits.

Pulkit Agrawal was involved in research discussions, suggested experiments related to dynamic programming, provided feedback on writing, positioning of the work and overall advising.

References

Appendix A Illustrative Examples

We empirically demonstrate the ability of Decision Diffuser to perform implicit dynamic programming in Maze2D-open environment from Fu et al. (2020). The task in Maze2D-open environment is to reach point C and the reward is negative distance from point C. The training dataset consists of 500500 trajectories from point A to point B and 500500 trajectories from point B to point C. The maximum trajectory length is 5050. During test time, the agent starts from point A and needs to reach point C as quickly as possible. As shown in Figure A1, Decision Diffuser can stitch trajectories in training dataset to form trajectories that goes from point A to point B in (near) straight lines.

A.2 Constraint Combination

In linear system robot navigation, Decision Diffuser is trained on 10001000 expert trajectories either satisfying the constraint ∥sT∥≤R\|s_{T}\|\leq R (R=1R=1) or the constraint ∥sT∥≥r\|s_{T}\|\geq r (r=0.7r=0.7). Here, sT=[xT,yT]s_{T}=[x_{T},y_{T}] represents the final robot state in a trajectory, specifying its final 2d position. The maximum trajectory length is 5050. During test time, Decision Diffuser is asked to generate trajectories satisfying ∥sT∥≤R\|s_{T}\|\leq R and ∥sT∥≥r\|s_{T}\|\geq r to test its ability to satisfy single constraints. Furthermore, Decision Diffuser is also asked to generate trajectories satisfying r≤∥sT∥≤Rr\leq\|s_{T}\|\leq R to test its ability to satisfy combined constraints.

Results

Figure 2 shows that Decision Diffuser learns to generate trajectories perfectly (i.e. with 100%100\% success rate) satisfying single constraints in linear system robot navigation. Furthermore, it learns to generate trajectories satisfying the composed constraint in linear system robot navigation with 91.3%(±2.6%)91.3\%(\pm 2.6\%) accuracy where the standard error is calculated over 55 random seeds.

Appendix B Hyperparameter and Architectural details

In this section, we describe various architectural and hyperparameter details:

We represent the noise model ϵθ\epsilon_{\theta} with a temporal U-Net (Janner et al., 2022), consisting of a U-Net structure with 6 repeated residual blocks. Each block consisted of two temporal convolutions, each followed by group norm (Wu & He, 2018), and a final Mish nonlinearity (Misra, 2019). Timestep and condition embeddings, both 128128-dimensional vectors, are produced by separate 2-layered MLP (with 256256 hidden units and Mish nonlinearity) and are concatenated together before getting added to the activations of the first temporal convolution within each block. We borrow the code for temporal U-Net from https://github.com/jannerm/diffuser.

We represent the inverse dynamics fϕf_{\phi} with a 2-layered MLP with 512512 hidden units and ReLU activations.

We represent the gait classifier with a 3-layered MLP with 10241024 hidden units and ReLU activations.

We train ϵθ\epsilon_{\theta} and fϕf_{\phi} using the Adam optimizer (Kingma & Ba, 2015) with a learning rate of 2e−42e-4 and batch size of 3232 for 2e62e6 train steps.

We train the gait classifier using the Adam optimizer with a learning rate of 2e−42e-4 and batch size of 6464 for 1e61e6 train steps.

We choose the probability pp of removing the conditioning information to be 0.250.25.

We use a planning horizon HH of 100100 in all the D4RL locomotion tasks, 5656 in D4RL kitchen tasks, 128128 in Kuka block stacking, 5656 in unitree-go-running tasks, 5050 in the illustrative example and 6060 in Block push tasks.

We use a guidance scale s∈{1.2,1.4,1.6,1.8}s\in\{1.2,1.4,1.6,1.8\} but the exact choice varies by task.

We choose α=0.5\alpha=0.5 for low temperature sampling.

Appendix C Importance of low temperature sampling

In Algorithm 1, we compute μk−1\mu_{k-1} and Σk−1\Sigma_{k-1} from a noisy sequence of states and predicted noise. We find that sampling xk−1∼N(μk−1,αΣk−1)x_{k-1}\sim{\mathcal{N}}(\mu_{k-1},\alpha\Sigma_{k-1}) (where α∈[0,1)\alpha\in[0,1)) with a reduced variance produces high-likelihood state sequences. We refer to this as low-temperature sampling. To empirically show its importance, we compare performances of Decision Diffuser with different values of α\alpha (Table A1). We show that low temperature sampling (α=0.5)(\alpha=0.5) gives the best average returns. However, reducing the α\alpha to eliminates the entropy in sampling and leads to lower returns. On the other hand, α=1.0\alpha=1.0 leads to a higher variance in terms of returns of the trajectories.

Appendix D Composing conditioning variables

In this section, we detail how Decision Diffuser trained with different conditioning variables {yi(τ)}i=1n\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n} composes these conditioning variables together. It learns the denoising model ϵθ(xk(τ),yi(τ),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\boldsymbol{y}^{i}(\tau),k) for a given conditioning variable yi(τ)\boldsymbol{y}^{i}(\tau). From the derivations outlined in prior works (Luo, 2022; Song et al., 2021), we know that ∇xk(τ)log⁡q(xk(τ)∣yi(τ))∝−ϵθ(xk(τ),yi(τ),k)\nabla_{\boldsymbol{x}_{k}(\tau)}\log q(\boldsymbol{x}_{k}(\tau)|\boldsymbol{y}^{i}(\tau))\propto-\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\boldsymbol{y}^{i}(\tau),k). Therefore, each conditional trajectory distribution {q(xk(τ)∣yi(τ))}i=1n\{q(\boldsymbol{x}_{k}(\tau)|\boldsymbol{y}^{i}(\tau))\}_{i=1}^{n} can be modelled with a single denoising model ϵθ\epsilon_{\theta} that conditions on the respective variable yi(τ)\boldsymbol{y}^{i}(\tau).

In order to compose nn different conditioning variables (i.e. skills or constraints), we would like to model q(xk(τ)∣{yi(τ)}i=1n)q(\boldsymbol{x}_{k}(\tau)|\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n}). We assume that {yi(τ)}i=1n\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n} are conditionally independent given xk(τ)\boldsymbol{x}_{k}(\tau). Thus, we can factorize as follows:

Using the above equations, we can sample from q(x0(τ)∣{yi(τ)}i=1n)q(\boldsymbol{x}_{0}(\tau)|\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n}) with classifier free guidance using the perturbed noise:

We use the perturbed noise to compose skills or combine constraints at test time. This derivation was borrowed from Liu et al. (2022) and is presented here for completeness.

While the composition of conditioning variables {yi(τ)}i=1n\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n} requires them to be conditionally independent given the state trajectory x0(τ)\boldsymbol{x}_{0}(\tau), we empirically observe that this condition doesn’t have to be strictly satisfied. However, we require composition of conditioning variables to be feasible (i.e. ∃  x0(τ)\exists\;\boldsymbol{x}_{0}(\tau) that satisfies all the conditioning variables). When the composition is infeasible, Decision Diffuser produces trajectories with incoherent behavior, as expected. This is best illustrated by videos viewable at https://anuragajay.github.io/decision-diffuser/.

First, the dataset should have a diverse set of demonstrations that shows different ways of satisfying each conditioning variable yi(τ)\boldsymbol{y}^{i}(\tau). This would allow Decision Diffuser to learn diverse ways of satisfying each conditioning variable yi(τ)\boldsymbol{y}^{i}(\tau). Since we use inverse dynamics to extract actions from the predicted state trajectory x0(τ)\boldsymbol{x}_{0}(\tau), we assume that the state trajectory x0(τ)\boldsymbol{x}_{0}(\tau) resulting from the composition of different conditioning variables contains consecutive state pairs (st,st+1)(s_{t},s_{t+1}) that come from the same distribution that generated the demonstration dataset. Otherwise, inverse dynamics can give erroneous predictions.

Appendix E Runtime characteristic of Decision Diffuser

We analyze the runtime characteristics of Decision Diffuser in this section. After training the Decision Diffuser on trajectories from the D4RL Hopper-Medium-Expert dataset, we plan in the corresponding environment according to Algorithm 1. Every action taken in the environment requires running 100 reverse diffusion steps to generate a state sequence taking on average 1.26s in wall-clock time. We can improve the run-time of planning by warm-starting the state diffusion as suggested in Janner et al. (2022). Here, we start with a generated state sequence (from the previous environment step), run forward diffusion for a fixed number of steps, and finally run the same number of reverse diffusion steps from the partially noised state sequence to generate another state sequence. Warm-starting in this way allows us to decrease the number of denoising steps to 40 (0.48s on average) without any loss in performance, to 20 (0.21s on average) with minimal loss in performance, and to 5 with less than 20%\% loss in performance (0.06s on average). We demonstrate the trade-off between performance, measured by normalized average return achieved in the environment, and planning time, measured in wall-clock time after warm-starting the reverse diffusion process, in Figure A2.

Appendix F When to use Inverse dynamics?

In this section, we try to analyze further when using inverse dynamics is better than diffusing over actions. Table 2 showed that Decision Diffuser outperformed CondDiffuser on 33 hopper environment, thereby suggesting that inverse dynamics is a better alternative to diffusing over actions. Our intuition was that sequences over actions, represented as joint torques in our environments, tend to be more high-frequency and less smooth, thus making it harder for the diffusion model to predict (Kingma et al., 2021). We now try to verify this intuition empirically.

Setup We choose Block Push environment adapted from Gupta et al. (2018) where the goal is to push the red cube to the green circle. When the red cube reaches the green circle, the agent gets a reward of +1. The state space is 1010-dimensional consisting of joint angles (33) and velocities (33) of the gripper, COM of the gripper (22) and position of the red cube (22). The green circle’s position is fixed and at an initial distance of 0.50.5 from COM of the gripper. The red cube (of size 0.030.03) is initially at a distance of 0.10.1 from COM of the gripper and at an angle θ\theta sampled from U(−π/4,π/4)\mathcal{U}(-\pi/4,\pi/4) at the start of every episode. The task horizon is 6060 timesteps.

There are 2 control types: (i) torque control, where the agent needs to specify joint torques (33 dimensional) and (ii) position control where the agent needs to specify the position change of COM of the gripper and the angular change in gripper’s orientation (Δx,Δy,Δϕ)(\Delta x,\Delta y,\Delta\phi) (33 dimensional). While action trajectories from position control are smooth, the action trajectories from torque control have higher frequency components.

Offline dataset collection To collect the offline data, we use Soft Actor-Critic (SAC) (Haarnoja et al., 2018) first to train an expert policy for 11 million environment steps. We then use 11 million environment transitions as our offline dataset, which contains expert trajectories collected towards the end of the training and random action trajectories collected at the beginning of the training. We collect 2 datasets, one for each control type.

Results Table LABEL:table:push shows that Decision Diffuser and CondDiffuser perform similarly when the agent uses position control. This is because action trajectories resulting from position control are smoother and hence easier to model with diffusion. However, when the agent uses torque control, CondDiffuser performs worse than Decision Diffuser, given the action trajectories have higher frequency components and hence are harder to model with diffusion.

Appendix G Robustness to stochastic dynamics

We empirically analyze robustness of Decision Diffuser to stochasticity in dynamics function.

We use Block Push environment, described in Appendix F, with torque control. However, we inject stochasticity into the environment dynamics. For every environment step, we either sample a random action from U(,)\mathcal{U}(,) with probability pp or execute the action given by the policy with probability (1−p)(1-p). We use p∈{0,0.05,0.1,0.15}p\in\{0,0.05,0.1,0.15\} in our experiments.

Offline dataset collection

We collect separate offline datasets for different block push environments, each characterized by a different value of pp. Each offline dataset consists of 1 million environment transitions collected using the method described in Appendix F.

Results

Table A3 characterizes how the performance of BC, Decision Diffuser, Diffuser, and CQL changes with increasing stochasticity in the environment dynamics. We observe that the Decision Diffuser outperforms Diffuser and CQL for p=0.05p=0.05, however all methods including the Decision Diffuser settle to a similar performance for larger values of pp.

Several works (Paster et al., 2022; Yang et al., 2022) have shown that the performance of return-conditioned policies suffers as the stochasticity in environment dynamics increases. This is because the return-conditioned policies aren’t able to distinguish between high returns from good actions and high returns from environment stochasticity. Hence, these return-conditioned policies can learn sub-optimal actions that got associated with high-return trajectories in the dataset due to environment stochasticity. Given Decision diffuser uses return conditioning to generate actions in offline RL, its performance also suffers when stochasticity in environment dynamics increases.

Some recent works (Yang et al., 2022; Villaflor et al., 2022) address the above issue by learning a latent model for future states and then conditioning the policy on predicted latent future states rather than returns. Conditioning Decision Diffuser on future state information, rather than returns, would make it more robust to stochastic dynamics and could be an interesting avenue for future works.

Appendix H Kuka Block Stacking

In the Kuka blocking stacking environment, the underlying goal is to stack a set of blocks on top of each other. Models have trained on a set of demonstration data, where a set of 4 blocks are sequentially stacked on top of each other to form a block tower.

We construct state-space plans of length 128. Following (Janner et al., 2022), we utilize a close-loop controller to generate actions for each state in our state-space plan (controlling the 7 degrees of freedom in joints). The total maximum trajectory length plan in Kuka block stacking is 384. We detail differences between the two consider conditional stacking environments below:

Stacking In the stacking environment, at test time we wish to again construct a tower of four blocks.

Rearrangement In the rearrangement environment, at test time wish to stack blocks in a configuration where a set of blocks are above a second set. This set of stack-place relations may not precisely correspond to a single block tower (can instead construct two block towers), making this environment an out-of-distribution challenge.

In addition to Diffuser (Janner et al., 2022), we used goal-conditioned variants of CQL (Kumar et al., 2020) and BCQ (Fujimoto et al., 2019) as baselines for the block stacking and rearrangement with single constraint. However, they get a success rate of 0.00.0.

Appendix I Unitree Go Running

We consider Unitree-go-running environment (Margolis & Agrawal, 2022) where a quadruped robot runs in 33 different gaits: bounding, pacing, and trotting. The state space is 5656 dimensional, the action space is 1212 dimensional, and the maximum trajectory length is 250250.

As described in Section 4.3, we train Decision Diffuser on expert trajectories demonstrating individual gaits. During testing, we compose the noise model of our reverse diffusion process according to equation 9. This allows us to sample trajectories of the quadruped robot with entirely new running behavior. Figures A4,A5,A6 shows the ability of Decision Diffuser to imitate bounding, trotting and pacing and their combinations.

We now try to quantitatively verify whether the trajectories resulting from composition of 22 gaits does indeed contain only those 22 gaits.

We learn a gait classifier that takes in a sub-sequence of states (of length 1010) and predicts the gait-ID. It is represented by a 33-layered MLP with 10241024 hidden units and ReLU activations that concatenates the sub-sequence of states (of length 1010) into a single vector of dimension 560560 before taking it in as an input. We train the gait classifier on the demonstration dataset. To ensure that the learned classifier can predict gait-ID on trajectories generated by the composition of skills, we use MixUp-style (Zhang et al., 2017) data augmentation during training. We create a synthetic sub-sequence of length 1010 by concatenating two sampled sub-sequence (from the demonstration dataset) of length lil_{i} and ljl_{j} (where li+lj=10l_{i}+l_{j}=10) from gaits with ID ii and jj and give it a label lili+ljone-hot(i)+ljli+ljone-hot(j)\frac{l_{i}}{l_{i}+l_{j}}\text{one-hot}(i)+\frac{l_{j}}{l_{i}+l_{j}}\text{one-hot}(j). During training, we sample a sub-sequence from the demonstration dataset with 70%70\% probability and a sythenthic sub-sequence with 30%30\% probability. We train the classifier for 2e62e6 train steps with a learning rate of 2e−42e-4 and a batch size of 6464.

Results

Figures A4,A5,A6 show that the classifier’s prediction is consistent with the visualized composed trajectories. Furthermore, we use Decision diffuser to act in the environment and generate 10001000 trott trajectories, 10001000 pace trajectories, 10001000 bound trajectories, and 10001000 composed trajectories for each possible pair of individual gaits. We then evaluate the learned gait classifier on these trajectories and compute the percentage of timesteps a particular gait has the highest probability. From Figures A4,A5,A6, we can see that if trajectories are generated by the composition of two gaits, then those two gaits will have the two highest probabilities across different timesteps in those trajectories.

I.2 A simple baseline for composition

Let one-hot(i)\text{one-hot}(i) and one-hot(j)\text{one-hot}(j) represent two different gaits that can be generated using noise models ϵθ(xk(τ),one-hot(i),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\text{one-hot}(i),k) and ϵθ(xk(τ),one-hot(j),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\text{one-hot}(j),k) respectively. To compose these gaits, we compose the above-mentioned noise models using equation 9. As an alternative, we see if the noise model ϵθ(xk(τ),one-hot(i)+one-hot(j),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\text{one-hot}(i)+\text{one-hot}(j),k) can lead to composed gaits. However, we observe that ϵθ(xk(τ),one-hot(i)+one-hot(j),k)\epsilon_{\theta}(\boldsymbol{x}_{k}(\tau),\text{one-hot}(i)+\text{one-hot}(j),k) catastrophically fail to generate any gait (see videos at https://anuragajay.github.io/decision-diffuser/). This happens because the condition variable one-hot(i)+one-hot(j)\text{one-hot}(i)+\text{one-hot}(j) was never seen by the noise model ϵθ\epsilon_{\theta} during training.

Appendix J NOT compositions with Decision Diffuser

Decision diffuser can also support ”NOT” composition. Suppose we wanted to sample from q(x0(τ)∣NOT yj(τ))q(\boldsymbol{x}_{0}(\tau)|\text{NOT }\boldsymbol{y}^{j}(\tau)). Let {yi(τ)}i=1n\{\boldsymbol{y}^{i}(\tau)\}_{i=1}^{n} be the set of all conditioning variables. Then, following derivations from Liu et al. (2022) and using β=1\beta=1, we can sample from q(x0(τ)∣NOT yj(τ))q(\boldsymbol{x}_{0}(\tau)|\text{NOT }\boldsymbol{y}^{j}(\tau)) using the perturbed noise:

We demonstrate the ability of Decision Diffuser to support ”NOT” composition by using it to satisfy constraint of type BlockHeight(i)>BlockHeight(j) AND (NOT BlockHeight(j)>BlockHeight(i))\texttt{BlockHeight}(i)>\texttt{BlockHeight}(j)\text{ AND }(\text{NOT }\texttt{BlockHeight}(j)>\texttt{BlockHeight}(i)) in Kuka block stacking task, as visualized in videos at https://anuragajay.github.io/decision-diffuser/. As the Decision Diffuser does not provide an explicit density estimate for each skill, it can’t natively support OR composition.

Appendix K Comparing Q-function guided diffusion and Classifier-free guided diffusion

Classifier-free guided diffusion and Q-value guided diffusion are theoretically equivalent. However, as noted in several works (Nichol et al., 2021; Ho & Salimans, 2022; Saharia et al., 2022), classifier-free guidance performs better than classifier guidance (i.e. Q function guidance in our case) in practice. This is due to following reasons:

Classifier-guided diffusion models learns an unconditional diffusion model along with a classifier (Q-function in our case) and uses gradients from the classifier to perform conditional sampling. However, the unconditional diffusion model doesn’t need to focus on conditional modeling during training and only cares about conditional generation during testing after it has been trained. In contrast, classifier-free guidance relies on conditional diffusion model to estimate gradients of the implicit classifier. Since the conditional diffusion model, learned when using classifier-free guidance, focuses on conditional modeling during train time, it performs better in conditional generation during test time.

Q function trained on an offline dataset can erroneously predict high Q values for out-of-distribution actions given any state. This problem has been extensively studied in offline RL literature (Kumar et al., 2020; Fujimoto et al., 2019; Levine et al., 2020). In online RL, this issue is automatically corrected when the policy acts in the environment, thinking an action to be good but then receives a low reward for it. In offline RL, this issue can’t be corrected easily; hence, the learned Q-function can often guide the diffusion model towards out-of-distribution actions that might be sub-optimal. In contrast, classifier-free guidance circumvents the issue of learning a Q-function and directly conditions the diffusion model on returns. Hence, classifier-free guidance doesn’t suffer due to errors in learned Q-functions and hence performs better than Q-function guided diffusion.

Appendix L Comparing Decision Transformer and Decision Diffuser

Decision Transformer models the likelihood of the next action given target return and sequence of past states, actions, and rewards. In contrast, Decision Diffuser models the score-function of future state trajectory given target return and past state trajectory. While both these methods learn a conditional probabilistic model of trajectories, Decision Diffuser allows for composition of conditioning variables by composing their respective score functions (see Appendix D for details). In theory, being a likelihood model, Decision Transformer can also model score function of trajectories. However, this might be computationally inefficient (Dieleman, 2022) and thus impractical.

Appendix M Limitations of Decision Diffuser

We summarize the limitations of Decision Diffuser:

No partial observability Decision Diffuser works with fully observable MDPs. Naive extensions to partially observed MDPs (POMDPs) may cause self-delusions (Ortega et al., 2021) in Decision Diffuser. Hence, extending Decision Diffuser to POMDPs could be an exciting avenue for future work.

Inability to explore the environment and update itself in online setting In this work, we focused on offline sequential decision making, thus circumventing the need for exploration. Using ideas from Zheng et al. (2022), future works could look into online fine-tuning of Decision Diffuser by leveraging entropy of the state-sequence model for exploration.

Experiments on only state-based environments While our work focused on state based environments, it can be extended to image based environments by performing the diffusion in latent space, rather than observation space, as done in Rombach et al. (2022).

Only AND and NOT compositions are supported Since Decision Diffuser does not provide an explicit density estimate for each condition variable, it can’t natively support OR composition.

Performance degradation in environments with stochastic dynamics In environments with highly stochastic dynamics, Decision Diffuser loses its advantage and performs similarly to Diffuser and CQL. To tackle environments with stochastic dynamics, recent works (Yang et al., 2022; Villaflor et al., 2022) propose learning a latent model for future states and then conditioning the policy on predicted latent future states rather than returns. Conditioning Decision Diffuser on future state information, rather than returns, would make it more robust to stochastic dynamics and could be an interesting avenue for future works.

Performance in limited data regime Since diffusion models are prone to overfitting in case of limited data, Decision Diffuser is also prone to overfitting in limited data regime.