Guided Flows for Generative Modeling and Decision Making

Qinqing Zheng, Matt Le, Neta Shaul, Yaron Lipman, Aditya Grover, Ricky T. Q. Chen

Introduction

Conditional generative modeling paves the way to numerous machine learning applications such as conditional image generation (Dhariwal and Nichol 2021; Rombach et al. 2022), text-to-speech synthesis (Wang et al. 2023; Le et al. 2023), and even solving decision making problems (Chen et al. 2021; Janner et al. 2021; Janner et al. 2022; Ajay et al. 2022). Models that appear ubiquitously across a variety of application domains are diffusion models (Sohl-Dickstein et al. 2015; Ho et al. 2020) and flow-based models (Song et al. 2020b; Lipman et al. 2023; Albergo and Vanden-Eijnden 2022). Majority of this development has been focused around diffusion models, where multiple forms of conditional guidance Dhariwal and Nichol 2021; Ho and Salimans 2022 have been introduced to place larger emphasis on the conditional information. While flow models have been shown to be more efficient alternatives than diffusion models (Lipman et al. 2023; Pooladian et al. 2023) in unconditional generation, requiring less computation to sample, their behavior in conditional generation tasks has not been explored as much. It also remains unclear whether conditional guidance can be applied to and help the performance of flow-based models.

In this work, we study the behavior of Flow Matching models for conditional generation. We introduce Guided Flows, an adaptation of classifier-free guidance (Ho and Salimans 2022) to Flow Matching models, showing that an analogous modification can be made to the velocity vector fields, including the optimal transport (Lipman et al. 2023) and cosine scheduling (Albergo and Vanden-Eijnden 2022) flows used by prior works.

We experimentally validate Guided Flows on a variety of applications, ranging from generative modeling over multiple modalities to offline reinforcement learning (RL), see Table 1.1. As we show in Section 4, for standard generative tasks including image synthesis and zero-shot text-to-speech generation, Guided Flows significantly improves the sampling quality over unguided counterparts (i.e., sample from the conditional distribution directly), attaining state-of-the-art (SOTA) performance.

Particularly, the integration of guidance enables us to apply flow-based models for return-conditioned plan generating in offline RL for the first time. The evaluation of offline trained RL agents is via sequential online interactions. The usage of conditional generative models for plan generation (Janner et al. 2021; Janner et al. 2022; Ajay et al. 2022) often requires the model to accurately model not only relations within the training data set, but to also generalize to unseen conditional signals during online evaluation. There, we find that guided flows generate reliable execution plans, given the current state and a target return values. Guided Flows also obtain notably higher returns than unguided flows, achieving SOTA performance as well, see Section 5.3.

In addition to its efficacy, for all these aforementioned tasks, Guided Flows also demonstrate favorable compute efficiency and performance tradeoffs. Particularly, for offline RL, Guided Flows enjoys a remarkable 10x speed up compared with diffusion models.

Related Work

Recent major developments in generative models are on building simple models that are highly-efficient to train while providing the means of using conditional inference for solving downstream applications (Janner et al. 2022; Ajay et al. 2022; Kawar et al. 2022; Pokle et al. 2023; Le et al. 2023). The predominant progress in this domain has primarily centered on diffusion models (Sohl-Dickstein et al. 2015; Ho et al. 2020; Song et al. 2020b). Within these models, two types of conditional guidance (Dhariwal and Nichol 2021; Ho and Salimans 2022) can be employed to promote the conditional information. Whether these form of conditional guidance can be integrated into and help other types of generative models, remains an open question.

As two concurrent works, Dao et al. 2023 derive an approach similar to ours but is mainly motivated heuristically for the conditional optimal transport probability path; Hu et al. 2023 develop another different approach by adding an offset to the learned vector field, without theoretical guarantees on the sample distribution. It is noteworthy that we also show the efficacy of guidance in flows across a much wider variety of domain settings. In particular, our offline RL use case is quite different from the generative modeling tasks considered in those works.

Reinforcement learning (RL) is a powerful paradigm that has been widely applied to solve complex sequential decision making tasks, such as playing games Silver et al. 2016, controlling robotics Kober et al. 2013, dialogue systems Li et al. 2016; Singh et al. 1999. The premise of offline RL is to learn effective policies solely from static datasets that consist of previous collected experiences, generated by certain unknown policies Levine et al. 2020. There, the agent is not allowed to interact with the environment, thus subsides the potential risk and cost of online interactions for high-risk domains such as healthcare.

Conditional generative models are becoming handy tools to help decision making. A rich body of work Chen et al. 2021; Janner et al. 2021; Janner et al. 2022; Ajay et al. 2022; Zhao and Grover 2023; Zheng et al. 2023 focuses on modeling trajectories as sequences of state, action, and reward tokens. Instead of optimizing the expected return as the classic RL methods, the training objective of all these methods is to simply maximize the likelihood of sequences in the offline dataset. In essence, these methods cast RL as supervised sequence modeling problems, and generative models such as Transformer Vaswani et al. 2017; Radford et al. 2018, Diffusion Models Ho et al. 2020 thus come into play. Solving RL from this new perspective improves its training stability, and further opens the door of multimodal and multitask pretraining Zheng et al. 2022; Lee et al. 2022; Reed et al. 2022, similar to other domains like language and vision Radford et al. 2018; Chen et al. 2020a; Brown et al. 2020; Lu et al. 2022. There, a generative model can be used as a policy to autoregressively generate actions Chen et al. 2021; Zheng et al. 2022; alternatively, it can generate imagined trajectories which serves as execution plans given current state and a target output such as return or goal location Janner et al. 2021; Janner et al. 2022; Ajay et al. 2022; Zhao and Grover 2023. There are other works that replace the parameterized Gaussian policies in classic RL methods by generative models to facilitate the modeling of multimodal action distributions Wang et al. 2022; Chi et al. 2023; Ward et al. 2019. Akimov et al. 2022 apply flow-based models to offline RL as conservative action encoders (and decoders), which decode actions from latent variables output by the policy.

Guided Flow Matching

where pt(⋅∣x1)p_{t}(\cdot|x_{1}) is a probability path interpolating between noise and a single data point x1x_{1}. That is, pt(⋅∣x1)p_{t}(\cdot|x_{1}) satisfies

where δx1(⋅)\delta_{x_{1}}(\cdot) is the delta probability that concentrates all its mass at x1x_{1}. Note that p0≡pp_{0}\equiv p is the noise distribution. An immediate consequence of the boundary conditions in equation (3) is that pt(⋅∣y)p_{t}(\cdot|y) indeed interpolates between noise and data, i.e., p0(⋅∣y)≡p(⋅)p_{0}(\cdot|y)\equiv p(\cdot) and p1(⋅∣y)≈q(⋅∣y)p_{1}(\cdot|y)\approx q(\cdot|y) for all yy.

Next, we assume ut(⋅∣x1)u_{t}(\cdot|x_{1}) is a velocity field that generates pt(⋅∣x1)p_{t}(\cdot|x_{1}) in the sense that solutions to equation (1), with ut(⋅∣x1)u_{t}(\cdot|x_{1}) as the velocity field and x0∼p(x0)x_{0}\sim p(x_{0}), satisfy xt∼pt(xt∣x1)x_{t}\sim p_{t}(x_{t}|x_{1}). The target velocity field for FM is then defined via

and can be proved to generate, in the sense described above, the marginal probability path pt(⋅∣y)p_{t}(\cdot|y) in equation (2). FM is trained by minimizing a tractable loss called the Conditional Flow Matching (CFM) loss (defined later), whose global minimizer is the target velocity field utu_{t}.

A popular instantiation of paths pt(x∣x1)p_{t}(x|x_{1}) are Gaussian paths defined by

By convention we denote by ∅\varnothing the null conditioning and set q(x):=q(x∣∅)q(x):=q(x|\varnothing), utθ(x):=utθ(x∣∅)u^{\theta}_{t}(x):=u^{\theta}_{t}(x|\varnothing), and ut(x):=ut(x∣∅)u_{t}(x):=u_{t}(x|\varnothing).

To justify this formula for velocity fields, we first relate ut(x∣y)u_{t}(x|y) to the score function ∇log⁡pt(x∣y)\nabla\log p_{t}(x|y) using the following lemma proved in Appendix A.1:

Let pt(x∣y)p_{t}(x|y) be a Gaussian Path defined by a scheduler (αt,σt)(\alpha_{t},\sigma_{t}), then its generating velocity field ut(x∣y)u_{t}(x|y) is related to the score function ∇log⁡pt(x∣y)\nabla\log p_{t}(x|y) by

Next, using this the lemma and plugging equation (7) for ut(x)=ut(x∣∅)u_{t}(x)=u_{t}(x|\varnothing) and ut(x∣y)u_{t}(x|y) in the r.h.s. of equation (6), we get

is the geometric weighted average of pt(x)p_{t}(x) and pt(x∣y)p_{t}(x|y).

Training guided flow follows the practice in CFG but replaces the Diffusion training loss with the Conditional Flow Matching (CFM) loss (Lipman et al. 2023). This leads to the following loss function:

Figure 3.1 shows a visualization of the effect of Guided Flows in a toy 2D example of a mixture of Gaussian distributions, where yy is the latent variable specifying the identity of the mixture component, and q(x1∣y)q(x_{1}|y) is one single Gaussian component. When the guidance weight is 1.01.0, there is no guidance performed, and we are sampling from the unconditional marginal q(x1)q(x_{1}). We see that as guidance weight increases, the samples move away from the unconditional distribution q(x1)q(x_{1}).

Conditional Generative Modeling

In this section, we perform experiments to test whether Guided Flows can help achieve better sample quality than unguided models. In particular, we consider two application settings: conditional image generation and zero-shot text-to-speech synthesis. The goal of these experiments to test Guided Flows on standard generative modeling settings, where the evaluation is purely based on sample quality, while exploring different data modalities, before we move on validating Guided Flows for more complex planning tasks in Section 5.

We downsample the official face-blurred ImageNet dataset to images of 64×64 pixels, using the open source preprocessing scripts from Chrabaszcz et al. 2017. We train Guided Flow Matching models pt(x∣y)p_{t}(x|y) where xx denotes an image and yy is the class label of that image. In particular, we consider two affine Gaussian probability paths, the optimal transport (FM-OT) path considered by Lipman et al. 2023 and the cosine scheduling (FM-CS) path considered by Albergo and Vanden-Eijnden 2022. As a baseline, we train diffusion models (DDPM; Ho et al. 2020; Song et al. 2020b) with classifier-free guidance (Ho and Salimans 2022). We report the results of both standard sampling of DDPM and also the deterministic DDIM sampling algorithm (Song et al. 2020a). All the models have the same U-Net architecture adopted from Dhariwal and Nichol 2021, and trained with the same hyperparameters and number of iterations, as listed in Table D.1.

Results are displayed in Figure 4.2. We see that guidance for Flow Matching models can drastically help increase sample quality (reducing FID from 2.542.54 to 1.681.68) using a midpoint solver with 200 number of function evaluations (NFE), i.e. 200 ODE steps. We do see that the optimal guidance weight can change depending on the compute cost (NFE), with NFE=10 having a higher optimal guidance weight. Additionally, in Figure 2(b) we show the Pareto front of the sample quality and efficiency tradeoff, plotted for each model. We note that the optimal guidance weight can be very different between each model, with DDPM noticeably requiring a larger guidance weight. Here we find that FM-OT models are slightly more efficient than the other models that we consider.

2 Zero-shot Text-to-Speech Synthesis

Given a target text and a transcribed reference audio as conditioning information yy, zero-shot text-to-speech (TTS) aims to synthesize speech resembling the audio style of the reference, which was never seen during training. As our Guided Flow model, we train a model on 60K hours ASR-transcribed English audiobooks, following the experiment setup in Le et al. 2023.

Specifically, we consider two main tasks. The first one is zero-shot TTS where the first 3 seconds of each utterance is provided and the model is requested to continue the speech. The second one is diverse speech generation, where only the text is provided to the model. In order to assess the accuracy of the generated results, we report the word error rate (WER) using automatic speech recognition (ASR) models following prior works (Wang et al. 2018).

Results of Guided Flows is provided in Table 4.1, where we also provide the results of Le et al. 2023 as reference. We also report our results when no guidance is used (weight equal to 1.0). With guidance, we see a marginal improvement for the continuation TTS task, and a much more sizable gain in the text-only TTS task. This likely due to the text-only TTS task being a much more diverse distribution.

Planning for Offline RL

We model our environment as a Markov decision process (MDP) (Bellman 1957) denoted by ⟨S,A,p,P,R,γ⟩\langle\mathcal{S},\mathcal{A},p,P,R,\gamma\rangle, where S\mathcal{S} is the state space, A\mathcal{A} is the action space, p(s0)p(s_{0}) is the distribution of the initial state, P(st+1∣st,at)P(s_{t+1}|s_{t},a_{t}) is the transition probability distribution, R(st,at)R(s_{t},a_{t}) is the deterministic reward function, and γ\gamma is the discount factor. At timestep tt, the agent observes a state st∈Ss_{t}\in\mathcal{S} and executes an action at∈Aa_{t}\in\mathcal{A}. The environment will provide the agent with a reward rt=R(st,at)r_{t}=R(s_{t},a_{t}), and also moves it to the next state st+1∼P(⋅∣st,at)s_{t+1}\sim P(\cdot|s_{t},a_{t}). Let τ\tau be a trajectory. For any length-HH subsequence τsub\tau_{\text{sub}} of τ\tau, e.g., from timestep tt to t+H−1t+H-1, we define the return-to-go (RTG) of τsub\tau_{\text{sub}} to be the sum of its discounted return g(τsub)=∑t′=tt+H−1γt′−trt′g(\tau_{\text{sub}})=\sum_{t^{\prime}=t}^{t+H-1}\gamma^{t^{\prime}-t}r_{t^{\prime}} This is slightly different from the standard RTG definition where the discounting factor is γt′−t\gamma^{t^{\prime}-t} rather than γt\gamma^{t}.. We also use s(τsub)\bm{s}(\tau_{\text{sub}}) to denote the state sequence extracted from the subsequence τsub\tau_{\text{sub}}. A deterministic inverse dynamics model (IDM) is a function f:S×S↦Af:\mathcal{S}\times\mathcal{S}\mapsto\mathcal{A} which predicts action using states: a^t=f(st,st+1)\widehat{a}_{t}=f(s_{t},s_{t+1}).

2 Our Setup

We consider the paradigm proposed by Ajay et al. 2022, where a generative model learns to predict a sequence of future states, conditioning on the current state and target output such as expected return. Intuitively, the predicted future states form a plan to reach the target output, and the predicted next state can be interpreted as the next intermediate goal on the roadmap. Based on the current state and the predicted next state, we use an inverse dynamics model to predict the action to execute. Algorithm 3 summarizes this framework. We emphasize that the target output gtg_{t} is the conditioning variable where guidance will perform, and the current state sts_{t} is a general conditioning variable. In this work, we consider the target return as our conditioning variable. Return-conditioned RL methods are widly used to solve standard RL problems where the environment produces dense rewards Srivastava et al. 2019; Kumar et al. 2019; Schmidhuber 2019; Emmons et al. 2021; Chen et al. 2021; Nguyen et al. 2022. Figure 5.1 and 5.2 plot the training and evaluation phases of our paradigm. Careful readers might notice that we resample the whole sequence s^t+1,…,s^t+H−1\widehat{s}_{t+1},\ldots,\widehat{s}_{t+H-1} at every timestep tt, but only use s^t+1\widehat{s}_{t+1} to predict a^t\widehat{a}_{t}. In fact, it is completely feasible to predict the actions for the next multiple steps using a single plan, which also saves computation. We note that replanning at every timestep is for the sake of planning accuracy, as the error will accumulate as the horizon expands, and there is a tradeoff between computational efficiency and agent performance. For diffusion models, previous works have used heuristics to improve the execution speed, e.g., reusing previously generated plans to warm-start the sampling of subsequence plans Janner et al. 2022; Ajay et al. 2022. This is beyond the scope of our paper, and we only consider replanning at every timestep for simplicity.

3 Experiments

Our experiments aim at answering the following questions:

Can conditional flows generate meaningful plans for RL problems given a target return?

How do flows compare to diffusion models, in terms of both downstream RL task performance and compute efficiency?

Compared with unguided flows, can guidance help planning?

We consider three Gym locomotions tasks, hopper, walker and halfcheetah, using offline datasets from the D4RL benchmark Fu et al. 2020. For all the experiments, we train 5 instances of each method with different seeds. For each instance, we run 20 evaluation episodes. We shall discuss the model architectures and important hyperparameters below, and we refer the readers to Appendix E for more details.

We compare our method to Decision Diffuser Ajay et al. 2022, which uses the same paradigm with diffusion models to model the state sequences.

For the locomotion tasks, we train flows for state sequences of length H=64H=64, parameterized through the velocity field uθu^{\theta}. Similar to the previous work Janner et al. 2022; Ajay et al. 2022, the velocity field uθu^{\theta} is modeled as a temporal U-Net consisting of repeated convolutional residual blocks. Both the time tt and the RTG g(s)g(\bm{s}) are projected to latent spaces via multilayer perceptrons, where tt is first transformed to its sinusoidal position encoding. For the baseline method, the diffusion model is training similarly with 200200 diffusion steps, where we use a temporal U-Net to predict the noise at each diffusion steps throughout the diffusion process.

For all the environments and datasets, we model the IDM by an MLP with 2 hidden layers and 1024 hidden units per layer. Among all the offline trajectories, we randomly sample 10% of them as the validation set. We train the IDM for 100k iterations, and use the one that yields the best validation performance.

To sample from the diffusion model, we use 200200 diffusion steps. For the sake of fair comparison, we also use 200200 ODE steps when sampling from the flow matching model. Following Ajay et al. 2022, we use the low temperature sampling technique for diffusion model, where at each diffusion step kk we sample the state sequence from N(μ^k,α2Σ^k)\mathcal{N}(\widehat{\mu}_{k},\alpha^{2}\widehat{\Sigma}_{k}) μ^k\widehat{\mu}_{k} and Σ^k\widehat{\Sigma}_{k} are the predicted mean and variance for sampling at the kkth diffusion step. We refer the readers to Ho et al. 2020 for more details. with a hand-selected temperature parameter α∈(0,1)\alpha\in(0,1). We sweep over 3 values of α\alpha for all our experiments: 0.1,0.250.1,0.25, and 0.50.5. For flow matching, we analogously set the initial distribution p0(x0)=N(0,ν2I)p_{0}(x_{0})=\mathcal{N}(0,\nu^{2}I) and we sweep over two values of ν\nu: 0.10.1 and 11.

3.1 Plan Generation (Q1)

To verify the capability of flows to generate meaningful plans that can guide the agent, we sample from a flow on the hopper-medium dataset. We randomly select a subtrajectory τsub\tau_{\text{sub}} from the dataset, and let flow condition on its first state and RTG 0.80.8. Figure 5.3 plots both state sequences. We can see that the generated state sequence is almost identical to the ground truth, demonstrating that guided flow is capable to generate meaningful plans to navigate the agent. We note that this RTG value 0.80.8 is out-of-distribution (OOD), as the maximum RTG value of the training dataset is 0.610.61. This suggests that guided flows might be even robust to OOD RTG values We note that the fundamental task for offline RL is to address the offline-to-online distribution shift. During online evaluation, an offline trained agent might encounter unseen data, potentially resulting in the generation of unreasonable actions or states that lead to poor performance. To address this issue, various notions of conservatism has been introduced into offline RL algorithms. The overarching objective of those diverse conservatism techniques is to maintain the output of the algorithm close to the training data distribution. , which we believe is an interesting property to understand and a potential direction to explore for future work, see related discussions in Chen et al. 2021; Emmons et al. 2021; Zheng et al. 2022; Nguyen et al. 2022. We refer the readers to Figure F.1 for more examples of generated plans.

3.2 Benchmark (Q2)

In this section, we conduct experiments to investigate the efficacy and efficiency of guided flows. We highlight the predominant trends from our findings:

Guided flows are on par with diffusion models with respect to the absolute performance, with a significant 10x speed up in sampling.

Throughout all our experiments, in addition to the temperature parameter for sampling, for both methods, we sweep over 3 values of RTG: 0.4, 0.5, and 0.6 for the medium-replay and medium datasets of halfcheetah, and 0.7, 0.8 and 0.9 for all the other datasets. We also sweep over 4 values of guidance parameters: 1.8,2.0,2.2,2.41.8,2.0,2.2,2.4 for diffusion models, and 1.0,1.5,2.0,2.51.0,1.5,2.0,2.5 for flows. We train both guided flows and guided diffusion models for 2 million iterations, and save checkpoints every 200k200k iteration. We evaluate the performance on all the saved checkpoints and report the best result in Table 5.1 This is because Ajay et al. 2022 report the best results over a collection of checkpoints.. The performances of flow matching and diffusion model are comparable, where flow matching performs marginally better.

A well-known pain-point of diffusion model is its computational inefficiency, since sampling from a diffusion model consists of iterative denoising steps. In our case, the diffusion model is trained with 200200 diffusion steps, thus it takes 200200 internal sampling steps to generate a sample with full quality. To accelerate sampling, many algorithms have been proposed to reduce the number of internal steps, including implicit models with deterministic sampling (DDIM, Song et al. 2020a), distillation Salimans and Ho 2022, noise schedule Chen et al. 2020b; Nichol and Dhariwal 2021; Lin et al. 2023. As a tradeoff, these methods all lead to loss in sample quality. Similarly, sampling from a flow model requires a number of internal steps to solve the ODE equation (1), and more internal steps leads to sample with better quality.

To understand the computation-vs-quality tradeoff for both flows and diffusion models, we run an ablation experiment for the hopper task, where we only sample with KK internal steps (K≤200)(K\leq 200) for both methods. In addition to the standard diffusion model (DDPM, Ho et al. 2020), we also compare with DDIM Song et al. 2020a, a deterministic sampling algorithm widely used in diverse domains. Figure 5.4 plots the normalized return we obtain versus the number of internal steps. For all 3 datasets, phase transitions occur for all three methods. Surprisingly, 1010 ODE steps are sufficient for flows to generate samples leading to the same return as 200200 ODE steps; whereas DDPM and DDIM both need 100100 diffusion steps. Again, the performance of flows is comparable to DDPM and is better than DDIM. Next, we compare the CPU time consumed by these methods. Figure 5.5 shows that the CPU time consumed by the diffusion model and the flow are roughly the same when the number of internal steps match, and it scales linearly as the number of internal steps increase. This means, compared with diffusion model, flows only need 10%10\% computing time to generate samples leading to the same downstream performance. The trend remains the same when the batch size increase, see Figure F.2.

3.3 Influence of Guidance (Q3)

Table 5.2 reports the normalized return obtained by our agents for the hopper and halfcheetah tasks, where the guidance weight varies from 1.0 to 3.0. The flows are trained on both medium and medium-replay datasets for hopper, and the medium dataset for halfcheetah. The guidance weight 1.01.0 yields unguided flows. The results show that the guided flows outperform the unguided ones on all three datasets.

Conclusion

We thoroughly explore the theory and effect of guidance for flow matching. We empirically validate the conditional generative capabilities of flow-based models trained through recently-proposed simulation-free algorithms (Lipman et al. 2023; Albergo and Vanden-Eijnden 2022) for a variety of applications, confirming its success across diverse domains. Our experiments show that conditional guidance can lead to better results for flow-based models. Moreover, guided flows excel in standard generative tasks like image synthesis and speech generation, achieving SOTA performance. Additionally, our experiments highlight both the efficacy and efficiency of guided flows in model-based planning: a significant 10x speedup in offline RL with performance on par with diffusion models. This underscores the great potential of flow matching in extending the application of generative models to planning problems, especially those demanding enhanced computational efficiency, such as online planning.

The authors thank Zihan Ding, Maryam Fazel-Zarandi, Brian Karrer, Maximilian Nickel, Mike Rabbat, Yuandong Tian, Amy Zhang, and Siyan Zhao for insightful discussions.

References

Appendix A Proofs

Proof (Lemma 1). The Gaussian probability path pt(x∣y)p_{t}(x|y) as in equation (2) is

where pt(x∣x1)=N(x∣αtx1,σt2I)p_{t}(x|x_{1})={\mathcal{N}}(x|\alpha_{t}x_{1},\sigma_{t}^{2}I). We express the score function as

The generating velocity field utu_{t} as in equation (4) is

where ut(x∣x1)=σ˙tσt(x−αtx1)+α˙tx1u_{t}(x|x_{1})=\frac{\dot{\sigma}_{t}}{\sigma_{t}}(x-\alpha_{t}x_{1})+\dot{\alpha}_{t}x_{1}. Hence, by linearity of integrals it is enough to show that

where in the last equality we used our assumption of Gaussian probability path that gives ∇log⁡pt(x∣x1)=−1σt2(x−αtx1)\nabla\log p_{t}(x|x_{1})=-\frac{1}{\sigma_{t}^{2}}(x-\alpha_{t}x_{1}). □\square

Appendix B Probability Flow ODE for Scheduler (αt,σt)(\alpha_{t},\sigma_{t})

We assume the marginal probability paths (see equation (2)) pt(x)p_{t}(x) and pt(x∣y)p_{t}(x|y) are defined with a scheduler (αt,σt)(\alpha_{t},\sigma_{t}) and data distribution q(x)q(x) and q(x∣y)q(x|y), respectively. CFG consider the probability path

Then the sampling is done with the Probability Flow ODE of diffusion models (Song et al. 2020b),

where ft=dlog⁡αtdtf_{t}=\frac{d\log\alpha_{t}}{dt}, gt2=dσt2dt−2dlog⁡αtdtσtg_{t}^{2}=\frac{d\sigma_{t}^{2}}{dt}-2\frac{d\log\alpha_{t}}{dt}\sigma_{t} (Kingma et al. 2023; Salimans and Ho 2022). Lastly,

plugging this in equation (24) is an ODE with a velocity field that coincides with the velocity field in equation (9).

Appendix C Flow Matching Sampling with Guidance for Offline RL

Comparing with standard generative modeling, the sequence model trained for RL needs to condition on the current state sts_{t}, see Section 5. Therefore, the sampling process is slightly different from Algorithm 2, as we need to zero out the vector fields corresponding to sts_{t}, as shown in Algorithm 4. Sampling for the goal-conditioned model can be done similarly.

Appendix D Image Generation Experiment Details

We train three models on ImageNet-64: DDPM (using noise prediction), FM-CS, and FM-OT.

FM-CS and FM-OT models are trained with the loss function in equation (11):

where tt is sampled uniformly in $,,b\sim\text{Bernoulli}(p_{\text{uncond}})isusedtoindicatewhetherwewillusenullcondition,is used to indicate whether we will use null condition,x_{0}isthenoise,is the noise,x_{1}andandyaresampledfromthetruedatadistribution,andare sampled from the true data distribution, andx_{t}=\alpha_{t}x_{1}+\sigma_{t}x_{0},,\dot{x}_{t}=u_{t}(x_{t}|x_{1})=\dot{\alpha}_{t}x_{1}+\dot{\sigma}_{t}x_{0}$. The noise scheduler of FM-CS is the cosine scheduler Albergo and Vanden-Eijnden 2022:

and the noise scheduler of FM-OT Lipman et al. 2023 is

DDPM models are trained with noise prediction loss as derived in Ho et al. 2020 and Song et al. 2020b:

We note that in our implementation, tt is sampled uniformly in $$. We use the VP scheduler

All three models have the same U-Net architecture adopted from Dhariwal and Nichol 2021, with hyperparameters listed below. For all the methods, we sweep the guidance weight across the range of 1.0 to 2.0 with a grid size of 0.05, and report the best results in Figure 2(b). In particular, we have reported DDPM, DDIM and FM-CS using guidance weight 0.2, and FM-OT using guidance weight 0.15.

Appendix E Offline RL Experiment Details

We summarize the architecture and other hyperparameters used for our experiments. For all the experiments, we use our own PyTorch implementation that is heavily influenced by the following codebases:

Decision Diffuser https://github.com/anuragajay/decision-diffuser Diffuser https://github.com/jannerm/diffuser/

We train both guided flows and guided diffusion models for state sequences of length H=64H=64. The probability of null conditioning puncondp_{\text{uncond}} is set to 0.250.25. The batch size is 6464. We normalize the discounted RTG by a task-specific reward scale, which is 400400 for hopper, 550550 for walker and 12001200 for halfcheetah. The final model parameter θˉ\bar{\theta} we consider is an exponential moving average (EMA) of the obtained parameters over the course of training. For every 1010 iteration, we update θˉ=βθˉ+(1−β)θ\bar{\theta}=\beta\bar{\theta}+(1-\beta)\theta, where the exponential decay parameter β=0.995\beta=0.995. We train the sequence model for 2×1062\times 10^{6} iterations, and checkpoint the EMA model every 200200k iteration.

We use a temporal U-net to model the velocity field uθu^{\theta}. It consists of 6 repeated residual blocks, where each block consists of 2 temporal convolutions followed by the group norm Wu and He 2018 and a final Mish nonlinearity activation Mish 2019. The time tt is first trainsformed to its sinusoidal position encoding and projected to a latent space via a 2-layer MLP, and the RTG g(s)g(\bm{s}) is transformed into its latent embedding via a 3-layer MLP. The model is optimized by the Adam optimzier Kingma and Ba 2014. The learning rate is 2×10−42\times 10^{-4} for hopper-medium-expert, 3×10−43\times 10^{-4} for walker-medium-replay and 10−410^{-4} for all the other datasets.

Guided Diffusion Models

We use the cosine noise schedule proposed by Nichol and Dhariwal 2021. We use a temporal U-net to model the noise ϵθ{\epsilon}_{\theta}, with the same architecture used for guided flows. The model is also optimized by the Adam optimzier, where the learning rate 2×10−42\times 10^{-4} for all the datasets.

Inverse Dynamics Model

The inverse dynamics model is modeled by an MLP with 2 hidden layers, 1024 hidden units per layer, and a 10%10\% dropout rate. We use the Adam optimizer with learning rate 10−410^{-4}. We randomly sample 10% of offline trajectories as the validation set. We train the IDM for 100k iterations, and use the one that yields the best validation performance.

Appendix F Additional Experiments

F.2 Computational Speed Comparison