Training-free Linear Image Inverses via Flows

Ashwini Pokle, Matthew J. Muckley, Ricky T. Q. Chen, Brian Karrer

Introduction

Solving an inverse problem involves recovering a clean signal from noisy measurements generated by a known degradation model. Many interesting image processing tasks can be cast as an inverse problem. Some instances of these problems are super-resolution, inpainting, deblurring, colorization, denoising etc. Diffusion models or score-based generative models (Sohl-Dickstein et al. 2015; Ho et al. 2020; Song & Ermon 2019; Song et al. 2021c) have emerged as a leading family of generative models for solving inverse problems for images (Saharia et al. 2022b; Saharia et al. 2022a; Wang et al. 2022; Chung et al. 2022a; Song et al. 2022; Mardani et al. 2023). However, sampling with diffusion models is known to be slow, and the quality of generated images is affected by the curvature of SDE/ODE solution trajectories (Karras et al. 2022). While Karras et al. 2022 observed ODE sampling for image generation could produce better results, sampling via SDE is still common for solving inverse problems, whereas ODE sampling has been rarely considered, perhaps due to the use of diffusion probability paths.

Continuous Normalizing Flow (CNF) (Chen et al. 2018b) trained with Flow Matching (Lipman et al. 2022) has been recently proposed as a powerful alternative to diffusion models. CNF (hereafter denoted flow model) has the ability to model arbitrary probability paths, and includes diffusion probability paths as a special case. Of particular interest to us are Gaussian probability paths that correspond to optimal transport (OT) displacement (McCann 1997). Recent works (Lipman et al. 2022; Albergo & Vanden-Eijnden 2022; Liu et al. 2022; Shaul et al. 2023) have shown that these conditional OT probability paths are straighter than diffusion paths, which results in faster training and sampling with these models. Due to these properties, conditional OT flow models are an appealing alternative to diffusion models for solving inverse problems.

In this work, we introduce a training-free method to utilize pretrained flow models for solving linear inverse problems. Our approach adds a correction term to the vector field that takes into account knowledge from the degradation model. Specifically, we introduce an algorithm that incorporates Π\PiGDM (Song et al. 2022) gradient correction to flow models. Given the wide availability of pretrained diffusion models, we also present a way to convert these models to arbitrary paths for our sampling procedure. Empirically, we observe images restored via a conditional OT path exhibit perceptual quality better than that achieved by the model’s original diffusion path, as well as recently proposed diffusion approaches such as Π\PiGDM (Song et al. 2022) and RED-Diff (Mardani et al. 2023), particularly in noisy settings. To summarize, our key contributions are:

We present a training-free approach for solving linear inverse problems with pretrained conditional OT flow models that adapts the Π\PiGDM correction, proposed for diffusion models, to flow sampling.

We offer a way to convert between flow models and diffusion models, enabling the use of pretrained continuous-time diffusion models for conditional OT sampling, and vice versa.

We demonstrate that images restored via our algorithm using conditional OT probability paths have perceptual quality that is largely on par with, or better than that achieved by diffusion probability paths, and other recent methods like Π\PiGDM and RED-Diff, without the need for problem-specific hyperparameter tuning.

Preliminaries

We introduce relevant background knowledge and notation from conditional diffusion and flow modeling, as well as training-free inference with diffusion models.

Conditional diffusion models.

Conditional diffusion uses a class of processes qq, and defines data to noise process pθp_{\theta} as a Markov process that in continuous time obeys a stochastic differential equation (SDE) (Song & Ermon 2019; Ho et al. 2020; Song et al. 2021c). The parameters of pθp_{\theta} are learned via minimizing a regression loss with respect to x1^\widehat{{\bm{x}}_{1}}

Conditional flow models

Alternatively, continuous normalizing flow models (Chen et al. 2018b) define the data generation process through an ODE. This leads to simpler formulations and does not introduce extra noise during intermediate steps of sample generation. Recently, simulation-free training algorithms have been designed specifically for such models (Lipman et al. 2022; Liu et al. 2022; Albergo & Vanden-Eijnden 2022), an example being the Conditional Flow Matching loss (Lipman et al. 2022),

where v^(xt,y)\widehat{{\bm{v}}}({\bm{x}}_{t},{\bm{y}}) denotes a parameterized vector field defining the ODE

If trained perfectly, the marginal distributions of xt{\bm{x}}_{t}, denoted pθ(xt∣y)p_{\theta}({\bm{x}}_{t}|{\bm{y}}), will match the marginal distributions of q(xt∣y)q({\bm{x}}_{t}|{\bm{y}}). Hence sampling from q(xt∣y)q({\bm{x}}_{t}|{\bm{y}}) involves starting from an initial value xt′∼q(xt′∣y){\bm{x}}_{t^{\prime}}\sim q({\bm{x}}_{t^{\prime}}|{\bm{y}}) and integrating the ODE from t′t^{\prime} to tt. Typically, one samples from t′=0t^{\prime}=0 since q(x0∣y)q({\bm{x}}_{0}|{\bm{y}}) is a tractable distribution. For general Gaussian probability paths q(xt∣x1,y)=N(μt(y,x1),σt(y,x1)2I)q({\bm{x}}_{t}|{\bm{x}}_{1},{\bm{y}})=\mathcal{N}(\mu_{t}({\bm{y}},{\bm{x}}_{1}),\sigma_{t}({\bm{y}},{\bm{x}}_{1})^{2}{\bm{I}}), one can set (Lipman et al. 2022)

Gaussian probability paths.

The time-dependent distributions q(xt∣y,x1)q({\bm{x}}_{t}|{\bm{y}},{\bm{x}}_{1}) are referred to as conditional probability paths. We focus on the class of affine Gaussian probability paths of the form

where non-negative αt\alpha_{t} and σt\sigma_{t} are monotonically increasing and decreasing respectively. This class includes the probability paths for conditional diffusion as well as the conditional Optimal Transport (OT) path (Lipman et al. 2022), where αt=t\alpha_{t}=t and σt=1−t\sigma_{t}=1-t. The conditional OT path used by flow models has been demonstrated to have good empirical properties, including faster inference and better sampling in practice, and has theoretical support in high-dimensions (Shaul et al. 2023). As emphasized in Lin et al. 2023, a desirable property for probability paths, obeyed by conditional OT but not commonly used diffusion paths, is to ensure q(x0∣y)q({\bm{x}}_{0}|{\bm{y}}) is known (i.e. N(0,I)\mathcal{N}(0,{\bm{I}})), as otherwise one cannot exactly sample x0{\bm{x}}_{0} which can add substantial error.

Training-free conditional inference using unconditional diffusion.

Following Eq. 6, past approaches (e.g., Chung et al. 2022a; Song et al. 2022) have used the pretrained model for the first term and approximated the second intractable term to produce an approximate x1^(xt,y)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t},{\bm{y}}). Diffusion posterior sampling (DPS) (Chung et al. 2022a) proposed to approximate q(y∣xt)q({\bm{y}}|{\bm{x}}_{t}) via q(y∣x1=x1^(xt))q({\bm{y}}|{\bm{x}}_{1}=\widehat{{\bm{x}}_{1}}({\bm{x}}_{t})). Later, Pseudo-inverse Guided Diffusion Models (Π\PiGDM) (Song et al. 2022) improved upon DPS for linear noisy observations where y=Ax+σyϵ{\bm{y}}={\bm{A}}{\bm{x}}+\sigma_{y}\epsilon, where A{\bm{A}} is some measurement matrix and ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,{\bm{I}}), by approximating q(y∣xt)q({\bm{y}}|{\bm{x}}_{t}) as N(Ax1^(xt),σy2I+rt2AAT)\mathcal{N}({\bm{A}}\widehat{{\bm{x}}_{1}}({\bm{x}}_{t}),\sigma_{\bm{y}}^{2}{\bm{I}}+r_{t}^{2}{\bm{A}}{\bm{A}}^{T}), derived via first approximating q(x1∣xt)q({\bm{x}}_{1}|{\bm{x}}_{t}) as N(x1^(xt),rt2I)\mathcal{N}(\widehat{{\bm{x}}_{1}}({\bm{x}}_{t}),r_{t}^{2}{\bm{I}}). Π\PiGDM also suggested adaptive weighting, replacing σt2/αt\sigma_{t}^{2}/\alpha_{t} with another function of time to account for the approximation.

Solving Linear Inverse Problems without Training via Flows

While our proposed algorithm is based on flow models, this poses no limitation on needing pretrained flow models. We first take a brief detour outside of inverse problems to demonstrate continuous-time diffusion models parameterized by x1^\widehat{{\bm{x}}_{1}} can be converted to flow models v^\widehat{{\bm{v}}} that have Gaussian probability paths described by Eq. 5. We separate training and sampling such that one can train with Gaussian probability path q′q^{\prime} and then perform sampling with a different Gaussian probability path qq. Equivalent expressions to this subsection have been derived in Karras et al. 2022, leveraging a more general conversion from SDE to probability flow ODE from Song et al. 2021c. Karras et al. 2022 similarly separate training and sampling for Gaussian probability paths, and demonstrate that using alternative sampling paths was beneficial. Our derivations avoid the SDE via a flow perspective and exposes a subtlety when swapping probability paths using pretrained models.

See proof on page .main-pratendconversion.tex

In particular, a diffusion model’s denoiser x1^(xt,y)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t},{\bm{y}}) trained using Gaussian probability path qq can be interchanged with a flow model’s v^(xt,y)\widehat{{\bm{v}}}({\bm{x}}_{t},{\bm{y}}) with the same qq via

Furthermore, x1^(xt,y)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t},{\bm{y}}) trained under q′q^{\prime} can be used for another Gaussian qq during sampling.

Consider two Gaussian probability paths qq and q′q^{\prime} defined by Eq. 5 with αt\alpha_{t}, σt\sigma_{t} and αt′\alpha^{\prime}_{t}, σt′\sigma^{\prime}_{t} respectively. Define t′(t)t^{\prime}(t) as the unique solution to σtαt=σt′′αt′′\frac{\sigma_{t}}{\alpha_{t}}=\frac{\sigma^{\prime}_{t^{\prime}}}{\alpha^{\prime}_{t^{\prime}}} when it exists for given tt. Then

See proof on page .main-pratendconversion.tex

So if x1^\widehat{{\bm{x}}_{1}} was trained under q′q^{\prime} it can be used for sampling under qq via evaluating at x1^(αt′(t)′xt/αt,t′(t),y)\widehat{{\bm{x}}_{1}}(\alpha^{\prime}_{t^{\prime}(t)}{\bm{x}}_{t}/\alpha_{t},t^{\prime}(t),{\bm{y}}) (with explicit time for clarity) whenever t′(t)t^{\prime}(t) exists. In particular, if pre-trained denoiser was trained with q′q^{\prime} and we perform conditional OT sampling, we utilize

where signal-to-noise ratio SNR(t)=αt/σt\text{SNR}(t)=\alpha_{t}/\sigma_{t}. The main avenue for non-existence for t′(t)t^{\prime}(t) is if the model under q′q^{\prime} is trained using a minimum SNR above zero, which induces a minimum tt for which t′(t)t^{\prime}(t) exists. When a minimum tt exists, we can only perform sampling with qq starting from xt∼q(xt∣y){\bm{x}}_{t}\sim q({\bm{x}}_{t}|{\bm{y}}). Approximating this sample is entirely analogous to approximating x0∼q′(x0∣y){\bm{x}}_{0}\sim q^{\prime}({\bm{x}}_{0}|{\bm{y}}). This error already exists for q′q^{\prime} because q′(x0∣y)q^{\prime}({\bm{x}}_{0}|{\bm{y}}) is not N(0,I)\mathcal{N}(0,{\bm{I}}) unless q′q^{\prime} is trained to zero SNR. An initialization problem cannot be avoided if q′q^{\prime} has limited SNR range by switching paths to qq.

2 Correcting the vector field of flow models

To solve linear inverse problems without any training via flow models, we derive an expression similar to Eq. 6 that relates conditional vector fields under Gaussian probability paths to unconditional vector fields.

Let qq be a Gaussian probability path described by Eq. 5. Assume we observe y∼q(y∣x1){\bm{y}}\sim q({\bm{y}}|{\bm{x}}_{1}) for arbitrary q(y∣x1)q({\bm{y}}|{\bm{x}}_{1}) and v(xt)v({\bm{x}}_{t}) is a vector field enabling sampling xt∼q(xt){\bm{x}}_{t}\sim q({\bm{x}}_{t}). Then a vector field v(xt,y)v({\bm{x}}_{t},{\bm{y}}) enabling sampling xt∼q(xt∣y){\bm{x}}_{t}\sim q({\bm{x}}_{t}|{\bm{y}}) can be written

See proof on page .main-pratendconversion.tex

We turn Theorem 1 into an training-free algorithm for solving linear inverse using flows by adapting Π{\Pi}GDM’s approximation. In particular, given v^(xt)\widehat{{\bm{v}}}({\bm{x}}_{t}) (or x1^(xt)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t})), our approximation will be

where following Π{\Pi}GDM terminology, we refer to γt=1\gamma_{t}=1 as unadaptive and other choices as adaptive weights. In general, we view adaptive weights γt≠1\gamma_{t}\neq 1 as an adjustment for error in qapp(y∣xt)q^{app}({\bm{y}}|{\bm{x}}_{t}).

For qapp(y∣xt)q^{app}({\bm{y}}|{\bm{x}}_{t}), we generalize Π{\Pi}GDM to any Gaussian probability path described by Eq. 5 via updating rt2r_{t}^{2}. We still have qapp(y∣xt)q^{app}({\bm{y}}|{\bm{x}}_{t}) is N(Ax1^(xt),σy2I+rt2AA⊤)\mathcal{N}({\bm{A}}\widehat{{\bm{x}}_{1}}({\bm{x}}_{t}),\sigma^{2}_{y}I+r^{2}_{t}{\bm{A}}{\bm{A}}^{\top}) when q(x1∣xt)q({\bm{x}}_{1}|{\bm{x}}_{t}) is approximated as N(x1^(xt),rt2I)\mathcal{N}(\widehat{{\bm{x}}_{1}}({\bm{x}}_{t}),r_{t}^{2}I). We choose rt2r_{t}^{2} by following Π{\Pi}GDM’s derivation, noting that if q(x1)q({\bm{x}}_{1}) is N(0,I)\mathcal{N}(0,{\bm{I}}) and q(xt∣x1)q({\bm{x}}_{t}|{\bm{x}}_{1}) is N(αtx1,σt2I)\mathcal{N}(\alpha_{t}{\bm{x}}_{1},\sigma_{t}^{2}I), then

When αt=1\alpha_{t}=1, we recover Π{\Pi}GDM’s rt2r_{t}^{2} as expected under their Variance-Exploding path specification.

Initializing conditional diffusion model sampling at t>0t>0 has been proposed by Chung et al. 2022c. For flows, we similarly want xt∼q(xt∣y){\bm{x}}_{t}\sim q({\bm{x}}_{t}|{\bm{y}}) at initialization time tt. In our experiments, we examine (approximately) initializing at different times t>0t>0 using

for ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,{\bm{I}}) when y{\bm{y}} is the correct shape. For super-resolution, we use nearest-neighbor interpolation on y{\bm{y}} instead. We also consider using A†y{\bm{A}}^{\dagger}y as an ablation in the Appendix C (where A†{\bm{A}}^{\dagger} is the pseudo-inverse of A{\bm{A}} (Song et al. 2022)). We may be forced to use this initialization for flow sampling due to converting a diffusion model not trained to zero SNR. However as shown in (Chung et al. 2022c) for diffusion, this initialization can improve results more generally. Conceptually, if the resulting xt{\bm{x}}_{t} is closer to xt∼q(xt∣y){\bm{x}}_{t}\sim q({\bm{x}}_{t}|{\bm{y}}) than achieved via starting from an earlier time t′t^{\prime} and integrating, then this initialization can result in less overall error.

Algorithm summary.

Putting this altogether, our proposed approach using flow sampling and conditional OT probability paths is succinctly summarized in Algorithm 1, derived via inserting αt=t\alpha_{t}=t and σt=1−t\sigma_{t}=1-t. Unlike Π\PiGDM, we propose unadaptive weights γt=1\gamma_{t}=1. By default, we set initialization time t=0.2t=0.2. The algorithm therefore has no additional hyperparameters to tune over traditional diffusion or flow sampling. In Appendix B, we detail our algorithm for other Gaussian probability paths, and the equivalent formulation when a pretrained vector field is available instead.

Experiments

We verify the effectiveness of our proposed approach on three datasets: face-blurred ImageNet 64×6464\times 64 and 128×128128\times 128 (Deng et al. 2009; Russakovsky et al. 2015; Yang et al. 2022), and AnimalFacesHQ (AFHQ) 256×256256\times 256 (Choi et al. 2020). We report our results on 1010K randomly sampled images from validation split of ImageNet, and 15001500 images from test split of AFHQ.

Tasks.

We report results on the following linear inverse problems: inpainting (center-crop), Gaussian deblurring, super-resolution, and denoising. The exact details of the measurement operators are: 1) For inpainting, we use centered mask of size 20×2020\times 20 for ImageNet-64, 40×4040\times 40 for ImageNet-128, and 80×8080\times 80 for AFHQ. In addition, for images of size 256×256256\times 256, we also use free-form masks simulating brush strokes similar to the ones used in Saharia et al. 2022a; Song et al. 2022. 2) For super-resolution, we apply bicubic interpolation to downsample images by 4×4\times for datasets that have images with resolution 256×256256\times 256 and downsample images by 2×2\times otherwise. 3) For Gaussian deblurring, we apply Gaussian blur kernel of size 61×6161\times 61 with intensity value 11 for ImageNet-6464 and ImageNet-128128, and 61×6161\times 61 with intensity value 33 for AFHQ. 4) For denoising, we add i.i.d. Gaussian noise with σy=0.05\sigma_{y}=0.05 to the images. For tasks besides denoising, we consider i.i.d. Gaussian noise with σy=0\sigma_{y}=0 and 0.050.05 to the images. Images x1{\bm{x}}_{1} are normalized to range $$.

Implementation details.

We trained our own continuous-time conditional VP-SDE model, and conditional Optimal Transport (conditional OT) flow model from scratch on the above datasets following the hyperparameters and training procedure outlined in Song et al. 2021c and Lipman et al. 2022. These models are conditioned on class labels, not noisy images. All derivations hold with class label cc since q(y∣c,x1)=q(y∣x1)q({\bm{y}}|c,{\bm{x}}_{1})=q({\bm{y}}|{\bm{x}}_{1}) (i.e. the noisy image is independent of class label given the image). We use the open-source implementation of the Euler method provided in torchdiffeq library (Chen 2018) to solve the ODE in our experiments. Our choice of Euler is intentionally simple, as we focus on flow sampling with the conditional OT path, and not on the choice of ODE solver.

Metrics.

We follow prior works (Chung et al. 2022a; Kawar et al. 2022) and report Fréchet Inception Distance (FID) (Heusel et al. 2017), Learned Perceptual Image Patch Similarity (LPIPS) (Zhang et al. 2018), peak signal-to-noise ratio (PSNR), and structural similarity index (SSIM). We use open-source implementations of these metrics in the TorchMetrics library (Detlefsen et al. 2022).

Methods and baselines.

We use our two pretrained model checkpoints— a conditional OT flow model and continuous VP-SDE diffusion model, and perform flow sampling with both conditional OT and Variance-Preserving (VP) paths, labeling our methods as OT-ODE and VP-ODE respectively. Because qualitative results are identical and quantitative results similar, we only include the VP-SDE diffusion model in the main text, and include the conditional OT flow model in Appendix D. We compare our OT-ODE and VP-ODE methods against Π\PiGDM (Song et al. 2022) and RED-Diff (Mardani et al. 2023) as relevant baselines. We selected these baselines because they achieve state-of-the-art performance in solving linear inverse problems using diffusion models. The code for both baseline methods is available on github, and we make minimal changes while reimplementing these methods in our codebase. A fair comparison between methods requires considering the number of function evaluations (NFEs) used during sampling. We utilize at most 100100 NFEs for our OT-ODE and VP-ODE sampling (see Appendix D), and utilize 100100 for Π\PiGDM as recommended in Song et al. 2022. We allow RED-Diff 10001000 NFEs since it does not require gradients of x1^\widehat{{\bm{x}}_{1}}. For OT-ODE following Algorithm 1, we use γt=1\gamma_{t}=1 and initial t=0.2t=0.2 for all datasets and tasks. For VP-ODE following Algorithm 2 in the Appendix, we use γt=αtαt2+σt2\gamma_{t}=\sqrt{\frac{\alpha_{t}}{\alpha_{t}^{2}+\sigma_{t}^{2}}} and initial t=0.4t=0.4 for all datasets and tasks. Ablations of these mildly tuned hyperparameters are shown in Appendix C. We extensively tuned hyperparameters for RED-Diff and Π\PiGDM as described in Appendix F, including different hyperparameters per dataset and task.

1 Experimental Results

We report quantitative results for the VP-SDE model, across all datasets and linear measurements, in Fig. 2 for σy=0.05\sigma_{y}=0.05, and in Fig. 9 within Appendix D for σy=0\sigma_{y}=0. Additionally, we report results for the conditional OT flow model in Fig. 11 and Fig. 10 for σy=0.05\sigma_{y}=0.05 and σy=0\sigma_{y}=0, respectively, in Appendix D. Exact numerical values for all the metrics across all datasets and tasks can also be found in Appendix D. Qualitative images have been selected for demonstration purposes.

We report qualitative noisy results for the VP-SDE model in Fig. 3 and for the conditional OT flow (cond-OT) model in Fig. 12. We observe that OT-ODE and VP-ODE outperforms Π\PiGDM and RED-Diff, both qualitatively and quantitatively, across all datasets for σy=0.05\sigma_{y}=0.05. As shown in these figures, Π\PiGDM tends to sharpen the images, which sometimes results in unnatural textures in the images. Further, we also observe some unnatural textures and background noise with RED-Diff for σy=0.05\sigma_{y}=0.05. For σy=0\sigma_{y}=0, OT-ODE has better FID and LPIPS, but Π\PiGDM shows improved PSNR and SSIM. Fig. 21 and Fig. 18 show qualitative examples for σy=0\sigma_{y}=0.

Super-resolution.

We report qualitative noisy results for the VP-SDE model in Fig. 4 and for the cond-OT model in Fig. 13. OT-ODE consistently achieves better FID, LPIPS and PSNR metrics compared to other methods for σy=0.05\sigma_{y}=0.05 (See Figs. 2 and 11). Similar to Gaussian deblurring, Π\PiGDM tends to produce sharper edges. This is certainly desirable to achieve good super-resolution, but sometimes this results in unnatural textures in the images (See Fig. 4). RED-Diff for σy=0.05\sigma_{y}=0.05 gives slightly blurry images. In our experiments, we observe RED-Diff is sensitive to the values of σy\sigma_{y}, and we get good quality results for smaller values of σy\sigma_{y}, but the performance deteriorates with increase in value of σy\sigma_{y}. For σy=0\sigma_{y}=0, as shown in Fig. 19 and Fig. 22, all the methods achieve comparable performance and the method declared best varies per metric and dataset.

Inpainting.

For centered mask inpainting, OT-ODE outperforms Π\PiGDM and RED-Diff in terms of LPIPS, PSNR and SSIM across all datasets at σy=0.05\sigma_{y}=0.05. Regarding FID, OT-ODE performs comparably to or better than VP-ODE (See Figs. 11 and 2). Similar observations hold true for inpainting with freeform mask on AFHQ. We present qualitative noisy results for the VP-SDE model in Fig. 5 and the cond-OT model in Fig. 14. As evident in these images, OT-ODE can result in more semantically meaningful inpainting (for instance, the shape of bird’s neck, and shape of hot-dog bread in Fig. 5). In contrast, the inpainted regions generated by RED-Diff tend be blurry and less semantically meaningful. However, we note that OT-ODE (and VP-ODE) inpainting occasionally produces artifacts in the inpainted region as the resolution of image increases. We show examples of such negative inpainting results in Section D.2. Empirically, we observe that performance of RED-Diff and Π\PiGDM improves as σy\sigma_{y} decreases. For σy=0\sigma_{y}=0, RED-Diff achieves higher PSNR and SSIM, but performs worse than OT-ODE in terms of FID and LPIPS (Refer to Fig. 9). OT-ODE’s tendency to produce inpainting artifacts for higher resolution images remains for σy=0\sigma_{y}=0, and can occur for the same images as σy=0.05\sigma_{y}=0.05. These artifacts can significantly degrade the pixel-based metrics PSNR and SSIM more than the perceptual metrics such as FID and LPIPS. We further note that noiseless inpainting for OT-ODE can be improved by incorporating null-space decomposition (Wang et al. 2022). We describe this adjustment in Appendix E.

Related Work

The challenge of solving noisy linear inverse problems without any training has been tackled in many ways, often with other solution concepts than posterior sampling (Elad et al. 2023). Utilizing a diffusion model has a host of recent research that we build upon. Our state-of-the-art baselines Π\PiGDM (Song et al. 2022) and RED-Diff (Mardani et al. 2023) correspond to lines of research in gradient-based corrections and variational inference.

RED-Diff (Mardani et al. 2023) approximates intractable q(x1∣y)q({\bm{x}}_{1}|{\bm{y}}) directly using variational inference, solving for parameters via optimization. RED-Diff was reported to have mode-seeking behavior confirmed by our results where RED-Diff performed better for noiseless inference. Another earlier variational inference method is Denoising Diffusion Restoration Models (DDRM) (Kawar et al. 2022). DDRM showed SVD can be memory-efficient for image applications, and we adapt their SVD implementations for super-resolution and blur. DDRM incorporates noiseless method ILVR (Choi et al. 2021), and leverages a measurement-dependent forward process (i.e. q(xt∣y,x1)≠q(xt∣x1)q({\bm{x}}_{t}|{\bm{y}},{\bm{x}}_{1})\neq q({\bm{x}}_{t}|{\bm{x}}_{1})) like earlier SNIPS (Kawar et al. 2021). SNIPS collapses in special cases to variants proposed in Song & Ermon 2019; Song et al. 2021c; Kadkhodaie & Simoncelli 2020 for linear inverse problems.

Discussion, Limitations, and Future work

We have presented a training-free approach to solve linear inverse problems using flows that can leverage either pretrained diffusion or flow models. The algorithm is simple, stable, and requires no hyperparameter tuning when used with conditional OT probability paths. Our method combines past ideas from diffusion including Π\PiGDM and early starting with the conditional OT probability path from flows, and our results demonstrate that this combination can solve inverse problems for both noisy and noiseless cases across a variety of datasets. Our algorithm using the conditional OT path (OT-ODE) produced results superior to the VP path (VP-ODE) and also to Π\PiGDM and REDDiff for noisy inverse problems. For the noiseless case, the perceptual quality from OT-ODE is on par with Π\PiGDM for super-resolution and gaussian deblurring, but lagging for inpainting due to image artifacts.

Another important limitation, shared with most of the past related research, is a restriction to linear observations with scalar variance. Our method can extend to arbitrary covariance, but non-linear observations are more complex. Non-linear observations occur with image inverse tasks when utilizing latent, not pixel-space, diffusion or flow models. Applying our approach to such measurements requires devising an alternative qapp(y∣xt)q^{app}({\bm{y}}|{\bm{x}}_{t}). Another shared limitation is that we consider the non-blind setting with known A{\bm{A}} and σy\sigma_{y}.

Future research could tackle these limitations. For non-linear observations in latent space, we could perhaps build upon Rout et al. 2023 that uses a latent diffusion model for linear inverses. For the blind setting, we might start from blind extensions to DPS and DDRM (Chung et al. 2023; Murata et al. 2023). As demonstrated here, we may be able to adapt and possibly improve these approaches via conversion to flow sampling using conditional OT paths.

References

Appendix A Proofs

For clarity, we restate Theorems and Lemmas from the main text before giving their proof.

Appendix B Our method for any Gaussian probability path

Algorithm 1 in the main text is specific to conditional OT probability paths. Here we provide Algorithm 2 for any Gaussian probability path specified by Eq. 5. Algorithm 1 and Algorithm 2 are written assuming a denoiser x1^(xt)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t}) is provided from a pretrained diffusion model. For completeness, we also include equivalent Algorithm 3 that assumes v^(xt)\widehat{{\bm{v}}}({\bm{x}}_{t}) is provided from a pretrained flow model. In all cases, the vector field or denoiser is evaluated only once per iteration.

Our VP-ODE sampling results correspond to αt\alpha_{t} and σt\sigma_{t} given from the Variance-Preserving path, which can be found in (Lipman et al. 2022).

Appendix C Ablation Study

We initialize the flow at time t>0t>0 as xt=αty+σtϵ{\bm{x}}_{t}=\alpha_{t}{\bm{y}}+\sigma_{t}\epsilon (y-init) where ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,{\bm{I}}). Another choice of initialization is to use xt=αtA†y+σtϵ{\bm{x}}_{t}=\alpha_{t}{\bm{A}}^{\dagger}{\bm{y}}+\sigma_{t}\epsilon. However, empirically we find that this initialization performs worse that y-init on cond-OT model with OT-ODE sampling. We summarize the results of our ablation study in Table 1. We find that on Gaussian deblurring, initialization with A†y{\bm{A}}^{\dagger}{\bm{y}} does worse than y-init, while the performance of both the initializations is comparable for super-resolution. In all our experiments, we use y-init, due to its better performance on Gaussian deblurring.

Ablation over γt\gamma_{t} for VP-ODE sampling.

We compare the performance of γt=1\gamma_{t}=1 against γt=αtαt2+σt2\gamma_{t}=\sqrt{\frac{\alpha_{t}}{\alpha_{t}^{2}+\sigma_{t}^{2}}}. We show results of VP-ODE sampling with VP-SDE model in Table 2 and Table 3. As seen our choice of γt\gamma_{t} outperform γt=1\gamma_{t}=1 across all the metrics on face-blurred ImageNet-128.

Variation of performance with NFEs.

We analyze the variation in performance of OT-ODE, VP-ODE and Π\PiGDM for solving linear inverse problems as NFEs are varied. The results have been summarized in Fig. 6. We observe that OT-ODE consistently outperforms VP-ODE and Π\PiGDM across all measurements in terms of FID and LPIPS metrics, even for NFEs as small as 2020. We also note that the choice of starting time matters to achieve good performance with OT-ODE. For instance, starting at t=0.4t=0.4 outperforms t=0.2t=0.2 when NFEs are small, but eventually as NFEs is increased, t=0.2t=0.2 performs better. We also note that Π\PiGDM achieves higher values of PSNR and SSIM at smaller NFEs for super-resolution but has inferior FID and LPIPS compared to OT-ODE.

Choice of starting time.

We plot the variation in performance of OT-ODE and VP-ODE sampling with change in start times for conditional OT model and VP-SDE model on AFHQ dataset in Fig. 7 and Fig. 8, respectively. We note that in general, OT-ODE sampling achieves optimal performance across all measurements and all metrics at t=0.2t=0.2 while VP-ODE sampling achieves optimal performance between start times of t=0.3t=0.3 and 0.40.4. In this work, for all the experiments, we use t=0.2t=0.2 for OT-ODE sampling and t=0.4t=0.4 for VP-ODE sampling.

Appendix D Additional Empirical Results

The main text includes Fig. 2 with σy=0.05\sigma_{y}=0.05 produced with the denoiser from the continuous-time VP-SDE diffusion model showing plots of various metrics across all datasets and tasks. Here we provide the same for our pre-trained conditional OT flow matching model using Algorithm 3 in Fig. 11, and noiseless figures for both models in Figs. 9 and 10. To save compute, for the flow model we only include our Π\PiGDM baseline as RED-Diff required extensive hyperparameter tuning. The qualitative results using the flow model instead of diffusion model checkpoint are identical.

This section also includes tables containing the numerical values of metrics across all datasets and tasks. The tables are hierarchically organized by noise, dataset, and task in consistent ordering. Noisy results with σy=0.05\sigma_{y}=0.05 are in Tables 4, 5, 6, 7, 8 and 9 and noiseless results with σy=0\sigma_{y}=0 are in Tables 10, 11, 12, 13, 8 and 9.

D.2 Negative results from Inpainting

Appendix E Noiseless null and range space decomposition

When σy2=0\sigma_{y}^{2}=0, we can produce a vector field approximation with even lower Conditional Flow Matching loss by applying a null-space and range-space decomposition motivated by DDNM (Wang et al. 2022). In particular, when y=Ax1{\bm{y}}={\bm{A}}{\bm{x}}_{1}, we have that A†y=A†Ax1{\bm{A}}^{\dagger}{\bm{y}}={\bm{A}}^{\dagger}{\bm{A}}{\bm{x}}_{1} (where A†{\bm{A}}^{\dagger} is the pseudo-inverse of A{\bm{A}}) and so

So when σy2=0\sigma_{\bm{y}}^{2}=0, it is only necessary to approximate the second term, as the first term is known through y{\bm{y}}. The regression loss is minimized for the first term automatically and x1^(xt,y)\widehat{{\bm{x}}_{1}}({\bm{x}}_{t},{\bm{y}}) is only responsible for predicting the second term.

In our experiments, we find that null space decomposition helps in inpainting but not other measurements. We summarize the results in Tables 16, 17, 19, 18, 20 and 21 and show qualitative results for inpainting in Figs. 28, 29 and 27.

Appendix F Baselines

We closely follow the official code available on github while implementing Π\PiGDM. For noisy case, we closely follow the Algorithm 1 in the appendix of Song et al. 2022. We use adaptive weighted guidance for both noiseless and noisy cases as in the original work. We always use uniform spacing while iterating the timestep over 100100 steps. We use ascending time from 00 to 11. Note that the original paper uses descending time from TT to 00. According to the notational convention used in this paper, this is equivalent to ascending time from 00 to 11. For the choice of rt2r_{t}^{2}, we consider the values derived from both variance exploding formulation and variance preserving formulation.

Value of rt2r_{t}^{2}.

Π\PiGDM sets the value of rt2=σ1−t21+σ1−t2r_{t}^{2}=\frac{\sigma_{1-t}^{2}}{1+\sigma_{1-t}^{2}} for VE-SDE, where q(xt∣x1)=N(x1,σ1−t2I)q({\bm{x}}_{t}|{\bm{x}}_{1})=\mathcal{N}({\bm{x}}_{1},\sigma_{1-t}^{2}{\bm{I}}). We can follow the same procedure as outlined in Song et al. 2022, and solve for rt2r_{t}^{2} in closed form for VP-SDE. We know for that VP-SDE, q(xt∣x1)=N(α1−tx1,(1−α1−t2)I)q({\bm{x}}_{t}|{\bm{x}}_{1})=\mathcal{N}(\alpha_{1-t}{\bm{x}}_{1},(1-\alpha_{1-t}^{2}){\bm{I}}), where αt=e−12T(t)\alpha_{t}=e^{-\frac{1}{2}T(t)}, T(t)=∫0tβ(s)dsT(t)=\int_{0}^{t}\beta(s)ds, and β(s)\beta(s) is the noise scale function. Using equation 13 for VP-SDE gives rt2=1−α1−t2r_{t}^{2}=1-\alpha_{1-t}^{2}. We can also obtain an alternate rt2r_{t}^{2} by plugging in value of σt2\sigma_{t}^{2} for VP-SDE into the expression of rt2r_{t}^{2} derived for VE-SDE, which evaluates to rt2=1−α1−t22−α1−t2r_{t}^{2}=\frac{1-\alpha_{1-t}^{2}}{2-\alpha_{1-t}^{2}}. Empirically, we find that rt2r_{t}^{2} for VE-SDE marginally outperforms VP-SDE. We report performance of Π\PiGDM with both choices of rt2r_{t}^{2} in Tables 22, 23 and 24.

Choice of starting time.

For OT-ODE sampling and VP-ODE sampling, we observe that starting at time t>0t>0 improves the performance. We therefore perform an ablation study on Π\PiGDM baseline, and vary the start time to verify whether starting at t>0t>0 helps to improve the performance. We plot the metrics for three different measurements in Fig. 30. We observe that starting later at time t>0t>0 consistently leads to worse performance compared to starting at time t=0t=0. Therefore, for all our experiments with Π\PiGDM, we always start at time t=0t=0.

F.2 RED-Diff

We use VP-SDE model for all experiments with RED-Diff. We closely follow the official code available on github while implementing RED-Diff. Similar to (Mardani et al. 2023), we always use uniform spacing while iterating the timestep over 10001000 steps. We use ascending time from 00 to 11. Note that the original paper uses descending time from TT to 00. According to the notational convention used in this paper, this is equivalent to ascending time from 00 to 11. We use Adam optimizer and use the momentum pair (0.9,0.99)(0.9,0.99) similar to the original work. Further, we use initial learning rate of 0.10.1 for AFHQ and ImageNet-128, as used in the original work, and learning rate of 0.010.01 for ImageNet-64. We use batch size of 11 for all the experiments. Finally, we extensively tuned the regularization hyperparameter λ\lambda to find the value that results in optimal performance across all metrics. We summarize the results of our experiments in Tables 25, 26, 27, 28, 29 and 30. We note that more extensive tuning may be able to find better performing hyperparameters but this goes against the intent of a training-free algorithm.

Appendix G Additional Background

In this section, we follow the notation used in the prior work by Lipman et al. 2022.

CNFs are usually trained by optimizing the maximum likelihood objective. As shown in Chen et al. 2018a, the exact likelihood computation can be done via relatively cheap operations despite the Jacobian term. However, this requires restricting the architecture of the neural network to constrain the Jacobian term. FFJORD (Grathwohl et al. 2018) improves upon this by proposing a method that uses Hutchinson’s trace estimator to compute log density, and allows CNFs with free-form Jacobians, thereby removing any restrictions on the architecture. This approach has difficulties for high-dimensional images where the trace estimator is noisy. Flow Matching provides an alternative, scalable approach to training CNFs with arbitrary architectures.

Flow Matching.

Suppose we have samples from an unknown data distribution x1∼q(x1){\bm{x}}_{1}\sim q({\bm{x}}_{1}). Let ptp_{t} denote a probability path from the prior distribution p0p_{0} to the data distribution p1p_{1} that is approximately equal to qq. Flow Matching loss is defined as

where ut(x){\bm{u}}_{t}({\bm{x}}) is a vector field that generates the probability path pt(x)p_{t}({\bm{x}}), and θ\theta denotes trainable parameters of the CNF. In practice, we usually do not have any prior knowledge on ptp_{t} and utu_{t}, and thus this objective is intractable. Inspired by diffusion models, Lipman et al. 2022 propose Conditional Flow Matching, where both the probability paths and the vector fields are conditioned on the sample x1∼q(x1){\bm{x}}_{1}\sim q({\bm{x}}_{1}). The exact objective for Conditional Flow matching is given by