The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation
Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, David J. Fleet
Introduction
Diffusion models have emerged as powerful generative models for high fidelity image synthesis, capturing rich knowledge about the visual world . However, at first glance, it is unclear whether these models can be as effective on many classical computer vision tasks. For example, consider two dense vision estimation tasks, namely, optical flow, which estimates frame-to-frame correspondences, and monocular depth perception, which makes depth predictions based on a single image. Both tasks are usually treated as regression problems and addressed with specialized architectures and task-specific loss functions, e.g., cost volumes, feature warps, or suitable losses for depth. Without these specialized components or the regression framework, general generative techniques may be ill-equipped and vulnerable to both generalization and performance issues.
In this paper, we show that these concerns, while valid, can be addressed and that, surprisingly, a generic, conventional diffusion model for image to image translation works impressively well on both tasks, often outperforming the state of the art. In addition, diffusion models provide valuable benefits over networks trained with regression; in particular, diffusion allows for approximate inference with multi-modal distributions, capturing uncertainty and ambiguity (e.g. see Figure 1).
One key barrier to training useful diffusion models for monocular depth and optical flow inference concerns the amount and quality of available training data. Given the limited availability of labelled training data, we propose a training pipeline comprising multi-task self-supervised pre-training followed by supervised pre-training using a combination of real and synthetic data. Multi-task self-supervised pre-training leverages the strong performance of diffusion models on tasks like colorization and inpainting [e.g., 54]. We also find that supervised (pre-)training with a combination of real and large-scale synthetic data improves performance significantly.
A further issue concerns the fact that many existing real datasets for depth and optical flow have noisy and incomplete ground truth annotations. This presents a challenge for the conventional training framework and iterative sampling in diffusion models, leading to a problematic distribution shift between training and inference. To mitigate these issues we propose the use of an loss for robustness, infilling missing depth values during training, and the introduction of step-unrolled denoising diffusion. These elements of the model are shown through ablations to be important for both depth and flow estimation.
We formulate optical flow and monocular depth estimation as image to image translation with generative diffusion models, without specialized loss functions and model architectures.
We identify and propose solutions to several important issues w.r.t data. For both tasks, to mitigate distribution shift between training and inference with noisy, incomplete data, we propose infilling, step-unrolling, and an loss during training. For flow, to improve generalization, we introduce a new dataset mixture for pre-training, yielding a RAFT baseline that outperforms all published methods in zero-shot performance on the Sintel and KITTI training benchmarks.
Our diffusion models is competitive with or surpasses SOTA for both tasks. For monocular depth estimation we achieve a SOTA relative error of 0.074 on the NYU dataset and perform competitively on KITTI. For flow, diffusion surpasses the stronger RAFT baseline by a large margin in pre-training and our fine-tuned model achieves an Fl-all outlier rate of 3.26% on the public KITTI test benchmark, 25% lower than the best published method .
Our diffusion model is also shown to capture flow and depth uncertainty, and the iterative denoising process enables zero-shot, coarse-to-fine refinement, and imputation.
Related work
Optical flow and depth estimation have been extensively studied. Here we briefly review only the most relevant work, and refer the interested readers to the references cited therein.
Optical flow. The predominant approach to optical flow is regression-based, with a focus on specialized network architectures to exploit domain knowledge, e.g., cost volume construction , coarse-to-fine estimation , occlusion handling , or iterative refinement , as evidenced by public benchmark datasets . Some recent work has also advocated for generic architectures: Perceiver IO introduces a generic transformer-based model that works for any modality, including optical flow and language modeling. Regression-based methods, however, only give a single prediction of the optical flow and do not readily capture uncertainty or ambiguity in the flow. Our work introduces a surprisingly simple, generic architecture for optical flow using a denoising diffusion model.
We find that this generic generative model is surprisingly effective for optical flow, recovering fine details on motion boundaries, while capturing multi-modality of the motion distribution.
Monocular depth. Monocular depth estimation has been a long-standing problem in computer vision with recent progress focusing on specialized loss functions and architectures such as the use of multi-scale networks , adaptive binning and weighted scale-shift invariant losses . Large-scale in-domain pre-training has also been effective for depth estimation , which we find to be the case here as well. We build on this rich literature, but with a simple, generic architecture, leveraging recent advances in generative models.
Diffusion models. Diffusion models are latent-variable generative models trained to transform a sample of a Gaussian noise into a sample from a data distribution . They comprise a forward process that gradually annihilates data by adding noise, as ‘time’ increases from 0 to 1, and a learned generative process that reverses the forward process, starting from a sample of random noise at and incrementally adding structure (attenuating noise) as decreases to 0. A conditional diffusion model conditions the steps of the reverse process (e.g., on labels, text, or an image).
Central to the model is a denoising network that is trained to take a noisy sample at some time-step , along with a conditioning signal , and predict a less noisy sample. Using Gaussian noise in the forward process, one can express the training objective over the sequence of transitions (as slowly decreases) as a sum of non-linear regression objectives, with the L2 loss (here with the -parameterization):
where , , and where is computed with a pre-determined noise schedule. For inference (i.e., sampling), one draws a random noise sample , and then iteratively uses to estimate the noise, from which one can compute the next latent sample , for .
Self-supervised pre-training. Prior work has shown that self-supervised tasks such as colorization and masked prediction serve as effective pre-training for downstream vision tasks. Our work also confirms the benefit of self-supervised pre-training for diffusion-based image-to-image translation, by establishing a new SOTA on optical flow and monocular depth estimation while also representing multi-modality and supporting zero-shot coarse-to-fine refinement and imputation.
Model Framework
In contrast to the conventional monocular depth and optical flow methods, with rich usage of specialized domain knowledge on their architecture designs, we introduce simple, generic architectures and loss functions. We replace the inductive bias in state-of-the-art architectures and losses with a powerful generative model along with a combination of self-supervised pre-training and supervised training on both real and synthetic data.
The denoising diffusion model (Figure 2) takes a noisy version of the target map (i.e., a depth or flow) as input, along with the conditioning signal (one RGB image for depth and two RGB images for flow). The denoiser effectively provides a noise-free estimate of the target map (i.e., ignoring the specific loss parameterization used). The training loss penalizes residual error in the denoised map, which is quite distinct from typical image reconstruction losses used in optical flow estimation.
Given that we train these models with a generic denoising objective, without task-specific inductive biases in the form of specialized architectures, the choice of training data becomes critical. Below we discuss the datasets used and their contributions in detail. Because training data with annotated ground truth is limited for many dense vision tasks, here we make extensive use of synthetic data in the hope that the geometric properties acquired from synthetic data during training will transfer to different domains, including natural images.
AutoFlow has recently emerged as a powerful synthetic dataset for training flow models. We were surprised to find that training on AutoFlow alone is insufficient, as the diffusion model appears to devote a significant fraction of its representation capacity to represent the shapes of AutoFlow regions, rather than solving for correspondence. As a result, models trained on AutoFlow alone exhibit a strong bias to generate flow fields with polygonal shaped regions, much like those in AutoFlow, often ignoring the shapes of boundaries in the two-frame RGB inputs (e.g. see Figure 3).
To mitigate bias induced by AutoFlow in training, we further mix in three synthetic datasets during training, namely, FlyingThings3D , Kubric and TartanAir . Given a model pre-trained on AutoFlow, for compute efficiency, we use a greedy mixing strategy where we fix the relative ratio of the previous mixture and tune the proportion of the newly added dataset. We leave further exploration of an optimal mixing strategy to future work. Zero-shot testing of the model on Sintel and KITTI (see Table 2 and Fig. 3) shows substantial performance gains with each additional synthetic dataset.
We find that pre-training is similarly important for depth estimation (see Table 7). We learn separate indoor and outdoor models. For the indoor model we pre-train on a mix of ScanNet and SceneNet RGB-D . The outdoor model is pre-trained on the Waymo Open Dataset .
2 Real data: Challenges with noisy, incomplete ground truth
Ground truth annotations for real-world depth or flow data are often sparse and noisy, due to highly reflective surfaces, light absorbing surfaces , dynamic objects , etc. While regression-based methods can simply compute the loss on pixels with valid ground truth, corruption of the training data is more challenging for diffusion models. Diffusion models perform inference through iterative refinement of the target map conditioned on RGB image data . It starts with a sample of Gaussian noise , and terminates with a sample from the predictive distribution . A refinement step from time to , with , proceeds by sampling from the parameterized distribution ; i.e., each step operates on the output from the previous step. During training, however, the denoising steps are decoupled (see Eqn. 1), where the denoising network operates on a noisy version of the ground truth depth map instead of the output of the previous iteration (reminiscent of teaching forcing in RNN training ). Thus there is a distribution shift between marginals over the noisy target maps during training and inference, because the ground truth maps have missing annotations and heavy-tailed sensor noise while the noisy maps obtained from the previous time step at inference time should not. This distribution shift has a very negative impact on model performance. Nevertheless, with the following modifications during training we find that the problems can be mitigated effectively.
Infilling. One way to reduce the distribution shift is to impute the missing ground truth. We explored several ways to do this, including simple interpolation schemes, and inference using our model (trained with nearest neighbor interpolation). We find that nearest neighbor interpolation is sufficient to impute missing values in the ground truth maps in the depth and flow field training data.
Despite the imputation of missing ground truth depth and flow values, note that the training loss is only computed and backpropagated from pixels with known (not infilled) ground truth depth. We refer to this as the masked denoising loss (see Figure 2).
Step-unrolled denoising diffusion training. A second way to mitigate distribution shift in the marginals in training and inference, is to construct from model outputs rather than ground truth maps. One can do this by slightly modifying the training procedure (see Algorithm 1) to run one forward pass of the model and build by adding noise to the model’s output rather than the training map. We do not propagate gradients for this forward pass. This process, called step-unrolled denoising diffusion, slows training only marginally (15% on a TPU v4). This step-unrolling is akin to the predictor-corrector sampler of which uses an extra Langevin step to improve the target marginal distribution of . Interestingly, the problem of training / inference distribution shift resembles that of exposure bias in autoregressive models, for which the mismatch is caused by teacher forcing during training . Several solutions have been proposed for this problem in the literature . Step-unrolled denoising diffusion also closely resembles the approach in for training denoising autoencoders on text.
We only perform step-unrolled denoising diffusion during model fine-tuning. Early in training the denoising predictions are inaccurate, so the latent marginals over the noisy target maps will be closer to the desired true marginals than those produced by adding noise to denoiser network outputs. One might consider the use of a curriculum for gradually introducing step-unrolled denoising diffusion in the later stages of supervised pre-training, but this introduces additional hyper-parameters, so we simply invoke step-unrolled denoising diffusion during fine-tuning, and leave an exploration of curricula to future work.
denoiser loss. While the loss in Eqn. 1 is ideal for Gaussian noise and noise-free ground truth maps, in practice, real ground truth depth and flow fields are noisy and heavy tailed; e.g., for distant objects, near object boundaries, and near pixels with missing annotations. We hypothesize that the robustness afforded by the loss may therefore be useful in training the neural denoising network. (See Tables 11 and 12 in the supplementary material for an ablation of the loss function for monocular depth estimation.)
3 Coarse-to-fine refinement
Training high resolution diffusion models is often slow and memory intensive but increasing the image resolution of the model has been shown to improve performance on vision tasks . One simple solution, yielding high-resolution output without increasing the training cost, is to perform inference in a coarse-to-fine manner, first estimating flow over the entire field of view at low resolution, and then refining the estimates in a patch-wise manner. For refinement, we first up-sample the low-resolution map to the target resolution using bicubic interpolation. Patches are cropped from the up-scaled map, denoted , along with the corresponding RGB inputs. Then we run diffusion model inference starting at time with a noisy map . For simplicity, is a fixed hyper-parameter, set based on a validation set. This process is carried out for multiple overlapping patches. Following Perceiver IO , the patch estimates are merged using weighted masks with lower weight near the patch boundaries since predictions at boundaries are more prone to errors. (See Section H.5 for more details.)
Experiments
As our denoiser backbone, we adopt the Efficient UNet architecture , pretrained with Palette style self-supervised pretraining, and slightly modified to have the appropriate input and output channels for each task. Since diffusion models expect inputs and generate outputs in the range , we normalize depths using max depth of 10 meters and 80 meters respectively for the indoor and outdoor models. We normalize the flow using the height and width of the ground truth. Refer to Section H for more details on the architecture, augmentations and other hyper-parameters.
Optical flow. We pre-train on the mixture described in Section 3.1 at a resolution of 320448 and report zero-shot results on the widely used Sintel and KITTI datasets. We further fine-tune this model on the standard mixture consisting of AutoFlow , FlyingThings , VIPER , HD1K , Sintel and KITTI at a resolution of 320768 and report results on the test set from the public benchmark. We use a standard average end-point error (AEPE) metric that calculates L2 distance between ground truth and prediction. On KITTI, we additionally use the outlier rate, Fl-all, which reports the outlier ratio in among all pixels with valid ground truth, where an estimate is considered as an outlier if its error exceeds 3 pixels and 5 w.r.t. the ground truth.
Depth. We separately pre-train indoor and outdoor models on the respective pre-training datasets described in Section 3.1. The indoor depth model is then finetuned and evaluated on the NYU depth v2 dataset and the outdoor model on the KITTI depth dataset . We follow the standard evaluation protocol used in prior work . For both NYU depth v2 and KITTI, we report the absolute relative error (REL), root mean squared error (RMS) and accuracy metrics ().
Depth. Table 3 reports the results on NYU depth v2 and KITTI (see Section D for more detailed results and Section B for qualitative comparison with DPT on NYU). We achieve a state-of-the-art absolute relative error of 0.074 on NYU depth v2. On KITTI, our method performs competitively with prior work. We report results with averaging depth maps from one or more samples. Note that most prior works use post processing that averages two samples, one from the input image, and the other based on its reflection about the vertical axis.
Flow. Table 2 reports the zero-shot results of our model on Sintel and KITTI Train datasets where ground truth are provided. The model is trained on our newly proposed pre-training mixtures (AutoFlow (AF), FlyingThings (FT), Kubric (KU), and TartanAir (TA)). We report results by averaging 8 samples at a coarse resolution and then refining them to the full resolution as described in Section 3.3. For a fair comparison, we re-train RAFT on this pre-training mixture; this new RAFT model significantly outperforms the original RAFT model. And our diffusion model outperforms the stronger RAFT baseline. It achieves the state-of-the-art zero-shot results on both the challenging Sintel Final and KITTI datasets.
Figure 4 provides a qualitative comparison of pre-trained models. Our method demonstrates finer details on both object and motion boundaries. Especially on KITTI, our model recovers fine details remarkably well, e.g. on trees and its layered motion between tree and background.
We further finetune our model on the mixture of the following datasets, AutoFlow, FlyingThings, HD1K, KITTI, Sintel, and VIPER. Table 2 reports the comparison to state-of-the-art optical flow methods on public benchmark datasets, Sintel and KITTI. On KITTI, our method outperforms all existing optical flow methods by a substantial margin (even most scene flow methods that use stereo inputs), and sets the new state of the art. On the challenging Sintel final, our method is competitive with other state of the art models. Except for methods using warm-start strategies, our method is only behind FlowFormer which adopts strong domain knowledge on optical flow (e.g. cost volume, iterative refinement, or attention layers for larger context) unlike our generic model. Interestingly, we find that our model outperforms FlowFormer on 11/12 Sintel test sequences and our overall worse performance can be attributed to a much higher AEPE on a single (possibly out-of-distribution) test sequence. We discuss this in more detail in Section 5. On KITTI, our diffusion model outperforms FlowFormer by a large margin (30.34).
2 Ablation study
Infilling and step-unrolling. We study the effect of infilling and step-unrolling in Table 5. For depth, we report results for fine-tuning our pre-trained model on the NYU and KITTI datasets with the same resolution and augmentations as our best results. For flow, we fine-tune on the KITTI train set alone (with nearest neighbor resizing to the target resolution being the only augmentation) at a resolution of 320448 and report metrics on the KITTI val set . We report results with a single sample and no coarse-to-fine refinement. We find that training on raw sparse data without infilling and step unrolling leads to poor results, especially on KITTI where the ground truth is quite sparse. Step-unrolling helps to stabilize training without requiring any extra data pre-processing. However, we find that most gains come from interpolating missing values in the sparse labels. Infilling and step-unrolling compose well as our best results use both; infilling (being an approximation) does not completely bridge the training-inference distribution shift of the noisy latent.
Coarse-to-fine refinement. Figure 6 shows that coarse-to-fine refinement (Section 3.3) substantially improves fine-grained details in estimated optical flow fields. It also improves the metrics for zero-shot optical flow estimation on both KITTI and Sintel, as shown in Table 5.
Datasets. When using different mixtures of datasets for pretraining, we find that diffusion models sometimes capture region boundaries and shape at the expense of local textural variation (eg see Figure 3). The model trained solely on AutoFlow tends to provide very coarse flow, and mimics the object shapes found in AutoFlow. The addition of FlyingThings, Kubric, and TartanAir removes this hallucination and significantly improves the fine details in the flow estimates (eg, shadows, trees, thin structure, and motion boundaries) together with a substantial boost in accuracy (cf. Table 7). Similarly, we find that mixing SceneNet RGB-D , a synthetic dataset, along with ScanNet provides a performance boost for fine-tuning results on NYU depth v2, shown in Table 7.
3 Interesting properties of diffusion models
Multimodality. One strength of diffusion models is their ability to capture complex multimodal distributions. This can be effective in representing uncertainty, especially where there may exist natural ambiguities and thus multiple predictions, e.g. in cases of transparent, translucent, or reflective cases. Figure 1 presents multiple samples on the NYU, KITTI, and Sintel datasets, showing that our model captures multimodality and provides plausible samples when ambiguities exist. More details and examples are available in Section A.
Imputation of missing labels. A diffusion model trained to model the conditional distribution can be zero-shot leveraged to sample from where is the partially known label. One approach for doing this, known as the replacement method for conditional inference , is to replace the known portion the latent at each inference step with the noisy latent built by applying the forward process to the known label. We qualitatively study the results of leveraging replacement guidance for depth completion and find it to be surprisingly effective. We illustrate this by building a pipeline for iteratively generating 3D scenes (conditioned on a text prompt) as shown in Figure 7 by leveraging existing models for text-to-image generation and text-conditional image inpainting. While a more thorough evaluation of depth completion and novel view synthesis against existing methods is warranted, we leave that exploration to future work. (See Section C for more details and examples.)
Limitations
Latency. We adopt standard practices from image-generation models, leading to larger models and slower running times than RAFT. However, we are excited by the recent progress on progressive distillation and consistency models to improve inference speed in diffusion models.
Sintel fine-tuning. Under the zero-shot setting, our method achieves state-of-the-art results on both Sintel Final and KITTI. Under the fine-tuning setting, ours is state-of-the-art on KITTI but is behind FlowFormer on Sintel Final. We discuss several possible reasons for why this may be the case.
We follow the fine-tuning procedure in . While their zero-shot RAFT results are comparable to FlowFormer on Sintel and KITTI, the fine-tuned RAFT-it is significantly better on KITTI but less accurate on Sintel than FlowFormer. It is possible that the fine-tuning procedure (e.g. dataset mixture or augmentations) developed in is more suited for KITTI than Sintel.
Another possible reason is that there is substantial domain gap between the training and test data on Sintel than KITTI. On Sintel test, there is a particular sequence “Ambush 1”, where the girl’s right arm moves out of the image boundary. Our method has an AEPE close to 30 while FlowFormer has lower than 10. It is likely that the attention on the cost volume mechanism by FlowFormer can better reason about the motion globally and handles this particular sequence well. This particular sequence may account for the major difference in the overall results; among 12 available results on the Sintel website, ours has lower AEPE on 11 sequences but a higher AEPE on the “Ambush 1” sequence, as shown in Table 8. Figure 16 in the appendix further provides visualization.
Conclusion
We introduced a simple denoising diffusion model for monocular depth and optical flow estimation using an image-to-image translation framework. Our generative approach obtains state-of-the-art results without task-specific architectures or loss functions. In particular, our model achieves an Fl-all score of 3.26% on KITTI, about 25% better than the best published method . Further, our model captures the multi-modality and uncertainty through multiple samples from the posterior. It also allows imputation of missing values, which enables iterative generation of 3D scenes conditioned on a text prompt. Our work suggests that diffusion models could be a simple and generic framework for dense vision tasks, and we hope to see more work in this direction.
We thank Ting Chen, Daniel Watson, Hugo Larochelle and the rest of Google DeepMind for feedback on this work. Thanks to Klaus Greff and Andrea Tagliasacchi for their help with the Kubric generator, and to Chitwan Saharia for help training the Palette model.
References
Appendix A Multimodal prediction
We provide more qualitative examples for multimodal prediction. Figures 8 and 9 illustrate multimodal depth predictions on NYU and KITTI respectively. Multimodality of the posterior distribution exists in regions where there are multiple plausible predictions. For example, this includes reflective and transparent surfaces (mirrors and glass surfaces in rows 1 to 5 of Figure 8 and windows of cars in Figure 9). We further find that the model captures uncertainty in depth estimates in the vicinity of object boundaries, some of which arise due to noise in ground truth measurements in the training data. This can be observed at the boundaries of cars in Figure 9 and around edges of objects in Figure 8 (most clearly visible in the last row).
Figure 10 illustrates different samples on KITTI from the optical flow diffusion model, also capturing multiple modes of the predictive posterior. Multimodality exists on transparent surfaces and near occlusions. As shown in Figure 11, on Sintel, multimodality also exists on occluded or out-of-bounds pixels where multiple predictions are plausible.
Appendix B Qualitative comparison of depth estimation with DPT
Figure 12 provides a qualitative comparison of our model with DPT-Hybrid finetuned on the NYU depth v2 dataset. The depth estimates of our diffusion model are more accurate both on coarse-scale scene structure (walls, floors, etc.) and on individual objects.
Appendix C More samples for zero-shot imputation of depth
Figure 13 provides samples generated using our iterative text-to-3D pipeline. We note that such pipelines for iteratively generating 3D scenes have been previously proposed in literature . However, these methods explicitly learn networks to refine the color and the depth map . In contrast, we propose leveraging the text-conditioned image prior from existing large scale text-to-image and text-conditional image completion models, and use our depth estimation model zero-shot for depth completion. One caveat with our current approach of using the replacement method for conditional inference for imputing depth, is that it does not enable one to fix errors in the depth predicted in the previous step. One approach to fix artifacts would be by noising-denoising, like that used for coarse-to-fine refinement. We leave further exploration into this to future work.
Appendix D Complete depth results on NYU and KITTI
Tables 9 and 10 provide detailed results on the val set of NYU depth v2 and KITTI depth datasets. We follow the standard evaluation protocol used in prior work . For both the NYU depth v2 and KITTI datasets we report the absolute relative error (REL), root mean squared error (RMS) and accuracy metrics ( for ). For NYU we also report absolute error of log depths (). For KITTI we additionally report the squared relative error (Sq-rel) and root mean squared error of log depths (RMS log). The predicted depth is up-sampled to the full resolution using bilinear interpolation before evaluation. For the indoor model we evaluate on the cropped region proposed by and for the outdoor model the cropped region proposed by as is standard in prior work.
Appendix E Ablations
Tables 11 and 12 show that an loss in training the diffusion model performs much better than an loss for monocular depth estimation on NYU and KITTI. Tables 13 and 14 show the effectiveness of Palette-style self-supervised pretraining for monocular depth estimation on NYU and KITTI respectively. All results use a single sample. Because these findings are reasonable and expected to generalize to other dense vision tasks, we do not further ablate them for optical flow estimation for compute efficiency.
Appendix F Coarse-to-fine refinement for depth
Figure 14 demonstrates performance of coarse-to-fine refinement on the NYU depth v2 dataset. While refinement improves fine-scale details in the estimated depth maps, the qualitative improvements are small and we do not find significant quantitative improvements. Hence the results reported in this work do not use coarse-to-fine refinement for depth estimation. Further work is needed to develop a coarse-to-fine algorithm capable of more robust gains in depth estimation.
Appendix G Coarse-to-fine optical flow refinement for RAFT
For a fair comparison with optical flow estimation, we also apply our coarse-to-fine refinement scheme to RAFT , to determine whether our performance gains translate to RAFT as well. We first estimate flow at a low resolution, , upsample the low-resolution flow to the original resolution, divide original-resolution input images into overlapping patches of size , then estimate flow on the cropped patches using the upsampled flow field as the initial guess for the recurrent refinement (12 steps in total) of RAFT . After estimating flow of each patch, we merge them using weighted masks . Table 15 reports the result. Unlike our diffusion-based method, the coarse-to-fine scheme actually hurts the accuracy of RAFT on Sintel Clean and KITTI and only marginally improves the accuracy on Sintel Final. Further exploration into better approaches for coarse-to-fine refinement for RAFT is warranted. We leave that to future work.
Appendix H Training and inference details
UNet. The predominant architecture for diffusion models is the U-Net developed for the DDPM model , and later improved in several respects . Here we adapt the Efficient U-Net architecture that was developed for Imagen . It is more efficient that the U-Nets used in prior work owing to the use of fewer self-attention layers, fewer parameters and less computation at higher resolutions, along with other adjustments that make it well suited to training medium resolution diffusion models.
Specifically we adopt the configuration for the 6464 256256 super-resolution model (see Figure 15 for an overview) with several changes. We drop the text cross-attention layers but preserve the self-attention in the lowest resolution layers dblock4 and ublock4 (see Figure 15). For supervised training for the flow model, we find it beneficial to additionally enable self-attention for the last-but-one layers dblock3 and ublock3. The number of input and output channels differ across self-supervised pre-training and supervised pre-training and are also different for flow and depth models. For self-supervised pre-training CH_IN=6 and CH_OUT=3 (see Figure 15) since the input consists of a 3-channel source RGB image and a 3-channel noisy target image concatenated along the channel dimension and the output is a RGB image. The supervised depth model has CH_IN=4 (RGB image + noisy depth) and CH_OUT=1. The supervised optical flow model has CH_IN=8 (2 RGB images + noisy flow along and ) and CH_OUT=2. Note that this means we need to reinitialize the input and output convolutional kernels and biases before the supervised pretraining stage. All other weights are re-used.
Resolution. Our self-supervised model was trained at a resolution of . The indoor depth model is trained at 240320. For Waymo we use 256384 and for KITTI depth 256832. Flow pretraining is done at a resolution of 320448, and finetuning at 320768.
H.2 Datasets and augmentation
For unsupervised pre-training, we use the ImageNet-1K and Places365 datasets and train on the self-supervised tasks of colorization, inpainting, uncropping, and JPEG decompression, following . Throughout, we mix datasets at the batch level.
Flow. For supervised flow pretraining we use a mix of AutoFlow (native resolution 448576), FlyingThings (540960), Kubric (512512) and TartanAir (480640) synthetic datasets. We finetune on the standard mixture consisting of AutoFlow, FlyingThings, Viper (540960), HD1K (5401280), Sintel (4361024), and KITTI (3751242).
We follow the same photometric and geometric augmentation schemes from , comprising random affine transformation, flipping, and cropping.
Depth. For supervised pre-training of the indoor model we mix the following datasets. ScanNet is a dataset of 2.5M examples captured using a Kinect v1-like sensor. It provides depth maps at 480640 and RGB images at 9681296. SceneNet RGB-D is a synthetic dataset of 5M images generated by rendering ShapeNet objects in scenes from SceneNet at a resolution of 240320.
For the outdoor model training we use the Waymo Open Dataset , a large-scale driving dataset consisting of about 200k frames. Each frame provides RGB images from 5 cameras and LiDAR maps. We use the RGB images from the FRONT, FRONT_LEFT and FRONT_RIGHT cameras and the TOP LiDAR only to build about 600k aligned RGB depth maps.
For indoor fine-tuning and evaluation we use NYU depth v2 , a commonly used dataset for evaluating indoor depth prediction models. It provides aligned image and depth maps at 480640 resolution. We use the official split comprising 50k images for training and 654 for evaluation.
For outdoor fine-tuning and evaluation, we use KITTI , an outdoor driving dataset which provides RGB images and LiDAR scans at resolutions close to 3701226. We use the training/test split proposed by , comprising 26k training images and 652 test images.
We use random horizontal flip data augmentation which is common in prior work. Where needed, images and dense depth maps are resized using bilinear interpolation to the model’s resolution for training and nearest neighbor interpolation is used for sparse maps.
H.3 Step-unrolling and interpolation of missing depth and flow
As discussed in Section 3.2 of the main paper, infilling and step-unrolling are used to mitigate distribution shift between training and inference with diffusion models. The problem arises due to the missing data in the training depth maps and flow fields.
Infilling. For indoor depth maps, we use nearest neighbor interpolation during training (see Section 3.2 in the main paper). For the outdoor depth data we use nearest neighbor interpolation except for sky regions, as they are often large and are much further from the camera than adjacent objects in the image. We use an off-the-shelf sky segmenter , and then set all sky pixels to be the maximum modeled depth (here, 80m). For missing optical flow ground truth we employ a simple sequence of 1D nearest neighbor interpolations first along rows, and then along columns.
Step-unrolling. By default we use a single unroll step in all results where step-unrolling is enabled. In Table 16, we show that using multiple unroll steps can further improve performance on the task of monocular depth estimation.
Finally, while we use infilling and step-unrolling, there are other ways in which one might try to mitigate the problem. One such approach was taken by , which faced a similar problem when training a vector-quantizer on depth data. Their approach was to synthetically add more holes following a carefully chosen masking ratio. We prefer our approach since nearest neighbor infilling is hyper-parameter free and step-unrolled denoising diffusion could be more generally applicable to other tasks with sparse data.
We also considered the approach of self-conditioning as an alternative to step-unrolling. However, as we show in Table 17, we find that self-conditioning is unable to bridge the train-inference distribution shift of the noisy latent for the task of monocular depth estimation. This is specially apparent in the results for KITTI without infilling where self conditioning leads to no improvement whereas step-unrolling substantially improves performance.
H.4 Hyper-parameters
Self-supervised. The self-supervised model is trained for 2.8M steps with an loss and a mini-batch size of 512. Other hyper-parameters are same as those in the original Palette paper .
Supervised. The supervised flow and depth models are trained with loss. Usually a constant learning rate of with a warm-up over 10k steps is used. However, for depth fine-tuning we find that a lower learning rate of achieves slightly better results. All models are trained with a mini-batch size of 64. The indoor depth model is pre-trained for 2M steps and then fine-tuned on NYU for 40k steps. The outdoor depth model is pre-trained for 0.9M steps and fine-tuned on KITTI for 40k steps. For flow, we pretrain for 3.7M steps, followed by finetuning for 50k steps. Other details, like the optimizer and the use of EMA are the same as .
H.5 Inference
Sampler. We use the DDPM ancestral sampler with 128 denoising steps for monocular depth models and 64 steps for optical flow models. Increasing the number of denoising steps further did not greatly improve performance.
Coarse-to-fine refinement. We use 25 overlapping patches ({top, bottom} {left, center-left, center, center-right, right}) for coarse-to-fine refinement. For Sintel we use and for KITTI .
Appendix I Limitations
Efficiency. Inference speed with diffusion models is a well-known issue, as multiple denoising steps are used to transform noise to a target signal. This can be prohibitive for vision tasks where near real-time latency is often desired. Table 18 compares the inference speed of our diffusion model for depth against DPT . Despite having an efficient denoiser backbone (8.5 ms per denoising step on a TPU v4), the diffusion model is considerably slower than DPT in total wall time. The most obvious way to reduce inference latency is to reduce the number of denoising steps. This can be done with only moderate reduction in performance. As shown in Table 18, we perform comparably with DPT with as few as 24 denoising steps. However, a more thorough study into optimizing the inference speed of these models while preserving the generation quality is warranted. With the use of progressive distillation it is likely possible to reduce latency even further, as this approach has been shown to successfully distill generative image models with over 1000 denoising steps into those with just 2-4 steps.
Fine-tuning on Sintel. In Section 5 we discuss possible reasons for why our model’s superior zero-shot performance compared to FlowFormer does not transfer to fine-tuning on Sintel. Figure 16 provides qualitative examples to further support the claims.
We observe certain cases where the model is uncertain about the depth estimates. Interestingly, this uncertainty appears to be well captured in the predictive posterior, as illustrated in Figure 17.