AutoFlow: Learning a Better Training Set for Optical Flow

Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T. Freeman, Ce Liu

Introduction

Datasets have been a driving force for the development of AI algorithms. Convolutional neural networks (CNNs) were proposed in the 1990’s but were not widely adopted for vision tasks until the early 2010’s, with the advent of AlexNet . One key ingredient for deep CNN models was the large amount of manually labeled images, \eg, from ImageNet . The performance gain by AlexNet over shallow models stimulated a paradigm shift in high-level vision tasks. Since then, new models have been invented in rapid succession, even achieving “superhuman” performance on image classification tasks .

Manual labeling, however, cannot provide reliable ground truth for a variety of low-level vision tasks like optical flow and stereo. Since these labels are either difficult or impossible to obtain, synthetic data play a key role in enabling deep models to perform well on such tasks. For example, all top-performing CNN models for optical flow are pre-trained on two large synthetic datasets, FlyingChairs and FlyingThings3D , before being fine-tuned on limited target datasets, \eg, Sintel and KITTI .

However, the success of FlyingChairs raises some interesting questions. For example, how realistic should the rendering be? Several new datasets have been developed to be more realistic than FlyingChairs, such as virtual KITTI , VIPER , and REFRESH , but none of them have proven more effective than FlyingChairs and FlyingThings3D at pre-training models. In fact, a comprehensive study has revealed that “realism is overrated” . There are some hypotheses for why FlyingChairs works, \eg, that it has been designed to match the motion statistics of Sintel, or that it has many thin structures and fine motion details. However, it still remains unclear what set of principles makes an effective optical flow dataset.

To address these questions, we argue that we should make explicit the objective function for rendering training data. We formulate the generation of training data as a joint optimization problem, which couples rendering the data with training the model. This generation process depends on a set of hyperparameters being optimized. The hyperparameters are evaluated by the performance of the trained model on a target dataset, as shown in Fig. 1.

To understand what matters, we ask: how simple can the rendering be? Thus, we start from an even simpler rendering pipeline than FlyingChairs, a 2D layered approach that requires neither manual labeling nor 3D models. The motion and shape of each layer are randomly generated according to hyperparameters, as shown in Fig. 2. We can then learn the rendering hyperparameters to optimize the performance of a model on a target dataset.

This simple rendering pipeline is surprisingly effective at generating training datasets for optical flow. Trained on its rendered data from scratch, both the recent RAFT model and the widely-used PWC-Net model obtain consistent improvements in accuracy on Sintel and KITTI over the same models trained on FlyingChairs (Fig. 1 and Table 1). Further, using 4 AutoFlow examples with augmentation results in lower errors on Sintel.final for RAFT than using 22,872 FlyingChairs examples with augmentation. More interestingly, the gap between PWC-Net and RAFT becomes small when trained on enough AutoFlow examples.

An analysis of the rendered data also suggests some interesting properties. For example, the motion statistics of the AutoFlow dataset and its augmented version do not resemble those of Sintel (Fig. 8) and underrepresent small motions. Though at first glance this distribution may seem abnormal, there may be a simple, intuitive explanation: tiny motion matters little in the overall error.

To summarize, our contributions are the following.

We have introduced, to our knowledge, the first learning approach to render training data for optical flow.

AutoFlow compares favorably against FlyingChairs and FlyingThings3D in pre-training RAFT.

AutoFlow also leads to a significant performance gain for PWC-Net, even competitive against RAFT.

We present a detailed analysis of what features are important to dataset generation for optical flow.

Related Work

Manually labeled datasets, such as ImageNet , PASCAL , MS-COCO , and CityScapes , have been widely adopted for high-level vision tasks. However, manual labeling is hard to scale, and quite a few synthetic datasets have been developed . Meta-Sim learns to minimize the distribution gap between the rendered and target datasets and can also optimize task performance. However, Meta-Sim can model only limited scenes because it relies on obtaining valid scene structures from a grammar.

RenderGAN learns to augment the dataset for handwriting classification. Differentiable rendering enables gradients to be passed to rendering parameters, which, however, do not directly relate to the scene distribution hyperparameters. Yang and Deng proposed a “hybrid gradient” approach to make use of analytical gradients whenever available. These methods focus on the generation of a single image and cannot directly apply to the generation of optical flow.

Datasets for optical flow

Similar to other vision tasks, datasets have been the driving force behind the development of optical flow. However, unlike high-level vision tasks, it is only possible to obtain ground truth under controlled lab environments or rigid scenes/objects . Early work relied on synthetic datasets for evaluation, such as the well-known “Yosemite” sequence . MPI-Sintel , one of the leading benchmark datasets for optical flow, was rendered using the Blender engine. Roth and Black used real depth data to render synthetic data, which is limited to static scenes. KITTI was created using LIDAR for static scenes and later extended to rigidly moving cars for autonomous driving applications.

Dosovitskiy et al. created a synthetic dataset, FlyingChairs. Mayer et al. further introduced a large dataset for optical flow and related tasks, FlyingThings3D. Ilg et al. found that sequentially training on FlyingChairs and then on FlyingThings3D obtains the best results; this has since become standard practice in the field. Efforts to improve these two datasets include the autonomous driving scenario , more realistic rendering , realistic backgrounds from SLAM , and human datasets . However, none have proven more effective than FlyingChairs and FlyingThings3D for pre-training.

Mayer et al. performed a comprehensive study of synthetic datasets for optical flow and disparity estimation. They developed each synthetic dataset heuristically, with no regard for target dataset performance. Our rendering pipeline is largely inspired by their 2D rendering techniques. But instead of designing each dataset by hand, we learn these parameters via jointly solving rendering and training to optimize the performance on a target dataset.

CNN models for optical flow

The seminal FlowNet paper pioneered the CNN-based approach for optical flow. Its follow-up, FlowNet2 , significantly improved FlowNet’s performance by stacking several sub-networks into one large model. Spy-Net , PWC-Net , and LiteFlowNet were designed using several well-established principles for optical flow. For the first time, PWC-Net obtained more accurate results on the Sintel and KITTI benchmarks than traditional approaches. Quite a few new network architectures were proposed based on the PWC-Net framework . Recently, Teed and Deng introduced the RAFT architecture, which used a recurrent architecture to obtain a significant performance gain over its predecessors on Sintel and KITTI.

The advances in these network architectures have significantly improved their performance on benchmark datasets. However, all these models follow nearly the same training procedures, \ie, pre-training on FlyingChairs and FlyingThings3D and then fine-tuning on limited training data on the target domain. In this paper, we focus on dataset generation and show that it is possible to achieve accuracy similar to or better than that of FlyingChairs and FlyingThings3D in pre-training by learning to render training data. We learn the rendering hyperparameters for the recent RAFT model and find that they also apply to PWC-Net.

Evaluating CNN models for optical flow

The improvement in accuracy comes from innovations on both the model architecture and the training procedures. Previous work shows that changes in training procedure result in significant performance boosts for FlowNetC and PWC-Net. Here we find that changing the datasets and incorporating recent practices in training significantly improves PWC-Net and narrows down its performance gap from RAFT.

Self-supervised and semi-supervised learning of optical flow

Significant progress has been made on self-supervised optical flow . However, state-of-the-art, self-supervised methods still lag behind supervised ones, e.g., models pre-trained on FlyingChairs and FlyingThings are more accurate on Sintel than models trained on Sintel image pairs using self-supervised loss.

Learning to learn

A recent trend in neural network research is learning to learn, which aims at automating the manual process of network design or hyperparameter selection. Existing methods mainly focus on learning hyperparameters for the architecture , loss function, optimization, and augmentation . In contrast, we focus on learning to render synthetic training data for optical flow.

Generating training data

We take a layered approach to rendering image pairs and their optical flow, as shown in Fig. 2. For the first frame, we randomly sample K images I1k\mathbf{I}_{1}^{k} from an image dataset and order them by depth, with the first layer being the background. Next, we sample an alpha mask M1k\mathbf{M}_{1}^{k} (section 3.1) and an optical flow field Wk\mathbf{W}^{k} (section 3.2) for each layer according to the rendering hyperparameters (section 3.5). The optical flow field is used to warp the image and the mask into the second frame:

where ff represents the forward warping function according to the flow field.

We composite the images and the flow with back-to-front alpha blending, starting with the background layer:

Finally, we apply certain visual effects (section 3.3) to the images to cover some natural variations in videos. Figure 6 shows examples of complete images and their flows.

The background layer has a fully opaque mask. For each foreground layer, we test two ways of generating the mask: random polygons and manual segmentation.

For each foreground layer, we generate a random polygon to serve as its alpha mask. Each polygon has a random number of sides, with vertices randomly sampled in angle and radius around a center. Each polygon can also have a hole, which itself is a smaller random polygon. Further, we can control polygon smoothness through subdivision. Finally, the mask can be blurred with a Gaussian filter in order to feather its boundary (this is applied to both polygon and manual object masks). Examples of random polygon masks are shown in Fig. 3.

Manual segmentation

To make foreground objects more semantically congruent, we use the images and manual labels from OpenImages . The location and size of each foreground object within the image are randomly sampled.

2 Motion Model

The motion of each layer is a combination of rigid transformation (scale, rotation, translation), perspective distortion, and a bilinear grid warp. A bilinear grid warp of size nn is a set of flow vectors defined on the vertices of a n×nn\times n grid, then bilinearly interpolated in the interior of the grid (the grid being uniformly distributed over the image) . This allows for more complex forward flow with a fast analytic solution for forward image warping (we invert the bilinear interpolating function within each grid cell, which boils down to solving a quadratic equation). In fact, all of our base motions can be modeled with a bilinear grid warp: rigid transform can be expressed by rigidly moving the corners of a grid, and perspective distortion by independently moving the corners of a 2×22\times 2 grid.

For the foreground layers, we employ all modalities of motion (rigid + grid), while for the background we only apply moderate perspective distortion. Figure 4 demonstrates our motion modalities on a sample foreground object.

3 Visual Effects

To generalize better to more realistic video data, we simulate common visual effects including motion blur and fog (Fig. 5). These effects only modify the image data and have no influence over the ground truth flow.

We approximate the motion blur of each layer by applying a filter to both the image and the mask. Standard deviations of the filter are computed by taking a proportion of the average absolute flow in each dimension over all the pixels within the mask. We apply the same motion blur filter to both the first and second images.

Fog

To simulate fog, we generate a white image with a random semi-transparent alpha mask and overlay it on top of the composited initial and final images. The fog does not move between the images, nor does it affect the ground truth flow. To compute the alpha mask, we generate several random normal images of various resolutions, with their standard deviations being inversely proportional to their resolutions. We then bicubically resample each to the desired fog resolution and sum them up. Finally, we adjust the resulting image so that its mean and standard deviation match controllable hyperparameters.

4 Data Augmentation

To increase the diversity of the training data, we apply data augmentation to the rendered data. Inspired by RandAugment , we randomly select several transformations among rotation, scale, squeeze, translation, and additive noise at each iteration. The number of transformations and their strength levels are hyperparameters to learn.

5 Hyperparameters

During training, we tune a number of hyperparameters that dictate data generation and augmentation, including the shape, size, and position of masks, the complexity and magnitude of motion, and the visual effects. Respective values are uniformly sampled from the specified ranges, and the ranges are hyperparameters to learn. Please refer to the appendix for the detailed list of hyperparameters.

Learning to Render Training Data

Given a target dataset for optical flow, \eg, Sintel or KITTI, we want to learn the hyperparameters to render training data so that a CNN model trained on the rendered data has optimal performance on the target dataset. Every set of hyperparameters corresponds to a rendered training dataset. In this section, we will first present the learning objective and then the search algorithm.

Given the rendering pipeline for generating training data and the range for the rendering hyperparameters Λ\Lambda, our goal is to search for the set of optimal hyperparameters λ\lambda that optimizes a metric Ω\Omega on a model θ\theta,

The model θ\theta minimizes a loss function L\mathcal{L} on the rendered datasets according to the set of hyperparameters λ\lambda

where the model θ\theta includes the parameters of a network ϕθ\phi_{\theta} that maps two input images to their optical flow. Be default, we use the sequence loss function proposed by RAFT as the loss function and the average end-point error (AEPE) as the metric on the target datasets unless stated otherwise.

2 Hyperparameter Search Algorithm

To learn the hyperparameters for rendering the dataset, we develop a hybrid algorithm based on the population-based training (PBT) algorithm and the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) algorithm . Specifically, we classify the hyperparameters into subgroups and use the CMA-ES algorithm to search the selected subgroups of hyperparameters to optimize the learning metric. CMA-ES maintains a sampling distribution over the search space. It samples a few points, evaluates them, and updates the distribution based on the ranking of the points \wrtto the learning metric. The sampling distribution is a multivariate Gaussian whose covariance matrix is adapted over time. Our algorithm takes N\mathcal{N} iterations, with each iteration training M\mathcal{M} in parallel. The time complexity grows linearly w.r.t. the number of search iterations and the number of training steps per search.

Experimental Results

We randomly sample images from different sequences of the Davis dataset as appearances for each layer. Our baseline is a TensorFlow implementation of RAFT , the performance of which is similar to that of the official PyTorch implementation. Throughout this section, we refer to our method or the data it generates as AutoFlow. By default, we use the average end-point error (AEPE) on the final pass of Sintel training dataset (Sintel.final) as the learning metric because it is currently the most challenging dataset.

Empirically, it takes about 7 days to finish 8 searching iterations using 48 NVIDIA P100 GPUs, with each iteration training 8 models in parallel. Hyperparameters about 5% less accurate are often found within 2 days. Alternatively, the time can also be reduced to less than 2 days by using fewer training steps (40k) and then reusing the searched hyperparameters for the full 200k steps. That has roughly a 3% drop in accuracy on Sintel.

1 AutoFlow Versus the State of the Art

We pre-trained RAFT and PWC-Net from scratch using different datasets. The hyperparameters for AutoFlow have been learned for RAFT. As summarized in Table 1, models trained from scratch using AutoFlow are comparable to or more accurate than models trained on FlyingChairs or FlyingChairs  ⁣→ ⁣\!\rightarrow\! FlyingThings3D. As shown in Fig. 7, RAFT trained on AutoFlow can successfully recover blurry objects under large motion in the final pass of Sintel. Table 2 summarizes the errors for regions with different motion magnitude. RAFT trained on AutoFlow performs better than RAFT trained on FlyingChairs, especially in regions with large motion. Using end-point error (epe) as the learning metric results in better accuracy in regions with large motion than using angular error (ae).

Generalization across datasets

We further compared with the recent DSMNet method that aims at narrowing down domain gaps. DSMNet reported an F-all score of 11.2% in non-occlusion regions on the KITTI 2015 training set for a modified PWC-Net trained on FlyingThings3D and Sintel. The F-all scores by RAFT and PWC-Net trained on AutoFlow that has been optimized for Sintel.final are 8.7% and 11.0%, respectively, suggesting that AutoFlow generalizes well across datasets.

Improving PWC-Net

We modified the pre-training procedure of PWC-Net using the one-cycle learning rate schedule and gradient clipping from RAFT , as summarized in Table 3. Applying gradient clipping not only improves accuracy but also makes training more stable: two out of eight runs diverged without gradient clipping.

Fine-tuning results

We followed the TF-RAFT procedure to fine-tune the model pre-trained by AutoFlow and denoted the method as RAFT-A. We applied the same fine-tuned model to Sintel and KITTI, as summarized in Table 4. RAFT-A is more accurate than TF-RAFT on the more challenging Sintel.final and KITTI benchmarks, demonstrating the benefits of pre-training on AutoFlow.

2 Ablation Study

To further analyze AutoFlow, we performed a series of ablation studies designed to determine how different design choices affect performance. Since it is computationally expensive to learn all the hyperparameters for each setup, we fixed the learned hyperparameters unless explicitly specified. For each experiment, we ran 8 independent trials, and Table 5 summarizes the most accurate one for each setup.

Removing the motion blur effect leads to a significant drop in performance on the final pass of Sintel and KITTI, despite the rough approximation used for simulating the motion blur. Gaussian and box filters have similar results. Removing the fog effect also results in a moderate performance drop in the final pass of Sintel and KITTI. Neither motion blur nor fog effects have a significant effect on the clean pass of Sintel.

Appearance

We tested three different image sources for the appearance image of each layer: Davis, OpenImages , and Sintel (\cfTable 5). Neither OpenImages nor Sintel achieves better results than Davis. By default, we downsample Davis images to 1280 ⁣× ⁣7201280\!\times\!720 (720p) resolution as appearance images for each layer. Downsampling to 960 ⁣× ⁣540960\!\times\!540 (540p) has similar results while 1920 ⁣× ⁣10801920\!\times\!1080 (1080p) has degraded performance, likely because the hyperparameters have been learned for the 720p resolution.

Foreground object masks

We tested three versions of masks for the foreground objects: random polygons with sharp edges, random polygons with smooth edges (default), and instance segmentation from the OpenImage dataset. Polygons with smooth edges perform consistently better than those with sharp edges. We also experimented with instance segmentation from the OpenImage dataset due to its diverse set of segmentation masks, but there we only observed a small improvement on Sintel.clean.

Number of foreground objects

The number of foreground objects determines the complexity of a scene. Using only a background layer, \ie, 0 foreground object, results in large errors. Adding one foreground object significantly improves the performance. Using three or four foreground objects tends to work best, while more than four foreground objects bring no further gain.

Motion model

Removing the bilinear grid warping results in a performance degradation on both Sintel and KITTI, suggesting that more complex and flexible motion than parametric motion is critical.

Number of training steps

We learned the hyperparameters using 200k training steps for RAFT. With the same hyperparameters, running more iterations to train RAFT, such as 800k, results in moderate gains on both Sintel and KITTI.

Target datasets

AutoFlow directly optimizes the performance on a target dataset. To test how well AutoFlow generalizes, we learned hyperparameters for Sintel.final and KITTI separately and found that the generalization gap is small. It is likely that the rendering pipeline and the small number of hyperparameters act as a form of regularization, which helps generalization.

Data augmentation

RandAugment leads to moderate improvement over applying the same augmentation at every training step, likely because RandAugment increases the diversity of training data. Turning off spatial augmentation results in a moderate drop in accuracy on both KITTI and Sintel. Turning off color augmentation results in severe performance degradation on KITTI, likely because KITTI data includes more lighting changes.

Motion statistics

We compared the statistics of motion magnitude for different datasets in Figure 8. The motion statistics of AutoFlow differ from those of Sintel and FlyingChairs. AutoFlow has little small motion, concentrates mainly in the middle-range motion, and does not exhibit an exponential falloff. We further analyzed the augmented data, as it is used to train models. The augmented AutoFlow also has little small motion and concentrates in the middle to high-range motion, probably because tiny motion matters little in the overall learning metric.

Number of pre-training examples

With the hyperparameters learned, we can render different numbers of pre-training examples, as shown in Fig. 1. The training of both RAFT and PWC-Net converge using one pre-training example with data augmentation, and more examples lead to better results. Four AutoFlow examples result in lower errors on Sintel.final for RAFT than 22,872 FlyingChairs examples. In this low-data regime, data augmentation plays a key role. Without spatial augmentation, the AEPE by RAFT trained on 4 AutoFlow examples drops from 3.57 to 7.66, more severe than the drop from 2.75 to 3.37 in Table 5. Further, as shown in Fig. 9, although the statistics of 4 AutoFlow examples differ significantly from those of the full AutoFlow, they are similar for augmented data.

Discussions

While AutoFlow empirically works better than FlyingChairs/FlyingThings3D, we should note that the comparisons are not strictly fair because of differences in implementations and hyperparameters. Although comparing motion statistics reveals some interesting properties, learning hyperparameters for a 3D rendering pipeline in the same setup would help identify key design choices.

Conclusions

We have introduced AutoFlow, a simple and effective method to learn pre-training data for optical flow. AutoFlow uses 2D rendering but achieves results comparable to or better than those obtained by FlyingChairs and FlyingThings3D that have been generated using 3D models. In particular, using as few as 4 AutoFlow examples with augmentation results in more accurate results on Sintel.final for RAFT than 22,872 FlyingChairs examples with augmentation. AutoFlow also significantly improves PWC-Net, even on par with RAFT. We hope that our approach will provide another option for pre-training optical flow and enable further progress and innovation in this direction.

We would like to thank Shuyang Cheng, Ekin Dogus Cubuk, Alex Dosovitskiy, Rico Jonschkowski, David Kao, Ang Li, Aaron Sarna, Austin Stone, and Barret Zoph for helpful discussions and support.

References

Appendix A Rendering Hyperparameters

During training, we tune a number of hyperparameters that dictate data generation, including the shape, size, and position of masks, the complexity and magnitude of motion, and the visual effects. Respective values are uniformly sampled from the specified ranges, and the ranges are hyperparameters to learn. U\mathcal{U} denotes a random number uniformly sampled from [−1[-1, 1]1].

maximum size of the hole’s bounding box diagonal, relative to the polygon’s

Hyperparameters for all foreground masks:

minimum and maximum size of the object’s bounding box diagonal, relative to the image diagonal

minimum and maximum object center location, relative to image dimensions

Hyperparameters for motion

scale strength psp_{s} (≥1\geq 1), with scale sampled as psU{p_{s}}^{\mathcal{U}}

rotation strength prp_{r}, with rotation angle sampled as π⋅pr⋅U\pi\cdot p_{r}\cdot\mathcal{U}

Hyperparameters for visual effects:

Hyperparameters for RangAugment:

Appendix B More Samples

Figures B.1-B.5 show some more samples of the image pairs and flow field. Please go to our webpage autoflow-google.github.io to see the gif images.