FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow Estimation
Xiaoyu Shi, Zhaoyang Huang, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li
Introduction
Optical flow is a long-standing vision task, targeting at estimating per-pixel displacement between consecutive video frames. It can provide motion and correspondence information in many downstream video problems, including video object detection , action recognition , and video restoration .
Recently, FlowFormer introduces a transformer architecture for optical flow estimation and achieves state-of-the-art performance. The core of its success lies on two aspects: the ImageNet-pretrained transformer-based image encoder and the transformer-based cost-volume encoder. Notably, adopting an ImageNet-pretrained visual backbone leads to considerable performance gain over the train-from-scratch counterpart, indicating that random weight initialization hinders the learning of correspondence estimation. This naturally begs the question: can we also pretrain the transformer-based cost-volume encoder and thus further unleash its power to achieve more accurate optical flow?
In this paper, we propose masked cost-volume autoencoding (MCVA), a self-supervised pretraining scheme to enhance the cost-volume encoding on top of the FlowFormer framework. We are inspired by the recent success of masked autoencoding, such as BERT in NLP and MAE in computer vision. The key idea of masked autoencoding is masking a portion of input data, and requiring networks to learn high-level representation for masked contents reconstruction. However, it is non-trivial to adapt the masked autoencoding strategy to learn a better cost volume encoder for optical flow estimation, because of the two following reasons. Firstly, the cost volume might contain redundancy and the cost maps (cost values between a source-image pixel to all target-image pixels) of neighboring source-image pixels are highly correlated. Randomly masking cost values, as done in other single-image pretraining methods , leads to information leakage and makes the model biased towards aggregating local information. Secondly, existing masked autoencoding methods target at reconstructing masked content randomly selected from fixed locations. This suffices to pretrain general-purpose single-image encoder in other fields. However, the cost-volume encoder of FlowFormer is deeply coupled with the follow-up recurrent decoder, which demands cost information of long range at flexible locations.
To tackle the aforementioned issues, we introduce two task-specific designs. Firstly, instead of randomly masking the cost volume, we partition source pixels into large varied-size blocks and let source pixels within the same block share a common mask pattern on their cost maps. This strategy, termed block-sharing masking, prevents the cost-volume encoder from reconstructing masked cost values by simply copying from neighboring source pixels’ cost maps Such design enfoces the cost-volume encoder to abstract useful cues from cost maps belonging to far-away source pixels, which encourages long-range information aggregation. Secondly, to mimic the decoding process in finetuning and thus avoid pretraining-finetuning discrepancy, we propose a novel pre-text reconstruction task as shown in Fig. 3: small cost patches (of shape ) are randomly cropped from the cost maps to retrieve features from the cost-volume encoder, aiming to reconstruct larger cost map patches (of shape ) centered at the same locations. This is in line with the decoding process of FlowFormer in the finetuning stage. This pre-text task explicitly encourages the cost-volume encoder to capture long-range information for cost-volume encoding, which is critical for optical flow estimation. Besides, we empirically show that the image encoder, upon which the cost volume is built, should be frozen during pretraining to avoid training collapse.
In essence, the proposed masked cost-volume autoencoding (MCVA) has unique designs compared with conventional MAE methods, which encourages the cost-volume encoder 1) to construct high-level holistic representation of the cost volume, more effectively encoding long-range information, 2) to reason about occluded (i.e., masked) information by aggregating faithful unmasked costs, and 3) to decode task-specific feature (i.e., larger cost patches at required locations) to better align the pretraining process with that of the finetuning. These designs contribute to better handling of hard cases, such as noises, large-displacement motion and occlusion, for more accurate flow estimation.
To conclude, the contributions of this work are three-fold: 1) We propose the masked cost-volume autoencoding scheme to better pretrain the cost-volume encoder of FlowFormer. 2) We propose task-specific masking strategy and reconstruction pre-text task to mitigate pretraining-finetuning discrepancy, fully taking advantage of the learned representations from pretraining. 3) With the proposed pretraining technique, our proposed FlowFormer++ obtains all-sided improvements over FlowFormer, setting new state-of-the-art performance on public benchmarks.
Related Work
Optical Flow. Compared with traditional optimization-based optical flow methods empirically formulating flow estimation, data-driven methods directly learn to estimate optical flow from labeled data. Since FlowNet , learning optical flow with neural networks presents superior performance and is still fast progressing where network architecture design becomes the key to improving optical flow accuracy. A series of excellent works are devoted to designing better network modules, which, indeed, introduced better inductive bias to the optical flow formulation. For example, encoding image feature with CNNs brings locality prior, and the all-pairs 4D cost volume outperforms the coarse-to-fine cost volumes in modeling small fast-motion objects. However, the empirical network design may always ignore some unintended cases. Due to the success of transformers in image recognition , the optical flow community also tries transformers to further weaken the network-determined bias and learn feature relationships from data. By replacing the handcrafted modules, i.e., the CNN image encoder, the cost pyramid, and the indexing-based costs retrieval, in RAFT with transformers, FlowFormer achieves state-of-the-art accuracy. However, transformers are known for requiring tremendous training data to capture feature relationships while collecting ground-truth flows for supervised optical flow learning is expensive. Inspired by the emerging pretraining-finetuning paradigm for vision transformers , we explore to pretrain FlowFormer to capture the feature relationship for optical flow.
Masked Autoencoding (MAE). As a self-supervised learning technique, MAE, e.g., BERT , achieved great success in NLP. Based on transformers, they mask a portion of the input tokens and require the models to predict the missing content from the reserved tokens. Pretraining with MAE encourages transformers to build effective long-range feature relationships. Recently, transformers also stream into the computer vision area, such as image recognition , video inpainting , optical flow , point cloud recognition . By breaking the limitations that convolution can only model local features, transformers present a significant performance gap compared to the previous counterparts. Pretraining with MAE is also introduced to these modalities, e.g., image , video , point cloud . These works show that MAE effectively releases the transformer power and do not require extra labeled data. FlowFormer presents a transformer-based cost volume encoder and achieves state-of-the-art accuracy. In this paper, we propose the masked cost-volume autoencoding to pretrain the cost volume encoder on a video dataset, which further unleashes the power of the transformer-based cost-volume encoder.
Method
As presented in Fig. 2, we propose a masked cost-volume autoencoding (MCVA) scheme to pretrain the cost-volume encoder of FlowFormer framework for better performance. The key of general masked autoencoding methods is to mask a portion of data and encourage the network to reconstruct the masked tokens from visible ones. Due to the redundant nature of the cost volume and the original FlowFormer architecture being incompatible with masks, naively adopting this paradigm to pretrain the cost-volume encoder leads to inferior performance. Our proposed MCVA tackles the challenge and conducts masked autoencoding with three key components: a proper masking strategy on the cost volume, modifying FlowFormer architecture to accommodate masks, and a novel pre-text reconstruction task supervising the pretraining process.
In this section, we first revisit the FlowFormer architecture, and then elaborate the proposed three key designs. We first introduce the masking strategy, dubbed as block-sharing masking, and then show the masked cost-volume tokenization that makes the cost-volume encoder compatible with masks. Coupling these two designs prevents the masked autoencoding from being hindered by information leakage in pretraining. Finally, we present the pre-text cost reconstruction task, mimicing the decoding process in finetuning to pretrain the cost-volume encoder .
FlowFormer is the first transformer architecture specifically designed for optical flow estimation, which enjoys the benefits of long-range information encoding via self-attention, but also encounters the similar problem to general vision transformers: it needs large-scale training data to model unbiased representations. The FlowFormer with the ImageNet-pretrained Twins-SVT backbone leads to boosted accuracy, while the same model with a train-from-scratch Twins-SVT or a shallow CNN achieve similar degraded performances, demonstrating the necessity of pretraining transformers for optical flow estimation. However, the ImageNet can only be used for pretraining the single-image encoder and the cost-volume encoder in FlowFormer is still trained from scratch and might not converge to the optimal point. To enable the pretraining of the cost-volume encoder to further enhance optical flow estimation, we propose the masked cost-volume autoencoding scheme.
2 Block-sharing Cost Volume Masking
To prevent such an over-simplified learning process, we propose a block-sharing masking strategy. We partition source pixels into non-overlapping blocks in each iteration. All source pixels belonging to the same block share a common mask for masked region reconstruction. In this way, neighboring source pixels are unlikely to copy each other’s cost maps to over-simplify the autoencoding process. Besides, the size of block is designed to be large (height and width of blocks are of pixels) and randomly changes in each iteration, and thus encouraging the cost-volume encoder to aggregate information from long-range context and to filter noises of cost values. The details of the mask generation algorithm are provided in supplementary.
Specifically, for each source pixel’s cost map, we first generate the mask map at resolution, and then up-sample it for three times to obtain a pyramid of mask maps , where , which are used for the down-sampling encoding process and will be discussed later in Sec. 3.3.
Another key design is that, in pretraining, we freeze the ImageNet-pretrained Twins-SVT backbone to build the cost volume from the pair of input images. Freezing the image encoder ensures the reconstruction targets (i.e., raw cost values) to maintain static and avoids training collapse.
3 Masked Cost-volume Tokenization
where indicates element-wise multiplication, , and is the raw cost map . The masked convolutions with the three binary mask maps remove all cost features in the masked regions in pretraining. Secondly, FlowFormer further projects the patchified cost-map features into the latent space via cross-attention. We thus remove the tokens in indicated by the mask map and then only project the remaining tokens into the latent space via the same cross-attention. During finetuning, the mask maps are removed to utilize all cost features, which converts the masked convolution to the vanilla convolution but the pretrained parameters in the convolution kernels and cross-attention layer are maintained.
The masked cost-volume tokenization completes two tasks. Firstly, it ensures the subsequent cost-volume encoder only processes visible features in pertaining. Secondly, the network structure is consistent with the standard tokenization of FlowFormer and can directly be used for finetuning so that the pretrained parameters have the same semantic meanings. After the masked cost-volume tokenization, the cost aggregation layers (i.e., AGT layers) take visible features as input which also don’t need to be modified in finetuning. The latent features interact with those of other source pixels in AGT layers and are transformed to the cost memory . We explain how to decode the cost memory to estimate flows in following section.
4 Reconstruction Target for Cost Memory Decoding
With the masked cost-volume tokenization, the cost encoder encodes the unmasked cost volume into the cost memory. The next step is decoding and reconstructing the masked regions from the cost memory.
In this section, we formulate the pre-text reconstruction targets, which supervises the decoding process as well as aformentioned embedding and aggregation layers. We start by revisiting the dynamic positional decoding scheme of FlowFormer, and present our reconstruction targets which are highly consistent with the finetuning tasks. FlowFormer adopts recurrent flow prediction. In each iteration of the recurrent process, the flow of source pixel is decoded from cost memory , conditioned on current predicted flow, to update the flow prediction. Specifically, current predicted corresponding location in the target image is computed as , where is current predicted flow. A local cost patch is then cropped from the window centered at on the raw cost map . FlowFormer utilizes this local cost patch (with positional encoding) as the query feature to retrieve aggregated cost feature via cross-attention operation:
Pre-text Reconstruction. Intuitively, should contain long-range cost information for better optical flow estimation and it is conditioned on local cost patch , which indicates the interested location on the cost map. We design a pre-text reconstruction task in line with these two characteristics to pretrain the cost-volume encoder as shown in Fig. 3: small cost-map patches are randomly cropped from the cost maps to retrieve cost features from the cost memory, targeting at reconstructing larger cost-map patches centered at the same locations.
where is the set of source pixels.
Discussion. The key of pretraining is to maintain consistent with finetuning, in terms of both network arthictecture and prediction target. To this end, we keep the cross-attention decoding layer unchanged and construct inputs with the same semantic meaning (e.g., replacing dynamically predicted with randomly sampled ); we supervise the extracted feature with long-range cost values to encourage the cost-volume encoder to aggregation global information for better optical flow estimation. What’s more, our scheme only takes an extra light-weight MLP as prediction head, which is unused in finetuning. Compared with previous methods that use a stack of self-attention layers, it is much more computationally efficient.
Experiments
We evaluate our FlowFormer++ on the Sintel and KITTI-2015 benchmarks. We pretrain FlowFormer++ using the proposed Masked Cost-volume Autoencoding on YouTube-VOS dataset. For the supervised finetuning, following previous works, we train FlowFormer++ on FlyingChairs and FlyingThings , and then respectively finetune it on the Sintel and KITTI-2015 benchmarks. FlowFormer++ obtains all-sided improvements over FlowFormer, ranking 1st on both benchmarks.
Experimental Setup. We adopt the commonly-used average end-point-error (AEPE) as the evaluation metric. It measures the average distance between predictions and ground truth. For the KITTI-2015 dataset, we additionally use the F1-all (%) metric, which refers to the percentage of pixels whose flow error is larger than 3 pixels or over 5% of the length of ground truth flows. YouTube-VOS is a large-scale dataset containing video clips from YouTube website. The Sintel dataset is rendered from the same movie in two passes: the clean pass is rendered with easier smooth shading and specular reflections, while the final pass includes motion blur, camera depth-of-field blur and atmospheric effects. The motions in the Sintel dataset are relatively large and complicated. The KITTI-2015 dataset constitutes of real-world driving scenarios with sparse ground truth.
Implementation Details. We use the same architecture of FlowFormer for fair comparison. The image feature encoder and context feature encoder are chosen as the first two stages of ImageNet-pretrained Twins-SVT, which are frozen in pretraining for better performance. We pretrain our model on YouTube-VOS for 50k iterations with a batch size of 24. The highest learning rate is set as . During finetuning, we follow the same training procedure of FlowFormer. We train our model on FlyingChairs for 120k iterations with a batch size of 8 and on FlyingThings with a batch size of 6 (denoted as ‘C+T’). Then, we train FlowFormer++ by combining data from Sintel, KITTI-2015 and HD1K (denoted as ‘C+T+S+K+H’) for another 120k iterations with a batch size of 6. This model is submitted to Sintel online test benchmark for evaluation. To obtain the best performance on the KITTI-2015 benchmark, we further train FlowFormer++ on the KITTI-2015 dataset for 50k iterations with a batch size of 6. The highest learning rate is set as for FlyingChairs and on other training sets. In both pretraining and finetuning, we choose AdamW optimizer and one-cycle learning rate scheduler. We crop images and tile predictions from all patches to obtain full-resolution flow predictions following Perceiver IO and FlowFormer.
As shown in Table 1, we evaluate FlowFormer++ on the Sintel and KITII-2015 benchmarks. Following previous methods, we evaluate the generalization performance of models on the training sets of Sintel and KITTI-2015 (denoted as ‘C+T’). We also compare the dataset-specific accuracy of optical flow models after dataset-specific finetuning (denoted as ‘C+T+S+K+H’). Autoflow is a synthetic dataset of complicated visual disturbance, while its training code is unreleased.
Generalization Performance. The ‘C+T’ setting in Table 1 reflects the generalization capacity of models. FlowFormer++ ranks 1st on both benchmarks among published methods. It achieves 0.90 and 2.30 on the clean and final pass of Sintel. Compared with FlowFormer, it achieves 4.26% error reduction on Sintel clean pass. On the KITTI-2015 training set, FlowFormer++ achieves 3.93 F1-epe and 14.13 F1-all, improving FlowFormer by 0.16 and 0.59, respectively. These results show that our proposed MCVA promotes the generalization capacity of FlowFormer.
Dataset-specific Performance. After training the FlowFormer++ in the ‘C+T+S+K+H’ setting, we evaluate its performance on the Sintel online benchmark. It achieves 1.07 and 2.09 on the clean and final passes, a 7.76% and 7.18% error reduction from previous best model FlowFormer.
We further finetune FlowFormer++ on the KITTI-2015 training set after the Sintel stage and evaluate its performance on the KITTI online benchmark. FlowFormer++ achieves 4.52 F1-all, improving FlowFormer by 0.16 while also outperforming the previous best model S-Flow by 0.12.
To conclude, FlowFormer++ shows greater optical flow estimation capacity for both naturalistic non-rigid motions (Sinel) and real-world rigid scenarios (KITTI-2015). This validates that our proposed MCVA improves the FlowFormer architecture by enhancing the cost-volume encoder.
2 Qualitative Experiments
We visualize flow predictions by our FlowFormer++ and FlowFormer on Sintel and KITTI test sets in Fig. 4 to qualitatively show how FlowFormer++ outperforms FlowFormer. The red arrows highlight that FlowFormer++ preserves clearer details than FlowFormer: in the first row, FlowFormer misses the flying bird while FlowFormer++ produces clear results; in the second row, FlowFormer++ keeps the boundaries of leaves while FlowFormer generates blurry prediction. FlowFormer++ also shows greater global aggregation capacity indicated by blue boxes. In the first row, FlowFormer produces obviously inconsistent prediction on the large-area sky, while FlowFormer++ yields consistent prediction. In the second row, the black car is partially occluded by the foreground tree, which challenges the optical flow model to aggregate information in long range. FlowFormer++ generates consistent prediction for the two separated parts of the car, while FlowFormer mixes the left part of the car with background and thus produces inconsistent optical flow prediction.
3 Ablation Study
We conduct a set of ablation studies to show the impact of designs in the Masked Cost-volume Autoencoding (MCVA). All models in the experiemnts are first pretrained and then finetuned on ‘C+T’. We report the test results on Sintel and KITTI training sets.
Masking Strategy. Masking strategy is one important design of our MCVA. As shown in Table 2, pretraining FlowFormer with random masking already improves the performance on three of the four metrics. But the proposed block-sharing masking strategy brings even larger gain, which demonstrates the effectiveness of this design. Besides, we observe higher pretraining loss with block-sharing masking than that with random masking, validating that the block-sharing masking makes the pretraining task harder.
Masking Ratio. Masking ratio influences the difficulty of the pre-text reconstruction task. We empirically find that the mask ratio of 50% yields the best overall performance.
Pre-text Reconstruction Design. The conventional MAE methods aim to reconstruct input data at fixed locations and use the positional encodings as query features to absorb information for reconstruction (the first row of Table 4). To keep consistent with the dynamic positional query of FlowFormer architecture, we propose to reconstruct contents at random locations (the second row of Table 4) and additionally use local patches as query features (the third row of Table 4). The results validate the necessity of ensuring semantic consistency between pretraining and finetuning.
Freezing Image and Context Encoders. The FlowFormer architecture has an image encoder to encode visual appearance features for constructing the cost volume, and a context encoder to encode context features for flow prediction. As shown in Table 5, freezing the image encoder is necessary, otherwise the model diverges. We hypothesize that the frozen image encoder ensures the reconstruction targets (i.e., raw cost values) to keep static. Freezing the context encoder leads to better overall performance.
Comparisons with Unsupervised Methods. We also use conventional unsupervised methods to pretrain FlowFormer with photometric loss and smooth loss following and then finetune it in the ‘C+T’ setting as FlowFormer++. As shown in Table 6, our MCVA outperforms the unsupervised counterpart for pretraining FlowFormer.
FlowFormer++ v.s. FlowFormer on FlyingChairs. We show the training and validating loss of the training process on FlyingChairs in Fig. 5. FlowFormer++ presents faster convergence during training and better validation loss at the end, which reveals that FlowFormer++ learns effective feature relationships during pretraining and benefits the supervised finetuning.
Conclusion
In this paper, we propose Masked Cost Volume Autoencoding (MCVA) to enhance the cost-volume encoder of FlowFormer by pretraining. We show that the naive adaptation of MAE scheme to cost volume does not work due to the redundant nature of cost volumes and the incurred pretraining-finetuning discrepancy. We tackle these issues with a specially designed block-sharing masking strategy and the novel pre-text reconstruction task. These designs ensure semantic integrity between pretraining and finetuning and encourage the cost-volume to aggregate information in a long range. Experiments demonstrate clear generalization and dataset-specific performance improvements.