Learning Pixel Trajectories with Multiscale Contrastive Random Walks

Zhangxing Bian, Allan Jabri, Alexei A. Efros, Andrew Owens

Introduction

Temporal correspondence underlies a range of video understanding tasks, from optical flow to object tracking. At the core, the challenge is to estimate the motion of some entity as it persists in the world, by searching in space and time. For historical reasons, the practicalities differ substantially across tasks: optical flow aims for dense correspondences but only between neighboring pairs of frames, whereas tracking cares about longer-range correspondences but is spatially sparse. We argue that the time might be right to try and re-unify these different takes on temporal correspondence.

An emerging line of work in self-supervised learning has shown that generic representations pretrained on unlabeled images and video can lead to strong performance across a range of tracking tasks . The key idea is that if tracking can be formulated as label propagation on a space-time graph, all that is needed is a good measure of similarity between nodes. Indeed, the recent contrastive random walk (CRW) formulation shows how such a similarity measure can be learned for temporal correspondence problems, suggesting a path towards a unified solution. However, scaling this perspective to pixel-level space-time graphs holds challenges. Since computing similarity between frames is quadratic in the number of nodes, estimating dense motion is prohibitively expensive. Moreover, there is no way of explicitly estimating the motion in ambiguous cases, like occlusion. In parallel, the unsupervised optical flow community has adopted highly effective methods for dense matching , which use multiscale representations to reduce the large search space, and smoothness priors to deal with ambiguity and occlusion. But, in contrast to the self-supervised tracking methods, they rely on hand-crafted distance functions, such as the Census Transform . Furthermore, because they focus on producing point estimates of motion, they may be less robust under long-term dynamics.

In this work, we take a step toward bridging the gap between tracking and optical flow by extending the contrastive random walk formulation to much denser, pixel-level space-time graphs. The main contribution is introducing hierarchy into the search problem, i.e., the multiscale contrastive random walk. By integrating local attention in a coarse-to-fine manner, the model can efficiently consider a distribution over pixel-level trajectories. Through experiments across optical flow and video label propagation benchmarks, we show:

This provides a unified technique for self-supervised optical flow, pose tracking, and video object segmentation.

For optical flow, the model is competitive with many recent unsupervised optical flow methods, despite using a novel loss function (without hand-crafted features).

For tracking, the model outperforms existing self-supervised approaches on pose tracking, and is competitive on video object segmentation.

Contrastive cycle-consistency provides a complementary learning signal to photo-consistency.

Multi-frame training improves two-frame performance.

Related work

Recent work has proposed methods for tracking objects in video through self-supervised representation learning. Vondrick et al. posed tracking as a colorization problem, by training a network to match pixels in grayscale images that have the same held-out colors. The assumption that matching pixels have the same color may break down over long time horizons, and the method is limited to grayscale inputs. This approach was extended to obtain higher-accuracy matches with two-stage matching . In contrast to our approach, the predictions are coarse, patch-level associations. Another line of work learns representations by maximizing cycle consistency. These methods track patches forward, then backward, in time and test whether they end up where they began. Wang et al. proposed a method based on hard attention and spatial transformers . Jabri et al. posed cycle-consistency as a random walk problem, allowing the model to obtain dense supervision from multi-frame video. Tang et al. proposed an extension that allowed for fully convolutional training. These approaches are trained on sparse patches and learn coarse-grained correspondences. In contrast, we learn pixel-to-pixel correspondences. Other work has encouraged patches in the same position in neighboring frames to be close in an embedding space .

Lucas and Kanade used Gauss-Newton optimization to minimize a brightness constancy objective. In their seminal work, Horn and Schunck combined brightness constancy with a spatial smoothness criteria, and estimated flow using variational methods. Later research improved flow estimation using robust penalties , coarse-to-fine estimation , discrete optimization , feature matching , bilateral filtering , and segmentation . In contrast, our model estimates flow using neural networks.

Early work used Boltzmann machines to learn transformations between images. More recently, Yu et al. train a neural network to minimize a loss very similar to that of optimization-based approaches. Later work extended this approach by adding edge-aware smoothing , hand-crafted features , occlusion handling , learned upsampling , and depth and camera pose . Another approach learns flow by matching augmented image pairs . Recently, Jonschkowski et al. surveyed previous literature and conducted an exhaustive search to find the best combination of methods and hyperparameters. In contrast to these works, our goal is to learn self-supervised representations for matching, in lieu of hand-crafted features. Moreover, we aim to produce a distribution over motion trajectories suitable for label transfer, rather than motion estimates alone. In very recent work, Stone et al. adapted unsupervised flow methods to the RAFT architecture and proposed new augmentation, self-distillation, and multi-frame occlusion inpainting methods. By contrast, we use PWC-net , since it is the standard architecture considered in prior work, and since it can obtain strong performance with careful training .

Cycle consistency has long been used to detect occlusions , and is used to discount the loss of occluded pixels in unsupervised flow . Zou et al. used a cycle consistency loss as part of a system that jointly estimated depth, pose, and flow. Recently, Huang et al. combined cycle-consistency with epipolar matching, but their method is weakly supervised with camera pose and assumes egomotion. In contrast, ours is entirely unsupervised and is capable of working solely with cycle-consistency and smoothness losses. Without extra constraints, their cycle consistency formulations have trivial solutions (e.g., all-zero flow). Other recent work learns to match images by ensuring that both an image and a warped variation of it match consistently with a second image, and Li et al. used random walks with fixed transition matrices to smooth scene flow on point clouds. A random walk formulation of cycle consistency also been used in semi-supervised learning , using labels to test consistency.

Many methods use a third frame to obtain more local evidence for matching. Classic methods assume approximately constant velocity or acceleration and measure the photo-consistency over the full set of frames. Recently, Janai et al. proposed an unsupervised multi-frame flow method that used a photometric loss with a low-acceleration assumption and explicit occlusion handling. In contrast to these approaches, our method can be deployed using two frames at test time. We use subsequent frames as a training signal. There have also been a variety of approaches that track over long time horizons, often using sparse (or semi-dense) keypoints . Other work chains together optical flow , typically after removing low-texture regions. In contrast, our method also learns “soft” per-pixel tracks, which convey the probability that pairs of pixels match.

Early work learned optical flow with probabilistic models, such as graphical models . Other work learns parameters for smoothness and brightness constancy or robust penalties . More recent methods has used neural networks. Fischer et al. proposed architectures with a built-in correlation layers. Sun et al. introduced a network with built-in coarse-to-fine matching. Recent work iteratively updates a flow with multiscale features, in lieu of coarse-to-fine matching.

Method

We first show how to learn dense space-time correspondences using mutiscale contrastive random walks, resulting in a model that obtains high quality motion estimates via simple nonparametric matching. We then describe how the learned representation can be combined with regression to handle occlusions and ambiguity, for improved optical flow.

We review the single-scale contrastive random walk, then extend the approach to multiscale estimation.

We build on the contrastive random walk (CRW) formulation of Jabri et al. . Given an input video with kk frames, we extract nn patches from each frame and assign each an embedding using a learned encoder ϕ\phi. These patches form the vertices of a graph that connects all patches in temporally adjacent frames. A random walker steps through the graph, forward in time from frames 1,2,...,k1,2,...,k, then backward in time from k−1,k−2,...,1k-1,k-2,...,1. The transition probabilities are determined by the similarity of the learned representations:

where the log⁡\log is elementwise and Aˉt,t+k\bar{A}_{t,t+k} are the transition probabilities from frame tt to t+kt+k: Aˉt,t+k=∏i=tt+k−1Ai,i+1\bar{A}_{t,t+k}=\prod_{i=t}^{t+k-1}A_{i,i+1}.

1.2 Optical flow as a random walk

In contrast to widely-used forward-backward cycle consistency formulations , which measure the deviation of a predicted motion from a starting point, there is no trivial solution (e.g., all-zero flow). This is because cycle consistency is measured in an embedding space defined solely from the visual content of image regions.

1.3 Multiscale random walk

As presented so far, this formulation is expensive to scale to high-resolutions because computing the transition matrix is quadratic in the number of nodes. We overcome this by introducing hierarchy into the search problem. Instead of comparing all pairs, we only attend on a local neighborhood. By integrating local search across scales in a coarse-to-fine manner, the model can efficiently consider a distribution over pixel-level trajectories.

Computing the transition matrix closely resembles cost volume estimation in optical flow . This inspires us to draw on the classic spatial pyramid commonly used for multiscale search in optical flow, by iteratively computing the dense transition matrix, from coarse to fine spatial scales l∈[1..L]l\in[1..L].

After computing the transition matrices between all pairs of adjacent frames, we sum the contrastive random walk loss over all levels:

where Aˉs,tl\bar{A}_{s,t}^{l} is defined as in Eq. 2 and nln_{l} is the number of nodes in level ll. In our experiments, we use L=5L=5 scales and consider k∈[2..4]k\in[2..4] length cycles.

1.4 Smooth random walks

Since natural motions tend to be smooth , we follow work in optical flow , and incorporate smoothness as an additional desiderata for our random walks. We use the edge-aware loss of Jonschkowski et al. , which penalizes spatial changes in flow near similarly-colored pixels:

where Id(p)=13∑c∣∂Ic∂d∣I_{d}(p)=\frac{1}{3}\sum_{c}\lvert\frac{\partial I_{c}}{\partial d}\rvert is the spatial derivative averaged over all color channels IcI_{c} in direction dd. The parameter λc\lambda_{c} controls the influence of similarly colored pixels. We apply this loss to each scale of the model.

2 Handling occlusion

While effective for most image content, nonparametric matching has no mechanism for estimating the motion of pixels that become occluded, since it requires a corresponding patch in the next frame. We propose a variation of the model that combines the multiscale contrastive random walk with a regression module that directly predicts the flow values at each pixel.

The architecture of our regressor closely follows the refinement module of PWC-net . We learn a function greg(⋅)g_{\mathtt{reg}}(\cdot) that regresses the flow at each pixel from the multiscale contrastive random walk cost-volume and convolutional features. These features are obtained from the same shared backbone that is used to compute the embeddings (model diagram provided in Figure 5).

where fs,t\mathbf{f}_{s,t} is the predicted flow and λa\lambda_{a} is a constant. To prevent the regression loss from influencing the learned embeddings that it is based on, we do not propagate gradients from the regressor to the embeddings XX during training. As in Eq. 7, we apply the loss to each scale and sum.

3 Training

The pure nonparametric model (Section 3.1) can be trained by simply minimizing the multiscale contrastive random walk loss with a smoothness penalty:

Adding the regressor results in the following loss:

We include weighting factors to control the relative importance of each loss (Section B).

We follow and include subcycles in our contrastive random walks: when training the model on kk-frame videos, we include losses for walks of length kk, k−1k-1, … 2. These losses can be estimated efficiently by reusing the transition matrices for the full walk.

When training with k>2k>2 frames, we use curriculum learning to speed up and stabilize training. We train the model to convergence with 2, 3, … kk frame cycles in succession.

To implement the contrastive random walk, we exploit the sparsity of our coarse-to-fine formulation, and represent the transition matrices As,tlA_{s,t}^{l} as sparse matrices. This significantly improved training times and reduced memory requirements, especially in the finest scales. It takes approximately 3 days to train the full model on one GTX2080 Ti. We train our network with PyTorch , using the Adam optimizer with a cyclic learning rate schedule with a base learning rate of 10−410^{-4} and a max learning rate of 5×10−45\times 10^{-4}. We provide training hyperparameters in Section B.

The contrastive random walk can potentially obtain shortcut solutions when it is trained with a fully convolutional network by exploiting positional information . While recent work has shown this can be solved through augmentation , we found that we avoided trivial shortcuts when using reflection padding in our network (for all convolution layers except for in the regressor). This may be because we simultaneously optimize multiple losses and use a limited search window, making the trivial solution harder to find.

Results

Our model produces two outputs: the optical flow fields and the pixel trajectories (which are captured in the transition matrices). We evaluate these predictions on label transfer and motion estimation tasks. We compare them to space-time correspondence learning methods, and with unsupervised optical flow methods.

For simple comparison with other methods, we train on standard optical flow datasets. We note that the training protocols used by unsupervised optical flow literature are not standardized. We therefore follow the evaluation setup of . We pretrain models on unlabeled videos from the Flying Chairs dataset . We then train on the KITTI-2015 multi-view extension and Sintel . To evaluate our model’s ability to learn from internet video, we also trained the model on YouTube-VOS , without pretraining on any other datasets.

We also evaluate our model on standard label transfer tasks. The JHMDB benchmark transfers 15 body parts to future frames over long time horizons. The DAVIS benchmark transfers object masks.

2 Label propagation

We evaluate our learned model’s ability to perform video label propagation, a task widely studied in video representation learning work. The goal is to propagate a label map provided in an initial video frame, which might describe keypoints or object segments, to the rest of the video.

We use variations of our model that was trained on the unlabeled Sintel and YouTube-VOS datasets, and use the transition matrix and flow fields at the penultimate level of the pyramid, i.e., level 44. Since the transition matrix describes residual motion, we warp (i.e., with ft,s4f^{4}_{t,s}) each label map before querying. Finally, since the features in level 44 have only 1616 channels, we stack features from levels 33 and 44 to obtain hypercolumns , before computing attention.

Our model is able to estimate motion solely through nonparametric matching, as can be seen qualitatively in Figure 4. Despite the model’s simplicity and the fact that it is based on very different principles than existing flow methods, it obtains strong performance on matching non-occluded pixels (Tab. 2). It outperforms many unsupervised optical flow models, such as SelFlow (Tab. 4) on KITTI noc metric (non-occluded endpoint error). We see that the full regression-based variation of the model obtains better results, particularly on the all metric.

To help understand the importance of our multiscale formulation, we compared to Jabri et al. on flow benchmarks, using their publicly released model (Tab. 2). This model resembles our nonparametric model, but with the random walk occurring at a single scale, with no smoothness prior. We evaluate the model using dense features, as in their approach. We found that our model significantly outperforms it. To control for training differences, we also tried removing scales from multiscale training (Tab. 6), by using untrained (random) embeddings for the fine scales. We found that this significantly reduced performance.

3.2 Effects of multi-frame cycles

We asked how the quality of the representation changes as we vary the number of frames kk used to train the random walk. We test all models on 2-frame optical flow. As seen in Table 6, the model obtains better performance on all metrics using 3-frame and 4-frame cycles.

3.3 Photometric feature learning

Next, we considered using a state-of-the-art hand-crafted feature, the Census transform , resulting in a model similar to the Census variation of UFlow. We found that our features obtained competitive performance on non-occluded pixels, but that there was a significant advantage to Census features on the all metric. This is understandable since the contrastive random walk does not have a way of learning features for occluded pixels. Interestingly, we found that combining the two features together improved performance, and that the gap improves further when multi-frame walks are used, obtaining the overall best results.

Moreover, the combined features show more robustness on image pairs with rapid exposure and hue changes. We evaluated models trained on hue- and brightness-jittered image pairs and found that the model with our learned features performed significantly better, and that the gap increased with the magnitude of the jittering (Tab. 4b). Please see Section C for details.

3.4 Motion estimation ablations

To help understand which properties of our model contribute to its performance, we perform an ablation study on KITTI-15 (Table 3). We asked how the different losses contribute to the performance. We ablated the smoothness loss (Eq. 8), the self-supervision loss (Sec. 3.2), and removing the constraint on the regressor in Eq. 9 by setting λa=0\lambda_{a}=0. We see that the smoothness loss significantly improves results. Similarly, we discarded the feature consistency term from Eq. 9, which reduces the quality of the results but outperforms the nonparametric model on the all metric.

We found that our model generalized well to benchmark datasets when training solely on YouTube-VOS (Tab. 7). For comparison, we also trained ARFlow on YouTube-VOS. Our model obtained better performance on KITTI, while ARFlow performed better on Sintel.

3.5 Comparison to recent optical flow methods

To help understand our model’s overall performance, we compare it to recent optical flow methods (Table 4). We include models that use different numbers of frames for the random walk, and a variation of ARFlow that uses our self-supervised features to augment its photometric loss.

We found that our models outperform many recent unsupervised optical flow methods, including the (3-frame) MFOccFlow and EPIFlow. In particular, we significantly outperform the recent SelFlow method , despite the fact that it takes 3 frames as input at test time and uses Census Transform features. By contrast, our model uses no hand-crafted image features. The highest performing method is the very recent, highly optimized SMURF model, which uses the RAFT architecture instead of PWC-net and which extends UFlow . This model uses a variety of additional training signals, such as extensive data augmentation, occlusion inpainting with multi-stage training, self-distillation, and hand-crafted features.

We have proposed a method for learning dense motion estimation using multiscale contrastive random walks. We see our work as a potential step toward unifying self-supervised tracking and optical flow. Moreover, the model can learn from internet video, which suggests that the emergent representations of such a hierarchical tracker may learn interesting part-whole structure at scale.

Motion analysis has many applications, such as in health monitoring, surveillance, and security. There is also a potential for the technology to be used for harmful purposes if weaponized. The released models are limited in scope to the datasets used in training.

References

We thank David Fouhey and Jeff Fessler for the helpful feedback. AO thanks Rick Szeliski for introducing him to multi-frame optical flow. This research was supported in part by Toyota Research Institute, Cisco Systems, and Berkeley Deep Drive.

Appendix A Architecture

We provide additional details about the network architecture, which closely resembles PWC-Net with the simplifications introduced by ARflow . We attach a 1×11\times 1 convolutional layer to the each scale to obtain the embeddings (of 32 channels) for contrastive random walk at each scale. We show a diagram for the model (with the regressor) in Fig. 5.

We train our network with PyTorch , using the Adam optimizer with a cyclic learning rate schedule with a base learning rate of 10−410^{-4} and a max learning rate of 5×10−45\times 10^{-4}. We use batch size of 4 for 2-cycle model and 2 for 3- and 4-cycle models (due to memory constraints). The total training takes approximately four days on two GTX 2080Ti, two days for training on Flying Chairs and two days for training on Sintel/KITTI. Experiments on Sintel and KITTI start from a model that was first trained on Flying Chairs, as in .

Appendix B Hyperparameters

We list the hyperparameters and ranges considered during our experiments. Weights for the boundary loss, learned photometric loss are hand-chosen. Parameters in bold are the ones that were systematically explored via ablations. For the image resolution, we follow the experimental setup of Jonschkowski et al. i.e., Flying Chairs: 384×512384\times 512, Sintel: 448×1024448\times 1024, KITTI: 640×640640\times 640. We use the same loss weight across different scales for a specific type of loss. RGB image values are scaled to $$ and augmented by randomly shifting the hue and brightness and randomly flipped left/right. The augmentation is kept the same across frame in a pair of images. In contrast to other work , we did not modify the model per dataset. For the optical flow baselines in label propagation (Table 3) we used the supervised RAFT trained on FlyingThings3D .

Appendix C Robustness on jittered images

To help understand the flexibility of our self-supervised model, we trained a variation of our model on images with large, simulated brightness and hue variations, inspired by the challenges of rapid exposure changes (e.g., in HDR photography). During training and testing, we randomly jitter the brightness and hue of the second image in KITTI by a factor of up to 0.6 and 0.3 respectively, using PyTorch’s built-in augmentation. We finetuned the variation of our model that combines our learned features with Census features (Tab. 5), since it obtained strong performance on KITTI. We also finetuned a model with only Census features. We found that the model with our learned features performed significantly better, and that the gap increased with the magnitude of the jittering (Tab. 4b). We show qualitative results in Fig. 6.

Appendix D Additional qualitative results

We provide additional qualitative results for optical flow in Figure 7.