Video Frame Synthesis using Deep Voxel Flow

Ziwei Liu, Raymond A. Yeh, Xiaoou Tang, Yiming Liu, Aseem Agarwala

Introduction

Videos of natural scenes observe a complicated set of phenomena; objects deform and move quickly, occlude and dis-occlude each other, scene lighting changes, and cameras move. Parametric models of video appearance are often too simple to accurately model, interpolate, or extrapolate video. None the less, video interpolation, i.e., synthesizing video frames between existing ones, is a common process in video and film production. The popular commercial plug-in Twixtorhttp://revisionfx.com/products/twixtor/ is used both to resample video into new frame-rates, and to produce a slow-motion effect from regular-speed video. A related problem is video extrapolation; predicting the future by synthesizing future video frames.

The traditional solution to these problems estimates optical flow between frames, and then interpolates or extrapolates along optical flow vectors. This approach is “optical-flow-complete”; it works well when optical flow is accurate, but generates significant artifacts when it is not. A new approach uses generative convolutional neural networks (CNNs) to directly hallucinate RGB pixel values of synthesized video frames. While these techniques are promising, directly synthesizing RGB values is not yet as successful as flow-based methods, and the results are often blurry.

In this paper we aim to combine the strengths of these two approaches. Most of the pixel patches in video are near-copies of patches in nearby existing frames, and copying pixels is much easier than hallucinating them from scratch. On the other hand, an end-to-end trained deep network is an incredibly powerful tool. This is especially true for video interpolation and extrapolation, since training data is nearly infinite; any video can be used to train an unsupervised deep network.

We therefore use existing videos to train a CNN in an unsupervised fashion. We drop frames from the training videos, and employ a loss function that measures similarity between generated pixels and the ground-truth dropped frames. However, like optical-flow approaches our network generates pixels by interpolating pixel values from nearby frames. The network includes a voxel flow layer — a per-pixel, 3D optical flow vector across space and time in the input video. The final pixel is generated by trilinear interpolation across the input video volume (which is typically just two frames). Thus, for video interpolation, the final output pixel can be a blend of pixels from the previous and next frames. This voxel flow layer is similar to an optical flow field. However, it is only an intermediate layer, and its correctness is never directly evaluated. Thus, our method requires no optical flow supervision, which is challenging to produce at scale.

We train our method on the public UCF-101 dataset, but test it on a wide variety of videos. Our method can be applied at any resolution, since it is fully convolutional, and produces remarkably high-quality results which are significantly better than both optical flow and CNN-based methods. While our results are quantitatively better than existing methods, this improvement is especially noticeable qualitatively when viewing output videos, since existing quantitative measures are poor at measuring perceptual quality.

Related Work

Video interpolation is commonly used for video re-timing, novel-view rendering, and motion-based video compression . Optical flow is the most common approach to video interpolation, and frame prediction is often used to evaluate optical flow accuracy . As such, the quality of flow-based interpolation depends entirely on the accuracy of flow, which is often challenged by large and fast motions. Mahajan et al. explore a variation on optical flow that computes paths in the source images and copies pixel gradients along them to the interpolated images, followed by a Poisson reconstruction. Meyer et al. employ a Eulerian, phase-based approach to interpolation, but the method is limited to smaller motions.

Convolutional neural networks have been used to make recent and dramatic improvements in image and video recognition . They can also be used to predict optical flow , which suggests that CNNs can understand temporal motion. However, these techniques require supervision, i.e., optical flow ground-truth. A related unsupervised approach uses a CNN to predict optical flow by synthesizing interpolated frames, and then inverting the CNN. However, they do not use an optical flow layer in the network, and their end-goal is to generate optical flow. They do not numerically evaluate the interpolated frames, themselves, and qualitatively the frames appear blurry.

There are a number of papers that use CNNs to directly generate images and videos . Blur is often a problem for these generative techniques, since natural images follow a multimodal distribution, while the loss functions used often assume a Gaussian distribution. Our approach can avoid this blurring problem by copying coherent regions of pixels from existing frames. Generative CNNs can also be used to generate new views of a scene from existing photos taken at nearby viewpoints . These methods reconstruct images by separately computing depth and color layers at each hypothesized depth. This approach cannot account for scene motion, however.

Our technical approach is inspired by recent techniques for including differentiable motion layers in CNNs . Optical flow layers have also been used to render novel views of objects and change eye gaze direction while videoconferencing . We apply this approach to video interpolation and extrapolation. LSTMs have been used to extrapolate video , but the results can be blurry. Mathieu et al. reduce blurriness by using adversarial training and unique loss functions, but the results still contain artifacts (we compare our results against this method). Finally, Finn et al. use LSTMs and differentiable motion models to better sample the multimodal distribution of video future predictions. However, their results are still blurry, and are trained to videos in very constrained scenarios (e.g., a robot arm, or human motion within a room from a fixed camera). Our method is able to produce sharp results for widely diverse videos. Also, we do not pre-align our input videos; other video prediction papers either assume a fixed camera, or pre-align the input.

Our Approach

We propose Deep Voxel Flow (DVF) — an end-to-end fully differentiable network for video frame synthesis. The only training data we need are triplets of consecutive video frames. During the training process, two frames are provided as inputs and the remaining frame is used as a reconstruction target. Our approach is self-supervised and learns to reconstruct a frame by borrowing voxels from nearby frames, which leads to more realistic and sharper results (Fig. 4) than techniques that hallucinate pixels from scratch. Furthermore, due to the flexible motion modeling of our approach, no pre-processing (e.g., pre-alignment or lighting adjustment) is needed for the input videos, which is a necessary component for most existing systems .

The spatial component of voxel flow F\mathbf{F} represents optical flow from the target frame to the next frame; the negative of this optical flow is used to identify the corresponding location in the previous frame. That is, we we assume optical flow is locally linear and temporally symmetric around the in-between frame. Specifically, we can define the absolute coordinates of the corresponding locations in the earlier and later frames as L0=(x−Δx,y−Δy)\mathbf{L}^{0}=(x-\Delta x,y-\Delta y) and L1=(x+Δx,y+Δy)\mathbf{L}^{1}=(x+\Delta x,y+\Delta y), respectively. The temporal component of voxel flow F\mathbf{F} is a linear blend weight between the previous and next frames to form a color in the target frame. We use this voxel flow to sample the original input video X\mathbf{X} with a volume sampling function Tx,y,t\mathcal{T}_{x,y,t} to form the final synthesized frame Y^\mathbf{\hat{Y}}:

The volume sampling function samples colors by interpolating within an optical-flow-aligned video volume computed from X\mathbf{X}. Given the corresponding locations (L0,L1)(\mathbf{L}^{0},\mathbf{L}^{1}), we construct a virtual voxel of this volume and use trilinear interpolation from the colors at the voxel’s corners to compute an output video color Y^(x,y)\mathbf{\hat{Y}}(x,y). We compute the integer locations of the eight vertices of the virtual voxel in the input video X\mathbf{X} as:

where ⌊⋅⌋\lfloor\cdot\rfloor is the floor function, and we define the temporal range for interpolation such that t=0t=0 for the first input frame and t=1t=1 for the second. Given this virtual voxel, the 3D voxel flow generates each target voxel Y^(x,y)\mathbf{\hat{Y}}(x,y) through trilinear interpolation:

where Wijk\mathbf{W}^{ijk} is the trilinear resampling weight. This 3D voxel flow can be understood as the joint modeling of a 2D motion field and a mask selecting between the earlier and later frame. Specifically, we can separate F\mathbf{F} into Fmotion=(Δx,Δy)\mathbf{F}_{motion}=(\Delta x,\Delta y) and Fmask=(Δt)\mathbf{F}_{mask}=(\Delta t), as illustrated in Fig. 2 (e-f). (These definitions are later used in Eqn. 6 to allow different weights for spatial and temporal regularization.)

Network Architecture. DVF adopts a fully-convolutional encoder-decoder architecture, containing three convolution layers, three deconvolution layers and one bottleneck layer. Therefore, arbitrary-sized videos can be used as inputs for DVF. The network hyperparamters (e.g., the size of feature maps, the number of channels and activation functions) are specified in Fig. 1.

For the encoder section of the network, each processing unit contains both convolution and max-pooling. The convolution kernel sizes here are 5×55\times 5, 5×55\times 5 and 3×33\times 3, respectively. The bottleneck layer is also connected by convolution with kernel size 3×33\times 3. For the decoder section, each processing unit contains bilinear upsampling and convolution. The convolution kernel sizes here are 3×33\times 3, 5×55\times 5 and 5×55\times 5, respectively. To better maintain spatial information we add skip connections between the corresponding convolution and deconvolution layers. Specifically, the corresponding deconvolution layers and convolution layers are concatenated together before being fed forward.

2 Learning

For our DVF training, we exploit the l1l_{1} reconstruction loss with spatial and temporal coherence regularizations to reduce visual artifacts. Total variation (TV) regularizations are used here to enforce coherence. Since these regularizers are imposed on the output of the network it can be easily incorporated into the back-propagation scheme. Our overall objective function that we minimize is:

where D\mathcal{D} is the training set of all frame triplets, NN is its cardinality and Y\mathbf{Y} is the target frame to be reconstructed. ∥∇Fmotion∥1\|\nabla\mathbf{F}_{motion}\|_{1} is the total variation term on the (x,y)(x,y) components of voxel flow, and λ1\lambda_{1} is the corresponding regularization weight; similarly, ∥∇Fmask∥1\|\nabla\mathbf{F}_{mask}\|_{1} is the regularizer on the temporal component of voxel flow, and λ2\lambda_{2} its weight. (We experimentally found it useful to weight the coherence of the spatial component of the flow more strongly than the temporal selection.) To optimize the l1l_{1} norm, we use the Charbonnier penalty function Φ(x)=(x2+ϵ2)1/2\Phi(x)=(x^{2}+\epsilon^{2})^{1/2} for approximation. Here we empirically set λ1=0.01\lambda_{1}=0.01, λ2=0.005\lambda_{2}=0.005 and ϵ=0.001\epsilon=0.001.

We initialize the weights in DVF using Gaussian distribution with standard deviation of 0.010.01. Learning the network is achieved via ADAM solver with learning rate of 0.00010.0001, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999 and batch size of 3232. Batch normalization is adopted for faster convergence.

Differentiable Volume Sampling. To make our DVF an end-to-end fully differentiable system, we define the gradients with respect to 3D voxel flow F=(Δx,Δy,Δt)\mathbf{F}=(\Delta x,\Delta y,\Delta t) so that the reconstruction error can be backpropagated through a volume sampling layer. Similar to , the partial derivative of the synthesized voxel color Y^(x,y)\mathbf{\hat{Y}}(x,y) w.r.t. Δx\Delta x is

where Eijk\mathbf{E}^{ijk} is the error reassignment weight w.r.t. Δx\Delta x. Similarly, we can compute ∂Y^(x,y)/∂(Δy)\partial\mathbf{\hat{Y}}(x,y)/\partial(\Delta y) and ∂Y^(x,y)/∂(Δt)\partial\mathbf{\hat{Y}}(x,y)/\partial(\Delta t). This gives us a sub-differentiable sampling mechanism, allowing loss gradients to flow back to the 3D voxel flow F\mathbf{F}. This sampling mechanism can be implemented very efficiently by just looking at the kernel support region for each output voxel.

3 Multi-scale Flow Fusion

As stated in Sec. 3.2, the gradients of reconstruction error are obtained by only looking at the kernel support region for each output voxel, which makes it hard to find large motions that fall outside the kernel. Therefore, we propose a multi-scale Deep Voxel Flow (multi-scale DVF) to better encode both large and small motions.

Specifically, we have a series of convolutional encoder-decoder HN,HN−1,⋯ ,H0\mathcal{H}_{N},\mathcal{H}_{N-1},\cdots,\mathcal{H}_{0} working on video frames from coarse scale sNs_{N} to fine scale s0s_{0}, respectively. Typically, in our experiments, we set s2=64×64s_{2}=64\times 64, s1=128×128s_{1}=128\times 128 and s0=256×256s_{0}=256\times 256. In each scale kk, the sub-network Hk\mathcal{H}_{k} predicts 3D voxel flow Fk\mathbf{F}_{k} at that resolution. Intuitively, large motions will have a relatively small offset vector Fk\mathbf{F}_{k} in coarse scale sNs_{N}. Thus, the sub-networks HN,⋯ ,H1\mathcal{H}_{N},\cdots,\mathcal{H}_{1} in coarser scales sN,⋯ ,s1s_{N},\cdots,s_{1} are capable of producing the correct multi-scale voxel flows FN,⋯ ,F1\mathbf{F}_{N},\cdots,\mathbf{F}_{1} even for large motions.

We fuse these multi-scale voxel flows to the finest network H0\mathcal{H}_{0} to get our final result. The fusion is conducted by upsampling and concatenating the multi-scale voxel flow Fkx,y\mathbf{F}^{x,y}_{k} (only the spatial components (Δx,Δy)(\Delta x,\Delta y) are retained) to the final decoder layer of H0\mathcal{H}_{0}, which has the desired spatial resolution s0s_{0}. Then, the fine-scale voxel flow F0\mathbf{F}_{0} is obtained by further convolution on the fused flow fields. The network architecture of multi-scale DVF is illustrated in Fig. 3. And it can be formulated as

Since each sub-network Hk\mathcal{H}_{k} is fully differentiable, the multi-scale DVF can also be trained end-to-end with reconstruction loss ∥Yk−T(Xk,Fk)∥1\|\mathbf{Y}_{k}-\mathcal{T}(\mathbf{X}_{k},\mathbf{F}_{k})\|_{1} for each scale sks_{k}.

4 Multi-step Prediction

Experiments

We trained Deep Voxel Flow (DVF) on videos from the UCF-101 training set . We sampled frame triplets with obvious motion, creating a training set of approximately 240,000240,000 triplets. Following and , both UCF-101 and THUMOS-15 test sets are used as benchmarks. The pixel values are normalized into the range of $$. We use both PSNR and SSIM (on motion regionsWe use the motion masks provided by .) to evaluate the image quality of video frame synthesis; higher values of PSNR and SSIM indicate better results. However, our goal is to synthesize pixels that look realistic and artifact-free to human viewers. It is well-known that existing numerical measures of visual quality are not good facsimiles of human perception, and temporal coherence cannot be evaluated in paper figures. We find that the visual difference in quality of our method and competing techniques is much more significant than the numerical difference, and we include a user study in Section 4.4 that supports this conclusion.

Competing Methods. We compare our approach against several methods, including the state-of-the-art optical flow technique EpicFlow . To synthesize the in-between images given the computed flow fields we apply the interpolation algorithm used in the Middlebury interpolation benchmark . For the CNN-based methods, we compare DVF to Beyond MSE , which achieved the best-performing results in video prediction. However, their method is trained using 44 input frames, whereas ours uses only 22. We therefore try both numbers of input frames. The comparisons are performed under two settings. First, we use their best-performing loss (ADV+GDL), and replace the backbone networks in Beyond MSE with ours and train using the same data and protocol as in DVF (i.e., 22 frames as input on UCF-101). Second, we adapt DVF to their setting (i.e., using 44 frames as input and adopting the same number of network parameters ) and directly benchmark against the numbers reported in .

Results. As shown in Table 1 (left), our method outperforms the baselines for video interpolation. Beyond MSE is a hallucination-based method and produces blurry predictions. EpicFlow outperforms Beyond MSE by 1.41.4dB because it copies pixel based on estimated flow fields. Our approach further improves the results by 1.61.6dB. Some qualitative comparisons are provided in Fig. 4 (a).

Video extrapolation results are shown in Table 1 (middle). The gap between Beyond MSE and EpicFlow shrinks to 0.70.7dB for video extrapolation since this task requires more semantic inference, which is a strength of deep models. Our approach combines the advantages of both, and achieves the best performance (32.732.7dB). Qualitative comparisons are provided in Fig. 4 (c).

Finally, we explore the possibility of multi-step prediction, i.e., interpolate/extrapolate three frames (step=1,2,3step=1,2,3) at a time instead of one. From Fig. 5, we can see that our approach consistently outperforms other alternatives along all time steps. The advantages become even larger when evaluating on long-range predictions (e.g., step=3step=3 in extrapolation). DVF is able to learn long-term temporal dependencies through large-scale unsupervised training. The qualitative illustrations are provided in Fig. 4 (b)(d).

In this section, we demonstrate the merits of Multi-scale Voxel Flow (Multi-scale VF); specifically, we examine results separately along two axes: appearance, and motion. For appearance modeling, we identify the texture regions by local edge magnitude. For motion modeling, we identify large motion regions according to the flow maps provided by . Fig. 6 compares the PSNR performance on UCF-101 test set without and with multi-scale voxel flow. The multi-scale architecture further enables DVF to deal with large motions, as shown in Fig. 6 (b). Large motions become small after downsampling, and these motion estimates are mixed with higher-resolution estimates at the final layer of our network. The plots show that the multi-scale architecture add the most benefit in large-motion regions.

We also validate the effectiveness of skip connections. Intuitively, concatenating feature maps from lower layers, which have larger spatial resolution, helps the network recover more spatial details in its output. To confirm this claim, we conducted an additional ablation study, showing that removing skip connections reduced the PSNR performance by 1.11.1dB.

2 Generalization to View Synthesis

Here we demonstrate that DVF can be readily generalized to view synthesis even without re-training. We directly apply the model trained on UCF-101 to the view synthesis task, with the caveat that we only produce half-way in-between views, whereas general view synthesis techniques can render arbitrary viewpoints. The KITTI odometry dataset is used here for evaluation, following .

Table 1 (right) lists the performance comparisons of different methods. Surprisingly, without fine-tuning, our approach already outperforms and by 0.1640.164 and 0.1350.135 respectively. We find that fine-tuning on the KIITI training set could further reduce the reconstruction error. Note that KITTI dataset exhibits large camera motion, which is much different from our original training data. (UCF-101 mainly focuses on human actions.) This observation implies that voxel flow has good generalization ability and can be used as a universal frame/view synthesizer. The qualitative comparisons are provided in Fig. 7.

3 Frame Synthesis as Self-Supervision

In addition to making progress on the quality of video interpolation/extrapolation, we demonstrate that video frame synthesis can serve as a self-supervision task for representation learning. Here, the internal representation learned by DVF is applied to unsupervised flow estimation and pre-training of action recognition.

As Unsupervised Flow Estimation. Recall that 3D voxel flow can be projected into a 2D motion field, which is illustrated in Fig. 2 (e). We quantitatively evaluate the flow estimation of DVF by comparing the projected 2D motion field to the ground truth optical flow field. The KITTI flow 2012 dataset is used as a test set. Table 2 (left) reports the average endpoint error (EPE) over all the labeled pixels. After fine-tuning, the unsupervised flow generated by DVF surpasses traditional methods and performs comparably to some of the supervised deep models . Learning to synthesize frames on a large-scale video corpus can encode essential motion information into our model.

As Unsupervised Representation Learning. Here we replace the reconstruction layers in DVF with classification layers (i.e., fully-connected layer + softmax loss). The model is fine-tuned and tested with an action recognition loss on the UCF-101 dataset (split-1) . This is equivalent to using frame synthesis by voxel flow as a pre-training task. As demonstrated in Table 2 (right), our approach outperforms random initialization by a large margin and also shows superior performance to other representation learning alternatives . To synthesize frames using voxel flow, DVF has to encode both appearance and motion information, which implicitly mimics a two-stream CNN .

4 Applications

DVF can be used to produce slow-motion effects on high-definition (HD) videos. We collect HD videos (1080×7201080\times 720, 3030fps) from the web with various content and motion types as our real-world benchmark. We drop every other frame to act as ground truth. Note that the model used here is trained on the UCF-101 dataset without any further adaptation. Since the DVF is fully-convolutional, it can be applied to videos of an arbitrary size. More video quality comparisons are available on our project pagehttps://liuziwei7.github.io/projects/VoxelFlow.

Visual Comparisons. Existing video slo-mo software relies on explicit optical flow estimation to generate in-between frames. Thus, we choose EpicFlow to serve as a strong baseline. Fig. 8 illustrates slo-mo effects on the “Throw” and “Street” sequences, respectively. Both techniques tend to produce spatially coherent results, though our method performs even better. For example, in the “Throw” sequence, DVF maintains the structure of the logo, while in the “Street” sequence, DVF can better handle the occlusion between the pedestrian and the advertisement. However, the advantage is much more obvious when the temporal axis is examined. We show this advantage in static form by showing xtxt slices of the interpolated videos (Fig. 8 (c)); the EpicFlow results are much more jagged across time. Our observation is that EpicFlow often produces zero-length flow vectors for confusing motions, leading to spatial coherence but temporal discontinuities. Deep learning is, in general, able to produce more temporally smooth results than linearly scaling optical flow vectors.

User Study. We conducted a user study on the final slo-mo video sequences to objectively compare the quality of different methods. We compare DVF against both EpicFlow and ground truth. For side-by-side comparisons, synthesized videos of the two different methods are stitched together using a diagonal split, as illustrated in Fig. 9 (a). The left/right positions are randomly placed. Twenty subjects were enrolled in this user study; they had no previous experience with computer vision. We asked participants to select their preferences on 1010 stitched video sequences, i.e., to determine whether the left-side or right-side videos were more visually pleasant. As Fig. 9 (b) shows, our approach is significantly preferred to EpicFlow among all testing sequences. For the null hypothesis: “there is no difference between EpicFlow results and our results”, the p-value is p<0.00001p<0.00001, and the hypothesis can be safely rejected. Moreover, for half of the sequences participants choose the result of our method roughly equally as often as the ground truth, which suggests that they are of equal visual quality. For the null hypothesis: “there is no difference between our results and ground truth”, the p-value is 0.8381930.838193; statistical significance is not reached to safely reject the null hypothesis in this case. Overall, we conclude that DVF is capable of generating high-quality slo-mo effects across a wide range of videos.

Failure Cases. The most typical failure mode of DVF is in scenes with repetitive patterns (e.g., the “Park” sequence). In these cases, it is ambiguous to determine the true source voxel to copy by just referring to RGB differences. Stronger regularization terms can be added to address this problem.

Discussion

In this paper, we propose an end-to-end deep network, Deep Voxel Flow (DVF), for video frame synthesis. Our method is able to copy pixels from existing video frames, rather than hallucinate them from scratch. On the other hand, our method can be trained in an unsupervised manner using any video. Our experiments show that this approach improves upon both optical flow and recent CNN techniques for interpolating and extrapolating video. In the future, it may useful to combine flow layers with pure synthesis layers to better predict pixels that cannot be copied from other video frames. Also, the way we extend our method to multi-frame prediction is fairly simple; there are a number of interesting alternatives, such as using the desired temporal step (e.g., t=.25t=.25 for the first out of three interpolated frames) as an input to the network. Compressing our network so that it may be run on a mobile device is also a direction we hope to explore.

References