Unsupervised convolutional neural networks for motion estimation
Aria Ahmadi, Ioannis Patras
Introduction
Motion fields, that is fields that describe how pixels move from a reference to a target frame, are rich source of information for the analysis of image sequences and beneficial for several applications such as video coding , medical image processing , segmentation and human action recognition . Traditionally, motion fields are estimated using the variational model proposed by Horn and Schunck and its variants such as . Very recently, inspired by the great success of Deep Neural Networks in several Computer Vision problems , a CNN has been proposed Fischer et al. in for motion estimation. The method showed performance that was close to the state-of-the-art in a number of synthetically generated image sequences.
A major problem with the method proposed in is that the proposed CNN needed to be trained in a supervised manner, that is, it required for training synthetic image sequences where ground truth motion fields were available. Furthermore, in order to generalize well to an unseen dataset, it needed fine tuning, also requiring ground truth data on samples from that dataset. Ground truth motion estimation are not easily available though. For this reason, the method proposed in was applied only on synthetic image sequences.
In this paper, we propose training a CNN for motion estimation in an unsupervised manner. We do so by designing a cost function that is differentiable with respect to the unknown motion field and, therefore, allows the backpropagation of the error and the end to end training of the CNN. The cost function builds on the widely used optical flow constraint - our major difference to Horn-Schunk based methods is that the cost function is used only during training and without regularization. Once trained, given a pair of frames as input the CNN gives at its output layer an estimate of the motion field. In order to deal with motions large in magnitude, we embed the proposed network in a classical iterative scheme, in which at the end of each iteration the reference image is warped towards the target image and in a classical coarse-to-fine multiscale framework. We train our CNN using randomly chosen pairs of consecutive frames from UCF101 dataset and test it on both the UCF101 where it performs similarly to the state-of-the-art methods and on the synthetic MPI-Sintel dataset where it outperforms them.
The remainder of the paper is organized as follows. In Section 2 we describe our method. In Section 3 we present our experimental results. Finally, in Section 4 we give some conclusion.
Method
At the heart of all motion estimation methods is the minimization of the difference between features extracted at a certain location , at the reference frame and its correspondence in the frame . The now classical Horn and Schunck method penalizes the deviation from the assumption of constant intensity, that states that the intensity at a pixel in the reference frame at time and the intensity at its correspondence at time are the same. Formally, the goal is the minimization of the motion compensated intensity differences, that is,
where is the intensity at pixel at frame , and F(x,y)\triangleq\left[\begin{array}[]{c}u(x,y)\\ v(x,y)\end{array}\right] is the unknown motion vector at pixel . Clearly, and are respectively the horizontal and vertical displacement of the pixel with coordinates . To arrive at a computationally tractable method, Horn and Schunck introduced a regularization term that penalized discontinuities in the motion field and linearized the cost by taking the first order Taylor expansion with respect to the horizontal and vertical displacements. By doing the latter, they arrived at the optical flow equation , where , and are the horizontal, vertical and temporal intensity derivatives respectively, and penalized deviations from it. In the equation above, we omit the pixel coordinates for notation simplicity.
In this work, we build on the constant intensity assumption. However, instead of using it at test time to estimate the motion field as the one that minimizes the deviations from it, we use it at training time only in order to train a CNN that takes as input a pair of images and outputs at its last layer a dense motion field between them. Clearly, the motion field is a function of the weights of the CNN and the images at its input. In order to reduce the influence of outliers we use a robust error norm, that is the Charbonnier penalty , a differentiable variant of the norm, the most robust convex function. Formally, during training we learn the CNN by optimizing the sum of costs that for a pair of images and are defined as follows:
where the image coordinates , are omitted for notation simplicity.
Clearly, the motion field is a function of the weights of the CNN and the images at its input. Therefore, the cost function in Eq. 2 is a function of the CNN weights. More importantly, our cost function allows us to calculate the derivatives of it with respect to the network weights. Specifically, using the chain rule,
The second part, that is , are the partial derivatives of the output of the CNN with respect to its weights . This can be calculated in a classical manner using the standard form of the backpropagation algorithm. The first term, that is , can be calculated in closed form as
The cost function that is used for training relies on the optical flow constraint. This is known that it does not hold when the motions are large in magnitude. Following the dominant paradigm in the field, we embed our method in a coarse-to-fine multiscale iterative scheme.
At test time, that is once trained, given a pair of frames as input the CNN gives as output an update on the motion field. The updated motion field is then used to warp the second frame towards the first one, and the new pair of images are given as input to the CNN to calculate another update on the motion field. Several iterations are performed at each scale of the muti-scale framework. After each update, and similarly to other methods in the literature , the calculated motion field in each iteration is filtered (by a median filter in our case). The proposed algorithm at test time is summarized in Algorithm 1.
We propose a fully convolutional neural network with 12 convolutional layers. The architecture could be imagined as two parts. The CNN makes a compact representation of motion information in the first part which involves 4 downsamplings. This compact representation is then used to reconstruct the motion field in the second part which involves 4 up-samplings. The up-sampled is performed by simply repeating the rows and columns of the feature maps. Since our proposed DNN is fully convolutional, the input could be of any size. Fig. 3 shows the two parts of the proposed CNN. To update the CNN weights during the training phase, we used ADAM and calculate that spatiotemporal intensity derivatives , and as proposed in .
Experiments
In order to evaluate the performance of our method, we report its results on 2 datasets, namely the real UCF101 and the synthetically generated MPI-Sintel, and compare them with the results of other state-of-the-art methods. We train the proposed Unsupervised trained CNN (USCNN) on 20K pairs of consecutive grayscale frames randomly selected from the about 1 million frames of the UCF101 dataset . We first test the method on the same dataset, using for the evaluation 10K pairs of frames that were randomly selected from the UCF101 dataset making sure that there is no overlap with the training set. Since there is no ground truth motion field for UCF101 dataset, we use as ground truth the output of the EpicFlow - to the best of our knowledge this is the state-of-the-art method for motion estimation. Table 1 reports the results of the proposed method and three other state-of-the-art methods in the field, with the notable exception of FlowNet for which, neither the training network nor the training dataset are available. As it can be seen, the proposed method has comparable performance for motions less than 5 pixels - for larger motions as it is largely expected from methods that rely on coarse-to-fine schemes that involve downsampling it has lower accuracy.
To evaluate how well our method can generalize to an unknown dataset, we have applied it on the MPI-Sintel Final which is one of the most realistic synthetic datasets for which ground truth is available. As reported in Table 2, the USCNN has a better performance than LDOF and has a comparable performance in comparison with the other state-of-the-art methods. In Table 2, the performance measures for EpicFlow, DeepFlow, LDOF, and FlowNet are from and the performance measure for HAOF is calculated using their publicly available code.
To have a fair comparison, we report the performance measure for the two architectures introduced by without finetuning. ’FlowNetS’ has a rather generic architecture which receives as input the two images stacked together. In ’FlowNetC’, first meaningful representations are made in two separate but identical streams each receiving one of the input images and then their combination is fed into another stream for motion estimation. Our network architecture is closest to ’FlowNetS’. The motion field of several samples from UCF101 and MPI-Sintel datasets, estimated by various methods are depicted in Fig. 1 and Fig. 2 respectively.
Conclusions
In this work, we propose estimating dense motion fields with CNNs. We show that, surprisingly perhaps, a simple cost function that relies on the optical flow equation can be used successfully for training a deep convolutional network in a completely unsupervised manner and without the need of any regularization or other constraints. Our CNN is trained on the UCF101 dataset. We show that it has a performance that is comparable to other state-of-the-art methods, especially for motions that are not large in magnitude, and that it can generalize very well to an unknown dataset, MPI-Sintel, without the need for refinement. The proposed method in this paper is among the very few studies conducted on the application of DNNs for motion estimation.
Acknowledgment
We gratefully acknowledge the support of NVIDIA Corporation with the donation of the Tesla K40 GPU used for this research.