Video Prediction Recalling Long-term Motion Context via Memory Alignment Learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, Yong Man Ro
Introduction
Video prediction in computer vision is to estimate upcoming future frames at pixel-level from given previous frames. Since predicting the future is an important basement for intelligent decision-making systems, the video prediction has attracted increasing attention in industry and research fields. It has the potential to be applied to various tasks such as weather forecasting , traffic situation prediction , and autonomous driving . However, the pixel-level video prediction is still challenging mainly due to the difficulties of capturing high-dimensionality and long-term motion dynamics .
Recently, several studies with deep neural networks (DNNs) have been proposed to capture the high-dimensionality and the long-term dynamics of video data in the video prediction field . The models considering the high-dimensionality of videos tried to simplify the problem by constraining motion and disentangling components . However, these methods did not consider the long-term frame dynamics, which leads to predicting blurry frames or wrong motion trajectories. Recurrent neural networks (RNNs) have been developed to capture the long-term dynamics with consideration for long-term dependencies in the video prediction . The long-term dependencies in the RNNs is about remembering past step inputs. The RNN-based methods exploited the memory cell states in the RNN unit. The cell states are recurrently changed according to the current input sequence to remember the previous steps of the sequence. However, it is difficult to capture the long-term motion dynamics for the input sequence with limited dynamics (i.e., short-term motion) because such cell states mainly depend on revealing relations within the current input sequence. For example, given short-length input frames for a walking motion, the leg movement from the input is limited itself. Therefore, it is difficult to grasp what will happen to the leg in the future through the cell states of the RNNs. In this case, the long-term motion context of the partial action may not be properly captured by the RNN-based methods.
Our work addresses long-term motion context issues for predicting future frames, which have not been properly dealt with in previous video prediction works. To predict the future precisely, it is required to capture which long-term motion context the input motion belongs to. For example, in order to predict the future of leg movement, we need to know such partial leg movement belongs to either walking or running (i.e., long-term motion context). The bottlenecks arising when dealing with long-term motion context are as follows: (i) how to predict the long-term motion context naturally matching input sequences with limited dynamics, (ii) how to predict the long-term motion context with high-dimensionality.
In this paper, we propose novel motion context-aware video prediction to address the aforementioned issues. To solve the bottleneck (i), we introduce a long-term motion context memory (LMC-Memory) with memory alignment learning. Contrary to the internal memory cells of the RNNs, the LMC-Memory externally exists with its own parameters to preserve various long-term motion contexts of training data, which are not limited to the current input. Memory alignment learning is proposed to effectively store long-term motion contexts into the LMC-Memory and recall them even with inputs having limited dynamics. Memory alignment learning contains two training phases to align long-term and short-term motions: storing long-term motion context from long-term sequences into the memory, matching input short-term sequences with the stored long-term motion contexts in the memory. As a result, the long-term motion context (e.g., long-term walking dynamics) can be recalled from the input short-term sequence alone (e.g., short-term walking clip).
Furthermore, to resolve the bottleneck (ii), we propose decomposition of a memory query that is used to store and recall the motion context. Even if various motion contexts of training data are stored in the LMC-Memory, it is difficult to capture the motion context that is exactly matched with the input. This is because motions of video sequences have high-dimensionality (e.g., complex motion with local motion components). The dimensionality indicates the number of pixels in a video sequence. Since each motion is slightly different from one another in a global manner even for the same category, the proposed memory query decomposition is useful in that it enables to store local context (i.e., low-dimensional dynamics) and recall the suitable local context for each local part of the input individually. It can boost the alignment effects between the input and the stored long-term motion context in the LMC-Memory.
The major contributions of the paper are as follows.
We introduce novel motion context-aware video prediction to solve the inherent problem of the RNN-based methods in capturing long-term motion context. We address the arising long-term motion context issues in the video prediction.
We propose the LMC-Memory with memory alignment learning to address storing and recalling long-term motion contexts. Through the learning, it is possible to recall long-term motion context corresponding to an input sequence even with limited dynamics.
To address the high-dimensionality of motions, we decompose memory query to separate an overall motion into local motions with low-dimensional dynamics. It makes it possible to recall suitable local motion context for each local part of the input individually.
Related Work
In video prediction, errors for predicting future frames can be divided into two factors . The first one is about systematic errors due to the lack of modeling capacity for deterministic changes. The second one is related to modeling the intrinsic uncertainty of the future. There have been several works to address the second factor . These methods utilized stochastic modeling to generate plausible multiple futures. Contrary to these, our paper addresses the video prediction focusing on the first factor.
Recently, deep learning-based video prediction methods have been proposed to deal with the first factor. They considered the problems leading to prediction difficulty such as capturing high-dimensionality and long-term dynamics in video data . Finn et al. incorporated appearance information in the previous frames with the predicted pixel motion information for long-range video prediction. Villegas et al. introduced a hierarchical prediction model that generates the future image from the predicted high-level structure. A predictive recurrent neural network (PredRNN) model was presented by Wang et al. to model and memorize both spatial and temporal representations simultaneously. Wang et al. further extended this model, named PredRNN++ to solve the vanishing gradient problem in deep-in-time prediction by building adaptive learning between long-term and short-term frame relation. Recently, eidetic 3D LSTM (E3D-LSTM) was proposed to integrate 3D convolutions into the RNNs for effectively addressing memories across long-term periods. Jin et al. introduced spatial-temporal multi-frequency analysis for high-fidelity video prediction with temporal-consistency . Su et al. proposed convolutional tensor-train decomposition to learn long-term spatio-temporal correlations. However, these works still have a limitation in encoding long-term dynamics in that they mainly rely on the input sequence to find frame relations. Therefore, it is difficult to capture the long-term motion context for predicting the future from the input sequence with limited dynamics.
2 Memory Network
Memory augmented networks have recently been introduced for solving various problems in computer vision tasks . Such computer vision tasks include anomaly detection , few-shot learning , image generation , and video summarization . Kaiser et al. presented a large scale long-term memory module for life-long learning. Memory-attended recurrent network was proposed by Pei et al. to capture the full-spectrum correspondence between the word and its visual contexts across video sequences in training data. To utilize the external memory network for our purposes, we introduce novel memory alignment learning that enables to store the long-term motion contexts into the memory and to recall them with limited input sequences. In addition, we separate overall motion into low-dimensional dynamics and utilize them as an individual memory query to recall proper long-term motion context for each local part of inputs.
Proposed Method
First, in the lower path of Figure 1, the differences between the consecutive frames (i.e., difference frames) are used as inputs of motion matching encoder . Then, a motion matching feature is extracted to recall the motion context memory feature from the external memory, named LMC-Memory. This LMC-Memory contains various long-term motion contexts of training data. Thus, from the memory can be considered as long-term information corresponding to the input sequence (described in Section 3.2 in detail). It is then embedded in the upper part . This long-term motion context embedding contains 2D-DeConvs to match the spatial size with the upper part, which results in .
The upper part of Figure 1 demonstrates long-term motion context-aware video prediction scheme. In this path, the required motion context is refined through attention-based encoding to effectively embed it in predicting future frames. Each frame of the input sequence is independently fed to spatial encoder with 2D-Convs to extract appearance characteristics. The ConvLSTM receives each extracted spatial feature as inputs in time step order. A cell state and an output state are obtained from recurrent processing of the ConvLSTM. Since contains the information from the past to the present of the input sequence, we use to refine for embedding the required motion context at the current step. and are concatenated and pass through fully connected layers to make channel-wise attention for . The channel-wise refined feature and output state from the ConvLSTM are concatenated to embed long-term context to the ConvLSTM output (i.e., spatio-temporal information of the input). The concatenated feature is fed to a frame decoder with 2D-DeConvs to generate corresponding next frame . The embedded motion context memory feature can provide the prior of long-term motion context for the current input sequence. Note that the generated next frame enters as a new input to create the further future frame.
2 LMC-Memory with Alignment Learning
The long-term motion context memory, named LMC-Memory is to provide the long-term motion context for current input sequences to predict future frames. To effectively recall the long-term motion context even for the input sequence with limited dynamics, we propose novel memory alignment learning. Figure 2 shows the training scheme of the LMC-Memory. The memory is trained alternately with two phases: storing long-term motion context into the memory and matching a limited sequence with the corresponding long-term context in the memory.
Optimization is performed with the prediction framework (see Figure 2). Only the short-term sequence is fed as a main input of . The memory path receives the long-term and short-term alternately. Two training phases are alternately performed in each iteration. In both phases, according to , we exploit a prediction loss function as follows
where denotes predicted future frames while denotes ground truth future frames. Note that the proposed method only takes short-term sequences at inference time as shown in Figure 1. Training procedure is further described in Algorithm 1.
Experiments
To validate the proposed method, we utilize both synthetic and natural video datasets. We use a synthetic Moving-MNIST dataset that is mainly used in video prediction. In addition, we use a KTH Action and a Human 3.6M datasets including natural videos with human action scenarios.
Moving-MNIST. The Moving-MNIST contains the moving of two randomly sampled digits from the original MNIST dataset. Each digit moves in a random direction within a 6464 size image with a gray scale. The constructed Moving-MNIST dataset consists of 10,000 sequences for training and 5,000 sequences for testing as .
KTH Action. KTH Action dataset consists of 6 types of action videos for 25 subjects. It includes indoor, outdoor, scale variations, and different clothes. Each frame is resized to 128128 with a gray scale. The videos of 1-16 subjects are used as the training set while the videos of 17-25 subjects are used as the test set. We follow the experimental setting of video prediction for the KTH Action dataset.
Human 3.6M. The Human 3.6M includes 17 human action scenarios with total 11 actors. It contains 4 different camera views. Each frame is resized to 6464 with RGB color channels. Videos of subjects 1, 5, 6, 7, and 8 are used to train the model while videos of subjects 9 and 11 are used to test the model. We follow the experimental setting .
2 Implementation
The video frames are normalized to intensity of and resized to 64 64 (MNIST and Human 3.6M) or 128 128 (KTH) as . The proposed model is trained by Adam optimizer with a learning rate of 0.0002. Memory slot size is fixed as 100 for all experiments. Input short-term sequence length is set as 10. Long-term sequence length is set as 30 (MNIST) or 40 (KTH and Human 3.6M). Our model is trained to predict corresponding future frames. We use 4-layer ConvLSTMs for frame prediction. The overall detailed network structures are described in the supplementary material.
3 Evaluation
We use MSE, PSNR, SSIM , and LPIPS to measure the performances. MSE and PSNR are calculated by the pixel-wise difference between the actual frame and the predicted frame. We also evaluate the performance using SSIM that considers the structural similarity between frames. Furthermore, we utilize LPIPS as a perceptual metric, which tends to be similar to the human recognition system . Higher values are better for PSNR and SSIM while lower values are better for MSE and LPIPS. LPIPS results are represented in scale. Single TITAN XP is used to evaluate computational costs for all models. Note that official source codes are used for other methods.
Results on Moving-MNIST. Table 1 shows the performance comparisons with the state-of-the-art methods on the Moving-MNIST. The left part of the table shows the experimental results of 10 frames prediction with the input 10 frames. The right part of the table shows the experimental results for 30 frames prediction with 10 input frames. The proposed method outperforms the other state-of-the-art methods. In particular, our method far surpasses the others in predicting 30 frames in terms of the LPIPS metric. In addition, the proposed method shows much better results on the computational cost compared to the other methods. Compared to other complex RNN-based methods, we adopt simple ConvLSTMs. Further, memory feature is extracted only once at the beginning, which is advantageous in computational cost. Figure 3 shows examples of frames predicted by the proposed method and other video prediction methods. As shown in the figure, the predicted frames of the proposed method show convincingly similar results to the ground truth. However, the prediction results by other methods show that they lose the trajectories or the shape of digits, especially in long-term condition.
Results on KTH Action. Table 2 shows the quantitative results of the proposed method and other state-of-the-art methods on the KTH Action dataset. The left part of the table shows the experimental results for predicting 20 frames with 10 input frames. The right part of the table indicates performances for predicting the next 40 frames. As shown in the table, the proposed method mostly surpasses the other state-of-the-art methods in predicting 40 frames. Especially, it is significant in the human perceptual metric (i.e., LPIPS). In addition, the proposed method shows a much faster inference speed compared to the other methods also on the KTH. Figure 4 shows qualitative long-term prediction results for input sequence with limited dynamics on the KTH. This input motion is limited because motion actually starts from the middle. As shown in the figure, the other methods fail to capture the detailed leg movement, especially in long-term condition. Whereas, our predicted frames are very similar to the ground truth frames. The proposed method maintains a clear shape of the legs while following the long-term trajectories even in such a challenging condition (i.e., limited dynamics).
Results on Human 3.6M. Table 3 shows the performance comparisons with the other methods on the Human 3.6M. The left part of the table shows the experimental results for predicting 40 frames with given 10 input frames. The right part of the table indicates performances for the last 10 frames among the future 40 frames. The proposed method outperforms other state-of-the-art video prediction methods both in predicting 40 future frames and predicting the last 10 frames. Figure 5 shows qualitative results for the long-term prediction on the Human 3.6M. The proposed method captures direction changing in the long-term while the other methods show the disappearance of a person at the corner. Compared to the others, the proposed method properly captures the long-term motion context with redirection.
4 Ablation Study
We analyze the effects of network designs by ablating them as shown in Table 4. In detail, we investigate the effectiveness of the LMC-Memory (i.e., memory alignment learning) and the local motion context (i.e., memory query decomposition). The baseline, ‘Model w/o LMC-Memory’ consists of the spatial encoder, the ConvLSTMs, and the frame decoder. The second one, ‘Model w/ LMC-Memory (Non-local motion context)’ contains LMC-Memory but it does not adopt memory query decomposition to use local motion context as memory queries. This model uses a globally pooled motion context feature as a query. The last one indicates our final proposed model of this paper. As shown in the table, each component contributes to the performances in predicting the last 10 among 40 frames. The final model outperforms the other models, especially in terms of perceptual metric LPIPS. These results show that locally manipulated query boosts the effects of the memory since it is more accessible to store and recall the motion context with low-dimensional dynamics. Further, the additional computational cost to use the LMC-Memory is marginal.
Figure 6 shows qualitative results for different network designs. The first model does not properly capture the long-term motion. The second one predicts long-term motion to some extent. However, the detailed local parts are distorted because it addresses the motion context in an only global manner. The final model effectively predicts future frames by properly capturing the context of long-term motion.
5 Memory Addressing
We analyze the memory addressing for different sequences. Figure 7 shows the cosine similarity values between addressing vectors from long-term and short-term sequences including different subjects. The areas addressed in the memory are more comparable (similarity between addressing vectors is high) in the case of the same action scenario than in the case of the different actions. It shows that the long-term and short-term features that belongs to the similar action are convincingly aligned in the memory.
Conclusion
The objective of the proposed work is to predict future frames being aware of the long-term motion context. To this end, we propose the LMC-Memory with the alignment learning scheme to effectively store abundant long-term contexts of training data and recall suitable motion context even from limited inputs. In addition, we utilize memory query decomposition to separate overall motion into low-dimensional dynamics. It enables to cope with the high-dimensionality in terms of utilizing motion contexts in the memory. As a result, the proposed method outperforms the state-of-the-art methods with sophisticated RNNs. In particular, it is significantly noticeable in long-term condition. Further, the effectiveness of the proposed method is analyzed in both quantitative and qualitative ways. Acknowledgement. This work was partly supported by the IITP grant (No. 2020-0-00004) and BK 21 Plus project.