VideoINR: Learning Video Implicit Neural Representation for Continuous Space-Time Super-Resolution

Zeyuan Chen, Yinbo Chen, Jingwen Liu, Xingqian Xu, Vidit Goel, Zhangyang Wang, Humphrey Shi, Xiaolong Wang

Introduction

We observe the visual world in the form of streaming and continuous data. However, when we record such data with a video camera in a computer, it is often stored with limited spatial resolutions and temporal frame rates. Because of the high cost on recording and storing large time-scales of video data, oftentimes our computer vision system will need to process low-resolution and low frame rate videos. This introduces challenges in recognition systems such as video object detection , and we are still struggling at learning to recognize motion and actions from discrete frames . When presenting the video back to humans (e.g., on a TV), it is essential to visualize it in high resolution and high frame rate for user experience. How to recover the low resolution video back to high resolution in space and time becomes an important problem and the first step for many downstream applications.

Space-Time Video Super-Resolution (STVSR) approaches are developed to increase the spatial resolution and frame rate at the same time given a low-resolution and low frame rate video as the input. Instead of performing super-resolution in space and time separately in two stages, researchers recently propose to simultaneously perform super-resolution in one stage . Intuitively, the aggregated information in time from multiple frames can reveal missing details for each frame when spatial scaling is applied, and the temporal interpolation can be more smooth and accurate given higher and richer spatial representation. The one-stage end-to-end training has shown to unify the benefits from both sides. While these results are encouraging, most approaches can only perform super-resolution to a fixed space and time scale ratio.

In this paper, instead of super-resolution in a fixed scale, we propose to learn a continuous video representation, which allows to sample and interpolate the video frames in arbitrary frame rate and spatial resolution at the same time. Our key idea is to learn an implicit neural representation, which is a neural function that takes a space-time coordinate as input, and outputs the corresponding RGB value. Since we can sample the coordinate continuously, the video can be decoded in any spatial resolution and frame rate. Our work is inspired by recent progress on implicit functions for 3D shape representations and image representations with Local Implicit Image Functions (LIIF) using a ConvNet . Different from images, where interpolation in space can be based on the gradients between pixels, pixel gradients across frames with low frame rates are hard to compute. The network will need to understand the motion of the pixels and objects to perform interpolation, which could be hard to model by 2D or 3D convolutions alone.

We propose a novel Video Implicit Neural Representation (VideoINR) as a continuous video representation. In the STVSR task, two low-resolution image frames are concatenated and forwarded to an encoder which generates a feature map with spatial dimensions. VideoINR then serves as a continuous video representation over the generated feature map. It first defines a spatial implicit neural representation for a continuous spatial feature domain, from which a high-resolution image feature is sampled according to all query coordinates. Instead of using convolutional operations to perform temporal interpolation, we learn a temporal implicit neural representation to first output a motion flow field given the high-resolution feature and the sampling time as inputs. This flow field will be applied back to warp the high-resolution feature which will be decoded to the target video frame. Since all the operations are differentiable, we can learn the motion in feature level end-to-end without any extra supervision besides the reconstruction error. To summarize, given the input frames, an encoder generates a feature map, which can be then decoded by VideoINR to arbitrary spatial resolution and frame rate.

In our experiments, we demonstrate that VideoINR can not only represent video in arbitrary space and time resolutions on the scales within the training distributions, but also extrapolate to out-of-distribution frame rates and spatial resolutions. Given the learned continuous function, instead of decoding the whole video each time, it allows the flexibility to decode only a certain region and time scale when needed. We conduct experiments with Vid4 , GoPro and Adobe240 datasets. We demonstrate that VideoINR achieves competitive performances with state-of-the-art STVSR methods on in-distribution spatial and temporal scales and significantly outperforms other methods on out-of-distribution scales.

We highlight our main contributions as follows:

We propose a novel Video Implicit Neural Representation as a continuous video representation.

The proposed approach allows for representing videos in arbitrary space and time resolution efficiently.

VideoINR achieves out-of-distribution generalization in space-time video super-resolution and outperforms baselines by a large margin.

Related Work

Implicit Neural Representation. Implicit neural representations have been demonstrated as compact yet powerful continuous representations for various tasks, including 3D reconstruction and generation . These representations typically represent signals as a neural function that maps coordinates to signed distance , occupancy , or density and RGB values in a neural radiance field (NeRF ). Recent works also show promising results of applying this idea for modeling 2D images . Our continuous video representation is inspired by this rapidly growing field and has specific designs for videos, where a learnable flow can exploit the correspondences in video frames with inductive bias.

Video Frame Interpolation. Video Frame Interpolation (VFI) aims to synthesize unseen frames between the input video frames. Meyer et al. proposed a phase-based method where information across levels of a multi-scale pyramid is combined for the synthesis of interpolated frames. Niklaus et al. introduced a series of kernel-based VFI algorithms in which they took pixel synthesis for the target frame as local convolution over input frames. Optical flow based VFI methods utilized optical flow prediction networks (e.g. PWC-Net ) to compute bidirectional flows between input frames, which served as the guidance for new frame synthesis. Additional information including occlusion masks , depth maps , and cycle consistency were also incorporated in the models for better performances.

Video Super-Resolution. Video Super-Resolution (VSR) aims at increasing the spatial resolutions of low-resolution videos. Earlier approaches were typically built on the sliding-window framework, where they predicted optical flows between input frames and performed spatial warping for explicit feature alignment. Later on, implicit alignment started a new trend in this task . For instance, TDAN adopts deformable convolutions (DCNs) to align different input frames at feature levels. EDVR further extends DCNs to a multi-scale fashion for more accurate alignment. Kelvin et al. introduced BasicVSR , in which they analyzed basic components for VSR models and suggested a bidirectional propagation scheme to maximize the gathered information from input.

Space-time Video Super-Resolution The target of Space-Time Video Super-Resolution (STVSR) is to simultaneously increase the spatial and temporal resolutions of the given low-resolution low frame rate videos. Shechtman et al. tackled this problem by combining information from multiple input video sequences and applying a directional space-time regularization. Mudenagudi et al. proposed a unified framework for STVSR in which videos are modeled as Markov random fields, and the maximum a posteriori estimates are taken as final solutions. Shahar et al. introduced an effective space-time patch recurrence prior for STVSR. Recently, with the advances in deep learning, researchers started to employ powerful convolutional neural networks to address the task . Xiang et al. proposed a unified neural network for synthesizing the feature of the missing frame and used a deformable ConvLSTM to align and aggregate extracted temporal information for reconstruction. STARNet leveraged mutually informative relationships between time and space with the assistance of additional optical flow inputs. TMNet proposed a temporal modulation block to modulate deformable convolution kernels for supporting frame interpolation at arbitrary time instances. All these STVSR methods are designed to perform super-resolution on a specific up-sampling space scale defined before training, and some of them can only infer intermediate frames at pre-defined times. Therefore, the application scopes of these methods are limited. VideoINR serves as a continuous video representation that supports frame interpolation at arbitrary spatial resolution and frame rate. VideoINR is more flexible during the application and can be employed in more circumstances, such as non-uniform interpolation and video zoom-in in local regions.

Video Implicit Neural Representation

Given a video with limited spatial resolution and frame rate, our goal is to find a continuous representation for the video. The representation interprets arbitrary space-time coordinate (xs,xt)(x_{s},x_{t}) into RGB values. To this end, we introduce Video Implicit Neural Representation (VideoINR), which enables continuous space-time super-resolution. It is parameterized by multi-layer perceptrons (MLPs) and takes the form

where ff is the proposed video representation defined by the encoded feature and network parameters. xsx_{s} is the 2D spatial coordinate, xtx_{t} is the temporal coordinate, and ss is the predicted RGB value. For learning such implicit neural representation, we propose to decouple space and time and learn a continuous representation for each of them.

Figure 2 illustrates an overview of our model. Given a space-time coordinate (xs,xt)(x_{s},x_{t}) and the feature extracted from input frames by an encoder, the Spatial Implicit Neural Representation (SpatialINR) decodes the spatial coordinate xsx_{s} and output a corresponding feature vector (Sec. 3.1). The feature is then forwarded to the Temporal Implicit Neural Representation (TemporalINR) for the motion flow at the query coordinate (Sec. 3.2). The flow is applied back to warp the continuous feature defined by SpatialINR for a new feature vector (Sec. 3.3) which is finally decoded to the target RGB value (Sec. 3.4).

Inspired by LIIF , we learn a Spatial Implicit Neural Representation (SpatialINR) that defines a continuous 2D feature domain by the discrete encoded feature map. This continuous domain decodes arbitrary 2D spatial coordinate into a corresponding feature vector. Specifically, the feature vectors generated by the encoder are evenly distributed in the 2D space. We sample the feature vector (the dark blue cuboid in Fig 2) nearest to the queried spatial coordinate xsx_{s}, concatenate it with the relative position information between query coordinate and feature vector, and input them into the a function fsf_{s} to output the continuous feature at xsx_{s} (the green cuboid in Fig 2). This process could be expressed as

where Fs\mathcal{F}_{s} is the continuous feature domain defined by SpatialINR, z∗z^{*} is the feature vector nearest to the query coordinate xsx_{s} and v∗v^{*} is the spatial coordinate of the feature vector z∗z^{*}.

The main difference between LIIF and SpatialINR is that LIIF is proposed for continuous image representation, while SpatialINR defines a continuous feature domain, which is supposed to be further utilized for modeling temporal information in videos.

2 Continuous Temporal Representation

The proposed SpatialINR defines a new continuous feature domain in 2D space. Our next step is to learn the continuous Temporal Implicit Neural Representation (TemporalINR) and extend the feature domain from 2D space to 3D space and time, which can be achieved by decoding the temporal coordinate xtx_{t}. Directly generating the target decoded feature by a network can be fairly difficult, as the network has to learn not only the motion patterns between input frames but also the context information. Instead, we propose to learn a continuous motion flow field for the continuous temporal representation.

Particularly, given a space-time coordinate (xs,xt)(x_{s},x_{t}) and two consecutive input frames I0I_{0} and I1I_{1}, TemporalINR maps the coordinate to a motion flow

where M\mathcal{M} is the continuous motion flow field and ftf_{t} is the function for TemporalINR. Benefiting from the 2D continuous feature domain provided by SpatialINR, we could replace I0I_{0}, I1I_{1}, and xsx_{s} by the continuous feature at xsx_{s}. Thus the equation can be written as

where Fs(xs)\mathcal{F}_{s}(x_{s}) is the feature domain defined in Eq 2.

3 Space-Time Continuous Representation

With two continuous representations for space and time, we aim at combining them into a unified space-time continuous representation for videos. Starting from a space-time coordinate (xs,xt)(x_{s},x_{t}), we first use SpatialINR to predict the continuous feature at xsx_{s}. TemporalINR is then utilized for generating the motion flow of the query coordinate. Based on these outputs, we obtain the space-time feature by warping the continuous feature domain. The warped feature at xsx_{s} corresponds to the continuous feature at xs′x_{s}^{\prime}. The relationship between two coordinates can be written as

where M(xs,xt)\mathcal{M}(x_{s},x_{t}) is the motion flow vector at (xs,xt)(x_{s},x_{t}).

We query this new spatial coordinate in the continuous 2D feature domain and obtain a new feature vector (the light green cuboid in Fig 2), which is treated as the feature of our continuous space-time representation at coordinate (xs,xt)(x_{s},x_{t}). Accordingly, the continuous space-time feature Fst\mathcal{F}_{st} can be formulated as

In practice, we generate two independent flows for the motion flow field, and concatenate corresponding warped features. Intuitively, TemporalINR may implicitly learn bi-directional correspondences between the target frame and input frames, without explicit supervision.

4 Feature Decoding

Based on the continuous space-time representation, we can get the feature corresponding to any space-time coordinate. The final step is to decode the feature as an RGB value. A straightforward design is to take the obtained space-time feature for decoding directly. However, due to the MLP-based network architecture, the RGB value of every predicted pixel depends on a single feature vector, leading to a limited size of the network receptive field. To alleviate the negative impact of this disadvantage, we enrich the input information of the decoding network by aggregating features of different scales. In detail, we incorporate the encoded feature as well as two input frames for decoding. Since these additional features are typically of low-resolution compared with the target resolution, we sample feature vectors corresponding to the query coordinate by bilinear interpolation. All features are then combined together for predicting the RGB output.

5 Frame synthesis

From Section 3.1 to 3.4, we focus on predicting the RGB value at a specific coordinate. To synthesize an entire frame, we need to query coordinates of all pixels of it. Given these coordinates, we can convert the continuous feature from SpatialINR into a high-resolution feature map. We can also generate a whole motion flow field for the latent high-resolution interpolated frame. Therefore, we do not have to forward SpatialINR twice before and after warping as in the situation of one input coordinate. Instead, we directly warp the whole high-resolution feature map based on the motion flow and input the warped feature into the decoding network to synthesize the target frame at one time.

Experiments

Dataset. We use Adobe240 dataset as the training set, which includes 133 videos in 720P taken by hand-held cameras. We follow to split these videos into the train, validation, and test subsets with 100, 16, and 17 videos. All videos are converted into image sequences for training and testing. Each sequence contains approximately 3000 frames which are treated as high-resolution frames in training. The low-resolution counterparts are then generated by imresize function in Matlab with the default setting of bicubic interpolation. We use a sliding window to select frames from the image sequences for training. The length of the sliding window is set to 9. We take the 1st1^{st} and 9th9^{th} frames as network inputs. The 2nd2^{nd} to 7th7^{th} frames serve as ground-truth frames, and we randomly select three of them as the supervision of our network in every iteration. VideoINR is trained by two stages. In the first stage, we fixed the down-sampling space scale to ×4\times 4. In the second stage, we randomly sample scales in a uniform distribution U(1,4)\mathcal{U}(1,4). We provide more discussion about this two-stage training strategy in Section 4.3.

Datasets including Vid4 , Adobe240 , and GoPro are used for evaluation. On Vid4, we only conduct experiments on single frame interpolation of STVSR. For Adobe240 and GoPro, we evaluate on their test set. The image sequences extracted from videos in the datasets are split into groups of 9-frame video clips. We feed the 1st1^{st} and 9th9^{th} frames down-sampled by scale ×4\times 4 in each clip into models to generate 9 high-resolution frames from 1st1^{st} to 9th9^{th}. We separately evaluate the average metrics of the center frames (i.e. the 1st1^{st}, 4th4^{th}, 9th9^{th} frames) and all 9 output frames. They are denoted as -Center and -Average in Table 1.

Implementation details. We use Adam optimizer with β1=0.9\beta_{1}=0.9 and β2\beta_{2}=0.999. The learning rate is initialized as 1×10−41\times 10^{-4} and is decayed to 1×10−71\times 10^{-7} with a cosine annealing for every 150,000 iterations. The model is trained in a total of 600,000 iterations with batch size 24. The first training stage includes 450,000 iterations while the second stage includes 150,000 iterations. The input frames in one batch are down-sampled by the same space scale and randomly cropped into patches with size 32×\times32. We perform data augmentation by randomly rotating 90∘90^{\circ}, 180∘180^{\circ} and 270∘270^{\circ}, and horizontal-flipping. We use Zooming SlowMo as the encoder. For the two functions incorporated in continuous space and time representations, we utilize two 3-layer SIRENs with hidden dimensions of 64,64,25664,64,256. For the decoding network, we employ a 4-layer SIREN with hidden dimensions of 64,64,256,25664,64,256,256. As suggested in , we select the Charbonnier loss function for optimization.

Evaluation. Peak-Signal-to-Noise Ratio (PSNR) and Structual Similarity Index (SSIM) are employed to evaluate model performances. We also compare the model size and inference time to measure the efficiency of models.

2 Comparison to State-of-the-arts

We compare VideoINR with state-of-the-art two-stage and one-stage STVSR methods. For two-stage methods, we employ SuperSloMo , QVI , and DAIN for video frame interpolation (VFI); Bicubic Interpolation, EDVR , and BasicVSR for video super-resolution (VSR). For one-stage methods, we compare VideoINR with recently developed Zooming SlowMo and TMNet . To perform fair comparisons, we train the three VFI methods and Zooming SlowMo from scratch on Adobe240 dataset. For TMNet, as mentioned in the original paper that a two-stage training scheme is needed for convergence, we pre-train the model on Vimeo90K dataset and fine-tune it on Adobe240 dataset . Therefore, TMNet is trained on more data compared with other methods, which may lead to some advantages in the comparison. To compare with Zooming SlowMo that only supports fixed frame interpolation, we train a new version of VideoINR named VideoINR-fixed of which the interpolation time is fixed to 0.5.

Quantitative results. We present in-distribution quantitative comparisons between VideoINR and other STVSR methods in Table 1. On single frame interpolation of STVSR including Vid4, GoPro-Center, and Adobe-Center, VideoINR-Fixed achieves competitive performance compared with other state-of-the-art models, while the performance of VideoINR slightly suffers. We attribute this observation to the difference of training targets between VideoINR and VideoINR-Fixed. The training settings of VideoINR-Fixed aim for synthesizing frames at pre-defined times. Therefore, it only learns fixed patterns between input frames instead of learning a continuous representation as VideoINR does, leading to advantages in performances. On Vid4, TMNet performs the best, and we assume this is because TMNet is trained with more data as we noted in Section 4.2. For multiple frame interpolation of STVSR including GoPro-Average and Adobe-Average, VideoINR achieves the best performance, which indicates that the proposed implicit neural representation provides advances on modeling the temporal information in videos.

In Table 2, we present comparisons of STVSR methods on out-of-distribution space and time scales. For two-stage STVSR methods, we select SuperSloMo and DAIN as VFI methods, and LIIF as the SR method since it can perform super-resolution on arbitrary up-sampling scales. We also take TMNet into the comparison as it could generalize on time scales. We produce experiments on GoPro dataset. We observe that VideoINR outperforms other methods by a large margin, demonstrating the advantage of our continuous video representation in out-of-distribution generalization. In addition, we further compare VideoINR with Zooming SlowMo (the encoder for VideoINR) in out-of-distribution scales. As Zooming SlowMo only supports interpolating fixed frames, we apply the model twice to achieve out-of-distribution inferences. In Table 3, we observe that while Zooming SlowMo performs slightly better on single frame interpolation (×4×2\times 4\times 2), VideoINR achieves better performance in out-of-distribution testing (×16×4\times 16\times 4).

We compare the inference time of STVSR methods in Figure 3. We observe that the efficiency of different methods is close at up-sampling time scale ×2\times 2, and VideoINR inferences faster than other models on multi-frame interpolation. We attribute this feature to the design of VideoINR, where all the latent frames between two input frames can be directly synthesized by MLPs after encoding.

Qualitative Results We demonstrate a qualitative comparison in Figure 4. We compare VideoINR with two STVSR methods, DAIN + BasicVSR and TMNet. The selected temporal coordinates of the first sample are in the training distribution, while the coordinates of the second sample are out-of-distribution. We find that the performance of DAIN + BasicVSR degrades in out-of-distribution circumstances (see the rider’s head in the second sample). TMNet fails to recover objects with large motion between two input frames (see the flowers in the first sample). The performance of VideoINR is steady across both in-distribution and out-of-distribution temporal coordinates, indicating that learning continuous video representations helps to improve model generalization in STVSR task.

3 Ablation Study

Motion Flow Field. Motion flow is one critical component of VideoINR. Previous video interpolation methods have already demonstrated that such a learnable flow helps to interpolate frames with sharp edges and clear details. We propose that the motion flow field brings two main advantages. First, the flow field could capture non-local information and temporal contexts of large motions. Second, we explicitly apply spatial warping on features, which works as an inductive bias for the training. In Table 4 between VideoINR and VideoINR (-f), we show that the performance degrades when the motion flow is not incorporated.

VideoINR trained with different data settings. In Table 5, we compare the performances of VideoINR trained on different data settings. As noted before, VideoINR follows a two-stage training strategy: fixed down-sampling space scale for the first stage and continuous space scales sampled from a uniform distribution for the second stage. VideoINR-×4\times 4 indicates that the space scale is fixed to ×4\times 4 throughout the training of VideoINR. VideoINR-continuous represents VideoINR trained with continuous down-sampling space scales from scratch. We find that the performance suffers a significant drop when we train VideoINR only on continuous scales. We hypothesize this is because the network needs to learn spatial and temporal representations at the same time, and it becomes extremely difficult to learn such temporal representation when the scale of spatial features keeps varying. Besides, we observe that training VideoINR with a fixed space scale achieves slightly better performance for that specific scale. However, its generalization performance is competed by VideoINR trained by two stages, which is demonstrated by the comparisons between VideoINR and VideoINR (-×4\times 4) on space scales other than ×4\times 4.

Other design choices. We provide more ablation studies in Table 4. By comparing VideoINR with VideoINR (-m), we find that the proposed multi-scale feature aggregation contributes to performance improvement. We also try to replace SpatialINR and TemporalINR by a single network, that is, we use one network only for generating the continuous motion flow, and apply spatial warping only on the encoded feature and input frames. The results between VideoINR and VideoINR (-s) indicate that using two functions for representing space and time outperforms only one network for them all.

Discussion

Conclusion. In this paper, we present Video Implicit Neural Representation (VideoINR). It can represent videos in arbitrary spatial and temporal resolution, which brings natural advantages for solving Space-Time Video Super-Resolution (STVSR) tasks. Extensive experiments show that VideoINR performs competitively with state-of-the-art STVSR methods on common up-sampling scales and outperforms prior works by a large margin on out-of-distribution scales.

Limitations and Future Work. We observe that there exist few cases for which VideoINR does not perform very well. These cases typically need to handle very large motions, which is still an open challenge for video interpolation.

Acknowledgements. This work was supported, in part, by gifts from Picsart.

References

Appendix A Implementation Details

All models included in experiments are trained from scratch to perform fair comparisons. For video frame interpolation methods incorporated in the experiments (i.e. Super SloMo , QVI , and DAIN ), we train them on Adobe240 dataset . We keep all the training settings the same as proposed in their original papers, including the optimizer, initial learning rate, learning rate decay strategy, and the number of training epochs. For data settings, 99 consecutive frames are selected from video clips for training in every iteration. Networks take the first and last frames as inputs and generate intermediate 77 frames. We calculate loss between generated frames and the original ground-truth frames. Each video frame is resized to have a shorter spatial dimension of 360, and a random crop of 352×\times352 is performed.

Zooming SlowMo and TMNet are two STVSR models included in our experiments. Zooming SlowMo only supports fixed frame interpolation, and the interpolation time is set to , 0.50.5, 11 in the original paper. Following their settings, we also train the model from scratch to interpolate the fixed time instances. To ensure that the input video frames of all models are of the same frame rate, we extract 99 consecutive frames from video clips and take the 1st1^{st} and 9th9^{th} frames as inputs. We then down-sample the input frames via Bicubic interpolation by a factor of 4 and use the network to predict the high-resolution versions of the 1st1^{st}, 5th5^{th} and 9th9^{th} frames. TMNet supports arbitrary frame interpolation. In its paper, the authors mention that TMNet needs a two-stage training process for convergence, and we follow their suggestions. In the first stage, we pre-train the network on the Vimeo90K dataset . The Vimeo90K dataset consists of 7-frame video sequences. We use the 1st1^{st}, 3rd3^{rd}, 5th5^{th}, and 7th7^{th} frames after down-sampling as the network inputs and predict the high-resolution results of all the 77 frames, which means that the interpolation time is set to , 0.50.5, 11 in this stage. In the second stage, we select 99 consecutive frames from video clips, and the 1st1^{st} and 9th9^{th} frames are taken as inputs. After down-sampling, we use the network to generate high-resolution predictions of all 99 frames and calculate the loss value with the original high-resolution frames. TMNet is trained with more data, which may lead to advantages in the experiments.

For the training of VideoINR, we select 99 consecutive frames and down-sample the 1st1^{st} and 9th9^{th} frames as model inputs. In each iteration, We randomly select three frames from the 9-frame video sequence and use the network to generate high-resolution predictions at the time instances of the three selected frames.

We keep the training settings unchanged for Zooming SlowMo, TMNet, and VideoINR. All three models are optimized with the Charbonnier loss function .

Appendix B Efficiency on Different Scales

To evaluate the efficiency of VideoINR on different up-sampling space scales, we provide more inference time comparisons in Figure 5. We select the two-stage method composed of SuperSlomo and LIIF as the baseline, as it supports arbitrary up-sampling scales on both space and time.

Appendix C Limitations

In some challenging cases, large motion and occlusion result in errors on the motion flow field, leading to blurred results with unclear boundaries. We show a failure cases of VideoINR in Figure 6.

Appendix D Additional Qualitative Results

We provide more qualitative results in Figures 7,8,9,10. We compare VideoINR with two STVSR methods, DAIN + BasicVSR and TMNet . The up-sampling space scale is set to 4 for all examples. In Figure 7, 8, we set the time scale for interpolation to 8, which is in our training distribution. We observe that DAIN + BasicVSR and TMNet tend to generate blurry regions or artifacts. In contrast, the results of VideoINR are consistent and aligned across two input frames, with sharp edges and clear details. In Figure 9, 10, we set the time scale to 12 and 16, which are out of the training distribution. We find that VideoINR well recovers objects with large motion and preserves better textural information compared with other methods. In summary, VideoINR shows the advantages of learning continuous representation for videos, and address the Space-Time Video Super-Resolution task. More visualization results can be found in the provided video.