GenRec: Unifying Video Generation and Recognition with Diffusion Models
Zejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu, Yu-Gang Jiang
Introduction
Diffusion models have achieved significant success in the field of image and video generation over the past few years. A variety of generative tasks have been revolutionized by using diffusion models trained on Internet-scale data, such as text-to-image generation , image editing , and more recently, text-to-video generation and text&image-to-video generation . The excellent generative capabilities of diffusion models suggest that informative representation is learned during the generative training and strong visual priors are captured by the backbone models . Therefore, recent work has explored leveraging the image diffusion models for image understanding tasks, including image recognition , object detection , segmentation and correspondence mining . However, the capability of video diffusion models to effectively capture spatial-temporal information is not fully understood, and their potential for downstream video understanding tasks remains under-explored.
In this paper, we study the potential of video diffusion models , particularly the unconditioned or image-conditioned models, for video understanding by addressing the three key problems: (a) Does the backbone model trained for video generation extract effective spatial-temporal representations for semantic video recognition? (b) Can we retain the video generation capability by jointly optimizing generation and recognition? (c) Will such a unified training framework further benefit video understanding, especially in noisy scenarios where only limited frames are available .
While conceptually appealing, unifying video generation and recognition into a diffusion framework is non-trivial. Prior work either views the diffusion models as frozen feature extractors , or deconstructs them for new tasks while sacrificing their original generation capability . One major challenge comes from their distinct training and inference processes. Diffusion models are typically optimized using corrupted inputs, optionally augmented with a single conditioning frame, to achieve unconditioned or image-conditioned generation during inference . In contrast, video recognition models require access to multiple frames to reason about temporal relationships and expect clean inputs during inference . Consequently, training a recognition model using corrupted videos and single-image conditions tends to suffer from inferior model optimization and a more significant training-inference gap.
To this end, we propose GenRec, a unified video diffusion model that enables joint optimization for video generation and recognition. Our model is built upon the open-source, image-conditioned Stable Video Diffusion model (SVD) , which encodes strong spatial-temporal priors by pretraining on large-scale image and video data. However, instead of conditioning on the same image across all video frames, we propose to condition on a random subset of frames while masking the remaining ones (see Figure 2). This simple random-frame conditioning process effectively bridges the gap between the learning processes of the two tasks. On the one hand, the generation capability of SVD is extended to handle arbitrary frame prediction, which provides more flexible and unambiguous video generation. On the other hand, conditioning on a random subset of frames allows the model to learn more discriminative and robust features for the recognition task. As shown in Figure 1, the model is jointly optimized using both generative supervision (i.e., noise prediction) and classification supervision.
We conduct extensive experiments to evaluate the performance of GenRec for both recognition and generation. Without sacrificing the generation capabilities, GenRec demonstrates competitive video recognition performance, offering 75.8% and 87.2% accuracy on SSV2 and K400, respectively. Furthermore, GenRec demonstrates extraordinary robustness in scenarios that only limited frames can be observed. For example, when only the front half of the video can be observed, GenRec achieves the 57.7% accuracy, which corresponds to 76.6% of the accuracy (75.3%) when the entire video is visible, emonstrating a higher accuracy retention ratio than other methods. By leveraging the recognition model for classifier guidance , GenRec also achieves superior class-conditioned image-to-video generation results, with FVD scores of 46.5 and 49.3 on the SSV2 and EK-100 datasets, respectively.
Preliminary
Representing the data distribution as with a standard deviation of , we can obtain a family of smoothed distributions by adding independent and identically distributed Gaussian noise with standard deviation . In the spirit of diffusion models, the generation process begins with a noise image and iteratively denoises it at decreasing noise levels . The final denoised result is thus distributed according to the original data.
In the EDM framework, the original will be diffused as:
where is the score function. Noise schedule is set as time step . The training objective is to minimize the loss with the denoiser network for different :
with the relation between and the score function as follows:
GenRec
We now introduce GenRec, a simple yet efficient framework, that can not only generate temporally-coherent videos conditioned on an arbitrary number of provided frames but also is able to recognize actions and events with the help of encoded spatial-temporal priors. To this end, GenRec explores the strong spatial-temporal priors learned by a video diffusion model. In this work, we instantiate GenRec with the powerful open-source Stable Video Diffusion model (SVD) , which is pretrained on large-scale video datasets and is able to produce a photo-realistic video when provided a single frame. Then, for generation, GenRec follows the classical EDM framework to learn noise reduction trajectories. For recognition, on the other hand, GenRec operates on intermediate decoded features using a recognition head. Furthermore, to generate videos in a more free fashion, i.e. an arbitrary collection of frames used as condition, we design a latent masking strategy that “interpolates” masked frames. Such a strategy also benefits recognition by easing the training process. More importantly, by doing so GenRec supports a multitude of downstream tasks, particularly when limited visual information is provided.
Unifying generation and understanding.
For the generation task, the UNet aims to reconstruct the original latent representation from the combined noisy and masked inputs. Representing UNet as the mapping function , its goal is to predict clean latent, which, according to the EDM framework, takes the form of a representation mapping as follows:
in which we set the same skip connection , scaling factor and as .
2 Optimization
We train GenRec with both generation and classification objectives, encouraging the model to learn high-quality video generation and accurate video understanding.
The generative loss uses a loss to measure the difference between the original latent representation and the reconstructed output produced by the UNet, and is defined as:
where denotes the ground truth labels, and represents the predicted labels referring to Equation Equation 6.
To balance the learning of generative and recognition tasks, we set a balancing weight to control the relative importance of each loss in the overall objective function. The total loss is given by:
3 Inference for Different Downstream Tasks
With the above training strategies, we now introduce how GenRec can flexibly support different types of generation and recognition tasks.
Once trained, GenRec is able to generate high-quality videos conditioned on an arbitrary number of given frames, thanks to the latent masking strategy. Particularly, following the EDM stochastic sampler framework and Equation 2, GenRec iteratively denoises the video conditioned on the masked latent , as shown below:
Video generation conditioned on classes.
When the number of visible frames is extremely limited, the motion trajectory becomes unpredictable and thus it would be hard to make a reliable prediction of the future. To mitigate this issue, GenRec supports adding category information to guide video generation in the expected desired direction.
Formally, we simplify Equation 11 with Equation 4, and obtain:
We can now replace the original score function with , in which denotes the conditional class, to get the conditional version of residual, denoted as :
Considering the scaling factor of : (following , and ), that would pre-scale the input as before model processing, the formulation can be further transferred as:
Following , we sharpen the distribution of by multiplying a scaling factor , shown as where is an arbitrary constant. Larger scaling value would bring more attention to the target category. Here, comes from the classification branch in GenRec. Finally, we can use the same EDM sampling procedure with the derived class information to generate samples.
Standard video recognition.
Based on Equation 6, GenRec can do the classical video recognition by setting constant no-mask, and thus is replaced by and the prediction follows:
Video recognition with partially observed frames.
Experiments
In our experiments, we use the following four datasets: Something-Something V2 (SSV2) , Kinetics-400 (K400) , UCF-101 and Epic-Kitchen-100 (EK-100) . SSV2 dataset is designed for fine-grained action recognition and it contains 174 action classes, 220,847 short video clips with an average duration of 4 seconds. K400 contains 400 action classes, 306,245 video clips with an average duration of 10 seconds. The UCF-101 dataset comprises 13,320 videos from 101 action categories and is widely utilized for human action recognition. The EK-100 dataset focuses on egocentric vision. It contains a total of 90,000 annotated action segments, encompassing 97 verb classes and 300 noun classes.
Evaluation protocols.
GenRec performs both generation and recognition tasks. For generation, we use the Fréchet Video Distance (FVD) metric to assess the quality of the generated videos. A lower FVD score indicates higher fidelity and realism. For recognition, we measure the top-1 accuracy that reflects the portion of correctly classified videos. We validate our model performance in formal video recognition, partial video recognition, class-conditioned image-to-video generation and frame completion with the above metrics.
Implementation details.
We initially set the learning rate to and set the total batch size as 32. Only generation loss will be retained for model adaptation on specific datasets. We train 200k steps on EK-100 and UCF, and 300k steps on SSV2 and K400, respectively. Subsequently, we finetune GenRec with both generation and recognition losses. The learning rate is set to and decayed to using a cosine decay scheduler. We warm up models with 5 epochs, during which the learning rate is initially set as and linearly increases to the initial learning rate . The loss balance ratio is set to 10, and the learning rate for the classifier head is ten times higher than the base learning rate. We drop out the conditions 10% of the time for supporting classifier-free guidance , and we finetune on K400 for 40 epochs and 30 epochs on other datasets. The training is executed on 8 A100s and each contains a batch of 8 samples. We sample 16 frames for each video.
2 Main Results
We compare with state-of-the-art methods in terms of their recognition accuracy and generation quality. The results are summarized in Table 1. The first two blocks of the table presents current advanced video recognition models, while the third block demonstrates the performance of the diffusion-based class-guided image-to-video generation.
As shown in the table, GenRec achieves optimal results or performs on par with the state-of-the-art approaches. In terms of video recognition, GenRec achieves 75.8% accuracy on SSV2 dataset, surpassing the majority of current state-of-the-art methods. On K400, GenRec achieves 87.2% accuracy, which is on par with the performance of MVD-H (87.2%) and Hiera (87.3%, 87.8%), and surpasses other advanced methods. These results indicate the effectiveness of our approach in video recognition. In addition, GenRec shows a slight performance gap compared to the methods in the second block. It is important to note that these advanced methods benefit significantly from pretraining on large-scale multimodal alignment datasets, which provide extensive cross-modal supervision that enhances their ability to capture semantic relationships across video frames.
We further construct two strong baselines. Baseline I adapts SVD to the respective dataset through generative fine-tuning, followed by attentive-probing for classification, where the backbone is frozen and all frames are used as input. Baseline II involves fully fine-tuning the original SVD model with classification supervision only, ensuring that all frames are visible during training. Compared with them, GenRec performs on par or better. GenRec performs good in supporting not only classification but also generation, demonstrating its comprehensive capability in handling both tasks effectively.
In terms of video generation, we evaluate the model on class-conditioned image-to-video generation task following . Comparing the FVD scores of SEER:112.9 and SEER:355.4, it can be inferred that generating longer videos with 16 frames is more difficult than generating 12 frames. GenRec generates videos with 16 frames and achieves much lower FVD scores than the other methods, demonstrating the effectiveness of our approach in video generation.
It is worth highlighting that, current research always treats video recognition and generation tasks in a separate manner, and most of the advanced methods focus primarily on either recognition or generation tasks. For instance, SEER method excels in class-conditioned image-to-video generation, but lacks the ability to do video recognition. While current research on representation learning, shown as the first and second blocks in Table 1, lacks the ability to do video generation tasks. In contrast, GenRec not only unifies these tasks, but also achieves competitive results compared to the specialized methods.
Comparison to state-of-the-art in video recognition with limited frames.
GenRec supports video recognition when only partial frames can be observed. We evaluate this capability on the SSV2 and EK-100 datasets. Our evaluation includes two tasks: an early prediction task, where the model has access only to previous continuous frames following the setting of , and a recognition task where videos are sparsely sampled, and the model is expected to make correct predictions. For fair comparisons, we construct two strong baselines. We first apply MVD to directly deal with the recognition task by constructing a dense video through nearest neighbor interpolation. We also construct another baseline similar to our training pipeline, where we apply frame dropout in the training process of MVD for better fitting on task with partial frames, and is named as MVD. In all settings, the number of fully observed frames is 16.
Table 2 shows the results under these settings, in which denotes the visible ratio (there are a total of 16 frames). In the early action prediction task, GenRec achieves the highest accuracy and ratio metrics at all observation levels. Notably, GenRec and MVD exhibit similar performance when all frames are observed, but as the number of observed frames decreases, GenRec demonstrates higher accuracy. GenRec also shows superior performance when videos are sparsely sampled, maintaining high accuracy even with fewer observed frames (e.g., 55.7% for 2 frames and 70.8% for 4 frames), indicating its robustness in handling sparse data. Moreover, we compute the ratio metric representing the percentage of maximum performance that the model can maintain at various frame rates, mitigating the unfairness caused by different backbone networks. In this scenario, GenRec still achieves the best performance.
We further investigate the contributions of generation supervision for recognition. As seen in Table 2, removing generation supervision results in noticeable performance degradation across various tasks, especially when the number of visible frames get less. For example, in the early prediction task, the accuracy decreases by 0.3% at and by 2.3% at . These results suggest that generation supervision is essential for maintaining high performance, particularly when the model has to make predictions with limited visual information. By incorporating generation supervision, the model can better handle scenarios with incomplete data, improving robustness and accuracy.
We also evaluate the early action prediction on EK-100 and UCF-101. EK-100 is a temporally sensitive dataset similar to SSV2, demanding in terms of the model’s temporal modeling capability, while UCF-101 demands more on appearance modeling. We conduct early prediction evaluation on them to further reveal the robustness of our GenRec. As shown in Figure 3, GenRec clearly outperforms TemPr and MVD. In particular, the improvement becomes more significant as the number of observed frames decreases. More evaluation results can be seen in Section A.1.
These results collectively demonstrate that GenRec effectively handles missing video frames. The robustness and high accuracy of GenRec across different datasets and observation ratios highlight its potential for real-world applications where video data might be incomplete or sparsely sampled.
The relationships between generation and recognition .
We further investigate the consistency between video generation and recognition, as shown in Table 4. We evaluate the performance of video recognition and generation with limited frames and find that the recognition accuracy not only depends on the number of visible frames but also significantly on the location of these frames. Interestingly, uniform sampling appears to facilitate video recognition better than dense sampling from the video prefix. Specifically, with the same number of frames, early prediction consistently shows lower accuracy compared to uniformly sampled frames (e.g., 28.9% vs. 55.7% with 2 frames) and worse FVD scores (e.g., 57.8 vs. 46.7 with 2 frames). When only three interpolated frames are visible, the 31.7 FVD score is comparable to that of an eight-frame prefix (30.3), while achieving much higher recognition accuracy. These results highlight the importance of complete state observation for action recognition and also suggest that video generation performance can potentially reflect task difficulty.
Choice of UNet layers.
As described in Section 3, the UNet mapping function is decoupled into , where serves as the feature extractor for video recognition. Our UNet model contains 4 main up-sampling blocks. We investigate which one is best suited for recognition. As shown in Table 4, using the second up-sampling block (Up Index 2) yields the best performance with an accuracy of 75.8%. The third block (Up Index 3) followed with 75.2%, while the first block (Up Index 1) has the lowest accuracy. As such, we choose the second block for feature extraction.
Explore the influence of the masking strategy.
We also conduct an ablation study on the masking schemes using different expected masking ratios, as shown in Table 5. The results show that the FVD scores remain similar across different ratios, and a larger masking ratio might be beneficial for generation, as it closely resembles our class-conditioned frame prediction scenario with one or two given frames. However, an excessively large masking ratio (87.5%) negatively impacts action recognition accuracy, leading to a 0.9% decrease compared to our selected ratio.
Related Work
The great success of diffusion models in image generation has led to rapid advancements in video generation, including text-to-video generation , image&text-to-video generation , and video editing . Many current works adapts the diffusion models from images to videos by incorporating temporal convolutions and attention mechanisms. One typical and excellent work, Stable Video Diffusion , follows the above description and has provided valuable foundations for generating high-quality, diverse, and temporally consistent videos. Different from the previous work, in our paper, we pursue not only the quality of generation, but also the unity of model generation capability and classification ability.
Diffusion models for visual understanding.
Recently, researchers start to uncover the significance of diffusion models for discrimination tasks. A notable approach involves utilizing pretrained visual diffusion models for various downstream tasks, such as image segmentation and visual content correspondence . Additionally, some studies treat diffusion learning as a self-supervised method to acquire valuable feature representations . However, most current works either use stable diffusion networks as pretrained backbones for downstream tasks or completely destroy their generative capabilities. Consequently, the potential benefits of integrating generation and classification abilities into a single model remain under-explored, which is the primary focus of our paper.
Conclusion
In this work, we presented GenRec, a unified video diffusion model that enables joint optimization for both video generation and recognition. GenRec exploits the significant temporal modeling power embedded in the diffusion model, allowing for mutual reinforcement between generation and recognition tasks. Extensive experiments were conducted to evaluate the performance of GenRec, demonstrate our approach contains strong generation and recognition capabilities at the same time in different kinds of scenarios, including normal or partial video recognition, video completion and class-conditioned image-to-video generation. Our findings highlight the potential of combining generation and classification tasks within a single unified model, providing valuable insights into the development of more sophisticated and versatile video analysis models. Future work will focus on further refining this integration and exploring its applications across various real-world scenarios.
References
Appendix A Appendix / supplemental material
More detailed evaluation results of the video recognition with limited frames on EK-100 and UCF-101 can be seen here.
A.2 Case Study for Class-Conditioned Image-to-Video Generation and for Video Interpolation
We show the generated visualization of the GenRec. The model can support video generation given various numbers of frames, as well as category-guided generation. We show two of the most difficult generative scenarios, which are: (1) given the first frame and different action categories to guide the video generation, and (2) given the start and end frames, the model is expected to complement the video. We also compare our methods with SEER with cases picked from its official website. The generation results can be seen in Figure 4, Figure 5 and Figure 6.
A.3 Limitations and Broader Impacts
The objective of our paper is to unify the tasks of generation and recognition, achieving or even surpassing the state-of-the-art experimental performance across various tasks. However, our method is based on fine-tuning a pretrained video diffusion model, using more pretraining data and having a larger number of parameters compared to previous methods. This is an issue we need to address in the future, and exploring the distillation of a well-pretrained video diffusion model into a smaller model is a worthwhile future endeavor.
Broader Impacts
The broader impact of the GenRec framework extends into various fields, enhancing capabilities in content creation, security, and accessibility. In the media industry, it allows for the automated generation of tailored, high-quality videos, reducing production costs and fostering creativity. For surveillance, its robustness in limited information scenarios improves monitoring effectiveness, particularly in challenging environments. Additionally, advancements of GenRec in video prediction can aid in developing assistive technologies, making digital content more accessible and interactive, particularly for individuals with visual impairments.