InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, Limin Wang, Yu Qiao

Introduction

Foundation models have been gaining increasing attention in the research community , since they give a practical paradigm for scaling to numerous perception tasks with surprisingly good results. Through simple adaption or zero/few-shot learning, foundation models greatly reduce downstream design and training costs using generic representations learned from web-scale data with a strong backbone of high capacity. It is expected that developing foundation models can cultivate cognition from perception, obtaining general vision capability.

Though a line of vision foundation models is proposed , video understanding and the corresponding tasks are less explored compared with image ones, mainly used for validating that the visual features from these models are also beneficial to spatiotemporal representations. We conjecture this relatively scarce focus from the academic community is caused by 1) a high computing burden from video processing, and 2) quite a few current video benchmarks can be handled by exploiting appearance features from image backbones with accordingly temporal modeling. Specifically, for efficiency, the additional time dimension in video processing makes at least one order of magnitude higher complexity than image processing when their spatial resolutions are close and the temporal sampling ratio is usually 16. For some current video datasets, image features alone or with lateral temporal modules are sufficient to give decent results, especially with the rise of the multimodal model CLIP . Its various temporal variants yield competitive or state-of-the-art performance in several core tasks . Regarding this, a simultaneous spatiotemporal learner does not seem like a sweet spot between research & development cost and payback.

Moreover, the transferability of current vision foundation models is somewhat narrow considering the wide spectrum of video applications. These models either concentrate on action understanding tasks (e.g. action recognition, spatiotemporal action localization, etc) or video-language alignment ones (e.g. video retrieval, video question answering, etc). We suppose this results from their learning schemes, as well as the lack of a comprehensive benchmark for measuring video understanding capabilities. Thus, these works focalize a few specific tasks to demonstrate their spatiotemporal perceptions. The community desires a general foundation model that enables a broader application domain.

In this paper, we advance video foundation model research with a cost-effective and versatile model InternVideo. To establish a feasible and effective spatiotemporal representation, we study both popular video masked modeling and multimodal contrastive learning . Note that video masking modeling specializes in action understanding, and it is still worth exploring regarding its limited model scale caused by the current decoder. For multimodal contrastive learning, it embeds rich semantics into video representation while ignoring concrete spatiotemporal modeling. To address these challenges, we make these two self-supervised approaches learn at scale efficiently in modular designs. To significantly broaden the generalization of the current video foundation models, we propose a unified representation learning with both two self-supervised training manners. To validate such a generalized representation, we propose a systematic video understanding benchmark. It involves evaluations of action understanding, video-language alignment, and open-world video applications, which we believe are three core abilities of generic video perception. Instead of introducing new data or annotations to this system, we initially choose ten representative video tasks with 39 public datasets, and categorize them into those three types. To our best knowledge, InternVideo is the first video foundation model which demonstrates promising transferability with state-of-the-art performance in all those three different types of video tasks.

In InternVideo, we design a unified video representation (UVR) learning paradigm. It explores both masked video modeling with autoencoders (MAE) and multimodal contrastive learning for two types of representations, strengthens them by supervised action classification, and generates a more general representation based on the cross-representation learning between them. UVR not only empirically shows video representation outperforms image one with temporal capturing significantly on core video tasks, but also is training-efficient. Its MAE exploits high redundancies in videos and trains with only a few visible tokens. Meanwhile, multimodal learning in InternVideo extends existing image-pretrained backbones for video contrastive training. After supervised training these two video encoders, we craft cross-model attention to conduct feature alignments between these two almost frozen encoders.

More than a unified video representation learning paradigm, we also make practices and guidelines for training large-scale video foundation models in a tractable and efficient manner. Our work contains and is not limited to 1) making VideoMAE scalable and exploring its scalability in model and data scale; 2) efficient and effective multimodal architecture design and training receipt about how to leverage existing image-pretrained backbones; 3) empirically finding features from VideoMAE and multimodal models are complementary and studying how to deduce more powerful video representations by coordinating different existing models. Specifically,

For the scalability study of VideoMAE, we show that the proper diversity and scaling-up size in training videos can improve the scalability of the used video encoder. With a new pretrained dataset in a masked autoencoder training setting, ViT lifts its action recognition performance on Kinetics-400 with finetuning from 81.01% to 85.35% from base to large, and further reaches 86.9% with the huge setup, surpassing the performance reported in with a notable margin. The scalability of VideoMAE enables its usage in video foundation model development.

For reusing existing foundation models for multimodal learning, we extend an image-pretrained vision transformer to video representation learning. This transfer learning requires substantial structure and optimization customizations or using local and global spatiotemporal modules for multimodal pretraining. The local module disentangles spatiotemporal modeling by consecutive and independent spatial and temporal attention computation. Meanwhile, the global module computes token interaction across space and time. Experiments show that this reuse design is effective in spatiotemporal representation learning.

More than self-supervised pretraining, we also employ supervised action recognition to further enhance video representation. Results demonstrate action recognition is a fine source task for transferring to various downstream applications.

To coordinate foundation models, we unify masked video encoder with multimodal one by cross-representation learning, instead of training them jointly in one formulation. Regarding the optimization of MAE and multimodal learning (MML) that may contradict each other, combining them without compromising their merits remains an open question . More importantly, MML with contrastive learning demands huge batches for better contrastive optimization. Adding MAE to it will inevitably lead to numerous headachy implementation issues. Considering their potential training adversaries, we train MAE and MML separately. After their training converges, we then dynamically combine their representations with the proposed cross-model attention (CMA) modules. It implements cross-attention between MAE and MML mid-level features, adaptively fusing their high-level features for prediction. In the model-level representation interaction phase, we freeze backbones trained by MAE and MML separately and only let CMA be updated in supervised learning with a few epochs. Experiments show it is a computationally tractable and efficient means to exploit both MAE and MML features.

We validate our proposed video foundation model in 10 tasks with 39 datasets (including core tasks e.g. action recognition, spatiotemporal action localization, video question answering, video retrieval, etc), and it outperforms all the state-of-the-art methods in each task non-trivially. We suppose these overall superior results obtained by our approach, along with observations and analysis, set up a new baseline for the video understanding community. The empirical evidence in this paper raises the confidence that video perceptive tasks and partial high-order tasks (formulated to perceptive forms) can be well-addressed by video foundation models, serving as a performance-critical method across a spectrum of applications.

In summary, we contribute to video foundation models in the following aspects:

We explore a general video representation paradigm with both masked and contrastive modeling, and realize this design by unifying their representation by lightweight model interaction learning in supervision. We confirm features learned by generative and contrastive training are complementary to experiments and can deliver better results than either of them trained independently.

We find masked video encoder can be scalable in model and data size with proper tuning. We devise pluggable local temporal and global spatiotemporal interaction modules to reuse pretrained ViT with image-text data for multimodal learning, easing the training burden and yielding better downstream performance.

We make a tentative attempt in constructing a systematic video understanding benchmark. Our general video foundation models achieve state-of-the-art performance on 39 datasets with several core tasks in this benchmark, e.g., Kinetics-400 and Something-Something v2 in action recognition. We empirically find our learned video representations outperform their rivals, dominating vision-language tasks by a large margin, especially for some image-based ones. It suggests general video representations will be a central player in video tasks. We believe the openness of our proposed methods and models will provide the research community with handy tools to foundation models and their features with easy access.

Related Work

Image Foundation Models. Most of the current vision models are only suitable for specific tasks and domains, and they require manually labeled datasets for training. Regarding this, recent works have proposed vision foundation models. CLIP and ALIGN prepare web-scale noisy image-text pairs to train dual-encoder models with contrastive learning, leading to robust image-text representations for powerful zero-shot transfer. INTERN expands the self-supervised pretraining into multiple learning stages, which use a large quantity of image-text pairs as well as manually annotated images. INTERN achieves a better linear probe performance compared with CLIP, and improves data efficiency in the downstream image tasks. Florence extends them with unified contrastive learning and elaborate adaptation models, which support a wide range of vision tasks in different transfer settings. SimVLM and OFA train encoder-decoder models with generative targets and show competitive performances on a series of multimodal tasks. Besides, CoCa unifies contrastive learning as CLIP and generative learning as SimVLM. Recently, BeiT-3 introduces Multiway Transformers with unified BeiT pretraining, achieving state-of-the-art transfer results on several vision and image-language tasks.

Video Foundation Models. Previous image foundation models only show promising performance for video recognition (especially on Kinetics). As for video multimodal tasks, VIOLET combines masked language and masked video modeling, All-in-one proposes unified video-language pretraining with a shared backbone, and LAVENDER unifies the tasks as masked language modeling. Though they perform well in multimodal benchmarks, they are trained with limited video-text data and struggle for video-only tasks, e.g. action recognition. In contrast, MERLOT Reserve collects 20M video-text-audio pairs to train the joint video representations with contrastive span matching, thus setting state-of-the-art video recognition and visual commonsense reasoning. Compared with image foundation models, current video foundation models support limited video and video-language tasks, especially for those fine-grained temporal discrimination tasks such as temporal localization.

Self-supervised Pretraining. Self-supervised learning has developed rapidly recently. It focuses on designing different pretext tasks for pretraining , which can be mainly divided into contrastive learning and masked modeling. Contrastive learning adopts various data augmentations to generate different views of an image, then pulls together the positive pairs and pushes apart the negative pairs. To maintain enough informative negative samples, previous methods depend on large memory banks or batch size . BYOL and SimSiam eliminate the requirement of negative samples, designing elaborate techniques to avoid model collapse. As for masked modeling, it learns rich visual representation via masked prediction based on visible context. iGPT firstly mentions Masked Image Modeling (MIM). BeiT propose visual token prediction with the pretrained tokenizer , MaskFeat predicts the hand-crafted image descriptor, and MAE directly reconstructs the raw pixels. For spatiotemporal representation learning, VideoMAE and BEVT respectively extend MAE and BeiT to spatiotemporal space.

Multimodal Pretraining. Starting from the development of image-text pretraining, large-scale video-text pretraining with specific downstream task finetuning has become the standard paradigm in the video-language area . The seminal methods use pretrained visual and language encoders to extract the offline video and text features, while the recent methods have demonstrated the feasibility of end-to-end training. Besides, the popular methods often include two or three pretraining tasks, e.g. masked language modeling , video-text matching , video-text contrastive learning and video-text masked modeling .

InternVideo

InternVideo is a general video foundation model along with its training and internal cooperation as given in Figure 2. In structure, InternVideo adopts the vision transformer (ViT) and its variant UniformerV2 , along with extra local spatiotemporal modeling modules for multi-level representation interaction. In learning, InternVideo improve its representation progressively, integrating both self-supervised (masked modeling and multimodal learning) and supervised training. Moreover, as we explore two types of self-supervised learning, we further integrate their merits. InternVideo dynamically derives new features from these two transformers via learnable interactions, getting the best of both worlds from generative and contrastive pertaining. Through the newly aggregated features, InternVideo sets new performance records on 34 benchmarks from 10 mainstream video tasks, and wins championships of five tracks in recent Ego4D competitions .

InternVideo conducts both masked and contrastive training without supervision for representation learning. According to , video masked modeling produces features that excel at action discrimination, e.g., action recognition and temporal action localization, and video-language contrastive learning is able to understand videos with semantics from text without annotations. We employ two transformers with different structures for better leveraging these two optimization targets. The final representation is constructed by adaptively aggregating these two types of representations.

We follow most rituals from our proposed VideoMAE work to train a vanilla Vision Transformer (ViT) as a video encoder for spatiotemporal modeling, as given in Figure 3 (a). VideoMAE conducts a video reconstruction task with highly masked video inputs, using an asymmetric encoder-decoder architecture. The used encoder and decoder are both ViTs. The channel number of the decoder is half of that of the encoder, with 4 blocks by default. Specifically, we divide the temporal strided downsampled video inputs into non-overlapping 3D patches and project them linearly into cube embeddings. Then we apply tube masking with notably high ratios (e.g. 90%) to these embeddings and input them into the asymmetric encoder-decoder architecture to perform the masked video modeling pretraining. To characterize spatiotemporal interaction globally, we employ joint space-time attention in ViT, making all visible tokens globally interact with each other. It is computationally tractable as only a few tokens are preserved for calculation.

1.2 Video-Language Contrastive Learning

We conduct both video/image-text contrastive learning and video captioning tasks for pretraining, as given in Figure 3 (b). For training efficiency, we build our multimodal structure based on the pretrained CLIP . Instead of directly employing a vanilla ViT, we use our proposed UniformerV2 as the video encoder for better and more efficient temporal modeling. Moreover, we adopt an extra transformer decoder for cross-modal learning. Specifically, we follow a typical align-before-fuse paradigm as given in . First, video and text are separately encoded. Then a contrastive loss is utilized to align the embedding space of video and text features. In the fusing stage, we apply a caption decoder as a cross-modality fuser, which uses cross attention for a captioning pretext. This align-before-fuse paradigm not only ensures the modalities can be aligned into the same single embedding space, which is beneficial for tasks like retrieval but also gifts the model with the ability to combine different modalities and can be beneficial for tasks like question answering. The introduction of a caption decoder both extends the potential of the original CLIP and improves the robustness of multimodality features.

2 Supervised Video Post-Pretraining

Empirically, action recognition acts well as a meta task in video downstream applications, widely validated in . Thus we train a masked video encoder and a multimodal one with supervised action classification separately as a post-pretraining step for better performance in diverse tasks. To promote the learning capacity of these encoders, we propose a unified video benchmark Kinetics-710 (K710, described in Section 4.1) for finetuning our video encoders.

Masked Video Encoder. We finetune the masked video encoder with 32 GPUs on K710. We adjust the learning rate linearly according to the base learning rate and batch size, lr=base learning rate×batch size256\textit{lr}=\textit{base learning rate}\times\frac{\textit{batch size}}{256}. We adopt DeepSpeedhttps://github.com/microsoft/DeepSpeed framework to save memory usage and speed up training. We set the base learning rate to 0.0010.001, the drop path rate to 0.20.2, the head dropout rate to 0.50.5, the repeated sampling to 2, the layer decay to 0.80.8, and trained for 4040 epochs.

Multimodal Video Encoder. We follow most of the training recipes in UniFormer . For the best result, we adopt CLIP-ViT as the backbone by default, due to its robust representation pretrained by vision-language contrastive learning. We insert the global UniBlocks in the last 4 layers of ViT-B/L to perform the multi-stage fusion. We set the base learning rate to 1e−51e-5, the repeated sampling to 1, the batch size to 512, and trained for 40 epochs. We adopt sparse sampling with a resolution of 224 for all the datasets. In post-pretraining, we use a UniformerV2 as the visual encoder and initialize additional parameters in a way that the output is identical to the original CLIP model which we find to be essential for good zero-shot performance. The video captioning module is a standard 6-layer transformer decoder with c=768c=768 followed by a two-layer MLP. Other setting leaves CLIP Large/14 untouched.

3 Cross-Model Interaction

To learn a unified video representation based on both video masked modeling and video-language contrastive learning, we conduct cross-representation learning with added cross-model attention modules, as shown in Figure 4.

Regarding optimizing both models at the same time is computing-intensive, we freeze both backbones except the classification layers and the query tokens in the multimodal video encoder, only updating newly added components. We add some elaborate learnable modules (cross-model attention) for aligning representations learned in different approaches. Cross-model attention (CMA) is formed by standard Multi-Head Cross Attention (MHCA) along with Feed-Forward Network (FFN). It employs intermediate tokens from a multimodal video encoder as keys and values while using these from a masked video encoder as queries. The new tokens computed from CMA are treated as a gradually aligned representation with that from the multimodal video encoder. This procedure mainly transfers multimodal knowledge to CMAs in the masked video encoder. One design exception is that for the last CMA module, its keys and values are from the tokens of the masked video encoder and the query is from the class token of the multimodal video encoder. Thus, the class token is updated based on tokens from the masked encoder. It transfers single-modal knowledge to CMA in the multimodal video encoder. From this perspective, the features in all stages of the masked video encoder and the ones in the final stage of the multimodal video encoder are enhanced to coordinate with each other, in the supervision by action recognition. Finally, we utilize a learnable linear combination to dynamically fuse the two prediction scores.

Experiments

We detail our experimental configurations first (Section 4.1), then we present the downstream performance of InternVideo on the proposed video understanding benchmark with three types of tasks (action understanding, video-language alignment, and open understanding) in Section 4.3.

General video foundation model pretraining requires data from various domains at scale. To achieve a data distribution with diversity, we employ 6 public datasets and our self-collected video clips as shown in Table 1.

We adopt a new customized kinetics action dataset Kinetics-710 for supervised training, both separate and joint ones. It has 650K videos with 710 unique action labels. It combines all the unique training data from Kinetics 400/600/700 . To avoid the training leak, some training data existing in the testing set from Kinetics of a specific version are abandoned.

The UnlabeledHybrid dataset is used for masked video pretraining, which is consist of Kinetics-710 , Something-Something V2 , AVA , WebVid2M , and our self-collected videos. For AVA, we cut the 15-minute training videos by 300 frames and get 21k video clips. We just randomly pick 250k videos from Self-collected videos and WebVid2M respectively. More details can be seen in Table. 1.

2 Implementations

With the initialization from CLIP, we post-pretrain our multi-modal model with WebVid2M, WebVid10M, and HowTo100M. Since the training corpus of video-text datasets is not as rich as CLIP-400M , we co-train the video model with image-text datasets, a subset of LAION-400M containing 100M image-text pairs. We alternate images and videos for each iteration. The batch size of video-text is 14,336, and the batch size of image-text is 86,016. We train for 400k steps on 128 NVIDIA A100 GPUs in 2 weeks, with a learning rate of 8×10−58\times 10^{-5}, weight decay of 0.2, cosine annealing schedule, and 4k warm-up steps.

2.2 Masked Video Training

We train the VideoMAE-Huge for 1200 epochs on the UnlabeledHybrid dataset with 64 80G-A100 GPUs. The model adapts the cosine annealing learning rate schedule and warmup 10% total epochs. The learning rate is set to 2.5e−42.5e-4. Only MultiScaleCrop is used for data augmentation.

2.3 Model Interaction

As shown in Figure 4, we freeze both backbones except the classification layers and the query tokens in the multimodal video encoder. To maintain the original output, we add tanh gating layers in the extra MHCA and FFN as in Flamingo , and the parameters in dynamic weighted sum are initialized as zero. We train the coordinated models with a batch size of 64, a learning rate of 5×1055\times 10^{5}, a weight decay of 0.001, a dropout rate of 0.9, and an EMA rate of 0.9999. Besides, we use a cosine annealing schedule for 5 epochs with 1 warmup epoch. All used data augmentations are the same as in UniFormerV2 .

3 Downstream Tasks

We conduct extensive experiments on a spectrum of downstream tasks to evaluate InternVideo. The employed tasks are of three categories that consider action understanding, video-language alignment, and open understanding. Since InternVideo contains masked video encoder specializing in spatiotemporal variation characterization and fused multi-modality video encoder, it can improve action understanding (Section 4.3.1) and video-language alignment (Section 4.3.2) tasks significantly. Its generalization brought by large-scale training data also enables its impressive zero-shot and open-set capabilities on the related tasks (Section 4.3.3). Even transferred to ego-centric tasks, InternVideo still gives an overwhelmingly favorable performance with simple heads . Details are given as follows.

Action Recognition. Actions derive spatiotemporal patterns. InternVideo aims to learn the representation of suitable spatiotemporal features, and the modeling of dynamical patterns. We evaluate InternVideo on 8 action recognition benchmarks, including popular Kinetics and Something-Something.

We evaluate VideoMAE and UniFormerV2 in InternVideo on Kinetics-400 , Kinetics-600 , Kinetics-700 , Something-in-Something-V1 , Something-in-Something-V2 , ActivityNet , HACS , and HMDB51 . We use the top-1 accuracy as a comparison indicator. In Table 2 and 3, InternVideo demonstrates exceedingly promising performance on all these action recognition benchmarks. Our InternVideo significantly surpasses previous SOTA methods on almost all benchmarks and matches the SOTA result on ActivityNet. The rised accuracy brought by extra fused model (InternVideo-D vs. InternVideo-T) demonstrates it is necessary to explore a broad technical roadmap as different lines benefits each other in performance.

Temporal Action Localization. This task (TAL) aims at localizing the start and end points of action clips from the entire untrimmed video with full observation. We evaluate our InternVideo on four classic TAL datasets: THUMOS-14 , ActivityNet-v1.3 , HACS Segment and FineAction . In line with the previous temporal action localization tasks, we use mean Average Precision (mAP) for quantitative evaluations. Average Precision (AP) is calculated for each action category, to evaluate the proposals on action categories. It is computed under different tIoU thresholds. We report the performance of the state-of-the-art TAL methods whose codes are publicly available, including ActionFormer method for THUMOS-14, ActivityNet-v1.3, and FineAction and TCANet method for HACS Segment.

We use ViT-H from our InternVideo as backbones for feature extraction. In our experiment, ViT-H models are pretrained from Hybrid datasets. As shown in Table 4, our InternVideo outperforms best than all the preview methods on these four TAL datasets. Note that, our InternVideo achieves huge improvements in temporal action localization, especially in fine-grained TAL datasets such as THUMOS-14 and FineAction.

Spatiotemporal Action Localization. This task (STAL) is to predict the frames and corresponding actions of people in video keyframes. We evaluate InternVideo on two classic STAL datasets AVA2.2 and AVA-Kinetics . In AVA2.2 , each video lasts 15 minutes, and it gives a keyframe every second. The annotation is provided for keyframes instead of all frames. Here we use a classic two-stage approach to handle this task. We apply a well-trained (on MS-COCO ) Mask-RCNN to detect humans on each keyframe, and the keyframe boxes are provided in the Alphaction project. In the second stage, centering around the key frame, a certain number of frames are extracted and fed into our video backbone. Similarly, in the training, we use the ground truth box for training , and the boxes predicted in the first stage for testing.

We used ViT-Huge in InternVideo for experiments. The specific results can be seen in Table 5. The classification head uses a simple linear head, achieving the SOTA performance on both datasets. Note using the ViT-H model, and training with the AVA-Kinetics dataset not only improves the overall mAP, but also significantly improves the mAP obtained by testing on AVA alone. It suggests the introduction of some Kinetics videos to AVA will improve the generalization of the model over AVA; on the other hand, observing the various distributions of the AVA dataset, we find that AVA presents a typical long-tailed distribution. The introduction of Kinetics video will alleviate this issue for better results. Due to the small number of models validated on the AVA-Kinetics dataset, only the results from the paperswithcode website are selected in the Table 5.

3.2 Video-Language Alignment Tasks

Video Retrieval. We evaluate InternVideoon the video retrieval task. Given a set of videos and related natural language captions, this task requires retrieving the matched video or caption corresponding to its inter-modality counterpart from candidates. We follow the common paradigm to capture visual and text semantics by a visual encoder fv(⋅)f_{v}(\cdot) and a text encoder ft(⋅)f_{t}(\cdot), then calculate the cross-modality similarity matrices as the retrieval guidance. We leverage the multimodal video encoder as fv(⋅)f_{v}(\cdot) and ft(⋅)f_{t}(\cdot) with pretrained ViT-L/14 as the basic CLIP architecture and finetune the entire model on each retrieval dataset. The training recipes and most of the hyperparameter settings follow CLIP4Clip , including training schedule, learning rate, batch size, video frames, maximum text length, etc. To boost model performance, we also adopt the dual softmax loss as the post-processing operation.

Our model is evaluated on six public benchmarks: MSR-VTT , MSVD , LSMDC , DiDeMo , ActivityNet , and VATEX , where we report the results on the standard split following previous works. We measure the retrieval results under the rank-1 (R@1) metric both on text-to-video and video-to-text tasks, which are shown in Table 6. Results show that our model significantly outperforms all previous methods by a large margin, showing the superiority of InternVideo on video-language related tasks. More detailed retrieval results, including rank-5 (R@5) and rank-10 (R@10), can be found in the supplementary materials.

Video Question Answering. To further demonstrate the vision-language capability of InternVideo , we evaluate InternVideo on video question answering (VQA). Given a video and question pair, VQA is to predict the answer to the question. Unlike the vanilla CLIP model without cross-modality fusion, our multimodal video encoder is able to capture the interaction between modalities with the proposed caption decoder. There are three potential ways to generate features required by the VQA classifier: concatenating the features of the video encoder and text encoder, utilizing the features of the caption decoder only, and concatenating all features from the video encoder, text encoder, and caption decoder. After comparison, we choose to use all three sources of features to boost the performance. The VQA classifier is a three-layer MLP.

We evaluate on three popular public benchmarks: MSR-VTT , MSVD , and TGIF . We mainly follow the practice in . The results are shown in Table 7 and our model outperforms all previous SOTA, which demonstrates the effectiveness of our cross-modality learner.

Visual Language Navigation. Visual-Language Navigation requires an agent to navigate in unknown photo-realistic environments based on its visual perceptions following natural language instructions. Navigation agents should be capable of capturing spatiotemporal information such as the relative motion of objects from navigation histories, especially when the agent navigates with short step size in continuous spaces. To verify the effectiveness of such ability of our model, we conduct our experiments on the VLN-CE benchmark , demanding the agent to function in a continuous environment.

We conduct our experiments using the method proposed in (CWP-HEP). The history-enhanced planner is a customized variant of HAMT which uses a concatenation of depth embedding and RGB embedding as the input embedding. Note that we don’t use the tryout controller here since VLN-CE setting allows sliding. This is a strong baseline that already outperforms the previous state-of-the-art method CWP-VLNBERT . In each decision loop, we collect the latest 16-frame observations to form panoramic navigation videos and then encode the video using ViT-L in InternVideo. The video embedding is concatenated with the RGB embedding and depth embedding as the final image embedding. For evaluation, we refer to for detailed metrics. InternVideo could improved our baseline from 50.2% to 52.9% in Success Rate(SR) (Table 8).

3.3 Video Open Understanding Tasks

Zero-shot Action Recognition. Zero-shot recognition is one of the extraordinary capabilities of the original CLIP model. With our designed multimodal video encoder, we can also achieve remarkable zero-shot action recognition performance without further optimization. We evaluate our model on Kinetics-400 dataset with 64.25% accuracy, which outperforms the previous SOTA 56.4% by a large margin. Zero-shot Video Retrieval. We compare InternVideo with CLIP on zero-shot text-to-video and video-to-text retrieval. For fair comparisons, we use the ViT-L/14 model with the pretrained weights https://github.com/openai/CLIP. of CLIP. Wise-finetuning and model ensemble are employed to further boost the model performance on zero-shot video retrieval. We empirically find that the optimal number of video frames for zero-shot retrieval is between 4 and 8, and the best-performed frame on each benchmark dataset is yielded via grid search. As illustrated in Table 9, InternVideo demonstrates superior retrieval ability across all six benchmark datasets. Besides, Florence used 900M image-text pair for pretraining and it achieved 37.6 R@1 text-to-video retrieval accuracy on MSR-VTT. In comparison, our model outperforms Florence by 4.1% with much less training data (14.35M video + 100M image v.s. 900M image). These results reveal the effectiveness of our method in learning the joint video-text feature space during pretraining.

Zero-shot Multiple Choice. Zero-shot multiple choice is another zero-shot task that can demonstrate the model’s generality. Multiple choice task aims to find the correct answer in the given choices, usually a small subset such as 5 words. We find that co-training with image-text pairs, wise-finetuning, and the ensemble is essential for the performance on zero-shot multiple choice. We report the zero-shot multiple-choice results in Table 10 on MSR-VTT and LSMDC datasets. We use the zero-shot performance as a handy indicator for generality in training, and the results show that our model is robust and effective.

Open-set Action Recognition. In open-set action recognition (OSAR), the model is required to recognize the known action samples from training classes and reject the unknown samples that are out of training classes. Compared with images, video actions are more challenging to be recognized in an OSR setting than images due to the uncertain temporal dynamics, and static bias of human actions . Our InternVideo generalizes well to unknown classes that are out of training classes and outperforms the existing method without any model calibration.

We use the ViT-H/16 model of InternVideo as a backbone, and finetune it with a simple linear classification head on UCF-101 training set. To enable InternVideo to “know unknown", we follow the method DEAR proposed in and formulate it as an uncertainty estimation problem by leveraging evidential deep learning (EDL), which provides a way to jointly formulate the multiclass classification and uncertainty modeling. Specifically, given a video as input, the Evidential Neural Network (ENN) head on top of a InternVideo backbone predicts the class-wise evidence, which formulates a Dirichlet distribution so that the multi-class probabilities and predictive uncertainty of the input can be determined. During the open-set inference, high-uncertainty videos can be regarded as unknown actions, while low-uncertainty videos are classified by the learned categorical probabilities.

InternVideo can not only recognize known action classes accurately but also identify the unknown. Table 11 reports the results of both closed-set (Closed Set Accuracy) and open-set (Open Set AUC) performance of InternVideo and other baselines. It shows that our InternVideo consistently and significantly outperforms other baselines on both two open-set datasets, where unknown samples are from HMDB-51 and MiT-v2 , respectively.

Concluding Remarks

In this paper, we propose a versatile and training-efficient video foundation model InternVideo. To our best knowledge, InternVideo is the first work to perform best among existing researches on all action understanding, video-language alignment, and video open understanding tasks. Compared with previous related work , it greatly lifts the generality of video foundation models to a new level, by achieving state-of-the-art performance on nearly 40 datasets covering 10 different tasks. The model exploits a unified video representation based on the cross-model learning between masked video learning (VideoMAE) and video-language contrastive modeling along with supervised training. Compared with previous foundation models, it is efficient in training. With simple ViT and its corresponding variants, we achieve generalized video representations with 64.5K GPU hours (A100-80G), while CoCa requires 245.76K TPU hours (v4). We validate such generalized spatiotemporal representation on a spectrum of applications. With simple task heads (even linear ones) and proper downstream adaption tuning, our video representation demonstrates record-breaking results in all used datasets. Even for zero-shot and open-set settings, our model spectrum still gives consistent and non-trivial performance increases, further proving its generalization and adaption.

Our study shows the effectiveness and feasibility of video foundation models instead of giving brand-new formulations or model designs. It focuses on the current popular video perception tasks and handles videos using clips. Its devise can hardly process long-term video tasks, as well as high-order ones, e.g. anticipating plots from the seen parts of a movie. Gaining the capacity to address these tasks is crucial to further push the generality of video representation learning.

2 Future Work

To further extend the generality of the video foundation models, we suppose embracing model coordination and cognition is necessary for its studies. Specifically, how to systematically coordinate foundation models trained from different modalities, pretraining tasks, and even varied architectures for a better representation remains open and challenging. There are multiple technical routes to address it, e.g. model distillation, unifying different pretraining objectives, and feature alignment, to name a few. By exploiting previously learned knowledge, we can accelerate video foundation model development sustainably.

In the long run, foundation models are expected to master cognitive capabilities more than perceivable ones. Considering its feasibility, we suppose one of its research trends is to achieve large-scale spatiotemporal analysis (long-term & big scene) from the foundational dynamic perception in the open world, leading to essential cognitive understanding. Moreover, it has raised a tide that combining foundation models with decision-making to form intelligent agents to explore new tasks. In this interaction, data collection and model training are also automated. The whole process enters a closed loop as the interactive results will adjust agent strategies and behaviors. Our initial experiments (Section 4.3.2) on vision-language navigation demonstrate the promising future of integrating video foundation models into Embodied AI.

Broader Impact

We give a video foundation model spectrum InternVideo. It is able to deliver the state-of-the-art performance on around 40 datasets, capable of action discrimination, video-language alignment, and open understanding. Besides of public data, we also exploit self-collected data from the Internet. The employed queries for gathering data are checked for ethic and legal issues and so are the curated data. The power consumption of training InternVideo is much lower than CoCa , only taking up 23.19 % of CoCa. For further impact studies, we need to explore the bias, risks, fairness, equality, and many more social topics.

References