InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, Conghui He, Ping Luo, Ziwei Liu, Yali Wang, Limin Wang, Yu Qiao

Introduction

Learning transferable video-text representations is both challenging and essential for video understanding in various real-world applications such as autonomous driving, intelligent surveillance, human-computer interaction, and visual searching. While multimodal contrastive learning using web-scale data has been successful in image-text representation, it remains underexplored in the video-language domain.

A key reason for this limited exploration is the lack of a high quality video-language dataset for pretraining at scale. Current research relies on datasets like HowTo100M , HD-VILA , and YT-Temporal , whose texts are generated using automatic speech recognition (ASR). Despite their large scale, these datasets often have low semantic correlations between the videos and corresponding textual descriptions . Empirical studies demonstrate that improving this correlation (e.g. aligning videos with subtitles to improve their matching) significantly benefits downstream tasks such as video retrieval and video question answering . Recent works have utilized WebVid10M , a dataset with higher-quality alt-texts, to address the low video-text correlation issue. However, its limited scale and dynamics hinder its use in current data and model scaling studies. Specifically, only 10M video-text pairs are provided, and the depicted scenes contain relatively few actions or activities.

We propose a large-scale video-centric dataset InternVid to address the challenge of scaling up video-language modeling while maintaining high video-text correspondence. Visual examples are given in Figure 1. Note the ASR transcripts barely depict visual elements in videos while the generated captions do. The dataset contains highly-correlated video-text pairs and includes over 7 million videos, totaling 760,000 hours and resulting in 234 million video clips, with various subsets for different needs. These videos cover 16 scenarios and around 6,000 motion descriptions. To improve video-text matching, we generate captions using a multiscale approach. In the coarse scale, we caption the middle frame of each video and use the description as the video caption. In the fine scale, we produce frame-by-frame captions and summarize them with a language model.

Leveraging InternVid, we scale a video-language transformer (ViT-L) in contrastive learning from a data perspective, and its experiments prove InternVid enables learning scalable video-text models. We introduce video masking to the model to accelerate the whole learning without compromising its effectiveness. The video and text encoders are initialized from the CLIP pretrained model with the same scale. With InternVid, we learn a video-text model for several epochs, achieving impressive zero-shot performance. Compared with previous Video CLIP variants, our proposed ViCLIP shows notable performance improvement, especially in zero-shot settings.

In addition to large-scale video-language contrastive pretraining, we discover its effectiveness in producing interleaved video-text data for learning a video-centric dialogue system like Flamingo , and advancing video generation. Since the text-annotated clips are extracted from videos, we naturally collect clips and their corresponding text based on the sampling locations. This results in approximately 7 million interleaved data pieces, suitable for instruction tuning as multi-turn video-centric dialogue. For video generation, we filter the core set and obtain 18 million video clips. Alongside WebVid-10M, InternVid can significantly improve a stable-diffusion based video generation model to new heights.

In summary, our contributions are threefold.

We introduce a new web-scale video-language dataset InternVid. This dataset, aimed at advancing video-related multimodal understanding and generation at scale, is created using a multi-scale video captioning approach powered by LLM, ensuring high-quality video-text data with minimal human intervention. InternVid has 7 million videos, corresponding to 234 million clips each with the generated captions. Spanning 16 scenes and about 6 thousand actions, the dataset includes computational features (video-text correlation and visual aesthetics) across the entirely of the dataset and gives way to diverse subsets to cater to varying training needs.

We learn a new video-language model, ViCLIP, which is trained on InternVid using ViT-L. It incorporates both constrastive learning and mask modeling techniques, allowing for efficient learning of transferrable video-language representation. This model achieves state-of-the-art zero-shot action recognition in Kinetics, scoring 75.7, 73.5, and 66.4 on K400, K600, and K700 with the average top1 and top5 accuracies, respectively. It also gets competitive performance on video retrieval, setting a new baseline for video-text understanding.

InternVid fosters the development of multimodal dialogue systems and text-to-video generation. The proposed ViCLIP learned on InternVid could serve as a vision backbone of video-centric dialogue systems, conducting tasks as action recognition, temporal understanding, reasoning, and creativity within an open-ended environment. Furthermore, we provide a subset, InternVid-Aesthetics, created using specific video-text relation and visual aesthetic filtering. This subset aids in generating high-resolution watermark-free videos. Utilizing InternVid-Aesthetics, both visual and quantitative outcomes of a simple text-to-video baseline can be noticeably enhanced (FVD: 705.3 -> 616.5).

Related Work

Vision-text data pairs are necessary to enable crossmodal learning. To learn vison-language representation effectively, these datasets should be large at scale and high at vision-text correlations. To this end, researches usually leverage existing web images with alt-text and videos with ASR transcriptions for scalable learning. With LAION-5B’s introduction , researchers now have access to hundreds or millions or billions of image-text pairs, opening up new avenues for research on large-scale image-language pretraining.

For video-centric multimodal datasets, HowTo100M collected instructional YouTube videos and exploited the corresponding ASR subtitles for learning joint representations. Zellers et al. and Xue et al. proposed YT-Temporal and HD-VILA for Audio-Visual-Language joint learning and high-resolution video crossmodal learning, respectively. On the other hand, Bain et al. found video-text alignment matters more than their quantities, so they produced WebVid where 10M videos with the corresponding alt-texts. This is frequently employed in recent video-language pretraining approaches . Similarly, based on CC3M, Nagrani et al. proposed VideoCC3M by transferring captions from image-text datasets to video ones. In this work, we target to present a large-scale video-language dataset with high-quality descriptions.

Video Understanding.

Pretraining large-scale video-text models and fine-tuning them for downstream tasks has become the norm in the video-language field . Early techniques used pretrained visual and language encoders to obtain offline video and text features, but recent methods highlight the advantages of end-to-end training. Common practices include two or three pretraining tasks, such as masked language modeling , video-text matching , video-text contrastive learning , masked video modeling , and video-text masked modeling .

In the multimodal video context, VIOLET combined masked language and video modeling, while All-in-one proposes a unified pretraining approach with a shared backbone, and LAVENDER unified tasks through masked language modeling. Despite their success in multimodal benchmarks, these methods’ reliance on limited video-text data hampers performance in video-only tasks like action recognition. Conversely, InternVideo and UMT combined masked modeling with crossmodal contrastive learning, leading to competitve performance in both video-only and video-language tasks. MERLOT Reserve exploited 20 million video-text-audio pairs for training joint video representations using contrastive matching, setting new standards in video recognition and visual commonsense reasoning. VALOR also employed different modality encoders for video, audio, and text processing, and introduces video-to-text and audio-to-text pretasks to improve vision-audio-language learning. To address modality entanglement in crossmodal learning, mPLUG-2 introduced a shared module across image, video, and text to encourage modality collaboration while reserving modality-specific modules for their differences. Similar to , VLAB adapted a CLIP-pretrained ViT to model spatiotemporal variations and blends it with CLIP ViT with cross attention for handling both images and videos.

InternVid: A Video-Centric Multimodal Dataset

A high-quality video-text dataset at scale is a premise to conduct large-scale video-language learning and associated tasks. We identify three crucial factors in constructing this dataset: substantial temporal dynamics, rich and diverse semantics, and strong video-text correlations. To ensure high temporal dynamics, we gather videos retrieved using action/activity-based query words. For rich and varied semantics, we not only crawl trending videos across various categories but also deliberately increase the proportion of data consciously collected from various countries and languages. To strengthen video-text correlations, we employ image captioning and language models to generate video descriptions from frame-specific annotations. Next, we elaborate the dataset construction process and discuss its statistics and characteristics.

We collect videos from YouTube considering the diversity and richness of its data, and its support for academic usage. Totally we obtain 7 million public YouTube videos with an average duration of 6.4 minutes, covering 16 topics. We ensure the uniqueness of our dataset by creating a database of YouTube video IDs and excluding any videos already present in publicly available datasets (released prior to April 2023). The data curation strategies are two-fold. On one hand, We select popular channels and the corresponding hot or high-rated videos from the categories e.g. news, gaming, etc., resulting in 2 million videos. On the other hand, we create a list of verbs related to actions/activities. With it, we also obtain 5.1 million videos by choosing the top retrieved ones.

We define around 6.1K action phrases from American Time Use Survey (ATUS), public video datasets, and text corpus. Then they are refined both manually and automatically. We employ actions from ATUS from 2017 to 2022 , merging them and removing the duplicates. For the referenced public video data, we leverage Kinetics , SomethingSomething series , UCF101 , and so on. This provides us with 1103 action labels. Moreover, we access several visual grounding corpus . A language model is employed to extract actions and their corresponding targets (if exist) to form phrases from the corpus, leading to 5001 actions with manual checking. Totally, we collect 6104 action queries for searching videos on YouTube.

Collection Strategies.

To ensure the quality of our dataset, we established specific crawling rules. We only collected videos that were between 10 seconds and 30 minutes in duration and had resolutions ranging from 360P to 720P. Videos with resolutions below 360P were excluded, and those above 720P were either downloaded in their 720P version or resized to 720P. In this process, we prioritize the highest available resolution. To provide a comprehensive mutimodal dataset, we gather videos along with their audio, subtitles, titles, and summaries. Captions for the videos were generated automatically using a video captioning pipeline described in Section 3.2.

In formation, the collected multimodal data contain videos V\mathbf{V}, their audios A\mathbf{A}, metadata (title Wtitle\mathbf{W}^{\text{title}}, video descriptions Wcontent\mathbf{W}^{\text{content}}, query words Wquery\mathbf{W}^{\text{query}}, tags Wtag\mathbf{W}^{\text{tag}}, etc), subtitles (user generated contents or auto-generated ones), and more. Each video V\mathbf{V} could be treated as a sequence of clips {Ci}i=1,2,...\{\mathbf{C}_{i}\}_{i=1,2,...}, and we can segment their corresponding audio as {Ai}i=1,2,...\{\mathbf{A}_{i}\}_{i=1,2,...} and ASR subtitles as {Wiasr}i=1,2,...\{\mathbf{W}_{i}^{\text{asr}}\}_{i=1,2,...}. For the metadata, we suppose clips share the same meta when they are sampled from the same video.

Trimming.

We segment videos (lasting an average of 5 minutes) into clips (for around 10 seconds) using scene variance. For starters, videos are cut into shorter ones based on their scene changes. We directly employ the corresponding filter in PySceneDetect https://github.com/Breakthrough/PySceneDetect with a threshold as 27. During this procedure, we also filter out clips in still or extreme dynamics (e.g. a browse of a photo gallery). After the filtering, we get total 234M video clips whose durations range from 2s to more than 30s.

2 Multiscale Video Captioning

To generate video captions that are scalable, rich, and diverse, we employ a multiscale method with two distinct captioning strategies, as depicted in Figure 2. On the finer scale, we simplify the video captioning process by concentrating on the common objects, actions, and scene descriptions within the video clip. We deliberately overlook intricate details such as subtle facial expressions & movements, and other nuanced elements. On the coarser scale, we adopt the single-frame bias assumption from and exclusively caption the central frame of the video. Given our focus on brief clips (around 10 seconds) filtered via scene segmentation, most videos predominantly display consistent objects without substantial appearance alterations. This circumvents the identity-preserving issue when dealing with videos from image perspectives. Technically, we employ the lightweight image captioning model Tag2Text for the finer scale, which describes videos at low fps in a frame-by-frame manner. These individual image captions are then synthesized into a comprehensive video description using a pretrained language model . At the coarser scale, we use BLIP2 to caption the middle frame of the clip.

3 Statistics and Features

We present the key statistics of InternVid with other popular video-language datasets in Table 1. More detailed ones are given below.

We collected videos from 16 popular categories with varying percentages, as illustrated in Figure 3. Unlike prior studies , we ensured diversity by selecting videos from countries with different languages instead of relying on a dominant language environment. The countries we sampled from include the UK, USA, Australia, Japan, Korea, China, Russia, and France, among others. In terms of duration, every video lasts 351.9s on average. Almost half (49%) of the videos are five minutes or less, while a quarter (26%) fall between five and ten minutes. Only 8% of the videos are over 20 minutes long. Among the curated videos, 85% were high-resolution (720P), while the remaining 15% had lower resolutions ranging from 360P to 720P. Although the lower-resolution videos may not perform as well as the high-resolution ones in content generation tasks, they can still be useful in video-language representation learning, provided that they have appropriate captions.

InternVid exhibits diverse clip durations and caption lengths in the segmented clip level. The aesthetic scores and clip-caption similarities are distributed uniformly, as shown in Figure 4. The majority of clips are 0-10 seconds in length, accounting for 85% of all clips (Figure 4: left). Approximately half of the clips have captions with 10-20 words, while one-third of the clip captions have fewer than 10 words. About 11% of clips have long captions with more than 20 words.

We measured the aesthetic scores of all clips using an open-source model . We uniformly sampled four frames of each clip, calculated their aesthetic scores, and took the maximum score as the video aesthetic score. For clip-caption similarity computation, we used a video-language model called UMT . We computed the cosine similarity between video embeddings and text embeddings, again using a uniform sampling of four frames for each clip. Most clips score around 4-6 in terms of aesthetics, accounting for approximately 75% of the data. For UMT-SIM, over 80% of the clips scored between 0.3-0.4, with the remaining clips scoring around 0.2-0.3 or 0.4-0.5. Based on these computed aesthetics and UMT-SIM scores, we can generate different versions of InternVid to meet various requirements.

Actionness.

In terms of actionness, the InternVid dataset contains about ten times more verbs than the WebVid10M dataset. To evaluate this, we used the NLTK toolkit to analyze the number of verbs in captions, focusing on extracting and tagging all unique verbs. We found a total of 109,485 verbs in the WebVid10M caption dataset, while the InternVid dataset contained 212,155 unique instances of verbs. While these counts may not be entirely accurate due to our simple counting method, we believe they provide a rough indication of the actionness of the two datasets.

4 Interleaved Video-Text Data Generation

Utilizing the created video captions, we can develop an integrated video-text dataset for in-context video learning, allowing video-based sequence models to perform new tasks without additional training. Previous research, such as Flamingo , Kosmos-1 , and Multimodal C4 , confirms that pretraining on the interleaved image-text sequences results in significant multimodal in-context abilities. To the best of our knowledge, a large-scale interleaved video-text dataset has not yet been established. Our work represents the initial step in creating and making it publicly available.

We create InternVid-ICL, containing 7.1M interleaved video-text data pairs. We propose three distinct methods for organizing clips and their captions:

∙\bullet Arrange clips and their descriptions sequentially based on their temporal order within the same video, as illustrated in Figure 5 (a).

∙\bullet Enhance diversity in interleaved video-text items by assigning ASR text to a used clip in addition to its caption, as demonstrated in Figure 5 (b).

∙\bullet Extend method 1 by concatenating two interleaved multimodal items, creating a video-centric dialogue simulating user queries involving multiple videos (Figure 5 (c)).

One visual example of these arrangements is provided in Table 9.

ViCLIP: Learning Video-Text Representation at Scale

Built upon CLIP , we make a simple video-text pretraining baseline ViCLIP. It consists of a video encoder (ViT) and a text encoder, as given in Figure 6. Both modules are initialized from the corresponding CLIP components. We update the native attention in the video encoder to spatiotemporal attention while maintaining other design elements. For efficient learning, we apply masking to videos in pre-training. The optimization target is the contrastive loss between input video and text embeddings.

Our video encoder uses a standard ViT with spatiotemporal attention. We apply random patch masking following MAE-based methods to the input videos. It significantly alleviates the computational burden. The used text encoder is also a transformer followed by .

Unmasked Video-Text Pretraining.

We feed all visual tokens into the video transformer instead of just the masked ones towards the end of the pretraining process. This helps bridge the gap between pretraining and downstream applications where the full video is used as input. We perform unmasked training for 0.5 epochs with a learning rate of 4e-6.

Training Objectives. Our framework optimizes video-text alignment. It minimizes InfoNCE loss using global video and text features, as

where fVf^{\mathbf{V}} and fTf^{\mathbf{T}} denote the learned video and text embeddings, respectively. sim(⋅)\text{sim}(\cdot) computes the cosine similarity between two features. τ\tau is the learnable temperature.

Implementation.

ViCLIP is learned with 64 NVIDIA A100 GPUs for 3 days with 50M video-text pairs. We introduce DeepSpeed and FlashAttention for training and inference acceleration.

We learn ViCLIP on five subsets of InternVid and evaluated its performance on popular video-related benchmarks using full-finetuned and zero-shot settings. We sample subsets InternVid-10M, InternVid-50M, and InternVid-200M randomly. For InternVid-10M-DIV, we prioritize to sample clips from different videos first, then we sample clips with varying probabilities according to the video length where they are extracted. The longer their source video is, the lower chance they are sampled. For InternVid-10M-FLT, we employ the sampling strategy of InternVid-10M-DIV and select clips with UMT-SIM scores ranking among the top 30% to ensure high quality.

1 Transferable Video Representation Performance

Action Recognition. In addition to OpenAI’s CLIP-L (CLIP400M ) and LAION (DataComp-1B ), we also include EVA-CLIP-L/14 and EVA-CLIP-E/14 for comparison. More experimental settings are given in App. E.1.

Zero-Shot. Table 3 shows that when trained on InternVid-10M-FLT, ViCLIP outperforms all other methods, including EVA-CLIP-E. This result validates InternVid’s effectiveness in learning video-text embeddings. Note that ViCLIP with InternVid-10M-FLT sets new records on zero-shot action recognition in Kinetics 400/600/700, demonstrating a significant performance boost compared to ViCLIP with WebVid10M or other models. Moreover, ViCLIP trained on InternVid-10M-FLT exceeds its performance on InternVid-200M. Normally, we would expect the model trained on InternVid-200M to perform better than those on -10M-DIV or -FLT, given that the latter two subsets derive from the former. Unless this discrepancy results from improper learning, we conjecture that false negative samples could severely impede video-text contrastive learning if we don’t purposefully reduce the number of clips taken from the same video. Specifically, we hypothesize that clips from the same video share similar representations and captions. Contrastive learning, however, assumes these clips to be different. This situation also undermines the significance of using a large batch size in current training since it increases the probability of encountering more false negatives. We believe this assumption is applicable to other video tasks as well and plan to explore this further in the future.

Fine-tuned. In Table 4, note when comparing ViCLIP trained on InternVid with image CLIP models or ViCLIP trained with WebVid, there is a clear increase in accuracy. Unlike the zero-shot results, when ViCLIP is pretrained with a larger number (200M) of video-text data pairs, it achieves higher accuracy in fine-tuned recognition tasks (87.9% in K400 and 73.6% in SthSthV2) compared to when pretrained (86.8% in K400 and 71.2% in SthSthV2) with fewer data (10M). This suggests that InternVid provides greater benefits for fine-tuned action-related tasks. The decrease in performance of ViCLIP with WebVid highlights the importance of addressing the distribution gap between WebVid and the action videos used for evaluation, emphasizing the need to collect videos with evident temporal dynamics.

Video-Text Retrieval. We evaluate the video retrieval performance of baselines and ViCLIP using different pretraining datasets on five popular benchmarks , as shown in Table 5 and 6. We uniformly sample eight frames from the input videos. For the CLIP models from OpenAI and LAION , we utilize their officially released ViT-L models and extract video embeddings by averaging the computed frame-wise image embeddings. Our ViCLIP directly predicts video embeddings. For evaluating retrieval performance, we report R@1 scores for both text-to-video (t2v) and video-to-text (v2t) tasks in 5 and 6.

Both Table 5 and 6 demonstrate that video-language pretraining is crucial for enhancing fine-tuned and zero-shot retrieval performance. This point is substantiated by the comparison between CLIP and ViCLIP using InternVid-50M. Table 5 exhibits a boost of nearly 4-10 points across different benchmarks in the zero-shot setting. Meanwhile, Table 6 shows an increase of approximately 10 points across all R@1 scores in the fine-tuned setting.

Zero-Shot. Table 5 reveals InternVid-10M outperforms WebVid when employing the same method, ViCLIP, with an average increase of 6.3% in R@1 across nearly all benchmarks. This improvement can be further amplified by diversifying the training clips used, as InternVid-10M-DIV and -FLT surpass WebVid on ViCLIP with gains in R@1 of 14.0% and 17.1%, respectively. These results underline, once again, the effectiveness of the correspondence between our generated video captions and their corresponding videos. Comparing CLIP4Clip using HowTo100M with ViCLIP using WebVid10M or InternVid-10M shows that the correlation between video and text influences performance more significantly than their quantity. Moreover, the zero-shot performance demonstrates that the video-text representation learned using InternVid is transferable. This claim is supported by its superior performance across multiple video retrieval benchmarks.

Fine-Tuned. Table 6 exhibits a noticeable improvement when transitioning from InternVid-10M to WebVid10M while using ViCLIP for both t2v and v2t retrieval across almost all datasets. On average, there is a 3.7% increase in t2v R@1 across all benchmarks, with particularly significant rise observed in ActivityNet (an increase of over 11.9%). However, ViCLIP using WebVid10M yields better v2t R@1 scores than when using InternVid-10M (81.2 vs. 80.0). We believe this does not alter the overall trend that InternVid-10M generally provides more advantage to ViCLIP than WebVid10M does.

The benefits of used video data become even more apparent when comparing InternVid-10M-DIV or InternVid-10M-FLT with WebVid10M. Their overall increases are 5.8% and 5.1%, respectively. Despite these improvements, issues related to data diversity persist.

Data Scaling and Issues. Figure 7 and 8 illustrate how ViCLIP’s performance changes in zero-shot and fine-tuning settings when varying the scale of InternVid. In both scenarios, increasing the data scale results in significant increases in performance. As shown in Figure 7, ViCLIP’s discriminative ability linearly increases with the increasing volume of training videos used (10M →\rightarrow 200M). Meanwhile, Figure 8 shows that the retrieval performance increase becomes marginal when scaling the training data beyond 50M. It’s vital to note our model is trained using only contrastive loss without employing popular designs such as matching head and its corresponding loss. Consequently, this retrieval result doesn’t allow for any definitive conclusions about whether there exists a turning point after which scaling up the training videos becomes less beneficial currently. More explorations are necessary in these retrieval experiments. However, these findings generally suggest that enhancing the scale of pretraining data can improve the transferability of the learned representation.

2 Text-to-Video Generation

Our InternVid dataset improves existing text-to-video generation models by providing video-text pairs with high correspondence. To establish a video generation baseline, we extend spatiotemporal modeling on the latent space of an open-source text-to-image diffusion model . We train the video generation approach with two settings: one using WebVid10M , and the other using InternVid-Aesthetics-18M in addition to WebVid10M . InternVid-Aesthetics-18M is a subset of InternVid consisting of clips with an aesthetic score of at least 4. Quantitative (Table 7) and qualitative (Figure 18) evaluations demonstrate the effectiveness of InternVid in video generation tasks. To evaluate our models quantitatively, we perform zero-shot text-to-video experiments and randomly sample 2,020 videos from the UCF-101 dataset and 2,990 videos from the MSRVTT dataset. Following the protocols in , we report CLIPSIM, IS, FID, and FVD metrics.

In Table 7, we observe that our t2v baseline trained on WebVid10M performs poorly in terms of IS, FID, and CLIPSIM when compared to other approaches. However, with the addition of InternVid-Aesthetics-18M, our t2v baseline demonstrates significant improvements in these metrics and outperforms other methods by a considerable margin. In Figure 18, we observe that the text-to-video (t2v) baseline using both WebVid10M and InternVid-Aesthetics-18M significantly outperforms other methods in terms of visual quality and temporal coherence. It is worth noting that the t2v baseline using InternVid does not contain watermarks, which is a data bias in WebVid10M. These results demonstrate the potential of InternVid for high-quality video generation.

3 Video-Centric Dialogue System

Inspired by recent vision-centric dialogue systems , we integrate our pretrained ViCLIP (with InternVid) into VideoChat to show how our data and model can empower multimodal dialogue methods with effective video modeling capability. In implementation, we inherit nearly all designs of VideoChat-Embed, just replacing its visual encoder with our ViCLIP (trained on InternVid). We evaluate VideoChat-ViCLIP in spatial understanding (Figure 10), action recognition (Figure 11), temporal understanding (Figure 12), video reasoning (Figure 13), and video creative (Figure 14) tasks. Our qualitative evaluations demonstrate its decent video-to-text capabilities, suggesting promising potential for improving video captioning further.

In terms of quantitative comparison, as shown in Table 8, VideoChat-ViCLIP significantly outperforms the vanilla VideoChat (using Eva-g as the vision encoder) and other systems across all evaluation aspects of the quantitative video conversation evaluation framework in . Specifically, the model shows remarkable improvements in the correctness of information (from 2.23 to 2.86), contextual understanding (from 2.53 to 3.08), and temporal understanding (from 1.94 to 2.36). The average score also increases from 2.29 to 2.64, showing an overall performance gain.

Conclusion

Our dataset, InternVid, is designed for multimodal research (both understanding and generation) focused on videos. It consists of over 230 million video clips sourced from 7 million high-resolution (720P) YouTube videos. We use existing models with a multiscale approach to generate clip-level descriptions. Our studies confirm the efficacy of captions, and the large volume of video-text data enables crossmodal learning and text-to-video generation at scale. By training with our data, we develop a video-text representation baseline ViCLIP using ViT-L and analyze briefly how the data scale affects learned crossmodal embeddings. In addition to perception tasks, we show that InternVid improves text-to-video generation performance when using a subset of clips based on their aesthetic scores. With its data, annotations, metadata, and computed scores, we believe InternVid can fuel a variety of studies and applications.

Appendix A Data Availability Statement

We are committed to maintaining transparency and compliance in our data collection and sharing methods. In accordance with these principles, please note the following:

Publicly Available Data: The data utilized in our studies is publicly available. We do not use any exclusive or private data sources.

Data Sharing Policy: Our data sharing policy builds upon the precedent set by prior works like Kinetics, HD-VILA, and others. Instead of providing the original raw data, we only supply the YouTube video IDs necessary for downloading the respective content.

Usage Rights: The data released by us is intended exclusively for research purposes. Any potential commercial usage is not sanctioned under this agreement.

Compliance with YouTube Policies: Our data collection and release practices are strictly in accord with YouTube’s data privacy policies. We ensure that no user data or privacy rights are violated during the process.

Data Licence: We employ the protocol of CC BY 4.0.

Appendix B Limitations & Societal Impact

All video data used in our research are downloaded from YouTube using Safe for Work (SFW) queries and channels. To ensure appropriate content, we employ a simple NSFW filter: a binary classifier designed to recognize and exclude non-ethical videos. For privacy considerations and in respect of data sharing practices, we share only the YouTube ID of the videos, similar to previous academic works. This approach aligns with YouTube’s data protocols and ensures no violation of privacy or data usage rules. Despite these precautions, our work has some limitations, primarily related to data diversity and representativeness. Although YouTube is an extensive source encompassing a wide range of video categories, certain specific types of footage may be excluded or scarcely collected, including: public area surveillance, sports competitions, movies, documentaries, etc. The exclusion of such categories is often due to copyright restrictions or other limits imposed by the platform. Therefore, while our dataset provides a broad view of everyday video content, its coverage does not extend to every possible category or type of video. These limitations should be taken into account when considering the generalizability of our results across all types of video data.

Appendix C More Statistics in InternVid

InternVid contains way more verbs than the WebVid10M. We used NLTK toolkit to analyze the number of verbs in captions, focusing on tagging all unique verbs. We found a total of 109,485 verbs in the WebVid10M, while InternVid contained 212,155 ones. While the counts may not be that accurate due to our simple counting, we believe they provide a rough indication of the actionness of the two datasets.

Video Caption and Transcript Distribution.

To analyze the word distribution of our generated captions and multilingual (ASR) transcripts, we compute their distributions. The resulting word distribution of the captions is presented in Figure 15, which includes objects (tv, car, door, plant, etc.), attributes (green, young, large, long, etc.), locations (middle, behind, south, next, etc.), scenes (room, stage, kitchen, office, etc.), actions/events (walking, eating, cutting, holding, etc.), and more.

We also include four word distributions of different languages in Figure 16, reflecting trends in different countries and offering potential data customization along with the provided metadata.

Appendix D InternVid-ICL: Interleaved Video-Text for In-Context Video Learning

As given in the paper, we provide examples video+text interleaved entries for in-cntext learning as Flamingo. Table 9 gives an example about format (a): arrange clips and their descriptions sequentially based on their temporal order within the same video. Note the videos are randomly dropped with a probability (0.3) for constructing richer text context compared with the original video-text pair combinations in sequential.

Appendix E Implementation Details

In the zero-shot action recognition, we sample 8 frames in each video. Following the settings in CLIP and EVA-CLIP, we report the mean of top-1 and top-5 accuracy for Kinetics-400 / -600 / -700. In Section 4.1, we show ViCLIP learnt on WebVid or InternVid is an effective zero-shot action recognition model.

In the full fine-tuned setting, we conduct two experiments with two receipts. In Table 4, for the experiments where the training data excluded K710, we followed the common practice of finetuning the pretrained ViCLIP with the training data from the evaluation dataset. On the other hand, for the experiments where the training data included K710, we adopted a training trick inspired by . We first finetuned the pretrained ViCLIP with K710 , and then proceeded with the common supervised finetuning setting. By incorporating the supervised finetuning with K710, ViCLIP demonstrated better performance in the fine-tuned tasks compared to experiments that did not include K710.

Video Retrieval.

In the full-finetuning setting, we tune the pretrained ViCLIP with not only video-text contrastive loss but also video-text matching loss on the training data of the evaluated benchmarks. During both training and testing, we sample 12 frames. Detailed hyper-parameters are given in Table 10. In the zero-shot setting, we sample only 8 frames for evaluations.

E.2 Video Generation Baseline

We used the spatiotemporal modeling approach from and built our text-to-video generation baseline on the work of . Our approach consists of a U-Net with a transformer that models its latents, using interleaved spatiotemporal attention (ST-Attn), cross-attention for visual-text, a feed-forward network (FFN), and temporal attention (T-Attn), as illustrated in Figure 17. To adapt the 2D convolutional layers in to 3D, we extended 3×33\times 3 kernels into 1×3×31\times 3\times 3 ones. We also extended the original spatial attentions to spatiotemporal ones. We initialized our baseline using all text-to-image diffusion model parameters, while the newly added temporal attention layers used default parameters.

For the ST-Attn implementation, we used frame embeddings from the U-Net encoder instead of video embeddings as in . We concatenated the embeddings of the previous and current frame for values and keys in attention, while using the current frame embedding alone as queries. The rest of the implementation remained the same as the original.

To evaluate our text-to-video model, we conducted zero-shot experiments on the UCF-101 and MSRVTT datasets, following the method from . For UCF-101, we used the class names as text prompts and generated 20 samples per class (total of 2,020 videos). For MSRVTT, we randomly selected one caption per video from the official test set (total of 2,990 videos). To ensure a fair comparison, we used the official implementation of VideoCrafter and VideoFusion to generate the same number of videos with the same text prompts. During video sampling and evaluation, we generated 16 frames per video.

We assess the overall quality of the synthesized results on UCF-101 using framewise-FID, FVD, and Inception Score (IS), and evaluate the text-video semantic similarity on MSRVTT using clip similarity (CLIPSIM). For framewise-FID and IS, we use the pretrained Inceptionv3 network weights as our image encoder. For FVD, we use the pretrained InceptionI3d model and followed the TATS method . To compute CLIPSIM, we calculate the clip text-image similarity for each frame with respect to the given text prompts and computed the average score. We use the ViT-B-32 clip model as the backbone, consistent with previous work .

Appendix F More Results

To further validate the effectiveness of our proposed captioning method, we establish a video caption baseline using the video multimodal model VideoChat for comparison. We input the video clip into the model with the prompt "Please describe the content in the given video." and apply it to InternVid-10M-FLT, resulting in 10 million new captions generated by VideoChat. Subsequently, we train two versions of ViCLIP-Base using InternVid-10M-FLT, each version trained with one of the two types of captions.

Table 11 demonstrates that ViCLIP-B trained using our captions outperforms the version trained using captions from VideoChat in both video retrieval (MSR-VTT) and action recognition (K400/600/700). These results are particularly noteworthy considering that the only difference in training lies in the captions generated by the two different approaches. Therefore, these findings further confirm the superior performance of our proposed captioning method compared to the baseline VideoChat.

F.2 Text-to-Video Generation

In Figure 18, we observe that the t2v baseline using both WebVid10M and InternVid-Aes-18M significantly outperforms others in visual quality and temporal coherence. Note that the t2v baseline using InternVid does not contain watermarks, which is a data bias in WebVid10M.

References