Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, Sergey Tulyakov
Introduction
We enter an era where the size of computing and data are indispensable for large-scale multimodal learning. Most breakthroughs are achieved by large-scale computing infrastructure, large-scale models, and large-scale data. Due to these integral components, we have powerful text-to-image and image-to-text models . Scaling the model size or the compute is challenging and expensive; however, it requires a finite amount of engineering time. Scaling the data is relatively more challenging, as it takes time for a human to analyze each sample.
Especially, compared to image-text pairs , video-text pairs are even harder to obtain. First, annotating videos is more time-consuming, as an annotator needs to watch the entire video before labeling. Second, videos often contain multiple scenes stitched together and consist of temporally varying content. Finally, meta-information, such as subtitles, video description, and voice-over, is often too broad or not correctly aligned in time or cannot precisely describe a video. For example, several 100M-scale datasets, such as HD-VILA-100M and HowTo100M , are annotated by automatic speech recognition (ASR). However, as shown in Figure 1, the subtitles usually fail to include the main content and action presented in the video. This limits the value of such datasets for multimodal training. We summarize the datasets available to the community in Table 1. Some are low-resolution, some are annotated by ASR, some contain data from a limited domain, some are small-scale, and some offer short captions.
In this work, we present a large-scale dataset containing 70M video clips with caption annotations. It includes high-resolution videos from an open domain with rich captions averaging 13.2 words per caption. While manually annotating 70M videos is prohibitively expensive, we opt for automatic annotation. Our key insight is that a video typically comes with information from several modalities that can assist automatic captioning. This includes the title, description, subtitles of the video, individual static frames, and the video itself. The value of this data cannot be fully maximized when only partially used. In comparison, we propose to utilize different combinations of multimodal data as inputs to various cross-modality captioning models. To substantiate this idea, we conduct a numerical analysis based on a human evaluation (the details are provided in Appendix B.3). If we use multiple cross-modality models to caption some video samples and evaluate the results by showing them to humans, we see that there is no single model able to generate good captions for more than of videos. However, if we jointly collect all the captions from different models, we observe that of videos can be annotated with at least one good caption.
To establish the dataset with this mindset, we begin by using 3.8M high-resolution long videos collected from HD-VILA-100M and process them through the following three steps. First, we design a semantics-aware video splitting algorithm to cut long videos into semantically consistent clips while striking the balance between semantics coherence and the duration of the video clips. Second, we use a range of cross-modality teacher models, including image captioning models and image/video visual-question-answering (VQA) models with additional text inputs, such as video description and subtitles, to predict several candidate captions for a clip. Lastly, we collect a 100K video subset, where human annotators act as an oracle to select the best caption for each video. We use this dataset to finetune a fine-grained video-to-text retrieval model which is then applied to the whole dataset to select the most precise caption as the annotation. Running multiple teacher models is computationally expensive and time-consuming. To pursue efficient video captioning at scale in the future, we train a student model to distill the knowledge from the teachers. The student model adopts a two-branch architecture which can take both visual and textual inputs to benefit the captioning from multimodal information.
Extensive experiments demonstrate that pretraining with the proposed Panda-70MWe call our dataset Panda, drawing an analogy to Panda Po, who learns from multiple martial arts teachers. can benefit several downstream tasks, including video captioning, video and text retrieval, and text-to-video generation. We also show that training a student model in a knowledge distillation manner facilitates learning a strong student model which can outperform any teacher model by more than preference ratio as in Table 3, where the performance can be further enhanced by additional text inputs, like video description and subtitles.
Related Work
Vision-Language Datasets. Training with millions or even billions of image-text pairs has been shown to be effective in learning powerful image foundation models . With this work, our goal is to build a large video-language dataset containing rich captions. We compare related datasets in Table 1. Several precedent video-language datasets contain data tackling various tasks, such as action recognition, video understanding, VQA, and retrieval. However, manually annotating data is costly and limits the scale of such datasets (typically they contain less than 120K samples). To alleviate the lack of data, the works of propose to automatically annotate data with subtitles, generated by ASR. While this approach significantly increases the dataset scale reaching 100M of samples, the subtitles, unfortunately, do not precisely describe the main video content, as shown in Figure 1. In comparison, in this work, we propose an automatic captioning pipeline with the inputs of multimodal data that enables us to scale up the dataset of high-quality video-caption pairs to a 70M scale.
Vision-Language Models learn the correlation between visual data (images or videos) and linguistic signals (words or sentences) and can be applied to several downstream applications, including text-driven image or video generation , captioning , VQA and retrieval . We utilize several vision-language models for the annotation of Panda-70M. BLIP-2 introduces an efficient vision-language pretraining that can facilitate image captioning. We use BLIP-2 as one of the teachers and input a randomly sampled video frame for captioning. MiniGPT-4 is an image VQA model that learns a projection layer to align a large language model (LLM) and a visual encoder. In addition to a video frame, we also input a prompt with extra text information, such as video description and subtitles, and ask the model to summarize all multimodal inputs. For the video modality, Video-LLaMA and VideoChat are both video VQA models and learn to extract LLM-compatible visual embeddings. We use both models and ask them to caption a video with prompt input. Besides, Unmasked Teacher is a video foundation model which can facilitate video understanding. We finetune it to implement fine-grained retrieval and use it to select the more precise caption as the annotation.
Video Annotation through Multi-modal Models. With the aforementioned development on vision-language models, some concurrent works also leverage these models for video captioning. VideoFactory employs BLIP-2 to caption video clips. However, as reported in Appendix B.3, the performance of a single BLIP-2 model is suboptimal. More similar to our captioning pipeline, InternVid and Stable Video Diffusion also use multiple captioning models which are followed by an LLM for summarization. In practice, we found the LLM would propagate errors from noisy outputs of vision-language models.
Methodology
To build Panda-70M, we utilize 3.8M high-resolution long videos collected from HD-VILA-100M . We then split them into 70.8M semantically coherent clips as described in Section 3.1. Section 3.2 shows how multiple cross-modality teacher models are used to generate a set of candidate caption annotations. Next, we finetune a fine-grained retrieval model to select the most accurate caption as detailed in Section 3.3. Finally, in Section 3.4, we describe our approach to training a student captioning model using Panda-70M. The high-level view of our approach is shown in Figure 2.
A desired video sample in a video-captioning dataset should have two somewhat contradictory characteristics. On the one hand, the video should be semantically consistent, so the video samples can better benefit the downstream tasks, such as action recognition, and the caption can also more accurately express its semantics content without ambiguity. On the other hand, the video cannot be too short or fragmentary to contain meaningful motion content, which is beneficial to tasks, like video generation.
To achieve both goals, we design a two-stage semantics-aware splitting algorithm to cut a long video into semantically coherent clips. In the first stage, we split the video based on shot boundary detection , as the semantics often change when a new scene starts. In the second stage, we stitch adjacent clips if they are incorrectly separated by the first stage, ensuring the videos do not end up being too short. To do so, we use ImageBind to extract embeddings of video frames and merge the adjacent clips if the frame embeddings from two clips are similar. We also implement additional procedures to handle: 1) long videos without any cut-scenes, 2) videos using complex transitions, such as fade-in and fade-out effects, which are not usually detected as cut-scenes, and 3) removal of redundant clips to increase the diversity of the dataset. More details of the splitting algorithm are in Appendix A. Notably, while our dataset focuses on fine-grained video-text pairs with consistent semantics, users can still acquire long videos with multiple cut-scenes by concatenating consecutive clips and captions, as these clips are split from the same long video.
To quantitatively verify the semantic consistency of a video clip, we introduce Max Running LPIPS, which highlights the most significant perceptual change within a video clip. Formally, given an -second video clip, we subsample the video frames each second and denote the keyframes as . The Max Running LPIPS is formulated as:
where LPIPS is the perceptual similarity of two images. As in Table 2, our splitting achieves a better semantics consistency than the splitting based on the alignment of subtitles sentences , while maintaining longer video length than the vanilla shot boundary detection .
2 Captioning with Cross-Modality Teachers
Videos in HD-VILA-100M contain rich multimodal information beneficial for captioning. Specifically, besides the video itself, there are also useful texts (e.g., video title, description, and subtitles) and images (e.g., individual video frames). Driven by this insight, we propose to use several captioning models with the inputs of different modalities.
We start with a large pool including 31 captioning models. The introduction of the model pool is in Appendix B.1. Since running the inference of all models on 70M video clips is computationally expensive, we construct a short list of eight well-performing models based on a user study. The list is shown in the y-axis of Figure 3. More details of this process are in Appendix B.3. Briefly, the models are composed of five base models with different pretraining weights and input information. The five base models include Video-LLaMA (video VQA), VideoChat (video VQA), VideoChat Text (natural language model which textualizes the video content), BLIP-2 (image captioning), and MiniGPT-4 (image VQA). To implement video captioning by cross-modality teacher models, we formulate distinct captioning processes tailored to each modality. For example, for the VQA models, in addition to visual data, we also input a prompt with additional text information and ask the models to summarize all multimodal inputs into one sentence. Details on the captioning process of each teacher model are described in Appendix B.2.
We hypothesize that teacher models using different modality data perform well on different kinds of videos. For example, video models can perform better on videos with complex dynamics due to the additional modules to handle temporal information. On the other hand, image models can accurately caption the videos with rare and uncommon objects, since they were trained using large-scale datasets of image-text pairs . Finally, for videos that are visually hard to understand, VQA models have leverage as they can employ additional textual clues.
This hypothesis can be supported by a numerical evaluation. Specifically, we conduct a user study where the participants are asked to select the best caption from eight candidates. We plot the selective rate of each teacher model in Figure 3 (blue bars). The results show that the best captions are generated by different teacher models. Moreover, the highest selective rate of an individual teacher model (i.e., BLIP-2 with opt6.7b ) is only . This fact expresses the limited captioning capability of a single model on a wide variety of videos.
3 Fine-grained Video-to-Text Retrieval
Given multiple candidate captions for a video, we seek the one that best aligns with the video content. An intuitive idea is to use the available generic video-to-text retrieval models to pick such a caption. Unfortunately, we find that they usually fail to pick the optimal result. One reason is that generic models are trained using contrastive learning objectives and learn to distinguish one sample from other completely unrelated samplesNegative samples for contrastive learning are usually randomly sampled from the within-batch data and show no association to the anchor.. In contrast, in our case, all candidate captions are highly relevant to the video sample and require the model to discern subtle distinctions within each caption for optimal performance.
To tailor the retrieval model to our “fine-grained” retrieval scenario, we collect a subset of 100K videos, for which human annotators select the caption containing the most correct and detailed information about the main content of the video. We then finetune Unmasked Teacher (UMT) on this dataset. We implement hard negative mining for contrastive loss, where the seven captions not selected by annotators compose the hard negative samples and are assigned a larger training weight. We describe the details of the dataset collection and finetuning of UMT in Appendix C.1 and C.2 respectively.
We quantitatively evaluate the retrieval performance of UMTs with and without finetuning on the validation set. The experiments indicate that a finetuned UMT can achieve R@1 accuracy which significantly outperforms a pretrained UMT which has R@1. Notably, we conducted a human agreement evaluation by asking two other persons to re-perform the annotation and comparing the results with the original annotations. The average human agreement score is only R@1 showing that the task is subjective when more than one caption is equally good. Alternatively, if we consider the captions selected by any of the three persons as good captions (i.e., a video might have multiple good captions), UMT achieves R@1. Besides, in Figure 3, we show that a finetuned UMT (green bars) can select the captions distributed similarly to human-selected captions (blue bars). We run the finetuned UMT on the whole dataset to select the best caption as the annotation as elaborated in Appendix C.3.
4 Multimodal Student Captioning Model
While the aforementioned captioning pipeline can generate promising captions, the heavy computational demands hinder its capability to expand the dataset to an even larger scale. Indeed, one needs to run different models to annotate a single video clip. To deal with this problem, we learn a student captioning model on Panda-70M to distill the knowledge from multiple teacher models.
As shown in Figure 4, the student model includes visual and text branches, leveraging multimodal inputs. For the vision branch, we use the same architecture as Video-LLaMA to extract LLM-compatible video representation. For the text branch, a straightforward design is to directly input text embedding into the LLM. However, this will lead to two problems: first, the text prompt with video description and subtitles can be too long, dominating the decision of the LLM and burdening heavy computation; second, the information from the description and subtitles is often noisy and not necessary to align with the content of the video. To tackle this, we add a text Q-former to extract the text representation with fixed length and better bridge the video and text representations. The Q-former has the same architecture as the Query Transformer in BLIP-2 . During training, we block the gradient propagation from the text branch to the vision branch and train the visual encoder only based on the video input. More details about the architecture and training of the student model are in Appendix D.
Experiments
We visualize the samples of Panda-70M in Appendix E. To quantitatively evaluate the effectiveness of Panda-70M, we test its pretraining performance on three downstream applications: video captioning in Section 4.1, video and text retrieval in Section 4.2, and video generation in Section 4.3. The training details of the downstream models adhere to the official codebases unless explicitly specified.
Experiment setup. To evaluate the performance of video captioning, we use Video-LLaMA with the vision branch only as the base model. We compare two pretraining weights: the official weight, which is jointly trained on 2.5M video-text pairs and 595K image-text pairs , and the weight trained on our Panda-2M from scratch. Panda-2M is a randomly sampled subset of Panda-70M and shares the same amount of training samples as the official weight. We also train our student model with both video and text branches on complete Panda-70M for better captioning performance. For all models, we use the same backbone, using Vicuna-7B as the large-language model, ViT and Q-Former as the video encoder, and the linear projection layer from MiniGPT-4 . For Panda-2M pretraining, we only use the video and caption data without using other textual information for a fair comparison. For the student model, in addition to the video, we also randomly input the metadata and subtitles into the model during training.
Downstream datasets and evaluation metrics. We test zero-shot video captioning on two benchmarks: MSR-VTT and MSVD . MSR-VTT contains 10K videos with 20 manually annotated captions for each video; we report the results on the 2,990 testing split. MSVD consists of 1,970 videos with a total of 80K descriptions; we report the numbers on the 670 testing videos. Note that we do not use any training or validation videos from the downstream datasets. To quantitatively evaluate the quality of output captions, we follow the common protocols and report BLEU-4 , ROGUE-L , METEOR , and CIDEr . All the metrics are computed using the pycocoevalcap package. We also compute BERTScore to evaluate the contextual similarity for each token in the ground truth and the predicted captions. The results are reported in Table 3. For a fair comparison, we do not input any additional text information to the student model during the inference on the downstream datasets. In Figure 5, we also showcase a video sample from the testing set of Panda-70M and the predicted captions for the qualitative comparison.
As in Table 3, Video-LLaMA with Panda-2M pretraining weight achieves significantly superior performance compared to the official weight. Numerically, our pretraining weight yields and improvement respectively on MSR-VTT and MSVD in terms of B-4. Besides, in Figure 5, we can find that the caption from the original Video-LLaMA contains irrelevant and generic information, such as date and location. In comparison, our prediction better aligns with the video content.
Can the student perform better than its teacher? In Section 3.4, we learn a student model in a knowledge distillation manner. To evaluate the performance of the student model, we conduct a user study where participants are asked to select the best caption from ten candidates for each video. Ten captions are predicted from eight teacher models and two student models (with and without text inputs). We collect the results from five participants to reduce the personal subjective bias. Each participant saw the same 200 videos, which were randomly sampled from the testing set and had not been seen during the training of the student model and UMT. We report the preference ratio of each model and the R@1 accuracy of the finetuned UMT (i.e., all teachers) in Table 4. We can observe that the student model outperforms any individual teacher model and achieves a comparable performance with all teacher models.
Can multimodal inputs leverage video captioning? Our student model supports both video and text inputs. In Table 4, we show that the student model with both video and text inputs outperforms the model with video input only by preference ratio. Qualitatively, we show the predictions with and without text inputs in Figure 5. While the prediction with pure video input can include partial content of the video, like “cactus”, the model with both video and text inputs can more comprehensively include keywords such as “succulents” and “different species” from the video title, description, and subtitles.
2 Video and Text Retrieval
Experiment setup. We use Unmasked Teacher as the base model to evaluate the performance on video and text retrieval. The standard protocols jointly use 3M images from CC3M and 2.5M videos as the pretraining datasets. Thus, we randomly sample a Panda-5M subset, which shares the same number of training samples as the standard pretraining dataset for a fair comparison. For both datasets, we use the same backbone composed of ViT-L/16 and BERTlarge . We use the official weights for the standard datasets pretraining and train the model from scratch for our Panda-5M.
Downstream datasets and evaluation metric. We test both zero-shot and finetune retrieval on three benchmarks: MSR-VTT , DiDeMo , and MSVD . For MSR-VTT, we follow the common protocol to evaluate on 1K testing split, which is not the same as the testing videos for captioning in Section 4.1. For DiDeMo , it contains 10K Flickr videos with a total of 40K dense captions. As in the previous standard , we evaluate paragraph-to-video retrieval by concatenating all sentence descriptions of one video into a single query. We report the results on the 1K testing set. For MSVD , we report the results on the 670 testing videos. We employ the standard metric and report R@1, R@5, and R@10 accuracy on both text-to-video and video-to-text retrieval in Table 5.
We can observe that pretraining with our Panda-5M outperforms the official weight in both zero-shot and finetune retrieval settings. Especially, our pretraining yields , , and lifts in terms of R@1 of zero-shot text-to-video retrieval on MSR-VTT , DiDeMo , and MSVD respectively. Besides, pretraining UMT with our Panda-5M also outperforms the existing state-of-the-art methods which are pretrained with much more vision-text data pairs (i.e., 100M).
3 Text-to-Video Generation
Panda-2M pretraining consistently shows superior performance on both metrics compared to the official weight. As highlighted, our pretraining yields lower FVD on UCF101 and outperforms state-of-the-art models pretrained on a dataset within a 10M scale in terms of FVD. Qualitatively, our pretraining weights can generate the video with a more meaningful motion and photorealistic appearance and do not include a watermark.
Conclusion and Limitations
This paper introduces Panda-70M, a large-scale video dataset with caption annotations. The dataset includes high-resolution and semantically coherent video samples. To caption 70M videos, we propose an automatic pipeline that can leverage multimodal information, such as video description, subtitles, and individual static video frames. We demonstrate that pretraining with Panda-70M can facilitate three downstream tasks: video captioning, video and text retrieval, and text-to-video generation.
Despite showing impressive results, the proposed dataset is still bound by a few limitations. First, we collect the videos from HD-VILA-100M , where most of the samples are vocal-intensive videos. Hence, the major categories of our dataset are news, television shows, documentary films, egocentric videos, and instructional and narrative videos. As our annotation pipeline does not require the presence of video subtitles, we list the collection of more unvocal videos as an important extension of this work.
Second, we focus on a fine-grained dataset where the video samples are semantically consistent so the caption can accurately express its semantics content without ambiguity. Nevertheless, it would limit the content diversity within a single video and also reduce average video duration, which might be hurtful to the downstream tasks, such as long video generation and dense video captioning . Future efforts in building datasets with long videos and dense captions can benefit these downstream applications.
Risk mitigation. Prior to the release of the dataset, we used the internal automatic pipeline to filter out the video samples with harmful or violent language and texts that include drugs or hateful speech. We also use the NLTK framework to replace all people’s names with “person”.
References
In Section 3.1, we propose a video splitting algorithm to cut a long video into several semantically coherent clips. The algorithm includes two stages, splitting and stitching, for which the details are described in Appendix A.1 and A.2.
To handle both cases, we propose creating artificial scene cuts each 5 seconds for clips without cut-scene. That is, if a video clip is longer than 5 seconds, we cut out the first 5 seconds as a new clip and recursively apply the same procedure to the remaining part. Since we are only interested in semantically consistent video clips, we extract the ImageBind features of the frames near the beginning or the end. If the features of these two frames are dramatically different we remove that clip. Specifically, given a -frame video clip , we extract the features and for the number and frames, denoted as and . We only keep the video clips if satisfying . As such, we can exclude video clips with transition effects or significant semantics changes within a clip.
A.2 Stage2: Stitching based on Semantics Similarity
The first stage introduces many short consecutive clips with the same semantic content. To this end, we propose an additional procedure to merge the clips with the same semantic content. Formally, given two adjacent clips and in sequence, we concatenate them into a clip if .
Finally, we perform a post-processing to stabilize the quality and diversity of the video clips with the following steps:
First, we exclude the clips shorter than seconds or clips that contain only slight motion (i.e., ). For the videos longer than seconds, we only retain the first seconds.
Next, we represent each clip by the average of ImageBind features extracted from stage1 (Section A.1) and only keep the video clips that are semantically different (i.e., Euclidean distance ) from the precedent clips to increase the diversity of the video samples.
Finally, we trim out the first and last of a video clip as we notice that the beginning and the ending of a clip usually contain unstable camera movement or transition effects.
With the proposed splitting algorithm, we split long videos into clips with an average clip duration of seconds. We plot the distribution of video length in Figure 7.
Appendix B Details of Teacher Captioning Models: Pool, Inference, and Selection
In Section 3.2, we propose to use multiple cross-modality teacher models for captioning. Specifically, we start with a large pool including 31 captioning models. We elaborate on the composition of the model pool and how we implement them for video captioning in Appendix B.1 and B.2 respectively. As running the inference of the models to 70M videos is computationally expensive, we select only 8 models as the representative, based on a human evaluation. We will describe more details about this process in Appendix B.3.
The primary reason to utilize cross-modality teacher models is to leverage multimodal data that would benefit video captioning. As such, we consider the base models, including image/video visual-question-answering (VQA) and image captioning models. Specifically, we employ Video-LLaMA , VideoChat , VideoChat Text , Video-ChatGPT , BLIP-2 , and MiniGPT-4 as the base models. Based on these models, we collect 31 captioning models in total using different weights and input information. We list the summary of all captioning models in Table 7.
B.2 Inference of Cross-Modality Teacher Model for Video Captioning
We list the inference details of each base model as follows:
Video-LLaMA is a video VQA model. We only use the vision branch and do not use the audio one. The model uses Vicuna-7B as the LLM to implement VQA. We use two official weights, including the pretraining weight, which is trained on 2.5M video-text pairs and LLaVA-CC3M , and the finetuning weight, which is further finetuned on instruction-tuning data from .
VideoChat and Video-ChatGPT are video VQA models. We use Vicuna-7B as the LLM and follow the official codebase for the rest of the configuration.
VideoChat Text is a natural-language processing (NLP)-based video VQA model. The model would textualize the video content into video tags, dense captions, and a general caption respectively by three models . As such, users can have a conversation with a chatbot and discuss the video based on the extracted textual content. The original codebase uses ChatGPT-4 as the chatbot, which is, however, not freely released to the public. Thus, we replace it with LLaMA for large-scale captioning.
BLIP-2 is a language-image pretraining model. We only use it for image captioning and do not input texts. We use the weights pretraining with different LLMs, including OPT (opt2.7b and opt6.7b) and FlanT5 (flant5xl).
MiniGPT-4 is an image VQA model. We use two variants respectively with Vicuna-7B and Vicuna-13B as LLMs.
To implement cross-modality teacher models for video captioning, we design the algorithms specifically for the models of different modalities. For an image model, given an -frame video clip, we randomly sample a video frame in-between number and frames as the input. For a VQA model, in addition to the visual data, we also input a text prompt that could include additional textual information, such as video title, description, and subtitles, to assist video captioning. Specifically, we use the prompt template in Figure 8 if we would like to include the information of either metadata or subtitles or both for captioning. In contrast, we use a dummy prompt: “Please faithfully summarize the video (or image) in one sentence.” if we only input the vision data for captioning.
B.3 Selecting 8 Captioning Models based on a Human Evaluation
Running 31 captioning models on 70M videos requires significant computation resources. Hence, we propose to find a well-performing subset of the models by a two-step algorithm, including a human evaluation and model selection algorithm.
Human evaluation. First, we conduct a user study by showing the output captions of each model to humans. Specifically, we randomly sample 1K video clips and perform the inference of 31 captioning models on each video. Next, the human annotators are asked to select “every good caption”, where a good caption is defined as: “the caption cannot contain any wrong information and needs to cover the main action OR all of the main objects presented in the video.” If none of the captions is a good caption, the annotators are asked to select the “All Bad” option. We randomly shuffle 31 captions to minimize the annotator’s bias on the order of the captions. Considering that a human is hard to focus on reading all 31 caption sentences at the same time, we split the captions into three groups. The annotator will see the same video three times with at most 11 captions once. We show the interface of this user study in Figure 9 and plot the results in Figure 10.
Algorithm of model selection. In the second step, we collect a list of 8 captioning models as the representative to reduce the computation for large-scale captioning. Intuitively, one may opt for the models exhibiting the top 8 performance. Nonetheless, such behavior does not align with the philosophy of our captioning algorithm. Specifically, our algorithm utilizes multiple cross-modality models to cover good captioning on various types of videos and only retrieves one best caption as the annotation for each video (as described in Section 3.3). Accordingly, we propose to use the set of models that can jointly cover a good caption for most video samples. The algorithm starts by selecting the best-performing model (i.e., BLIP-2 with opt6.7b). Next, we only consider the videos that the previously selected model(s) cannot generate a good caption and then greedily find the model that performs best on those videos. We recursively collect the models under this mindset until we make the list of 8 captioning models. The 8 selected models are highlighted in Figure 10.
Additional findings. From Figure 10, we can also observe that a single captioning model can predict a good caption for at most of the videos. In comparison, all 31 captioning can jointly predict at least one good caption for of the videos (based on the “All Bad” ratio of ). This fact supports our motivation to use multiple cross-modality teacher models to jointly predict the captions for a video. Last but not least, according to our statistics, using 8 selected teacher captioning models can jointly predict a good caption for of the videos which shows comparable performance with all 31 models while significantly reducing the computational requirements.
Appendix C Details of Fine-Grained Video-to-Text Retrieval: Dataset, Training, and Inference
In Section 3.3, we mention that the available generic retrieval models cannot pick the best caption from 8 candidates predicted by our teacher models. The main reason is that all of the candidate captions are highly relevant to the video sample and require the model to discern subtle distinctions within each caption for optimal performance. To better perform our “fine-grained” retrieval task, we first annotate a subset of video samples by manually selecting the best caption as detailed in Appendix C.1. Next, we finetune Unmasked Teacher (UMT) and run the inference of the model on all video samples respectively in Appendix C.2 and C.3.
We randomly sample 100K video samples from our dataset and ask human annotators to select “the best caption” for each video. At the beginning of the task, the annotator will read the task description as follows:
“You are presented with a short video clip and a set of textual summaries that describe this clip. Choose the textual summary that is the most faithful and descriptive of the content of the video clip. Imagine you are talking on the phone with your friend and you need to describe the video to him.”
Note that this task is different from the user study in Appendix B.3, where a human is asked to select “every good caption”. But, we also randomly shuffle the captions and provide an “All Bad” option if all of the captions contain wrong information. We filter out videos with the “All Bad” option selected and split the dataset into and videos for training and validation. We plot the selective rate of each teacher model on the validation set in Figure 3 (blue bar).
C.2 Finetuning of Retrieval Model
C.3 Inference of Retrieval Model on Panda-70M
With the finetuned UMT, we automatically retrieve the best caption as the annotation for all 70M videos. We illustrate the distribution of the finetuned UMT’s selection in Figure 11 and the caption length in Figure 12. We also plot the word cloud of the randomly sampled 100K caption annotations in Figure 13 to highlight the rich content within the annotated captions.
In addition to the retrieval result, UMT also predicts a matching score for the video-text pair. In practice, we find the score is highly correlated to the alignment of the contents within the video-text pair. A score higher than usually represents a strong association between the video and the caption. Numerically, of the samples in Panda-70M have matching scores higher than .
Appendix D Details of Student Captioning Model: Architecture and Training
Figure 4 shows the architecture of the student captioning model. The model includes a vision branch and a text branch for additional subtitle and metadata inputs.
For the text branch, given a prompt with an arbitrary length, the model first tokenizes the prompt and embeds each token into a feature vector with length by a pretrained embedding layer . Considering that the number of token embedding might be large for a longer prompt and the information of the prompt might not well align with the video content, we then design a text Q-Former to extract a fixed and shorter length of text embedding and at the same time, better bridge the feature of the input video and text prompt. Specifically, the text Q-Former takes the inputs of the video representation as the queries and multiple token embedding as the key and value. The module then outputs a text representation. Finally, we combine the multimodal inputs by concatenating the text and video representations in sequence to get a feature and input it to the LLM to predict the video caption.
D.2 Training Details
The training data includes a video-caption pair and additional text information (i.e., the metadata and subtitles). For the video data, we randomly sample 8 frames and apply the same video reading algorithm as in Appendix C.2. For the text branch, we embed the extra text information into the prompt. To learn a captioning model that can take both video-only and video-text inputs, we drop part of the text inputs at random. Formally, we use the prompt template in Figure 8 and employ the metadata or/and subtitles information with the probability of (the sampling for metadata and subtitles are independent).
We use the AdamW optimizer. The learning rate is initialized as and linearly warmed up to within the first 2,500 steps and gradually decreased to based on cosine annealing strategy . We set and use a weight decay of . We train the model on the whole Panda-70M with a batch size of 48 and last the training for 300K steps. The model is trained on 48 Nvidia A100 GPUs (80GB).
Appendix E Visualization of Panda-70M Dataset
In the following subsections, we visualize video-text pairs in Panda-70M by category.