mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou
Introduction
Large-scale pre-trained foundation models have been an emerging paradigm for a wide range of artificial intelligence (AI) fields, across language (Devlin et al., 2018; Brown et al., 2020), vision (Dosovitskiy et al., 2020; Liu et al., 2021b) and multi-modality (Radford et al., 2021; Yu et al., 2022; Wang et al., 2022e). With the broad success of Transformer architecture (Vaswani et al., 2017), recent years have featured a trend toward the big convergence of language, vision and multimodal pre-training (Yu et al., 2022; Wang et al., 2022e; Alayrac et al., 2022). One line along this trend proposes to unify the tasks and modalities with a unified sequence-to-sequence generation framework such as T5 (Raffel et al., 2020), OFA (Wang et al., 2022d) and Flamingo (Alayrac et al., 2022). On the other hand, BERT (Devlin et al., 2018), Florence (Yuan et al., 2021) and BEIT-3 (Wang et al., 2022e) models all the tasks as instance discrimination, and adopt the pure encoder-based architecture.
The predominant foundation models propose to share the same single network for multi-modality (Alayrac et al., 2022) to leverage the information from modality collaboration. However, the strategy will suffer from the issue of modality entanglement due to the large variance of different modality tasks. The challenge is that multiple modalities may interfere with each other (Huang et al., 2022b), especially when there are many modalities and tasks. It is difficult for a single-module foundation model to balance the gain of modality collaboration and the influence of modality entanglement on a large number of downstream tasks across multiple modalities.
To alleviate the challenge, in this work, we introduce a new unified paradigm of multi-modal foundation models, as shown in Figure 1. It features a module-based network design considering both the modality collaboration and modality entanglement, where mPLUG-2 designs certain shared functional modules to encourage the modality collaboration, while reserving modality-specific modules to tackle the problem of modality entanglement. Different modules are then jointly trained effectively on both the uni-modal and multi-modal datasets according to the task’s module design. As a result, different modules can be flexibly selected and combined for the large number of uni-modal and cross-modal understanding and generation tasks accordingly. The details of the supported downstream tasks are given in Table 1. To the best of our knowledge, the proposed method tackles the largest number of different kinds of downstream tasks across text, image and video.
Specifically, we design a unified dual-vision encoder module by disentangling spatial and temporal representations, where video inputs share the standard Transformer module with image inputs for modeling spatial information and an extra local temporal modeling module is used for temporal relation modeling on video-related tasks. Then a novel universal layers module is introduced to serve as a pivot across different modalities, where vision and language modalities are projected to the common language-guided semantic space by sharing self-attention modules. Besides, an extra cross-attention module is used to fuse the universal vision representation with the original fine-grained vision representation. The detailed module design is shown in Figure 2. Finally, different modules of mPLUG-2 are jointly pre-trained with task and modality instructions (Wang et al., 2022d) on both uni-modal and cross-modal tasks. During inference, mPLUG-2 can select different modules for various uni-modal and cross-modal tasks with the modularized Transformer architecture. The selected modules for different tasks can be found in Table 2 in Appendix.
We evaluate the new unified paradigm of mPLUG-2 on over 30 challenging uni-modal and cross-modal understanding and generation benchmarks and it achieves state-of-the-art or competitive results with a similar model size and data scale. Equipping with the module-based network design, mPLUG-2 can be also easily extended to additional tasks by selecting and adding modules. Notably, mPLUG-2 shows new state-of-the-art results of 48.0 top-1 accuracy and 80.3 CIDEr on the challenging MSRVTT video QA and video caption tasks, respectively. mPLUG-2 also demonstrates strong zero-shot transferability on vision-language and video-language tasks.
Related Work
ConvNets (Szegedy et al., 2017; He et al., 2015) have long been the main stream visual architecture before the emergence of vision transformer (a.k.a. ViT) (Dosovitskiy et al., 2020). Due to the superior capacity of Transformer network, ViT stands out in various downstream tasks (Carion et al., 2020; Xu et al., 2022). Apart from scaling up the naive ViT architecture with large-scale dataset such as JFT-3B (Zhai et al., 2021), SwinV2-G (Liu et al., 2021a) extends the original ViT with hierarchical architectures. In addition, EVA (Fang et al., 2022a) distills the multi-modal knowledge to scale up ViT by leveraging unlabeled images with the large-scale pre-trained image-text model (e.g. CLIP (Radford et al., 2021)). Recently, InternImage (Wang et al., 2022f) revitalizes the convolutional neural networks with deformable convolution and achieves the state-of-the-art performance on various vision downstream tasks. Besides, InternVideo (Wang et al., 2022g) extends to video tasks by assembling two large video models with both generative and discriminative self-supervised video learning.
Inspired by the successful practice of the BERT (Devlin et al., 2018) in natural language understanding, a massive large-scale language foundation models are proposed for natural language processing. BART (Lewis et al., 2020) is a denoising autoencoder like BERT but with encoder-decoder architecture which shows effectiveness for both text generation and comprehension tasks. Apart from BERT-series methods (Devlin et al., 2018; Lewis et al., 2020; Liu et al., 2019), there are numerous other effective architectures and pre-training objectives. T5 (Raffel et al., 2020) introduce a unified framework that covers all text-based language tasks into a text-to-text format. GPT-3 (Brown et al., 2020) is an auto-regressive language foundation model which includes 175 billion parameters, and shows strong performance on many NLP tasks under the few-shot and zero-shot settings.
Benefiting from a large number of image/video-text pairs in the Internet, the emergence of vision-language foundation models can subsume vision-language pre-training. The success of CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021) indicates that the model pre-trained with simple contrastive objectives on noisy image-text pairs can generate powerful vision-language representation. Moreover, ALBEF (Li et al., 2021b), BLIP (Li et al., 2022c) and mPLUG (Li et al., 2022a) extend the task with multi-modal text completion and text generation for auxiliary learning. On the other hand, some foundation models are built through task unification. For instance, Florence (Yuan et al., 2021) unifies the contrastive objectives that can leverage both vision and vision-language data. BEiT-3 (Wang et al., 2022e) ascribe the pre-training task to mask data modeling in terms of text, vision, and vision-language. SimVLM (Wang et al., 2021b), OFA (Wang et al., 2022d), and CoCa (Yu et al., 2022) perform the generative pre-training for vision-language understanding and generation. Different from predominant foundation models, mPLUG-2 introduces a new modularized transformer framework, which can leverage different compositions of modules for both uni-modal and cross-modal tasks by both sharing common universal modules and disentangling modality-specific ones to address the problem of modality entanglement.
Method
As shown in Figure 2, mPLUG-2 consists of a dual-vision encoder module for image and video, a text encoder module, a universal layers module that serves as a multi-modal pivot shared by all tasks, a multi-modal fusion module and a shared decoder module for uni-modal and cross-modal generation. We first use two uni-modal encoders which encode image/video and text separately to represent the inherent information of the individual modality. For image/video, we adopt the dual-vision encoder to encode visual features with spatial modeling and local temporal modeling. Then, the visual and linguistic representations are fed into the universal module separately, which consists of multiple universal layers. Each universal layer projects different modalities to shared semantic space for cross-modal alignment while preserving the original representation of different modalities. The output of universal layers is applied to conduct uni-modal discrimination tasks. For cross-modal tasks, an additional fusion module will be applied to produce cross-modal representations. Finally, the uni-modal and cross-modal representations can be incorporated as input to a shared Transformer decoder for various generation tasks, which facilitates multi-task pre-training and transfer learning. The modules for different downstream tasks are summarized in Table 2.
To capture the visual information of various vision modalities, we propose dual-vision encoder to model image and video simultaneously. Specially, we split the image and video frames into a sequence of non-overlapping visual tokens. Every sequence of visual tokens with learnable spatial position embeddings and an extra [CLS] token constitute an input visual sequence. However, modeling the completed visual sequences leads to difficulty in spatio-temporal learning without large-scale video pre-training (Li et al., 2022d; Wang et al., 2022a, b). To alleviate this problem, we decouple the visual representation into the spatial and temporal representation separately by introducing temporal locality. As illustrated in Figure 2(b), we leverage the self-attention (SA) layer and feed-forward layer (FFN) in the Transformer block for spatial modeling, and propose a novel local temporal modeling module (LT) to model the temporal dependency among the spatial representation as:
where LN is short for layer normalization. The local temporal modeling module captures the correlation among patches with the same spatial locations through multi-group fusion formulated as:
where and are linear transformation functions. is the learnable temporal relation parameter, which is instantiated as a convolution kernel. and are number of frames and size of hidden state. indicates the number of groups, and denotes concatenation function. By using multi-group fusion, the model is able to learn rich temporal information from distinctive representation subspaces at different temporal locations. As a result, except the local temporal module, the dual-vision encoder module enables weight sharing for images and videos, which effectively and efficiently learns the spatial and temporal representation.
For the text encoder module, we use BERT (Devlin et al., 2018) as the text encoder, which transforms the input text and an extra [CLS] token into a sequence of text embeddings. The embedding of [CLS] token is used to summarize the input text.
To benefit from modality collaboration, we propose the universal layers to model the vision and language modalities in the shared semantic space while preserving the original representation of the different modalities.
Before the universal module, we take a variable number of image or video features from the dual-vision encoders as input to produce a fixed number of visual tokens to reduce the computational complexity of universal layers. In the universal layer, the visual tokens and the text representation are fed to the shared self-attention layers to align semantics, and then the visual tokens are injected into the original visual feature space by the cross-attention layer to keep the original representation.
Then is fed into the next universal layer repeatedly to get the final common image and text representation. Finally, the output of the universal layers are combined with the original representations by the cross-attention layer for the text-aware visual and visual-aware text representation, where are the layers of universal module, dual-vision encoder and text encoder respectively.
To effectively capture the cross-modal interaction between vision and language modalities, we use the fusion module as in ALBEF (Li et al., 2021b), which is composed of a stack of Transformer blocks with cross-attention layers. Specifically, the fusion module takes the text embeddings from the universal layers module as the input. Then, the text-aware vision embedding cross-attends to the visual-aware text embeddings in language-shared common space. By cascading the Transformer blocks with cross-attention layers, fusion module is able to yield multi-modal vision-language representations.
To empower the model with the capability of generation, a shared decoder module is introduced to enable the model to generate text with both uni-modal and multi-modal information. In detail, the shared decoder module is a Transformer decoder with arbitrary inputs. For example, image captioning only requires the visual features, while the multi-modal features are used for visual question answering. By taking different types of input, our shared decoder module can adapt to a variety of tasks with text generation. The shared decoder module facilitates multi-task pre-training and transfer learning.
2 Unified Pre-training Objectives
We jointly train the multiple modules of mPLUG-2 with the following three objectives.
For the text encoder module, we use Masked Language Modeling (MLM) as in BERT (Devlin et al., 2018) to learn the text representation. We randomly mask 15% tokens in the text and the model is asked to predict these masked tokens with the context representations.
For the cross-modal module, we employ the Cross-modal Matching Losses (CML) as in ALBEF (Li et al., 2021b), which consists of Vision-language Matching (VLM) and Vision-language Contrastive Learning (VLC).
Following Flamingo (Alayrac et al., 2022) and OFA (Wang et al., 2022d), we adopt the Instruction-based Language Model Loss to unify various generation tasks. We use handcrafted instructions to discriminate tasks and modalities, which include Video/Image-Text Pairs, Video/Image Captioning, Video/Image Question Answering, Text Generation, etc.
Experiment
Following previous works (Li et al., 2021b, 2022a), we pre-train our model with the same popular image-text datasets with 14M images including MS COCO (Lin et al., 2014), Visual Genome (Krishna et al., 2017), Conceptual Captions 3M (Sharma et al., 2018), Conceptual Captions 12M (Changpinyo et al., 2021), and SBU Captions (Ordonez et al., 2011). For video-text datasets, we adopt the web-sourced video dataset WebVid-2M (Bain et al., 2021a) with 2.5M video-text pairs. The text datasets consists of WikiCorpus (Devlin et al., 2018) (about 20GB) and cleaned common crawl (about 350GB). The collection and cleaning method of the latter is generally the same as that used in c4 (Raffel et al., 2020). The implementation details of pre-training can be found in the Appendix.
2 Main Results
We evaluate the new unified paradigm of mPLUG-2 on over 30 benchmarks including vision-language tasks (e.g. multi-modal retrieval, question answering and captioning) (Xu et al., 2016, 2017; Chen & Dolan, 2011), language-only tasks (e.g. text classification, question answering and summarization) (Wang et al., 2018; Rush et al., 2015a), and vision-only tasks (e.g. image classification and video action recognition) (Deng et al., 2009; Kay et al., 2017). Specially, the vision-language benchmarks can be categorized as image-text parts and video-text parts. Details of these datasets can be found in the Appendix.
We compare mPLUG-2 with several state-of-the-art methods on MSRVTT (Xu et al., 2016), DiDeMo (Anne Hendricks et al., 2017) and LSMDC (Rohrbach et al., 2015) datasets. The results are summarized in Table 3. We can observe that mPLUG-2 outperforms the previous SoTA methods on most of the datasets. In particular, our method yields 5.7% lift in terms of R@1 on LSMDC datasets compared with HiTeA, which indicates that the proposed model can leverage the temporal information presented in fruitful movie clips through the proposed local temporal modeling module in the dual-vision encoder.
Table 4 summarizes the video question answering results on MSRVTT-QA (Xu et al., 2017), MSVD-QA (Xu et al., 2017), and TGIF-FrameQA (Jang et al., 2017). It can be observed that mPLUG-2 outperforms all the existing foundation models on MSRVTT-QA and TGIF-FrameQA by a large margin, and it also attains the comparable result with big foundation models GIT2 (Wang et al., 2022c) on MSVD-QA even using significantly smaller amount of pre-trained data. In particular, mPLUG-2 achieves absolute improvement 0.6% on MSRVTT and 0.5% on TGIF-FrameQA. Furthermore, mPLUG-2 achieves the comparable results compared to the large models (i.e., VideoCoCa and GIT2) with smaller model size.
Table 55 compares mPLUG-2 with existing methods on video captioning datasets MSRVTT and MSVD. As shown in the table, although pre-trained on less data, mPLUG-2 derives the significant improvement on MSRVTT dataset, and comparable performance on MSVD dataset. On MSRVTT Caption, our method surpasses SoTA method VideoCoCa (Yan et al., 2022) and GIT2 (Wang et al., 2022c) by 4.4% on CIDEr and 3.0% on BLEU@4. Moreover, we can notice mPLUG-2 outperforms HiTeA with the same amount of pre-training data, which shows that mPLUG-2 is able to generate stronger video-language representation.
We compare mPLUG-2 with existing state-of-the-art methods on visual grounding datasets including RefCOCO (Yu et al., 2016), RefCOCO+ (Yu et al., 2016) and RefCOCOg (Mao et al., 2016). Table 7 shows that mPLUG-2 achieves comparable performance to the state-of-the-art methods. Our method achieve 0.97% absolute improvement compared with the second best method on RefCOCO “testB” split without using object detection data for pre-training. Queries in “testB” split can refer to various visual concepts but only people in “testA”. The improvement demonstrates that the introduction of universal layers can help model the visual concepts in the image.
We evaluate mPLUG-2 on image-text retrieval datasets MSCOCO and Flickr30k. As shown in Table 6, both mPLUG-2 and mPLUG-2 achieves comparable or better performance than state-of-the-art methods. Florence (Yuan et al., 2021) and BLIP (Li et al., 2022c) use 0.9B and 129M data for pre-train respectively. In contrast, our mPLUG-2 only requires 17M data. It demonstrate that mPLUG-2 is data-efficient.
We report the performance of mPLUG-2 on visual question answering test sets. mPLUG-2 surpasses state-of-the-art method Florence (Yuan et al., 2021) 0.95% on test-dev and 0.77% on test-std. The scale of the pre-trained data used in our model is 89.11% less than that in Florence. It shows that our mPLUG-2 can learn multi-modal represent efficiently and effectively.
We compare mPLUG-2 with existing state-of-the-art methods on MSCOCO (Lin et al., 2014). Following (Li et al., 2020b), we train mPLUG-2 on the COCO Caption with cross-entropy loss and test on the same Karpathy split. As shown in Table 9, our mPLUG-2 achieves new SoTA results on COCO Caption. Moreover, our method achieves competitive results with big foundation models, such as LEMON (Hu et al., 2021) and BLIP (Li et al., 2022c) which use more than nearly 10x amount of pre-training data. Specifically, our mPLUG-2 outperforms BLIP on COCO caption by an obvious 1.2 point margin on BLEU@4, and 1 point on CIDEr.
2.2 Language Only Tasks
We evaluate mPLUG-2 on 6 tasks of the GLUE benchmark (Wang et al., 2018) for natural language understanding. Table 10 shows that mPLUG-2 achieves comparable performance to the state-of-the-art natural language and multimodal pretrained models including RoBERTa (Liu et al., 2019), DeBERTa (He et al., 2021b). Our method with DeBERTa achieves improvement compared with DeBERTa (He et al., 2021b) on three tasks, which also demonstrate the effectiveness of universal modules for modality collaboration.
We evaluate mPLUG-2 on Gigaword abstractive summarization (Rush et al., 2015b) for natural language generation. As shown in Table 1111, mPLUG-2 achieves the comparable result with the state-of-the-art models.
2.3 Vision Only Tasks
Video action recognition is the most representative for video understanding since it requires the model to understand the spatio-temporal cues revealed in the video. Table 12 summarizes the performance of different approaches on Kinetics 400, Kinetics 600, and Kinetics 700 datasets. Our mPLUG-2 surpasses the most of SoTA methods. For example, comapred with Florence pre-trained on 900M vision-text pairs, mPLUG-2 improves the Top-1 accuracy by 1.9% on Kinetics 600 and 0.6% on Kinetics 400. Meanwhile, we can notice that the performance of mPLUG-2 is better than OmniVL with similar amount of pre-training data, which shows the effectiveness of the dual-vision encoder module for video representation learning.
We further evaluate the performance of mPLUG-2 in terms of image classification on ImageNet-1K. As we can see in the Table 13, We can see that mPLUG-2 achieves comparable results or even surpass the SoTA methods on ImageNet-1K without using the ImageNet data for pre-training. Besides, to effectively evaluate the robustness and generalization ability of mPLUG-2 , we perform the evaluation on 5 ImageNet variants (i.e. IN-V2, IN-Real., IN-Adversarial, IN-Rendition, and IN-Sketch). Following standard evaluation procedure (Fang et al., 2022a), all these models are first fine-tuned on the original ImageNet-1K training set and directly tested on the 6 variants without further fine-tuning. As shown in Table 13, mPLUG-2 not only achieves the highest accuracy on ImageNet-1K validation set but also obtains the relative small gap (i.e., ), which reflects the excellent robustness and generalization capability of mPLUG-2 with the help of the universal layer module by learning language-shared representation.
3 Discussion
The instructional-based learning is able to distinguish different types of tasks with specific instructions. Table 14 demonstrates the effectiveness of instructional-based learning. In the table, we can observe that the instructional-based learning improves the performance of retrieval and question answering by at least 0.7% and 1.2% in Average Recall and accuracy respectively. With the help of instructional-based learning, mPLUG-2 is capable of utilizing the different modules when different instructions are used to boost the performance.
To validate the effectiveness of our proposed local temporal modeling module in the dual-vision encoder, we conduct experiments with the different temporal modeling structures. Specially, we have tried out the temporal self-attention and temporal convolution for comparison. The results are summarized in Table 15. We can notice that the local temporal modeling module outperforms temporal self-attention module by introducing modeling temporal locality. Meanwhile, with the help of the multi-group fusion mechanism, the local temporal modeling module can learn the diverse temporal representations in distinctive representation subspaces while the temporal convolution is restricted in the same temporal representation spaces, thus leading to the better performance.
To validate the effectiveness of our proposed universal layer module, we ablation this module for all uni-modal and multi-modal tasks. As shown in Table 16 and Table 17, we set Row 1/2/2 as the baseline of the vision/language/vision-language task in this experiment, respectively. We can find that compared with the baseline the shared universal layer is beneficial for all modality tasks by encouraging collaboration between modalities.
In Figure 3, we visualize the Grad-CAM on the cross-attention map in the first universal layer. For each sample, we present two cross-attention maps that attend to different visual concepts. The results show that the universal layer can encourage modality collaboration and modality entanglement between visual patch features and language features by attending the areas of various visual concepts in the image.
Here we investigate the influence of universal layer in terms of modality collaboration. We randomly sample some vision-language pairs, and sketch the UMAP visualization of the generated embeddings from pre-trained mPLUG-2 in the Figure 4. We can observe that with the help of universal layer, the distance between vision and text samples are more closer instead of solely two concentrated clusters. Besides, we quantitatively compute the modality gap (Liang et al., 2022), where the is the difference between the center of vision embeddings and text embeddings. It can be observed that the model with universal layer would encourage the collaboration between vision and language modalities thus yielding lower modality gap compared to the model without universal layer.
Conclusion
This paper presents mPLUG-2 , a new unified paradigm with modularized design for building multi-modal foundation models. mPLUG-2 introduces a module-based network design that shares common universal modules for modality collaboration and disentangles modality-specific modules to address the problem of modality entanglement. Experimental results show that the new unified paradigm of mPLUG-2 can achieve strong performances on a broad range of over 30 tasks across the text, image and video modalities. It is also easy to extend mPLUG-2 to more tasks by selecting and adding modules.
References
Appendix A More Results
We evaluate the object detection and instance segmentation performance of mPLUG-2 on COCO dataset (Lin et al., 2014), which is widely used for object-level detection and segmentation with 80 common categories. Table 18 reports the results on COCO dataset. We observe that mPLUG-2 outperform typical state-of-the-art resnet-based detection methods (e.g., DETR (Carion et al., 2020) and Pix2seq (Chen et al., 2021)). There is a performance gap between foundation model optimized for computer vision (e.g., Florence (Yuan et al., 2021) and Swin-Transformer (Liu et al., 2021b)) and mPLUG-2 . Note that mPLUG-2 does not pre-trained with vision only task and data. Lower performance than models pre-trained on ImageNet is to be expected.
A.2 Zero-Shot Transferability
For testing the transferability of pre-trained mPLUG-2 , we conduct the zero-shot evaluation on Text-to-Video Retrieval and the results are summarized in Table 19. We can find that mPLUG-2 obtains SoTA results on both MSRVTT, DiDeMo and LSMDC datasets, and outperforms previous methods by a large margin, such as 5.1 point of R@1 on the MSRVTT dataset. The results prove that our mPLUG-2 has excellent zero-shot transferability.
We testing the transferability of pre-trained mPLUG-2 on Video QA and the results are summarized in the Table 20. It can be observed that mPLUG-2 achieves the best zero-shot performance on both MSRVTT-QA and MSVD-QA datasets, which demonstrates the strong zero-shot transferability of our model under the help of universal module and instructional-based learning.
A.3 Visual Grounding
We visualize several cases of visual grounding task in Figure 5. The first row shows that our mPLUG-2 can understand various visual concepts and their relationships. It also can make fine-grained alignment between vision and language. The second row presents several failure cases. In the first sample, “trunk” is an ambiguous which result in a incorrect prediction. In the second sample, mPLUG-2 fail to recognize the blurred “donut”. In the thrid sample, mPLUG-2 does not realize the left and right are reversed in a mirror and predict the “left” item.
Appendix B Implementation Details
Our models are implemented in the PyTorch framework (Paszke et al., 2019). In detail, we instantiate the text encoder with BERT (Devlin et al., 2018) model pre-traiend on Wikipedia and Bookcorpus (Zhu et al., 2015). The visual encoder is initialized from CLIP-ViT (Radford et al., 2021) pre-trained on 400M noisy image-text pairs. For the base size of model namely mPLUG-2 , we use the ViT-B/16 for vision encoder and BERT-Base (Devlin et al., 2018) as the text encoder as well as the text decoder. For mPLUG-2 , we scale up the vision and text encoders with ViT-L/14 (Dosovitskiy et al., 2020) and BERT-Large (Devlin et al., 2018) respectively. and for mPLUG-2 and mPLUG-2 . We set for universal layers for the good empirical performance, and choose for multi-group mechanism in the local temporal modeling module empirically. The number of layers for fusion module is set to 3 for mPLUG-2 and 6 for mPLUG-2 , while the number of shared decoder layer is set to 12 for both mPLUG-2 and mPLUG-2 . We pre-train the model for 30 epochs with the total batch size of 1024 on 8 NVIDIA A100 GPUs for mPLUG-2 and batch size of 512 on 16 NVIDIA A100 GPUs. We use AdamW (Loshchilov & Hutter, 2019) optimizer with the weight decay factor 0.02 and betas (0.9, 0.98) for stabilizing the learning. The learning rate is firstly warmed up to in the first 5000 iterations then decays following the cosine annealing schedule. is set to 1e-4 for mPLUG-2 and 5e-5 for mPLUG-2 . During the pre-training, we randomly crop the images and video frames into resolution and sparsely sample 4 frames for each video while preserving their order in-between. For vision-text contrastive learning, the queue size and the momentum coefficient are set to 65,536 and 0.995 respectively.
B.2 Downstream Tasks
We first train mPLUG-2 on the Kinetics-710 dataset for 40 epochs which is the combination of Kinetics-400, Kinetics-600 and Kinetics-700 by removing the videos represented in the validation and test sets. Specially, the base learning rate is set to 1e-5 for mPLUG-2 and 5e-6 for mPLUG-2 with batch size 256 and 128 respectively. Then fine-tuning on Kinetics-400, Kinetics-600, and Kinetics-700 individually for 5 epochs with same learning rate and batch size.
We finetune mPLUG-2 for 30 epochs with the learning rate of 6e-5 and a batch size of 4096. we use the RandomCrop, HorizontalFlip, RandAug and RandErase transformations for data augmentation.
We keep the same setting as EVA (Fang et al., 2022b) to train mPLUG-2 on object detection and segmentation tasks. The different is that we do not pre-train mPLUG-2 on Object365 (Shao et al., 2019) before fine-tuning on MSCOCO.
B.2.2 Language Only Tasks
Following (Wang et al., 2022d), we select the best hyperparameters in a suitable range for fine-tuning. We tune the training epochs among 5, 7, 10, learning rate among 3e-5, 5e-5, 6e-5, 7e-5, 1e-4, batch size among 32, 64, 128. We report the best performance on the development set for each task.
Following (Wang et al., 2022d), we finetune mPLUG-2 for 50,000 steps with a learning rate of 3e-5 and a batch size of 256. During reference, we beam size with 5 and max generation length with 512.
B.2.3 Video-Text Multi-modal Tasks
For all video-language downstream tasks, we resize video frames to 224 224. During fine-tuning, we randomly sample 12 frames for text-to-video size video frames, 16 frames for video question answering and video captions. We perform uniform sampling during inference. We use RandomCrop with minimum ratio 0.5 and HorizontalFlip with 0.5 probability for data augmentation.
We train mPLUG-2 and mPLUG-2 on the training set of MSRVTT/DiDeMo/LSMDC for 10 epochs with a learning rate of 2e-5 and batch size of 192.
We train mPLUG-2 and mPLUG-2 on the training set of MSRVTT-QA/MSVD-QA/TGIF-FrameQA for 10 epochs with a learning rate of 2e-5 and batch size of 128.
For the video caption task, we use a prefix prompt ”What does the video describe?” to improve the quality of generated captions. We set the same training parameters for both the MSRVTT and MSVD datasets. Specifically, we fine-tune mPLUG-2 and mPLUG-2 with cross-entropy loss on the training set for 10 epochs with a learning rate of 2e-5 and a batch size of 128. Then, we perform CIDEr optimization for extra 5 epochs with a learning rate of 1e-6 and a batch size of 16. Finally, we evaluate the test set with a beam size of 5 and max generation length of 25.
B.2.4 Image-Text Multi-modal Tasks
We resize image frames to 336/576/384/336 for the retrieval/vqa/captioning/grounding tasks. We use ResizedCrop with a minimum ratio of 0.5 and HorizontalFlip with 0.5 probability for data augmentation. We perform center crop during inference.
We train mPLUG-2 on the training set of MSCOCO/Flickr30K for 8 epochs with a learning rate of 1e-5 and batch size of 512.
We train mPLUG-2 on the VQA dataset for 8 epochs with a learning rate of 3e-5 and batch size of 512.
For the image caption task, we use a prefix prompt ”What does the image describe ?” to improve the quality of generated captions. we first fine-tune mPLUG-2 with cross-entropy loss on COCO training set for 5 epochs with a learning rate of 1e-5 and a batch size of 256. Then we evaluate on the COCO Caption Karpathy validation split and reuse it to predict the Nocaps validation set directly. During inference, we use beam search with a beam size of 5 and set the maximum generation length as 25.
We first train the model with RefCOCO series datasets with a learning rate of 2e-5 for 120 epochs. Then we continue fine-tuning the model on each dataset with a learning rate of 2e-6 epochs for 30 epochs. We limit the query length to 20/40 for RefCOCO and RefCOCOg, respectively.
B.3 Dataset Description
We evaluate mPLUG-2 on three popular text-to-video retrieval datasets including MSRVTT (Xu et al., 2016), DiDeMo (Anne Hendricks et al., 2017), and LSMDC (Rohrbach et al., 2015).
MSRVTT consists of 10K YouTube sourced videos with 200K text descriptions. Following (Li et al., 2022d; Luo et al., 2022; Huang et al., 2022a), the dataset is divided into 9K and 1K videos for training and testing.
DiDeMo consists of 10K videos from Flickr and each video with 4 descriptions. Following (Li et al., 2022b; Ma et al., 2022; Li et al., 2022d), we concatenate all descriptions of a video as a paragraph, and evaluate the paragraph-to-video retrieval performance. The dataset is separated into 8K for training, 1K for validation and 1K videos for test.
LSMDC consists of 118,081 video clips from 202 movies. Following the standard splits from (Rohrbach et al., 2015), the dataset is divided into 101K and 1K videos for training and testing.
We evaluate mPLUG-2 on three popular video question answering datasets including MSRVTT-QA (Xu et al., 2017) MSVD-QA (Xu et al., 2017), and TGIF-FrameQA (Jang et al., 2017).
MSRVTT-QA is based on the MSRVTT dataset (Xu et al., 2016). The QA pairs are automatically generated by from the descriptions. This benchmark composed of 243K open-ended questions over 10K videos.
MSVD-QA is based on the MSVD datasets (Chen & Dolan, 2011) with automatically generated QA pairs. It consists 2K videos with 47K questions.
TGIF-FrameQA collects the answerable with just a single frame in the video, and is divided into training set with 35K questions and test set with 14K questions.
We use MSRVTT (Xu et al., 2016) and MSVD (Chen & Dolan, 2011) for video captioning evaluation.
MSRVTT is composed of 10K videos with 20 captions per video as described above. We take the same data split as text-to-video retrieval task.
MSVD contains 1970 YouTube short video clips. Following the standard splits from (Lin et al., 2022; Li et al., 2022d), we separate the dataset into 1,200 train, 100 validation and 670 test videos.
We evaluate our method on the VQA 2.0 dataset (Agrawal et al., 2017).
VQA 2.0 is a dataset containing open-ended questions about images and at least 3 questions (5.4 questions on average) per image. It contains 83k/41k/81k images for training/validation/test.
Two popular image-text retrieval benchmarks, COCO (Lin et al., 2014) and Flickr30K (Plummer et al., 2015) are used to evaluate the model. We adopt the widely-used Karpathy split (Karpathy & Fei-Fei, 2015) for both COCO and Flickr30K.
COCO has over 330k images and 5 independent human generated captions are be provided for each image. It contains 113k/5k/5k images for training/validation/testing.
Flickr30K contains 31k images from Flickr, each image with 5 human annotated sentences. It contains 29k/1k/1k images for training/validation/testing.
We evaluate our method on COCO (Lin et al., 2014) datasets.
COCO takes the same data split as the image-text retrieval task.
To verify the natural language understanding ability of our mPLUG-2 , we select 6 language understanding datasets from GLUE (Wang et al., 2018) benchmark, including both single-sentence classification tasks and sentence-pair classification tasks.
SST-2 The Stanford Sentiment Treebank consists of sentences from movie reviews and human-annotated sentiment. The task is to predict the sentiment of a given sentence.
RTE The Recognizing Textual Entailment dataset comes from a series of annual textual entailment challenges.
MRPC The Microsoft Research Paraphrase Corpus consists of a corpus of sentence pairs collected from online news sources, with human annotations for whether the sentences in the pair are semantically equivalent.
QQP The Quora Question Pairs dataset is a collection of question pairs from the community question-answering website Quora. The task is to predict whether a pair of questions are semantically equivalent.
MNLI The Multi-Genre Natural Language Inference Corpus consists of sentence pairs (premise, hypothesis) with textual entailment annotations. The task is to predict the entailment between the premise and the hypothesis.
QNLI The Stanford Question Answering Dataset is a question-answering dataset, where one of the sentences in the paragraph (drawn from Wikipedia) contains the answer to the corresponding question (written by an annotator). The task is to determine whether the context sentence contains the answer to the question.
We use Gigaword dataset (Rush et al., 2015a) for text summarization task to verify the natural language generation ability of our mPLUG-2 .
Gigaword Headline-generation on a corpus of article pairs from Gigaword consisting of around 4 million articles. It contrains 3803957, 189651 and 1951 samples for training/validation/testing.
We adopt three popular benchmarks Kinetics 400/600/700 dataset (Kay et al., 2017) to evaluate our model.
The videos in these three benchmarks are collected from YouTube. Each video clip lasts around 10 seconds and is labeled with a single action class. The videos include human-object interactions such as playing instruments, as well as human-human interactions such as shaking hands and hugging.
Kinetics 400 consists of 240K training videos and 20K validation videos that span 400 human action categories.
Kinetics 600 consists of 392K training videos and 30K validation videos spanning 600 action categories.
Kinetics 700 consists of 545K training videos and 35K validation videos spanning 700 action categories.
We evaluate performance of mPLUG-2 in terms of image classification on ImageNet-1K (Deng et al., 2009).
ImageNet-1K contains 1.28M training images and 50K validation images from 1,000 classes.