BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Ruyang Liu, Chen Li, Yixiao Ge, Ying Shan, Thomas H. Li, Ge Li
Introduction
Since the past year, Large Language Model (LLM) chatbots (Touvron et al., 2023; Wang et al., 2022c; Ouyang et al., 2022; Zeng et al., 2022) have emerged as one of the most remarkable advancements within the AI community. Excitingly, the LLM-driven AI assistants have showcased impressive abilities in comprehension, reasoning, and engaging in conversations. The success of text-only dialogue systems has also sparked the development of image-language conversation agents (Liu et al., 2023a; Dai et al., 2023; Zhu et al., 2023; Gao et al., 2023). These agents combine pretrained image models with LLMs, followed by instruction tuning, to create a fusion of visual and textual understanding for enhanced conversation capabilities.
In comparison to images, videos offer a more comprehensive representation of how humans perceive and interpret the world. Nevertheless, constructing a video-centric dialogue system is a more complex endeavor compared to the image-based one. Firstly, a typically multimodal agent, comprising LLM, a pretrained image encoder (usually CLIP-L (Radford et al., 2021)), and additional learnable modules, is already demanding in terms of GPU memory. The incorporation of video dialogue introduces even higher costs, as multiple frames need to be fed as inputs. Secondly, a pretrained encoder as potent and knowledge-enriched as CLIP is currently lacking in the realm of videos, leading to inferior visual understanding. Lastly, gathering video-text instruction data of comparable scale, quality, and diversity to images poses a significant challenge. Rather than building the video agent from scratch, an alternative and promising strategy involves adapting existing pretrained image-centric dialogue models to the video domain (Li et al., 2023b; Maaz et al., 2023; Zhang et al., 2023; Luo et al., 2023), which can leverage the abundant knowledge embedded within image models.
Expanding pretrained image models to the video domain is an extensively explored area in computer vision. The crux of the matter lies in enhancing the capability of 2D models to model temporal dynamics (Bertasius et al., 2021; Carreira & Zisserman, 2017; Tran et al., 2015). However, implementing effective temporal modeling within image dialogue models presents notable challenges, primarily due to the trade-off between efficiency and effectiveness in CLIP-based temporal modeling techniques. Methods with strong temporal modeling capabilities, e.g., joint-ST modeling (Bertasius et al., 2021; Wang et al., 2022b; Li et al., 2023c; Xue et al., 2022) and interpolate-style modeling (Ni et al., 2022; Pan et al., 2022), often necessitate finetuning of the entire visual encoder, thereby exacerbating the already significant GPU memory consumption of video conversation models. Moreover, full-finetuning would lead to the loss of knowledge encapsulated in pretrained models, resulting in reduced performance for image-based conversations In contrast, methods that focus on parameter-efficient temporal modeling aim to keep the image encoder frozen and introduce only a few trainable parameters, which are embraced by most video dialogue models. However, these approaches typically provide limited temporal modeling and struggle to capture crucial spatial-temporal features.
To tackle the problems in the aforementioned methods, we proposed Branching Temporal Adapter (BT-Adapter), a novel framework to migrate image-text pretrained models to the video domain. As the name implies, BT-Adapter incorporates a branching spatial-temporal module for temporal modeling. Alongside the pretrained image model, BT-Adapter inherits the parameter efficiency advantage from conventional adapters (Houlsby et al., 2019) while concurrently achieving effective temporal modeling. Notably, unlike the plug-style adapter, BT-Adapter does not disrupt the forward progression of the pretrained models, thereby safeguarding the integrity of the pretrained multimodal knowledge. After being pretrained with any version of CLIP, BT-Adapter can be seamlessly integrated with all image conversation models using this version of CLIP to activate video conversation, without necessitating video instructions, e.g., openai-CLIP for LLaVa, Eva-CLIP for MiniGPT4 and InstructionBLIP. Additionally, we have devised a unique asymmetric masking mechanism that exclusively implements tube token masking within the BT-Adapter. Building upon this, we formulate two custom training objectives for the BT-Adapter: Masked Branching Token Alignment (MBTA) and Masked Branching Cross-modal Alignment (MBCA). This approach not only reduces computational demands and accelerates convergence but also yields improved outcomes.
To validate the effectiveness and efficiency of our BT-Adapter, we provide a detailed analysis of various temporal modeling strategies for video conversation in Sec. 3.2. Furthermore, in Sec. 4, we carry out extensive qualitative and quantitative experiments on BT-Adapter, encompassing both traditional video tasks and video conversations. As depicted in Fig. 1(left), a series of designs ensures that BT-Adapter is highly resource-efficient: our post-pretraining demands just 8 V100(32G) GPUs in a mere 3 hours, leading to a reduction in carbon emissions by more than 10240 and 2687 compared to CoCa (Yu et al., 2022) and InternVideo (Wang et al., 2022b) respectively. Building upon this, we still achieve state-of-the-art results in zero-shot video-text retrieval. Regarding video conversation, as shown in Fig. 1(right), we clearly demonstrate that fine-tuned BT-Adapter surpasses the previous state-of-the-art by a significant margin across all benchmarks, and BT-Adapter without instruction tuning has better average performance than fine-tuned SOTAs.
Related Work
Temporal Modeling on Image-Language Pretrained Models. With the wide application of pretrained image-language models (Radford et al., 2021; Yu et al., 2022), how to extend these pretrained multimodal models into the video domain has emerged as a novel yet critical research topic. Existing temporal modeling methods can be broadly categorized into several types. The most straightforward and commonly used one is joint-spatial-temporal modeling (Bertasius et al., 2021; Wang et al., 2022b; Li et al., 2023c; Xue et al., 2022). By inputting all video tokens into the image encoder, this approach can effectively model temporal dependencies without other techniques. However, joint-ST modeling necessitates fine-tuning the entire encoder, resulting in significant computational costs and the degradation of pretrained knowledge. Another representative type is the interpolate-style temporal modeling, including separated-ST modeling (Bertasius et al., 2021; Zeng et al., 2023), message token (Ni et al., 2022), and ST-Adapter (Pan et al., 2022). Inserting temporal modules between the pretrained spatial layers, interpolate-style modeling shares similar cons and pros with joint-ST modeling. Specifically, ST-Adapter is also known as the parameter-efficient temporal module, where the distinctions with our methods are detailed in Sec. D in the Appendix. Different from the former two types, concatenate-style temporal modeling (Luo et al., 2022; Fang et al., 2021) attaches temporal modules following the pretrained encoder, which allows for the freezing of the backbone and the preservation of knowledge. Nevertheless, this style offers limited temporal modeling, as it hardly captures crucial low-level spatial-temporal features. In Sec. 3.2, we will experimentally compare the different temporal modeling methods for video dialogue.
In contrast to the preceding types, BT-Adapter adopts a branching structure for temporal modeling. Thanks to the intricate design, we circumvent the issues outlined above, enabling proficient temporal modeling, efficient parameter fine-tuning, and multimodal affinity simultaneously. Most similar to our methods, STAN (Liu et al., 2023b) also introduced branch-style temporal modeling. Nonetheless, there are three key distinctions between our work and STAN: (1) Our emphasis is on zero-shot image-to-video transfer and video conversations, areas that have never been explored by STAN. (2) Our performance is notably superior when the backbone is frozen, which is attributed to our unique design like temporal CLS token, zero-initialized temporal projection, and temporal selection. (3) We have introduced the novel training strategy and objectives tailored for the branching temporal structure. In Sec. D in the Appendix, we will also provide an experimental comparison between BT-Adapter and its similar competitors, including ST-Adapter and STAN.
Video Conversation. Recent advancements in multimodal learning have been predominantly propelled by the fusion of visual models with LLM. Yuan et al. (2021) initially demonstrated the potential of combining visual models with LLM. Blip-2 (Li et al., 2023a) further proposed Q-former that maps visual tokens into the text embedding space. Subsequently, several methods presented visual instruction tuning (Liu et al., 2023a; Zhu et al., 2023; Dai et al., 2023) for image-LLM to enable visual conversation. Most closely related to our topic, there have been several developments in the field of video-centric dialogue models over the past few months (Li et al., 2023b; Zhang et al., 2023; Maaz et al., 2023; Luo et al., 2023). Generally, these models consist of the visual encoder, LLM, and temporal module, tuned with video instruction data to realize video conversation. However, in pursuit of efficient training, the temporal modules utilized by these models, such as temporal position embedding (Zhang et al., 2023) and temporal pooling (Maaz et al., 2023; Luo et al., 2023), often exhibit limited temporal modeling capabilities. In contrast, our approach maintains parameter-efficient training while simultaneously delivering effective temporal modeling. Thanks to the benchmark initially raised by Maaz et al. (2023), we can quantitatively illustrate the superiority of our methods in Sec. 4. Moreover, we introduce a novel setting of zero-shot conversation by integrating pretrained temporal modules with image-centric dialogue models, demonstrating the feasibility of video conversation without any video instruction tuning.
Methodology
In this section, we elaborate on how to efficiently enhance pretrained image-language models (e.g.,, CLIP (Radford et al., 2021) and LLaVa (Liu et al., 2023a)) with temporal modeling capabilities while preserving the multimodal pretrained knowledge. The main framework is depicted in Fig. 2 . We will introduce the model architectures of BT-Adapter in Sec. 3.1, the exploration of temporal modeling in video dialogue models in Sec. 3.2, and the training strategy and objectives in Sec. 3.3.
As the pioneer of contrastive image-text pretraining, CLIP (Radford et al., 2021) has founded widespread application in various domains. Meanwhile, fundamental image dialogue models like Blip-2 (Li et al., 2023a) and LLaVa (Liu et al., 2023a) all employ CLIP as visual encoder. Hence, without losing the generalizability, we focus on the adaption of CLIP.
Model Architecture. To enable video input, within CLIP, we treat each frame as an individual image. Given a video with T frames, we divide each frame into N non-overlapping patches, represented as and , where denotes the [CLS] token. Then, tokens in each frame are added with spatial position and fed independently into the CLIP layers:
In contrast to the plug-style adapter that consists of multiple independent modules, BT-Adapter is a continuous network operating as a branch alongside the main backbone. Inside BT-Adapter, we adopt divided space-time attention (Bertasius et al., 2021). Given all video tokens, we first gather patch tokens in the same position across different frames, obtaining and . Then, tokens in each position are fed individually into the temporal layers:
Backbone-Branch Interaction. Unlike traditional adapters, which are added to every pretrained layer, our branching adapter has significantly fewer layers compared to the main backbone. Assuming we have K layers in the branch, BT-Adapter takes the output of the last K+1 CLIP layers as input, each layer on both sides corresponding pairwise. To construct the input of the first BT-Adapter layer from the CLIP layer, we first develop a new learnable video [CLS] token to represent the entire video. Then, we concatenate with and update the patch embeddings with frozen spatial positional embeddings and learnable temporal positional embeddings:
where is the spatial position embedding shared with CLIP while is the temporal position embedding. For any other layer of BT-Adapter, we construct its input from the previous branching layer and the CLIP layer at the same level with weight selecting as follows:
2 Temporal Modeling for Video Conversation
In this section, we conduct an empirical study to explore potential temporal modeling strategies for video dialogue models, demonstrating the advantages of our approach. We argue that an ideal temporal modeling approach for video-LLM models should meet several criteria: it should be parameter-efficient, as dialogue models are already quite large; it should be multimodal-friendly, preserving as much multimodal alignment knowledge as possible; and it should be temporal-sensitive, delivering strong performance in time-sensitive scenarios. To measure the “multimodal-friendly” and “temporal-sensitive”, we employ the “Correctness of Information” and “Temporal Understanding” metrics from the VideoChatGPT benchmark (Maaz et al., 2023) respectively. “parameter-efficient” is simply decided by whether the CLIP can be frozen. Next, we will explain how we implement these methods for video conversation.
Spatial-Temporal Pooling. Following Maaz et al. (2023), we do not integrate any module in CLIP and encode each frame independently. Then, all patch tokens are pooled along the time () and spatial () dimensions, resulting in a total of tokens. These tokens are subsequently input into the LLM, which has much fewer tokens compared to joint or separate temporal modeling.
Joint-ST modeling. Following Xue et al. (2022), we incorporate spatiotemporal positional embeddings to 2D patches and feed all video tokens into CLIP and LLM simultaneously. In this way, tokens in any position or frame can attend to each other, providing a straightforward yet effective method for modeling temporal dependencies. To make training feasible, we utilize the FSDP (Zhao et al., 2023) to facilitate parameter and gradient sharing between GPUs.
Separate-ST modeling. Unlike joint-ST modeling, separate-ST modeling retains and leverages CLIP’s spatial layer. All patch tokens are fed in LLM. Following Zeng et al. (2023), we insert temporal attention before each CLIP layer and add temporal position embeddings to the input. Regarding the backbone, separate-ST modeling offers flexibility, allowing us to experiment with both frozen and unfrozen CLIP layers.
Branch Temporal Adapter. We directly equip CLIP-L/14 with BT-Adapter in Sec. 3.1 to achieve video encoder. In LLava, LLM takes the output from the second-to-last layer of the visual encoder as its input. Hence, we also take the second-to-last output from both CLIP and BT-Adapter and combine them with learnable balance weight. Finally, combined patch tokens are pooled along the time () dimension, resulting in only tokens for inputting into LLM.
Results. We replace CLIP-L/14 in LLaVa with the implemented video encoders and then proceed with video instruction tuning. In all methods, we keep LLM frozen while opening the linear projection between the visual encoder and LLM. No pretraining is included for fair comparison. Results are presented in Table 1. It can be observed that methods with a frozen CLIP tend to be parameter-efficient but perform poorly in terms of temporal modeling. On the other hand, methods with good temporal understanding suffer from high computation costs and knowledge loss. In contrast, our BT-Adapter achieves a balance by being parameter-efficient, multimodal-friendly, and temporal-sensitive simultaneously. This gives it a clear advantage over other temporal modeling methods for video conversation.
3 Pretraining with Asymmetric Masking
As proved by previous studies (Zeng et al., 2023; Wang et al., 2022b; Xue et al., 2022), CLIP-based video encoder can harvest stronger performance on downstream video tasks after being post-pretrained on large-scale video-text data. Hence, we also involved our BT-Adapter with video-text pretraining. However, the expensive computation cost has always been a bottleneck when scaling up the training. Inspired by the recent success of masked modeling in visual-language pretraining (Li et al., 2023d; Tong et al., 2022), we develop a unique asymmetric masking strategy for BT-Adapter. Specifically, we maintain the tokens in the frozen CLIP unchanged while applying a tubular mask to the tokens in BT-Adapter. This mask randomly masks a certain percentage () of patch tokens in the same position across different frames. This approach allows us to maximize the retention of pretraining knowledge from CLIP while reducing spatial-temporal redundancy. Thanks to the asymmetric mask, we can maintain a high mask ratio () without compromising performance, leading to a reduction of at least half of the computational budget. As a result, with the assistance of frozen backbone and token masking, we can accomplish the costy video-text pretraining in just a few hours. Furthermore, based on asymmetric masking, we have devised two tailor-made training objectives for the branching temporal structure, in addition to the Video-Text Contrastive.
Video-Text Contrastive (VTC). VTC is the most widely-used basic objective for cross-modal alignment. Given the global video feature in Eq. 5 and global text feature , we formulate as:
where is the batch size and is the index in a batch and is the temperature scale.
Masked Branching Token Alignment (MBTA). Masked video modeling has been proven beneficial for spatial-temporal representation learning (Tong et al., 2022), but pixel reconstruction in Tong et al. (2022) is computationally expensive. In our approach, we have a frozen CLIP and the masked BT-Adapter, where the unmasked CLIP is naturally a teacher for the masked branch. Hence, we can conduct the token alignment in an end-to-end pretraining without extra forward propagation. Specifically, we compute the mean squared error (MSE) between the unmasked tokens in BT-Adapter and the corresponding tokens in CLIP in the last layer:
Masked Branching Cross-Modal Alignment (MBCA). In our model, the backbone and the branch handle different aspects: CLIP encodes static spatial features, while BT-Adapter captures dynamic information. Therefore, we additionally align the branching patch tokens with the text embedding to enhance temporal learning, which is formulated as:
Different from , we utilize the patch tokens for the alignment. This choice is made because the information is highly centralized in the CLS token for CLIP, whereas downstream applications like multimodal conversation or generation typically rely on the patch tokens as input.
Experiments
Tasks and Datasets. We have evaluated our BT-Adapter on two main aspects: traditional video tasks and video-centric dialogue. To begin with, we pretrained the BT-Adapter on WebVid-2M (Bain et al., 2021). For traditional video tasks, we consider zero-shot text-to-video retrieval and zero-shot video recognition, covering seven benchmarks: (a)MSR-VTT (Xu et al., 2016). (b)DiDeMo (Anne Hendricks et al., 2017). (c)LSMDC (Rohrbach et al., 2017). (d)ActivityNet (Caba Heilbron et al., 2015). (e)Kinetic-400 (Carreira & Zisserman, 2017). (f)HMDB-51 (Kuehne et al., 2011). (g)UCF-101 (Soomro et al., 2012). We used Recall@K (R@K) for retrieval tasks and Top-K Accuracy (A@K) for recognition tasks as the evaluation metrics.
For video dialogue, we conducted evaluations for the zero-shot video conversation (without instruction tuning) and instruction-tuned video conversation. The VideoChatGPT-100K (Maaz et al., 2023) is employed for supervised instruction tuning. To assess the quality of the responses quantitatively, we employ the VideoChatGPT benchmark, which consists of five metrics for video-based generative performance benchmarking and six metrics for zero-shot question-answer evaluation. Additional details and settings of each dataset can be found in Sec. A in the Appendix.
Implementation Details. We employ openai-CLIP-L/14 as the backbone for its wide application. We adopt 4 layers of BT-Adapter in default. During pretraining, the masking ratio is set as 70%. The temperature scale is fixed as 0.01 for contrastive loss. The weight of the three losses is . It takes 3 hours to train one epoch on WebVid on 8 V100-32G GPUs. For video conversation, we implement the instruction tuning based on BT-Adapter-LLaVA. InstructionBLIP and miniGPT4 are also included for zero-shot evaluation. During instruction tuning, we update the linear projection and BT-Adapter, while keeping the rest architecture frozen. It takes 3 hours to train three epochs on 8 A100 40GB GPUs. More training details are listed in Appendix.
2 Quantitative evaluation
Traditional Video Tasks. The results of zero-shot text-to-video retrieval are presented in Table 2. Compared to the previous SOTAs, BT-Adapter achieves competitive performance across all datasets, obtaining the best results on most metrics. For instance, in comparison to UMT and Singularity, although there is a slight lag in terms of R@1 on DiDeMo and LSMDC respectively, we have surpassed them by more than 5% on MSRVTT and ActivityNet. Furthermore, we achieve superior results with significantly fewer pretraining scales and GPU hours than all the mentioned methods. For example, we use 130 fewer GPU hours than UMT, 560 fewer than TVTSv2, and 2687 fewer than InterVideo while still outperforming them. The results of zero-shot action recognition are posted in Sec. C and Table 9 in Appendix.
Video Dialogue. Thanks to the benchmarks introduced by Maaz et al. (2023), we can quantitatively compare the performance of various video conversation models. The results are presented in Table 3 and Table 5, where we compare BT-Adapter with all existing video-centric dialogue models, including VideoLLaMA (Zhang et al., 2023), LLaMA-Adapter (Gao et al., 2023), VideoChat (Li et al., 2023b), and VideoChatGPT (Maaz et al., 2023). It is observed that our zero-shot model, even without any instruction tuning, outperforms methods that require instruction tuning on average. When fine-tuned using video instruction data, the superiority of our approach becomes even more pronounced. These results underscore the effectiveness of the BT-Adapter as a superior method for temporal modeling in video conversation models compared to existing approaches. Notably, our BT-Adapter exhibits a significantly larger performance margin over other methods on ActivityNet in both Table 5 and 2, highlighting its particular strength in handling long video sequences. Moreover, we have integrated the BT-Adapter with various pretrained image conversation models. As illustrated in Table 5, pretrained BT-Adapter yields consistent advancement on all image-language chatbots without extra instruction tuning, underscoring the broad applicability of our method.
3 Ablation Study
Model Structure. To investigate the components within the BT-Adapter, we initiated our study by exclusively employing VTC for pretraining and subsequently evaluating the zero-shot performance across retrieval tasks and video conversation, as presented in Table 7. Initially, we explored the implementation of separate-ST modeling within the last 4 layers. However, the resulting improvements were marginal. Subsequently, we transitioned the separate-ST network into the branch, where we witnessed notable progress. This shift underscores the efficacy of branching modeling. In the final step, we studied the interaction module, i.e., multi-level selective combination between backbone and branch. This addition led to further enhancements across all benchmark datasets. The more fine-grained ablations on model components can be found in Sec. E in the appendix.
Training Objectives. We investigated the influence of two innovative training objectives in Table 7, where all experiments were conducted using BT-Adapter with a mask rate of 70%. Our observations reveal that both MBTA and MBCA yield improvements in downstream results, with MBCA demonstrating more substantial progress. This outcome suggests that individually aligning the temporal module and the textual output effectively mitigates the limitations inherent in pretrained image-text models. Moreover, the combination of both objectives results in further advancements.
Number of Branching Layers and Masking Rate. The number of branching layers and the masking rate in our model inherently involve a trade-off between computational resources and performance. To determine the optimal settings, we conducted ablation experiments. Firstly, in Fig. 4, we present the results of zero-shot retrieval and conversation as we vary the number of layers. Notably, we observed that the performance improvement plateaus at around 4-6 layers. Secondly, in Fig. 4, we provide results for zero-shot retrieval and the corresponding convergence GPU time across different masking rates. Our findings indicate that the results remain stable when the masking rate is set at or below 70%, while higher masking rates significantly lower the training time. Therefore, we have chosen to maintain a masking rate of 70%.
4 Qualitative Results
In Figure 5, we present a qualitative example of video conversation. Unlike the general description from VideoChatGPT, our model provides an informative and accurate response to a video-related question, highlighting the effectiveness of the BT-Adapter in video comprehension. Additional examples of video dialogues covering various aspects can be found in the Appendix.
Conclusion
This paper presents a novel approach to achieve parameter-efficient yet effective image-to-video adaptation and video conversation. Our proposed solution, named Branching Temporal Adapter (BT-Adapter), is a branching separate-ST network for temporal modeling. Building upon the BT-Adapter, we introduce an asymmetric masking technique along with two novel training objectives. Extensive experiments demonstrate the superiority of our model over other temporal modeling methods for video conversation. By seamlessly integrating the BT-Adapter with any pretrained image conversation model, we achieve video dialogues without necessitating manual video instruction tuning. With significantly lower computation costs, we attain state-of-the-art results across two zero-shot video tasks and video conversations.
References
Appendix A Datasets
Test-to-Video Retrieval. The settings of the four zero-shot retrieval benchmarks are presented as follows: (1) MSRVTT (Xu et al., 2016), a widely used video-text retrieval benchmark, comprises 10,000 YouTube videos, each accompanied by 20 captions. Our reported results are based on the 1K-A split, which consists of 9,000 training videos and 1,000 testing videos. For MSRVTT, we sample 12 frames for each video and set max token length as 12. (2) DiDemo (Anne Hendricks et al., 2017) comprises 10,611 videos gathered from Flickr, along with 40,000 sentences. To form queries, we concatenate all captions associated with a video. We use a frame number of 64 and a mask token length of 64, consistent with prior research. (3) LSMDC (Rohrbach et al., 2017) consists of 118,081 videos extracted from 202 movies. We configure it with a frame number of 12 and a maximum token length of 32. (4) ActivityNet (Caba Heilbron et al., 2015) comprises 20,000 YouTube videos. To create queries, we concatenate all video descriptions into paragraphs. Our evaluation focuses on video-paragraph retrieval using the ’val1’ split. We set the frame number and maximum token length to 64.
Action Recognition. In all video recognition datasets, we refrain from utilizing templates such as “a video of **” and instead employ the tag itself as the textual query. The parameters for frame number and maximum token length are consistently configured at 16 and 12 respectively. The statistics pertaining to the three zero-shot action recognition benchmarks are provided below: (1) Kinetics-400 (Carreira & Zisserman, 2017) is a widely recognized dataset for video action recognition. It comprises a substantial collection of 260,000 videos, each with an average duration of approximately 300 frames. The dataset encompasses a diverse set of 400 action classes. (2) HMDB-51 (Kuehne et al., 2011) includes a total of 5,000 videos spanning 51 distinct action categories. The dataset is partitioned into training and test sets, with 3,500 videos allocated for training and 1,500 videos for testing. (3) UCF-101 (Soomro et al., 2012) comprises a comprehensive collection of 13,000 videos, representing 101 unique action categories. Within this dataset, the training set consists of 9,500 videos, while the test set contains 3,500 videos.
Video-Text pretraining. We adopt the WebVid2M (Bain et al., 2021) for pretraining, laying the foundation for the BT-Adapter’s video encoding capabilities. WebVid2M is a substantial video-text pretraining dataset composed of short videos paired with textual descriptions, sourced from stock footage sites. This dataset is characterized by its vast scale, encompassing approximately 2.5 million video-caption pairs and totaling 12,000 video hours. The videos within WebVid2M exhibit a rich diversity of content. During the pretraining, we configured the frame number and maximum token length to be 8 and 32 respectively.
Video Conversation. VideoChatGPT benchmark (Maaz et al., 2023) s the first benchmark designed for the quantitative evaluation of video conversation models. It was collaboratively annotated by ChatGPT and human annotators using the ActivityNet dataset, resulting in a dataset containing 100k video-text instruction pairs. For the video-based text generation benchmark, a test set was curated based on ActivityNet, which included captions and associated question-answer pairs obtained from human annotations. The evaluation pipeline used the GPT-3.5 model and assessed the model’s performance in various aspects, including Correctness of Information, Detail Orientation, Contextual Understanding, Temporal Understanding, and Consistency. The pipeline assigns a relative score to the generated predictions on a scale of 1 to 5 for each of these aspects. For zero-shot question-answer evaluation, three open-source video QA datasets were employed: MSRVTT-QA, MSVD-QA, and ActivityNet-QA. Also, GPT was used as the zero-shot evaluation assistor to assign relative scores on a scale of 1 to 5 for generated answers.
Appendix B Implementation Details
All experiments were conducted using PyTorch (Paszke et al., 2019). The pretraining and zero-shot inference processes were implemented based on mmaction2.0 (Contributors, 2020). Our configuration settings are detailed in Table 8, with the exception of specific cases where alternate configurations were used. It is noteworthy that our data augmentation techniques are notably simpler in comparison to those employed by other methods.
Appendix C Zero-Shot Results on Action Recognition
The results of zero-shot video recognition are reported in Table 9. Despite being pretrained solely on video-language datasets, BT-Adapter consistently contributes to the video-only task, achieving state-of-the-art zero-shot results among CLIP-based methods. Notably, even when compared to InternVideo, which employed self-supervised reconstruction during pretraining (proven to be more effective on single-modality tasks than contrastive learning), BT-Adapter still outperforms it, underscoring the effectiveness of BT-Adapter in video encoding and spatial-temporal modeling.
Appendix D Experimental Comparison With Similar Methods
In this section, we conduct an experimental comparison between two closely related works, ST-Adapter (Pan et al., 2022) and STAN (Liu et al., 2023b). ST-Adapter is also notably recognized as a parameter-efficient method for temporal modeling, while STAN also employs the branching temporal modeling strategy. We pretrain the three methods on MSRVTT for one epoch first, and the results of zero-shot performance on MSRVTT retrieval and video conversation are presented in Table 10. Initially, it is evident that ST-Adapter exhibits suboptimal results across all metrics. This outcome may be attributed to the fact that ST-Adapter is a single-modality temporal adapter, where the insertion of 3-D convolutions between transformer layers may lead to the rapid degradation of the pretrained multimodal knowledge. Next, we assess STAN under two conditions: with frozen CLIP and without. The results reveal that STAN, when used with an open CLIP, performs admirably in zero-shot retrieval tasks. However, it exhibits poorer outcomes in video conversation tasks, and it requires significantly longer pretraining hours. Conversely, when STAN is employed with a frozen CLIP, it shows improvements across all metrics, although it still falls short of BT-Adapter in all aspects. In contrast, BT-Adapter achieves both efficiency and effectiveness simultaneously, underscoring the superiority of our design over ST-Adapter and STAN in the context of zero-shot video encoding and video conversation.
Appendix E More Ablation Results
Temporal Projection and Initialization. We examine the appropriate way for instantiating the temporal projection. As demonstrated in Table 11(above), random initialization for the projection yields performance results similar to those obtained without projection. In contrast, zero initialization outperforms them by a significant margin. This suggests that building temporal reasoning capability from scratch, as opposed to random initialization, mitigates adverse effects on the well-established spatial prior. Consequently, zero initialization is better suited for knowledge transfer from images to videos. Backbone-Branch Combination. We further perform an ablation study to explore the most effective method for combining the output from the backbone and the branch, considering three approaches: direct addition, weighted selection, and concatenation with subsequent linear projection. As illustrated in Table 11(below), weighted selection yields the most favorable results. This observation suggests that different layers and samples require distinct degrees of information from the backbone and the branch.
Appendix F More Qualitative Results
In Figures 6, 7, 8, and 9, we present a comprehensive overview of the qualitative results obtained in video dialogues, encompassing diverse aspects. These visualizations vividly illustrate the capacity of our BT-Adapter to provide contextually appropriate responses in a variety of scenarios where temporal sensitivity is paramount. These results serve to underscore the efficacy of the BT-Adapter in video understanding. In Figure 10, we present a notable outlier case in which our method encounters challenges, where the BT-Adapter struggles to recognize the text content within the frames. This particular instance sheds light on the fact that, while the BT-Adapter diligently strives to preserve pretraining information to the greatest extent possible, it may still introduce some disruption to the pretraining knowledge compared to the fully concatenation-based modeling of VideoChatGPT.