Long Context Transfer from Language to Vision
Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, Ziwei Liu
Introduction
Driven by the progress of Large Language Models (LLMs) , multiple studies are conducted to extend their capability to understand images and videos . With modality alignment and visual instruction tuning, these Large Multimodal Models (LMMs) have shown impressive abilities such as captioning and visual question-answering. While current LMMs have demonstrated promising performance on tasks involving single images and short videos , effectively processing and understanding extremely long videos remains a significant challenge .
One primary reason for this challenge is the excessive number of visual tokens generated by the vision encoder. For instance, LLaVA-1.6 can produce 576 to 2880 visual tokens for a single image. The number of visual tokens increases significantly with the addition of more frames. To address this problem, numerous methods have been proposed to reduce the number of visual tokens. One popular direction is to modify the visual resampler that connects the vision encoder and LLM, aiming to extract fewer tokens . Alternative approaches employ heuristic techniques to prune or merge the visual features. However, despite these efforts, Table 1 demonstrates that the majority of current LMMs are still limited in their ability to process a large number of frames effectively.
Another issue hindering the development of high-performance long video LMMs is the lack of high-quality long video datasets. In Table 3, we list the average video length of existing video instruction tuning data. Most datasets consist of video clips within 1 minute. Even if some datasets do contain longer videos, the corresponding text pairs are generated by annotating only several frames within that video, lacking long and dense supervision signals.
Given the circumstance, in this paper, instead of reducing the visual tokens, we identify the more critical issue limiting the visual context length in existing LMMs: the context length of the language model backbone. Given a language model, we first extend its context length by training on longer text data. We then use this context-extended LM as the backbone to perform modality alignment and visual instruction tuning without any long video text pairs. By training this way, the context length of the language model is directly transferred to that of the LMMs. We further proposed UniRes, a unified encoding scheme that represents videos as extended images, enhancing the capability fusion between images and videos. To facilitate benchmarking and accurately assess the context length in the visual domain, we created V-NIAH, a synthetic visual benchmark based on the Needle-in-a-haystack test used in language models. Our model, Long Video Assistant (LongVA), is capable of accurately retrieving visual information from 2000 frames or more than 200K visual tokens. Experiments show that additional frames during inference lead to improved performance on long video question-answering benchmarks, and LongVA achieves state-of-the-art performance among 7B models on the Video-MME and MLVU dataset. In summary, our paper makes the following contributions:
(1) Long Context Transfer: We discovered the long context transfer phenomenon where the context of the language model can be directly transferred to the modality-aligned multi-modal models.
(2) Visual Needle-In-A-Haystack (V-NIAH): We proposed the V-NIAH benchmark to test LMMs ability in locating and retrieving visual information over extremely long contexts.
(3) Long Video Assistant (LongVA): With long context transfer and UniRes, we developed LongVA that can perceive more than 200K visual tokens, achieving SoTA performance on the Video-MME and MLVU dataset.
Related Work
Existing studies explore different architectures to extract and inject visual features into LLMs. One line of work , pioneered by Flamingo , adopts a resampler to compress the visual feature and inserts cross-gated attention layers into the LLM. Some other works still use a reampler while directly feeding the image feature into the input layer of the language model. The LLaVA series use a simple and scalable design to directly project the image features into language model without any pooling or resampling. When the field moves from image-only models to include multi-image and video inputs, more modifications to the visual language connector were proposed. and use a simple average pooling. dynamically drop the visual tokens. adopt a spatial-temporal convolution to better capture the dynamics of video data and reduce feature size. Our proposed context transfer from text to image is orthogonal to those works and can further enable LMMs to understand more frames.
Context Extrapolation in Transformer
Transformer does not directly work on sequences longer than its training length. To alleviate that, various RoPE-based extension techniques have been proposed to allow for training-free context extrapolation. Efforts have also been made on data curation and system optimization during long context training. There has been limited exploration of the context extrapolation in the domain of LMMs. are closest to our work and train LMM with long context language models, but they do not benchmark the effective visual context length of their model.
Video Language Benchmarks
Recent years have witnessed significant progress in Video Question-Answering. To accurately measure the progress of the video LMMs’ performance, researchers have developed various benchmarks encompassing a broad spectrum of tasks. These range from fundamental visual perception tasks such as activity recognition, concept detection , and counting , to more complex visual reasoning tasks including compositional , causal , and situated reasoning . However, most of those benchmarks focus on short videos, lacking data and metrics to test LMMs’ capability over a long context. Inspired by the NIAH test in the language model community, we proposed V-NIAH to benchmark LMMs’ ability over long visual inputs with the minimum overhead of data collection and human annotation. Several concurrent works also developed multimodal versions of the Needle-in-a-haystack test . However, they only measure on several hundreds of frames and lack a strong baseline to properly analyze the properties of visual context length.
Long Video Assistant
As in Figure 1, this paper centers around the hypothesis that if the modality of vision and language can be truly aligned, the capability to handle long contexts could also transfer from text to vision, and this could happen even without explicit long video training. Our methodology is thus very straightforward. Given a language model, we first perform long context training purely on language to extend its text context (Section 3.1). We then detailed how we augment this language model with long visual capabilities by training solely on short image data in Section 3.2.
We use Qwen2-7B-Instruct as the backbone language model and perform continued pretraining with a context length of 224K224K is the maximum we can fit with 8A100-80G for Qwen-2-7B. We find that the embedding size significantly impacts the maximum sequence length in our optimized codebase. Qwen2 has a huge vocabulary of 152K tokens. For LLaMA2 with 32K vocabulary, we can train it with 700K context length. over a total of 900M tokens. We follow to increase RoPE base frequency during the continued pertaining and specifically set it to 1B. A constant learning rate of 1e-5 is maintained for a batch size of one million tokens across 1,000 training steps. Following , we construct the dataset used for long context training from Slimpajama by upsampling documents longer than 4096 and keeping the domain mixture ratio unchanged. Multiple documents are packed into a single sequence separated by a BOS token.
We employed several optimization strategies to perform training on such long sequences. These includes FlashAttention-2 , Ring Attention , activation checkpointing, and parameter offload . To balance the load across different GPUs, we shard the sequence in a zigzag way in ring attention. The resulting training framework is memory efficient and maintains very high GPU occupancy. Note that we do not use any parameter-efficient methods such as LoRA or approximate attention . With those optimizations, the compute used in long context training is minimal compared to that of language model pretraining, making it feasible for academic budgets. The long context training can finish in 2 days with 8 A100 GPUs.
In Figure 7, we evaluate the extended Qwen2 with the Needle-in-a-haystack (NIAH) test . It achieves perfect results within the training context length (224K) and generalizes even further. We find the vanilla NIAH to be a relatively trivial benchmark and further test it with 5 distractors randomly inserted into the documents. The detailed configuration can be found in Appendix A.
2 Aligning Long Language Model Using Short Vision Data
Inspired by the AnyRes encoding scheme in LLaVA-NeXT , we designed UniRes that provides a unified encoding scheme for both images and videos, as shown in Figure 2. Unlike AnyRes which retains a small base image and flattens ViT patches across the grids, UniRes removes the base image, flattens patches within each grid, and 2x2 pool the visual features by default (Appendix B). This approach allows us to maintain consistent representation when extending image data into videos where multiple frames are viewed as multiple grids in a row.
Specifically, UniRes divides an input image of resolution into smaller grids, each with a resolution of pixels. This results in grids. For very high-resolution images, we limit the maximum number of grids to 49, resizing images larger than this threshold. Each grid is separately encoded using CLIP-ViT-L-336px and then projected through a 2-layer MLP to match the LM’s input dimension, resulting in 576 features per grid. We then apply 2x2 average pooling, finally converting an image into tokens. During inference, this visual encoding scheme allows videos to be represented as very long images (even though we do not train on videos). An -frame video is treated as an image of size , divided into grids where each grid corresponds to a video frame. Using CLIP encoding, MLP projection, and average pooling, an -frame video is encoded into visual tokens.
To clearly ablate the long context transfer phenomenon from language to vision, we adopt a train short, test long protocol where we only use image-text data during training, but test on long videos. We trained our model using the same data recipe and two-stage training approach as LLaVA-1.6. Our experiments show that compared to AnyRes, UniRes has slightly lower scores on low-resolution image benchmarks (Table 7) but performs better on V-NIAH (Figure 4) and Video-MME (Table 4). We believe the unified encoding scheme for images and videos is crucial, thus choosing this as the encoding scheme of LongVA. The image-text alignment can be finished in 1.5 days. With 2 days for long context training on text, the total training cost of LongVA is 3.5 days on 8A100-80G.
It is worth noting previous work largely inspired the design choice of LongVA. For example, first demonstrates the effectiveness of long context continued pretraining with increased RoPE base frequency (thus decreasing the rotation angles). We sample the long text data following the guidance of . We adopt the same vision encoder and training data as that of LLaVA-1.6 . We try to keep our methods as simple as possible to clearly show the phenomenon of long context transfer without other confounders.
V-NIAH
To measure the context length of language models on extremely long input, earlier works calculate perplexity scores over long documents. Recently, many have started using the Needle-in-a-Haystack (NIAH) test to benchmark LLMs’ ability to retrieve long context information precisely. We note that there is so far no benchmark to measure the visual context length of LMMs. To evaluate LongVA’s capacity to locate and retrieve long-range visual information, we extend the NIAH test from text to video and propose V-NIAH.
As shown in Table 9, we designed 5 video question-answering problems as the needle and inserted each as a single frame into hours-long videos. We sampled the videos at 1 FPS as the visual input. The image of the needle is sourced from existing VQA benchmarks or AI-generated to avoid any contamination. The AI-generated images and questions are purposely chosen to be "counterfactual" or "counter-commonsense", ensuring the model cannot answer based on language knowledge alone. Each question includes a "locating prompt" so that a capable system or human can locate the needle frame from the video haystack and answer the question.
When testing LongVA with visual inputs of up to 3000 frames, one difficulty we encountered was that processing a 200K-token input requires up to 100GB of GPU memory for the KV cache for a 7B LM like LLaMA. Even with advanced LM serving systems like vLLM with tensor parallelism to shard the KV cache across multiple GPUs, the sampling process remains extremely slow due to limited memory and batchsize. To address this, we used "perplexity-based" evaluation to measure the correctness of the model output. We first encode all frames and save their corresponding visual embeddings. During the evaluation, we only load the language model from LongVA and concatenate the visual embeddings, question tokens, and answer tokens for a single forward pass with ring attention. This approach makes the workload compute-bound and eliminates the need to cache the KV state. The model’s output is considered correct only if the highest output logits index of all tokens in the answer span matches the correct answer.
Experiments
We primarily assess the long visual capability of LongVA on two benchmarks: V-NIAH (Section 5.1 and Video-MME (Section 5.2). V-NIAH provides quick signals about the visual context length of LongVA. However, it only tests the model’s ability to retrieve information and does not cover other abilities necessary for a real-world long video assistant. Therefore, we also include LongVA’s performance on Video-MME, a comprehensive evaluation suite for video LMMs that includes diverse data types and qualitative annotations. Video-MME is an ideal benchmark for assessing LMMs’ ability to handle long videos in real-world scenarios, given its average video duration of 1017 seconds and the inclusion of short, medium, and long subsets. We further include the benchmark results on MLVU in Appendix C.
We mainly compare LongVA against other image and video LMMs. To validate the phenomenon of long context transfer, we trained LLaVA-Next-Qwen2, a baseline model based on Qwen2-7B-Instruct using the LLaVA-NeXT training recipe. Additionally, we trained LongVA (AnyRes) to showcase the advantages of our UniRes encoding scheme. The difference between LongVA and our baselines can be found in Table 5.
Long context transfers from language to vision Figure 4 shows the V-NIAH performance of LongVA and other LMMs. Specifically, Figure 4 (iii) demonstrates that the visual context length of LLaVA-NeXT-Video-32K is constrained by the 32K context length of its language backbone, Mistral-7B-Instruct-v0.2 , equivalent to approximately 200 frames. Beyond this limit, the V-NIAH accuracy drops significantly. As a stronger baseline, we include the results of LLaVA-NeXT-Video-32K enhanced with a training-free length extrapolation algorithm by increasing its RoPE base frequency. We empirically determine the optimal extrapolation frequency by choosing from [3M, 10M, 30M, 100M, 300M, 1B]. As indicated in Figure 4 (iv), although this training-free extrapolation allows the model to process information across an extended context, the improvement is marginal. These findings led us to develop LongVA, a model that unlocks the visual context by extending the language model purely on text. As shown in Figure 4 (i), LongVA can almost perfectly retrieve information and answer the needle question for input frames fewer than 2000. Although we only trained LongVA’s language backbone on a context length of 224K (equivalent to 1555 frames), it generalizes well beyond that, maintaining satisfactory performance within 3000 frames. Those results clearly corroborate of hypothesis of long context transfer.
Unified encoding enables better visual context extrapolation We also present the V-NIAH heatmap of LongVA trained with AnyRes encoding scheme, keeping all other factors unchanged in Figure 4 (ii). LongVA-AnyRes demonstrates strong retrieval capabilities. However, its performance still lags behind LongVA trained with UniRes. We believe that the unified representation of images and videos in UniRes, where a video is encoded in the same way as a long image, enhances the long context transfer from language to vision. This approach also facilitates effective training with short vision data (images) and enables zero-shot understanding of long videos during inference.
2 Video Evaluation
On Video-MME (Table 4), LongVA achieves state-of-the-art performance among LMMs under 10B parameters, rivaling much larger ones such as LLaVA-NeXT-Video-34B and InternVL-Chat-V1.5 . Notably, LongVA is trained without any video data, so its performance on video can be considered zero-shot. As the number of sampled frames increases, LongVA shows improved performance on the long subset, handling up to 384 framesWe limited our analysis to 384 frames due to computational and memory constraints as detailed in Section 4.. Even though LongVA’s score slightly drops when we upsample from 128 to 384 frames, it maintains a competitive performance. To our knowledge, LongVA is the only open-source model that can handle such large input frames on Video-MME. These findings highlight the long context transfer effect, where LongVA, originating from a long context language model, can process significantly more frames than its baseline, despite being trained on the same multimodal data.
We also tested LongVA on shorter benchmarks with average video durations under 120 seconds. As indicated in Table 6, although LongVA scores higher with more densely sampled frames on datasets such as NeXTQA and ActivityNetQA , the gains quickly plateau and are not as significant as those observed in Video-MME, which can be attributed to the shorter duration of these datasets. On the VideoChatGPT and Video Detailed Description (Video-DD) benchmarks, increasing frames does not lead to better performance, and LongVA generally achieves lower scores compared to LLaVA-NeXT-Video-7B. Since both benchmarks use OpenAI’s GPT API as a judge, we believe their metrics are closely related to the answering format. To address this, we perform a lightweight Direct Preference Optimization (DPO) on the LLaVA-Hound-DPO dataset. We observe significantly improved performance for LongVA-DPO, confirming the findings in .
3 Image Evaluation
We further evaluate our model on various image benchmarks to investigate the image performance of LongVA (Table 7). Compared to the LongVA (AnyRes) baseline, LongVA with UniRes achieves significantly increased performance on InfoVQA , while the scores drop to some extent on AI2D and ChartQA . To better understand this phenomenon, we recorded and analyzed the image size of those datasets, as shown in Figure 5. We found that InfoVQA consists of higher-resolution images, while many images in AI2D and ChartQA are smaller than 768768. Compared to Anyres, UniRes operate 22 average pooling on each image, reducing to visual tokens per image grid. However, the grid upper bound is set to 49 for UniRes while 4 for AnyRes, so UniRes may produce more image grids if the input images are of higher resolution. By using more grids per image, UniRes allocates more visual tokens on datasets such as InfoVQA, achieving superior performance compared to the previous 7B LLaVA model. However, most of the images in ChartQA and AI2D require fewer than 4 grids to represent. This may explain why the image performance decreases on those benchmarks.
Qualitative Results
The qualitative results of LongVA-DPO are illustrated in Figure 6. The short video example comes from and the two long videos are sourced from link1 and link2, respectively. In the figure, LongVA accurately describes the short, humorous video involving individuals playfully interacting with condiments. It also identifies specific details in long videos, such as the color of a train and the colors of umbrellas used in a scene, showcasing its proficiency in retrieving and interpreting visual information over extended video contexts. These capabilities highlight LongVA’s potential to overcome the challenges associated with processing and understanding extremely long videos.
Conclusion
This work addresses the challenges of understanding long videos in Large Multimodal Models. By extending the language model on text and then aligning this extended model with visual inputs, we significantly improved the capability of LMMs to handle long videos thanks to the long context transfer phenomenon. Our model, LongVA, shows improved performance with more input frames and achieves state-of-the-art results on Video-MME. Additionally, we introduce a synthetic benchmark, V-NIAH, to effectively measure the visual context length of video LMMs. We hope this work inspires further research in the field of long video LMMs and multimodal agents.
Acknowledgements
This research/project is supported by the National Research Foundation, Singapore under its AI Singapore Programme (AISG Award No: AISG2-PhD-2022-01-029). Besides, this project is supported by NTU NAP, MOE AcRF Tier 2 (MOE-T2EP20221-0012), and under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as advise from the industry partner(s).
References
Appendix
Appendix A Needle In A Haystack Test
When evaluating the Needle In A Haystack task , we focus specifically on an easier-to-evaluate variant that involves identifying and retrieving random numbers associated with various randomly assigned cities from the context. The input to the language model has below template:
We insert a needle with the key Singapore and a 7-digit randomly sampled magic number as the value into the haystack of Paul Graham’s Essays. The needle has the following format:
We iterate over various document depths (where the needle is placed) and context lengths to measure the performance. For each depth and context length, we conducted the test 5 times, each time with a different 7-digit needle. We also come up with a harder version where we also insert several (3 or 5) other needles with the same format but different city name as distractors. The results are shown in Figure 7.
Appendix B UniRes Encoding Scheme
Figure 8 indicates the difference between AnyRes and UniRes. Given a high-resolution image and assuming we use CLIP-ViT-L-336px as the vision encoder, both AnyRes and UniRes will divide it into multiple grids, each with the size 336x336. However, AnyRes will have a smaller version of the full image as the base image and prepended before the high-resolution image grids. Additionally, UniRes flattens the encoded image feature in a raster-order within each grid, while AnyRes combines all the grids as a big feature map and flattens them across the border of the grid. UniRes also apply 2x2 average pooling on the image feature. As shown in the rightmost part of Figure 8, the design of UniRes allows us to unifiedly encode videos as well. A video is treated as an extended image where each frame is considered as an image grid.
Appendix C MLVU Results
Table 8 includes the evaluation results by the authors of MLVU on their benchmark. LongVA achieves state-of-the-art results among open-source models and is only second to GPT-4o.
Appendix D Visual Needle In A Haystack Test
Table 9 lists the five VQA needles we used for V-NIAH. The 5 visual questions and answers are the only places where human annotation is involved in the construction of V-NIAH, making it an ideal testbed to benchmark LMMs’ long context capability.